AI
Clinical Trial Sites Need Queryable Data Before Custom AI
Most clinical trial sites should warehouse eSource data before custom AI, a split that now tracks ICH audit rules.
Clinical trial sites racing to buy or build AI tools got a colder order on 16 July 2026: fix the data stack first. CRIO hosted the panel that day, then posted a follow-up playbook on 2 September 2026 after attendee questions piled up.
Operators already running models said warehouse access, not the model brand, decides whether custom AI is even worth attempting.
Site Operators Put Buy-Versus-Build Last
Mike Wenger, CRIO’s chief innovation officer, moderated. Aneesh Vaze, managing director of Clinical Research Philadelphia, Sam Stein, CEO of ALSA Research, and Nick Spittal, chief operating officer of Velocity Clinical Research, joined him. Velocity runs 70-plus sites with 200-plus investigators, which matters later, because one voice on the call already works from a shared stack.
Attendees kept asking whether to build tools in house or buy them. The panel treated that as a maturity test. AI shows up in three forms: features inside platforms sites already pay for, point products for recruitment or scheduling, and homegrown scripts on top of local data. All three have a job. Homegrown work only pays once records are clean and staff already use a model every day.
Operators mapped a path in which paper-based sites should digitize first, write a basic AI policy, and hand staff enterprise LLM accounts before anyone codes a custom workflow. Skip that order and the cleanup usually exceeds the time you thought you would save.
THE PATH THE PANEL MAPPED FOR PAPER SITES
- Electronic records: Move source and regulatory files off paper into eSource and eReg before any model work.
- A written policy: Set a basic AI governance rule so staff know what they may paste into a prompt.
- Enterprise accounts: Issue company LLM seats with training protections, not personal logins.
- Small internal tools: Automate one painful bottleneck, then iterate with a human reading every draft.
- Backend access last: Open warehouse queries and heavier workflows only after those steps hold.
That sequence is dull on purpose. The failure mode the panel kept circling was the opposite: one automation that swallows documents, chat logs, and system data, then spits out emails, reports, and dashboards in a single pass, which no one can trust or repair when it breaks.
The Warehouse Layer Comes Before the Chat Window
Wenger had already made the data case in an 7 April 2026 post, updated the same month after compliance readers pushed back on how consumer AI plans handle training data. His claim was blunt, and it is the second-order split hiding under the buy-versus-build talk.
Most site technology platforms don’t give you a way to connect your data to AI tools. Your data is locked inside proprietary systems with no clean path out.
Mike Wenger, Chief Innovation Officer, CRIO, 7 April 2026 post
CRIO says visit data already flows into BigQuery and Looker for its customers, along with SiteApp activity, enrollment metrics, and other operational tables. Looker is the path Wenger recommends for operations leads. BigQuery is the path for teams that can write SQL and will sign a business associate agreement before they point a model at the full account.
LOOKER VERSUS BIGQUERY FOR SITE DATA
| Path | Best for | Setup time | What gets exposed |
|---|---|---|---|
| Looker | Operations leads and site managers | Minutes | Only the Looks, or reports, you choose to share |
| BigQuery | Analysts and custom reporting | Hours to days | The full dataset in the account |
The Looker route is a scheduled export to Google Drive, then a project in a tool such as Claude that always reads the newest file. Sites are already using that pattern for revenue forecasts from visit data, enrollment trends, data-quality flags, and plain-English questions that would otherwise need a report build. That is not a protocol engine. It is operational data the site already owns, which is why it is the layer custom AI actually needs.
Plan tier is the compliance decision people skip. On Anthropic’s commercial and enterprise Claude seats, the company says it does not train on customer data. Free, Pro, and Max consumer plans may use inputs for training unless the user opts out. The same product name, different legal facts. Wenger’s updated post told sites to confirm training rules, retention, and whether a BAA exists before any file that could touch PHI leaves the building.
Platforms that never expose a warehouse leave staff copying spreadsheets into a chat window. That is not a model problem. It is a plumbing problem, and it is why “build our own AI” so often becomes a second data-entry job.
What ICH E6(R3) Now Requires of Site Systems
The revised Good Clinical Practice guideline does not publish a chapter titled AI. It still decides how far a site can go. Computerised systems used in trials must be fit for purpose, with controls in proportion to risk to participants and to the reliability of results. Investigators who deploy their own systems, including an AI algorithm used to screen patients or measure endpoints, have to keep those systems validated and access-controlled.
That is why a site that turns on a homegrown screener before it can show which data entered the model is already behind the guideline, even if the tool feels helpful on a busy clinic day.
THE RULES THAT CAUGHT UP WITH SITE AI
- 6 January 2025: ICH endorses E6(R3) at Step 4.
- January 2025: FDA issues draft guidance on AI used to support regulatory decisions for drugs and biologics.
- 23 July 2025: E6(R3) principles and Annex 1 take effect in the EU after ICH and CHMP adoption.
- 14 January 2026: FDA and EMA release a joint set of 10 principles for AI in drug development.
- 16 July 2026: CRIO’s site-operator panel tells clinics to sequence data work ahead of custom tools.
- 2 September 2026: CRIO publishes the playbook answering the leftover attendee questions.
- 15 January 2027: E6(R3) Annex 2, covering pragmatic, decentralised, and real-world-data designs, takes effect.
The ICH E6(R3) principles and Annex 1 have been in force since 23 July 2025. A consolidated text that folds in Annex 2 was posted on 15 July 2026. Annex 2 itself takes effect on 15 January 2027. Sites that treat AI as a side experiment still have to document it as a computerised system if they deploy it on trial work. The documentation burden sits with the people using the tool, which is the site, not the model vendor.
CV Drafts Work, Eligibility Screens Do Not
Where the panel said AI is already earning its keep, the jobs are narrow. Drafting and updating investigator CVs was the concrete example. Staff used an enterprise LLM account, described the current manual process in plain language, and iterated from a rough draft instead of trying to finish the whole file in one shot. A human read every version. Claude came up most often in the room. Gemini and NotebookLM showed up for other tasks. No one argued there was a single correct model.
That pattern, small scope plus a person in the loop, is the whole practical recommendation. It also shows what the tools are good at: turning a known, repetitive document into a faster first draft. It does not show that a site is ready to let a model decide who qualifies for a study.
Inclusion and exclusion criteria, and prohibited medications, were flagged as a poor fit. Those lists change from protocol to protocol and do not follow one pattern. Feed that variability in and the output is a confident misread. The failure looks a lot like chatbots that stumble on clinical rules after they handle simpler health questions. Recruitment is in a similar bucket. Most sites on the call were not wiring agents straight into patient files. They were plugging point products into vendor ecosystems for outreach and pre-screening, then reviewing the queue themselves.
A monitoring visit will not grade your prompt style. It will ask how a patient got past screening. If the answer is a model that cannot show its inputs, the file is already a finding waiting to be written.
A January Rulebook Covers Site-Built Tools
FDA’s Center for Drug Evaluation and Research spent 2016 to 2023 looking at 500-plus submissions with AI components, then took 800-plus comments on a May 2023 discussion paper. The January 2025 draft guidance that followed, still described as draft on CDER’s AI page, is aimed at AI used to produce information for regulatory decisions on safety, effectiveness, or quality. It is written for sponsors more than for a two-coordinator clinic. The joint FDA and EMA note from 14 January 2026 is broader, and it is the text that maps onto site tools.
The agencies’ ten guiding principles for AI practice put human oversight, risk-based validation, and data provenance in writing. Three of those ten are the ones a site actually has to live with when a coordinator pastes a visit log into a chat window.
THE THREE PRINCIPLES THAT BIND A SITE TOOL
- Human-centric by design: The model has to sit inside a process a person can overrule, not replace the person who signs the record.
- Risk-based approach: A CV draft needs lighter controls than a screener that affects who enters a trial.
- Data governance and documentation: Provenance, processing steps, and analytic choices have to be traceable and verifiable, in line with GxP.
Principle 6 is the one that turns a warehouse from a nice dashboard into a compliance asset. If you cannot say which table fed the prompt, you cannot defend the output. An AI result that cannot be tied to a named reviewer, a timestamp, and the data that produced it will not survive a monitoring visit. That is true in a hospital ward and it is true in a research clinic, and it is cheaper to design now than to reconstruct after a finding.
Independent Sites Still Retrace Data by Hand
eSource adoption is wide in vendor surveys and thinner when you count all US clinics. That gap is the industry’s quiet split, and it is why “just use AI” lands so differently at a 70-plus-site network than at a single independent shop.
WHERE SITE DATA STANDS
- Survey eSource use: 74% of 255 professionals in RealTime eClinical Solutions’ October 2025 survey use eSource in trials.
- Full-study use: Nearly 40% of those respondents use it for 100% of active studies.
- Re-entry tax: Over 60% still spend more than 10 minutes per visit transcribing eSource into EDC.
- US-site estimate: CRIO puts eSource adoption at 35% of US-based research sites, concentrated among professional networks.
Those figures can sit together because they measure different groups. RealTime asked people who answered a vendor survey, more than half of them from independent sites. CRIO estimated the US site universe. Both point the same way. Electronic source has spread. Queryable, governed, AI-ready data has not. A clinic can be “on eSource” and still copy values by hand for more than 10 minutes after each visit, which is not a dataset a model should touch.
Velocity already scores protocols against its own longitudinal data, investigator history, and enrollment metrics, because those sites share a stack. Andrew Reina, the network’s chief revenue officer, described that tool on 30 January 2026 as a move from feasibility-as-opinion to feasibility-as-evidence. Independent clinics do not have that corpus. They have charts, a CTMS login, and a coordinator who is already behind on queries. Telling them to wait on custom AI is protective. It is also how networks that already warehouse their visits pull further ahead on forecasting, staffing, and study pickup, while paper sites are still told to buy an enterprise chatbot and draft CVs.
The Audit Trail Has to Name a Person
The panel’s test is simpler than a vendor bake-off. For any AI-assisted file a site produces, someone has to name the data that went in and the person who checked what came out. ICH E6(R3) already asks for that kind of record on computerised systems. The January 2026 FDA and EMA principles ask for it on AI. Consumer chat windows do not produce it. A warehouse plus an enterprise contract plus a human signature can.
Sites that still run paper should digitize. Sites that already capture eSource and still retype it should stop treating that export as a model feed. Sites that can query their own visits can try narrow jobs with a reviewer attached. Custom platforms come after that, if they come at all.
-
AI3 months agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
AI2 months agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI3 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
GAMING3 months agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
CRYPTO3 months agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
NEWS3 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
APPS3 months agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
AI3 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
