Connect with us

AI

Clinical Trial Sites Need Queryable Data Before Custom AI

Most clinical trial sites should warehouse eSource data before custom AI, a split that now tracks ICH audit rules.

Published

on

Clinical trial sites racing to buy or build AI tools got a colder order on 16 July 2026: fix the data stack first. CRIO hosted the panel that day, then posted a follow-up playbook on 2 September 2026 after attendee questions piled up.

Operators already running models said warehouse access, not the model brand, decides whether custom AI is even worth attempting.

Site Operators Put Buy-Versus-Build Last

Mike Wenger, CRIO’s chief innovation officer, moderated. Aneesh Vaze, managing director of Clinical Research Philadelphia, Sam Stein, CEO of ALSA Research, and Nick Spittal, chief operating officer of Velocity Clinical Research, joined him. Velocity runs 70-plus sites with 200-plus investigators, which matters later, because one voice on the call already works from a shared stack.

Attendees kept asking whether to build tools in house or buy them. The panel treated that as a maturity test. AI shows up in three forms: features inside platforms sites already pay for, point products for recruitment or scheduling, and homegrown scripts on top of local data. All three have a job. Homegrown work only pays once records are clean and staff already use a model every day.

Operators mapped a path in which paper-based sites should digitize first, write a basic AI policy, and hand staff enterprise LLM accounts before anyone codes a custom workflow. Skip that order and the cleanup usually exceeds the time you thought you would save.

THE PATH THE PANEL MAPPED FOR PAPER SITES

  • Electronic records: Move source and regulatory files off paper into eSource and eReg before any model work.
  • A written policy: Set a basic AI governance rule so staff know what they may paste into a prompt.
  • Enterprise accounts: Issue company LLM seats with training protections, not personal logins.
  • Small internal tools: Automate one painful bottleneck, then iterate with a human reading every draft.
  • Backend access last: Open warehouse queries and heavier workflows only after those steps hold.

That sequence is dull on purpose. The failure mode the panel kept circling was the opposite: one automation that swallows documents, chat logs, and system data, then spits out emails, reports, and dashboards in a single pass, which no one can trust or repair when it breaks.

The Warehouse Layer Comes Before the Chat Window

Wenger had already made the data case in an 7 April 2026 post, updated the same month after compliance readers pushed back on how consumer AI plans handle training data. His claim was blunt, and it is the second-order split hiding under the buy-versus-build talk.

Most site technology platforms don’t give you a way to connect your data to AI tools. Your data is locked inside proprietary systems with no clean path out.

Mike Wenger, Chief Innovation Officer, CRIO, 7 April 2026 post

CRIO says visit data already flows into BigQuery and Looker for its customers, along with SiteApp activity, enrollment metrics, and other operational tables. Looker is the path Wenger recommends for operations leads. BigQuery is the path for teams that can write SQL and will sign a business associate agreement before they point a model at the full account.

LOOKER VERSUS BIGQUERY FOR SITE DATA

Path Best for Setup time What gets exposed
Looker Operations leads and site managers Minutes Only the Looks, or reports, you choose to share
BigQuery Analysts and custom reporting Hours to days The full dataset in the account

The Looker route is a scheduled export to Google Drive, then a project in a tool such as Claude that always reads the newest file. Sites are already using that pattern for revenue forecasts from visit data, enrollment trends, data-quality flags, and plain-English questions that would otherwise need a report build. That is not a protocol engine. It is operational data the site already owns, which is why it is the layer custom AI actually needs.

Plan tier is the compliance decision people skip. On Anthropic’s commercial and enterprise Claude seats, the company says it does not train on customer data. Free, Pro, and Max consumer plans may use inputs for training unless the user opts out. The same product name, different legal facts. Wenger’s updated post told sites to confirm training rules, retention, and whether a BAA exists before any file that could touch PHI leaves the building.

Platforms that never expose a warehouse leave staff copying spreadsheets into a chat window. That is not a model problem. It is a plumbing problem, and it is why “build our own AI” so often becomes a second data-entry job.

What ICH E6(R3) Now Requires of Site Systems

The revised Good Clinical Practice guideline does not publish a chapter titled AI. It still decides how far a site can go. Computerised systems used in trials must be fit for purpose, with controls in proportion to risk to participants and to the reliability of results. Investigators who deploy their own systems, including an AI algorithm used to screen patients or measure endpoints, have to keep those systems validated and access-controlled.

That is why a site that turns on a homegrown screener before it can show which data entered the model is already behind the guideline, even if the tool feels helpful on a busy clinic day.

THE RULES THAT CAUGHT UP WITH SITE AI

  1. 6 January 2025: ICH endorses E6(R3) at Step 4.
  2. January 2025: FDA issues draft guidance on AI used to support regulatory decisions for drugs and biologics.
  3. 23 July 2025: E6(R3) principles and Annex 1 take effect in the EU after ICH and CHMP adoption.
  4. 14 January 2026: FDA and EMA release a joint set of 10 principles for AI in drug development.
  5. 16 July 2026: CRIO’s site-operator panel tells clinics to sequence data work ahead of custom tools.
  6. 2 September 2026: CRIO publishes the playbook answering the leftover attendee questions.
  7. 15 January 2027: E6(R3) Annex 2, covering pragmatic, decentralised, and real-world-data designs, takes effect.

The ICH E6(R3) principles and Annex 1 have been in force since 23 July 2025. A consolidated text that folds in Annex 2 was posted on 15 July 2026. Annex 2 itself takes effect on 15 January 2027. Sites that treat AI as a side experiment still have to document it as a computerised system if they deploy it on trial work. The documentation burden sits with the people using the tool, which is the site, not the model vendor.

CV Drafts Work, Eligibility Screens Do Not

Where the panel said AI is already earning its keep, the jobs are narrow. Drafting and updating investigator CVs was the concrete example. Staff used an enterprise LLM account, described the current manual process in plain language, and iterated from a rough draft instead of trying to finish the whole file in one shot. A human read every version. Claude came up most often in the room. Gemini and NotebookLM showed up for other tasks. No one argued there was a single correct model.

That pattern, small scope plus a person in the loop, is the whole practical recommendation. It also shows what the tools are good at: turning a known, repetitive document into a faster first draft. It does not show that a site is ready to let a model decide who qualifies for a study.

Inclusion and exclusion criteria, and prohibited medications, were flagged as a poor fit. Those lists change from protocol to protocol and do not follow one pattern. Feed that variability in and the output is a confident misread. The failure looks a lot like chatbots that stumble on clinical rules after they handle simpler health questions. Recruitment is in a similar bucket. Most sites on the call were not wiring agents straight into patient files. They were plugging point products into vendor ecosystems for outreach and pre-screening, then reviewing the queue themselves.

A monitoring visit will not grade your prompt style. It will ask how a patient got past screening. If the answer is a model that cannot show its inputs, the file is already a finding waiting to be written.

A January Rulebook Covers Site-Built Tools

FDA’s Center for Drug Evaluation and Research spent 2016 to 2023 looking at 500-plus submissions with AI components, then took 800-plus comments on a May 2023 discussion paper. The January 2025 draft guidance that followed, still described as draft on CDER’s AI page, is aimed at AI used to produce information for regulatory decisions on safety, effectiveness, or quality. It is written for sponsors more than for a two-coordinator clinic. The joint FDA and EMA note from 14 January 2026 is broader, and it is the text that maps onto site tools.

The agencies’ ten guiding principles for AI practice put human oversight, risk-based validation, and data provenance in writing. Three of those ten are the ones a site actually has to live with when a coordinator pastes a visit log into a chat window.

THE THREE PRINCIPLES THAT BIND A SITE TOOL

  • Human-centric by design: The model has to sit inside a process a person can overrule, not replace the person who signs the record.
  • Risk-based approach: A CV draft needs lighter controls than a screener that affects who enters a trial.
  • Data governance and documentation: Provenance, processing steps, and analytic choices have to be traceable and verifiable, in line with GxP.

Principle 6 is the one that turns a warehouse from a nice dashboard into a compliance asset. If you cannot say which table fed the prompt, you cannot defend the output. An AI result that cannot be tied to a named reviewer, a timestamp, and the data that produced it will not survive a monitoring visit. That is true in a hospital ward and it is true in a research clinic, and it is cheaper to design now than to reconstruct after a finding.

Independent Sites Still Retrace Data by Hand

eSource adoption is wide in vendor surveys and thinner when you count all US clinics. That gap is the industry’s quiet split, and it is why “just use AI” lands so differently at a 70-plus-site network than at a single independent shop.

WHERE SITE DATA STANDS

  • Survey eSource use: 74% of 255 professionals in RealTime eClinical Solutions’ October 2025 survey use eSource in trials.
  • Full-study use: Nearly 40% of those respondents use it for 100% of active studies.
  • Re-entry tax: Over 60% still spend more than 10 minutes per visit transcribing eSource into EDC.
  • US-site estimate: CRIO puts eSource adoption at 35% of US-based research sites, concentrated among professional networks.

Those figures can sit together because they measure different groups. RealTime asked people who answered a vendor survey, more than half of them from independent sites. CRIO estimated the US site universe. Both point the same way. Electronic source has spread. Queryable, governed, AI-ready data has not. A clinic can be “on eSource” and still copy values by hand for more than 10 minutes after each visit, which is not a dataset a model should touch.

Velocity already scores protocols against its own longitudinal data, investigator history, and enrollment metrics, because those sites share a stack. Andrew Reina, the network’s chief revenue officer, described that tool on 30 January 2026 as a move from feasibility-as-opinion to feasibility-as-evidence. Independent clinics do not have that corpus. They have charts, a CTMS login, and a coordinator who is already behind on queries. Telling them to wait on custom AI is protective. It is also how networks that already warehouse their visits pull further ahead on forecasting, staffing, and study pickup, while paper sites are still told to buy an enterprise chatbot and draft CVs.

The Audit Trail Has to Name a Person

The panel’s test is simpler than a vendor bake-off. For any AI-assisted file a site produces, someone has to name the data that went in and the person who checked what came out. ICH E6(R3) already asks for that kind of record on computerised systems. The January 2026 FDA and EMA principles ask for it on AI. Consumer chat windows do not produce it. A warehouse plus an enterprise contract plus a human signature can.

Sites that still run paper should digitize. Sites that already capture eSource and still retype it should stop treating that export as a model feed. Sites that can query their own visits can try narrow jobs with a reviewer attached. Custom platforms come after that, if they come at all.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending