Process Data in Clinical Trials: The Questions Your Stack Can’t Answer

Why your trial systems can’t tell you how a decision was made.

Six Questions Most Sponsors and CROs Can’t Answer

  1. How often do your two independent adjudicators disagree on first read, and is that number moving in the wrong direction?
  2. Which of your eligibility criteria generate the most back and forth before a patient is confirmed, and at which sites?
  3. How long does site qualification actually take, from first contact to activation, and where inside that span does the time go?
  4. Which protocol deviations recur, in what pattern, and how long does classification sit before someone acts on it?
  5. Are your reviewers, monitors, and medical leads applying the protocol consistently with one another, and with themselves six months into enrollment?
  6. On the fifty cases a regulator is most likely to question, what was decided, by whom, and on what basis?

Every one of these questions is answerable in principle. In practice, almost no sponsor or CRO can answer any of them today.

Every System in the Trial Records the Verdict

Clinical trial systems were built to record outcomes. EDC holds the value of the lab result. eCOA holds the patient’s response. IRT holds the assignment. CTMS holds the milestone, the visit, the payment. Each of these does exactly what it was designed to do, and each of them stores a conclusion.

What none of them stores is how the conclusion was reached.

The reaching happened somewhere else. It happened in an email thread with eleven replies, on a call where two physicians talked through a discordant read, in a tracking spreadsheet maintained by a project manager who left the study last year. That work was performed by expert people and it produced the most consequential outputs in the trial. Then it evaporated. What survives is the answer, stripped of everything that would let you evaluate it.

“What survives is the answer, stripped of everything that would let you evaluate it.”

– Abraham Gutman, Founder and CEO, AG Mednet

Two kinds of data are in play here. Outcome data records what was concluded. Process data records how the conclusion was produced: who reviewed, in what sequence, with what information in front of them, how many rounds it took, what came back and why, who disagreed with whom and how the disagreement resolved.

CTMS measures the logistics of running a trial, which is enrollment pace, monitoring visits, site payments. Process data measures the medical and scientific judgment itself, which sits far closer to the science and carries considerably more weight.

Capturing it requires a system that governs the decision while the decision is being made. That system is process management, and it is the layer the clinical trial stack has never had.

One Event, From Trigger to Verdict

Consider a suspected myocardial infarction in a cardiovascular outcomes trial.

EDC triggers the event, and the stack has done its job. Everything after that happens in the space between systems. The site is asked to gather source documents and send them to the CRO. The CRO reviews what arrives, finds a discharge summary missing and an ECG unreadable, and goes back to the site twice before the dossier can be assembled. The completed dossier goes to two adjudicators, who evaluate it independently. They disagree. The case escalates to committee, which resolves it eleven days later.

The final endpoint lands in the analysis dataset carrying none of this. Not the two document requests, not the eleven days, not the fact that two qualified physicians read the same dossier and reached different conclusions.

The verdict and the path to the verdict tell you different things. Only one of them is stored.

The same shape holds across every journey in a trial:

  • Eligibility: A patient’s eligibility gets confirmed, and nothing records how many clarifications it took or which criterion caused them.
  • Site activation: A site gets activated, and nothing records where the eight weeks went.
  • Deviations: A deviation gets classified, and nothing records how long it sat unclassified or who escalated it.

When the Questions Become Answerable

Once these processes run inside a governed system, the six questions stop being rhetorical.

Suppose first-read discordance between your two adjudicators runs at forty percent. The instinct is to treat that as a reviewer problem and address it through retraining. More often it points at the endpoint definition, because criteria that produce consistent disagreement among qualified physicians are criteria that were written ambiguously. That is a protocol finding, surfaced while the trial is still enrolling, and it is invisible in every dataset you hold today.

Suppose the same eligibility criterion generates clarification requests on a third of screened patients. The criterion is either ambiguous or misaligned with local practice, and both are correctable in an amendment rather than absorbed as screen failures for the length of the study.

Suppose document requests to sites cluster at four sites out of ninety, or classification of deviations stretches from three days to nineteen over the course of a year. Both are visible in process data long before database lock, and both are cheap to correct while the trial is still running.

Optimization in most conversations means shortening cycle time. Process data supports something more valuable, which is improving the design of the science itself.

“Process data supports something more valuable, which is improving the design of the science itself.”

– Abraham Gutman, Founder and CEO, AG Mednet

The Same Data, Read Across Trials

Because journeys are structured consistently, process data is comparable across studies. What the comparison yields depends on who holds it.

What sponsors see

A sponsor running one indication through three CROs has never had a way to see how those CROs actually execute. Contracted metrics confirm that milestones were hit. Process data shows whose dossier preparation ran clean, whose site qualification stalled and where, whose turnaround drifted as the trial matured. Vendor selection stops resting on reputation and reference calls.

What CROs see

A CRO working across twenty sponsors accumulates something different and arguably more valuable: a view of which protocol designs execute cleanly and which generate friction. Today that knowledge lives in the heads of a handful of senior project managers, and it walks out the door when they do. Captured as data, it becomes an asset the CRO brings to the next design conversation, backed by evidence from hundreds of trials rather than by recollection.

Process metadata, not clinical data

Both views rest on process metadata rather than clinical data. No sponsor’s trial data becomes visible to another. What travels is the shape of the work: cycle times, disagreement rates, rework patterns, handoff behavior.

This is also the only structured, attributable record of expert clinical judgment anywhere in the trial stack, which makes it the only part of it genuinely worth training a model on. Judi Insights is where that record becomes legible. The record itself matters more, and it exists only because the process ran somewhere it could be observed.

Capture systems taught the industry what happened in a trial. Process management is how it learns how the trial actually ran.

bio_abraham_gutman

Abraham Gutman Founder and CEO, founded AG Mednet in 2005 after selling his previous company to AT&T. He is a technologist with 35 years of experience in the conception, implementation and sales of advanced software systems and platforms in fields including Telecommunications and Life Sciences. He holds a B.A. in Computer Science from Cornell, and an M.S. in Computer Science (AI) from Yale.

Don't forget to share this post

Related Posts