Data quality tells us whether information can be trusted. Data relevance asks a different question: whether that information adequately represents the outcome we are trying to understand.
Clinical development is entering a more demanding phase of AI adoption. The question is no longer confined to what a technology can do in a controlled pilot. Organizations must also decide whether a use case addresses a sufficiently important problem, whether it can operate within clinical-development workflows and whether the available evidence creates enough confidence to change established practice.
These questions sit at the centre of From potential to practice: What will it take to make AI work in clinical trials?, a recent white paper with contributors from Fortrea and Cognivia. The paper examines why promising AI capabilities do not automatically translate into broader operational adoption. Among the conditions it identifies are a clear purpose, relevant data, operational integration, repeatable evidence and stakeholder alignment.
Each condition merits attention. For Cognivia, however, the discussion of data relevance raises a particularly consequential question: How should relevance be defined when the outcome under examination depends partly on patient decisions?
Data can be reliable without answering the right question
Data quality and data relevance are closely related, but they are not interchangeable.
Data quality concerns the integrity of the information available. Is it accurate, complete, consistent and sufficiently reliable for its intended use? Data relevance concerns the relationship between that information and the question an organization is trying to answer.
A dataset may be accurate and methodologically robust while still offering only a partial representation of the outcome under investigation. This does not make the dataset poor quality. It means that the limits of what it represents must be understood before conclusions are drawn from it.
The distinction becomes especially important when AI-supported approaches are used to predict, interpret or influence results. The analytical sophistication of a model cannot, by itself, establish that all materially relevant dimensions of the problem are represented in its inputs. A model may identify patterns within the evidence available to it, but the usefulness of those patterns still depends on whether the evidence corresponds to the decision that needs to be supported.
Relevance, in other words, is not simply a characteristic of the data. It is a relationship between the data, the outcome and the intended decision.
Observable outcomes do not always reveal their drivers
Consider patient participation in clinical trials.
Clinical and operational data can provide essential information about eligibility, protocol characteristics, site activity, study milestones and patterns of participation. They can show what happened, when it happened and where differences occurred. These observations are indispensable to trial design and execution.
Yet patient participation is not produced by an operational process alone. It also reflects decisions made by individuals who interpret information, develop expectations, assess burden and respond within personal circumstances that may evolve over time. The white paper specifically identifies cognitive, emotional and contextual factors as influences on how patients engage with clinical trials. It also observes that algorithms relying solely on clinical or operational data may not fully represent how patients perceive, engage with and ultimately participate in a study.
This creates an important distinction between observing an outcome and understanding the factors that contributed to it.
Operational evidence may reveal that a patient did not enrol, disengaged or discontinued participation. Clinical evidence may describe the patient’s medical profile and treatment experience. Neither necessarily explains how the patient interpreted the study, which expectations shaped the decision, how the burden was evaluated or what changed in the patient’s personal context.
This is not an argument that behavioral information should be introduced into every clinical-development model. Nor does it suggest that clinical and operational evidence should be replaced or treated as inherently incomplete. The relevant question is more disciplined:
If the outcome being examined is materially influenced by patient decisions, are the factors influencing those decisions sufficiently represented in the evidence being used?
The answer will depend on the use case. In some cases, the available clinical and operational data may be entirely appropriate. In others, an important part of the decision context may remain unobserved.
The problem is not necessarily missing data
It would be tempting to interpret this as another demand for more data. That would miss the point.
More data do not automatically create a more relevant representation of a problem. Additional variables can increase complexity without improving decision quality if they are not connected to a clearly defined question.
The more useful starting point is purpose. What decision is the organization trying to support? What outcome matters? Which factors could materially affect that outcome? Which of those factors are currently observable, and which remain inferred or absent?
These questions transform data relevance from a generic technical requirement into a use-case-specific judgment.
For an operational forecasting question, historical site and study-performance data may be central. For a clinical question, biological and treatment-related evidence may be most relevant. For a question involving patient participation, continued engagement or another patient-dependent outcome, the way individuals perceive and evaluate the trial may also warrant consideration.
The objective is not to assemble a complete representation of every patient reality. That would be neither realistic nor necessarily useful. It is to identify whether an omitted dimension is material enough to limit the interpretation or application of the analysis.
Better representation must still meet the conditions for adoption
Recognizing the potential relevance of patient behavior does not resolve the broader adoption challenge. Behavioral understanding must meet the same standards expected of any other source of evidence.
Its purpose must be specific. “Understanding patients better” is not, by itself, an operational objective. An organization must define which decision could improve, which outcome matters and how the information would be used.
The information must also be credible, interpretable and appropriate for that purpose. If it cannot be integrated into an existing decision, role or workflow, it risks becoming another source of insight that remains adjacent to clinical-development practice rather than influencing it.
Stakeholders must also agree on what the evidence means and what conclusions it can reasonably support. The introduction of a behavioral dimension should not create false confidence, extend an analysis beyond its evidentiary limits or imply that complex patient decisions can be reduced to a single explanatory factor.
This brings the discussion back to adoption. The objective is not to add behavioral information simply because it is available. It is to determine where such information is decision-relevant, what evidence is needed to support its use and how it can contribute without creating unnecessary complexity.
A more demanding definition of evidence
As AI becomes more embedded in clinical development, organizations will need to make increasingly deliberate decisions about which applications deserve to progress beyond experimentation.
Technical performance will remain important, but it will not be sufficient. Decision-makers must also consider whether a use case addresses a meaningful problem, whether its outputs can be integrated into practice and whether the underlying evidence adequately represents the outcome being examined.
For patient-dependent outcomes, this introduces a more demanding test of relevance. Clinical and operational evidence may show what occurred, while leaving some of the factors influencing the patient’s decision less visible.
The implication is not that every patient-related outcome requires a behavioral model. It is that organizations should resist treating data relevance as self-evident.
Before asking whether an AI-supported approach can predict an outcome, a more fundamental question may be required:
Does the evidence adequately represent the decision we are trying to understand?That question extends beyond model development. It speaks to the quality of the decisions clinical-development organizations will make about where, why and how AI should be adopted.
Explore the thought-leadership paper.


