The AI narrative in life sciences has been relentlessly optimistic. More data. Smarter models. Faster insights. And yet the actual output, therapies reaching patients, hasn’t improved at the same speed. If anything, the industry has gotten less productive during the AI era. Why? Because AI in pharma has largely been deployed to answer a question nobody actually asked. Teams stand up pilots that impress in a demo and fail the moment they hit a real workflow. They build prototypes that run ten times and return ten different, uncitable answers. They clear the tech bar and fail the decision bar.
The sequence error is almost always the same: most AI investments start with a model, and almost none start with the question that matters. Do I know what a good answer looks like? Do I have the right foundation to inform a model? Do I understand how this decision actually gets made? If you can’t answer those questions first, you’re not building AI for decisions. You’re building AI for theater.
The Echo Chamber Problem
There’s a more insidious issue underneath the pilot problem: the more content AI generates, the more that content feeds back into itself. For example, a team asks a generic LLM to map a competitive landscape. The model returns a confident, well-formatted answer — and it’s wrong. Drugs that were retired are listed as active. Drugs that just received approval are missing entirely. The model has no data scope, no transparency on what it actually knows, and no mechanism for a subject matter expert to validate the output efficiently. And an SME has to know enough to spot what’s wrong before they can trust anything that’s right.
This isn’t a failure of the underlying model’s raw capability. It’s a failure of context. The AI doesn’t know what it doesn’t know and has no grounding in the actual decision workflow. It wasn’t built to make this call; it was built to generate plausible text.
The result is an echo chamber: confident-sounding content that lacks the context, currency, and transparency that a high-stakes decision actually requires.
High-Stakes Decisions Require High-Stakes Clarity
The tools that move from demo to decision are the ones built with what I call the Clarity Criteria, which includes four questions:
- Cited: Can you trace every output to a source? Every time, not sometimes.
- Consistent: Does the grounding hold across queries? Answers can and will vary, but the context layer underneath them must not.
- Contextual: Is the answer grounded in how this decision actually gets made, in this organization, in this workflow, or is it a generic RAG dump?
- Consequential: Does the output move you toward a decision, or does it just sound confident?
Fail any one of those four, and you have a demo. Not a decision.
The Real Infrastructure Problem
The organizations seeing real productivity gains from AI aren’t the ones with the biggest models. They’re the ones with the smartest context underneath them: structured, longitudinal, domain-specific data that’s been trained not just on what’s publicly available, but on what the actual job requires.
A competitive intelligence analyst doesn’t just need to know what’s in the market. They need to know what regulatory events are pending, which trials are active or discontinued, what payers are doing, and how all of it connects to the specific decision on their desk today. That’s not something a general-purpose model has been trained to do. It requires expertise embedded in the system itself, not just retrieved from it. The agent should understand the goal, and what’s needed to clear the hurdle for that goal. We need to understand what good looks like from both a data/infrastructure standpoint and from an output and decision criteria standpoint.
We can’t expect AI to replicate the sophistication of a competitive intelligence analyst, a market access specialist, or a forecaster designing a launch model, unless the AI has been built against the right data, by people who actually do the job.
So, before the next model or pilot, run the Clarity Criteria test on what you already have:
- Can you trace every output to a source?
- Does the grounding hold across queries?
- Does it match how the decision actually gets made?
- Does it move you toward a decision?
If you can answer yes to all four, you’re not just building AI. You’re building something you can actually scale, deploy, and trust. The paradox isn’t a model problem. It’s a decision problem. And the clock the patient is on is the only clock that matters.
Norstella Atlas combines Norstella’s proprietary data assets with purpose-built AI agents to deliver decision-ready outputs across the drug development lifecycle. Want to learn more? Contact us.