All articles

Biopharma Data in an Age of AI

AI has removed the constraints that forced every team to interrogate the world through the same datasets.

Andrew Pannu
August 17, 2026

Back when I was an analyst manually assembling pipeline, trial or deal data, databases were a helpful way to interrogate the world. The alternative was haphazard googling, or just not getting any data at all. But AI has removed the constraints that forced every team to interrogate the world through the same datasets.

It used to be that aggregating a first-pass landscape was labor intensive. So a database with standardized fields (drug, target, modality, indication, etc.) was a much better starting point. Biopharma BD, strategy and CI teams across the industry adopted several of these tools to answer basic landscape questions like "which clinical assets target Alzheimer's disease?"

The data served to these customers was the same across the board - it had to be. Data construction was expensive, so the dimensions captured needed to be broadly reusable.

But the issue was that nearly all strategic work quickly outstripped this general-purpose data, with questions like:

  • How should we position our asset as the treatment setting evolves?
  • Which companies are the most plausible partners for this program?
  • Which competitors actually matter clinically?

These questions require specific datasets whose fields are designed around the decision that needs to be made - not just what is available from a database's existing menu of filters. Those tend to just define the universe, but the most useful dimensions are those a database was never designed to capture.

For asset positioning, this might mean structuring data around which patient archetype has the most compelling benefit-risk tradeoff, whereas for partner selection it might be around signals that suggest a company is actually prepared to prioritize a transaction now.

AI has dramatically lowered the cost of creating these specific datasets. Models can help decide and populate the parameters a question requires in real time, allowing the schema to reshape itself around the decision.

This flexibility creates new demands of what modern data systems need to do really well: define new fields precisely, retrieve the right evidence for each one and continuously evaluate their outputs as questions, sources and models change. Meanwhile, the foundational requirements such as comprehensive data collection, entity resolution and source-level provenance never went away.

The result is a significantly faster and more precise decision-making loop for life science teams:

  • Instead of asking "what can this database tell us?", teams can begin with "what would the perfect dataset to help inform this decision contain?"
  • Teams can iterate on the question without rebuilding the research from scratch
  • The right dataset can become the shared ground truth for a quarter of strategic work

At Sleuth we build the data foundation and systems that let teams move from an important strategic question to a trusted dataset shaped around it - and keep refining both as the question, evidence and market evolve.

This is the approach behind our intelligence reports: each one starts from a strategic question and builds the dataset around it, rather than filtering a general-purpose database. If there's a dataset your team needs that doesn't exist yet, that's exactly what our custom datasets work is for.

See what your team has been missing