AI Assistants Find Far More of the Pharma Landscape with Sleuth MCP
We ran five pharma BD and strategy questions through Claude, ChatGPT and Perplexity, with and without Sleuth MCP. Sleuth lifted entity coverage in every case.
Ask a frontier AI assistant like ChatGPT a typical pharma BD or S&E question:
"What is the current competitive landscape (including every program, all phases of development, actively developed assets only) for drugs targeting IL-23 or synonym targets and subtypes."
The AI assistant uses what it has been trained on and all the tools at its disposal to deliver a fluent, well-organized, and confident answer. It will name the obvious programs easily found on the web. It will also miss some assets, and misattribute or mislabel a few more.
We demonstrated that providing access to a curated, high-quality, and comprehensive pharma and biotech knowledge base dramatically improves the performance of AI assistants in answering complex questions. We ran five questions through three AI assistants (Claude, ChatGPT and Perplexity), twice each: once as any user would use them, and once with access to the Sleuth knowledge base through the Sleuth MCP connector.
Assessing the 15 pairs of responses, we observed that:
- More assets in every class. In all cases, Sleuth increased the unique entity* coverage in the AI assistant responses, by an average of 63% and a median of 40%.
- Entirely new classes surfaced. Sleuth not only added more of the same type of entities, it surfaced classes that the AI assistants missed. For example, Sleuth added the full TL1A × IL-23p19 bispecific field and the oral IL-23R challengers, increasing Claude’s output quantity by 97% when answering the IL-23 landscape question.
- Analytical depth. AI assistants generally returned the right column headers; however, the analytically challenging fields arrived empty or as placeholders. Sleuth populated those challenging datapoints. For example, in the response to the I&I deals question, Claude’s royalty terms column said "not disclosed" on 52% of rows, while Claude + Sleuth carried real content on 89% of rows.
- No data-access dead ends. Sleuth-enabled responses included credible data elements in all cases, whereas AI assistants reported difficulty finding the right data in a few cases. For example, in response to the I&I deals question, Perplexity started with: "I can’t truthfully provide an exhaustive, every-deal dataset from public web search alone. Comprehensive partnering databases are generally paywalled," and returned 18 rows based on a "source-supported set of qualifying deals located in public announcements." With Sleuth, Perplexity’s deals coverage increased by 150%.
- Improved data integrity. When Sleuth was used, AI assistants performed better at entity resolution and deduplication, and at timeline calculations in forecasting. For example, ChatGPT kept CPO102 separate from SYSA1801/EO-3021 despite them being the same CSPC ADC. With Sleuth, they were folded into one row (as was LCAR-C18S into LB1908), and unresolved identity was stated in-cell rather than left empty.
* Entity refers to the primary grain or dimension of the generated dataset used to answer the question; e.g., "programs" in the case of the IL-23 question above.
Experiment design
The goal of this experiment was to measure the impact of access to the Sleuth knowledge base on the output of three frontier AI assistant applications, with the following configuration:
- ChatGPT: GPT-5.6 Sol, Effort: High
- Claude: Opus 5, Effort: High
- Perplexity: Best
To do that:
- We chose five questions representing common themes explored by Sleuth users:
- Scouting assets: Identifying all China-origin assets whose primary molecular target is CLDN18.2, across all modalities and stages.
- Listing deals: Finding all BD deals announced from January 1, 2020 through today where the primary therapeutic area is immunology & inflammation and the disclosed upfront consideration exceeds US$50M.
- Market analysis: Modeling the market share and erosion path of BTK inhibitors in CLL over the next three years.
- Identifying whitespace: Presenting the top oncology modalities primed to take over in the next five years given market access dynamics, with a focus on EGFR targets.
- Comprehensive competitive landscape: Finding all IL-23 (and synonym/subtype) programs across all phases of development, focusing on actively developed assets only.
- We varied complexity and specificity across the questions to represent the different styles and formats of typical user prompts. For example, the I&I deals question listed a comprehensive set of columns to return, while the IL-23 question remained high level.
- The questions were given to the AI assistants in two settings: with and without Sleuth MCP connectivity. The output of each run, including all generated artifacts such as datasets, text, and visuals, was captured and anonymized so that nothing linked the output to the specific application.
- The outputs were compared quantitatively on entity coverage.
Note that to simulate the user experience and leverage the tools these AI assistant applications have access to (e.g., web search), we used the applications themselves rather than calling the LLMs via their platform APIs.

Experiment results
Before we dive into the results, a few key observations:
- The Sleuth knowledge base was exposed to the AI assistants via an MCP server. Our experiment did not tell the assistants whether or how to use the MCP. In fact, we observed that each AI assistant crafts its own prompt for Sleuth based on what it believes it needs. Some questions were passed to Sleuth as is; for others, the AI assistant generated a detailed prompt, and in some cases adjusted the scope of the question. For this reason, Sleuth’s response varied for the same user question.
- Engagement with Sleuth varied across questions and AI assistants. In some cases the AI assistant was satisfied with a single call; in others it took multiple calls, as many as six, before wrapping up.
- The Sleuth agent behind the MCP returns a full analysis, including the dataset and its own view of the prompt it was given. The AI assistant then decides what portion of Sleuth’s data and analysis to include in the final answer. In some cases, the AI assistant independently verified Sleuth’s data and analysis. The models’ reasoning indicated that multiple factors influence what Sleuth content makes it into the final answer, including prior skills and memories, LLM variation, and corroboration of specific data elements by what the AI assistant finds on its own.
We focused on unique entity coverage, which measures the comprehensiveness of the recalled data used to answer the user’s question. Without exception, Sleuth improved the entity coverage of all three products across all five questions. The relative improvement ranged from 6.2% to 194.7%.

We calculated summary statistics for the lift contributed by Sleuth per question and per AI assistant.
Discussion
The fundamental challenge AI assistants face in answering a question like the global IL-23 landscape is that the required information is scattered across hundreds of heterogeneous data sources of uneven quality, exposed through different interfaces, and in several languages. Some assets have limited presence in the public domain (e.g., a preclinical candidate may be mentioned only in a single publication or patent), while others appear many times under scientific names, brand names, trade names, aliases, and more.
Finding the data is only the first step. The insight comes from connecting the pieces: target to asset, asset to trial, trial to owner, owner to deal, deal to region, and so on. Resolving syntactic and semantic conflicts and duplicates takes fine-tuned algorithms, advanced domain-specific classifiers, a large amount of compute, and a rigorous quality control process. AI assistants have seconds to minutes, and a limited budget, to perform these tasks while answering a complex question in real time. So the quality of their answer depends on which content they happened to retrieve via their tools, the capability of the underlying LLM, and the phrasing of the user prompt.
A curated knowledge base like Sleuth’s removes that dependency. The entities are already resolved, the relationships already extracted, and every claim carries its source, date, and calculated confidence score. Given structured, cited evidence, today’s LLMs reason over it very well, and they do so consistently whether you’re using Claude, ChatGPT or Perplexity. The insight was already in the data delivered to the AI assistant.
Want these results in your own AI assistant? Talk to us about connecting Sleuth MCP.
