How We Build Our Benchmarks
Methodology
Our benchmarks are calibrated against 1,600+ verified biopharma transactions sourced from regulatory filings, public disclosures, and proprietary intelligence. Here's how we turn raw data into actionable deal intelligence.
Data Foundation
Every benchmark in our platform is grounded in real, publicly disclosed transactions. Our database encompasses 1,600+ biopharma deals spanning 2017 through 2026, covering 12 therapeutic areas, 5 deal structures, and 15+ modalities.
Primary Sources
- SEC Regulatory Filings — 8-K material definitive agreements, the filings companies are legally required to submit when entering material licensing, collaboration, or acquisition agreements
- FTC Premerger Filings — Hart-Scott-Rodino Act filings and Federal Trade Commission merger review actions that capture deal activity above reporting thresholds
- Press Release Wires — Deal announcements from global press distribution networks, filtered for biopharma relevance and financial term disclosure
- Regulatory Agency Databases — FDA, EMA, and other regulatory body approval and authorization records that signal commercial-stage deal activity
- Clinical Trial Registries — Partnership and sponsor changes in registered clinical trials that indicate underlying licensing or collaboration agreements
- Proprietary Intelligence Feeds — Web-wide deal monitoring that captures announcements from industry conferences, investor presentations, and non-US company disclosures
Benchmarking Engine
Raw transaction data is transformed into actionable benchmarks through a multi-factor quantitative model that adjusts for the variables that matter most in deal structuring.
Comparable Matching
Comparable transactions are ranked with an additive match score, not a regression. Each candidate deal earns points for every dimension it shares with your asset, and the top-scoring deals become the comparable set. The weights below are read directly from the scoring code:
| Dimension | Points |
|---|---|
| Same therapeutic area | 3 |
| Same clinical phase | 4 |
| Adjacent phase (±1 step) | 2 |
| Same modality | 3 |
| Same indication | 3 |
| Same deal structure | 2 |
| Recency (current year full points, prior year half) | 2 |
| Verifier-confirmed against a primary source | 1 |
| Maximum score | 18 |
A deal must share your therapeutic area and at least one of phase, adjacent phase, or indication to qualify. If fewer than 5 deals qualify, the filter relaxes to therapeutic area + modality, then therapeutic area alone, so you always see the closest available comparables and the panel tells you which rung was used.
Benchmark statistics computed from those comparables (medians and percentile ranges) are recency-weighted with an exponential decay: each deal's weight is 0.5 raised to (years since signing ÷ 2.5). A deal signed 24 months ago therefore carries 0.57× the weight of a deal signed this year (current deals have roughly 1.7× the influence), and a deal from five years ago carries 0.25×. Older deals still count, they just count less.
Monte Carlo Simulation
Every calculation runs 10,000 Monte Carlo iterations. Each iteration first draws a macro scenario (bear, base, or bull) and then samples probability of success, peak sales, discount rate, and timing around that scenario. Scenario weights depend on development stage, because late-stage assets have tighter outcome distributions. The table below is read from the engine at render time:
| Stage | Bear | Base | Bull |
|---|---|---|---|
Early stage Discovery, preclinical, Phase 1 | 25% | 40% | 35% |
Mid stage Phase 1/2, Phase 2, Phase 2/3 | 20% | 50% | 30% |
Late stage Phase 3, NDA/BLA filed, approved | 15% | 60% | 25% |
This produces probability-adjusted ranges rather than single-point estimates, reflecting the inherent variability across different market conditions and negotiation outcomes.
Risk-Adjusted NPV
rNPV analysis incorporates phase-specific probability of success rates, indication-specific modifiers (biomarker validation, regulatory precedent, competitive density, modality risk), and scenario bridges that quantify dollar-impact risks like competitor entry, payer restrictions, and label expansion potential.
Semantic Deal Matching
Traditional deal databases match on keywords — "oncology" finds oncology deals. Our platform goes further with semantic matching technology that understands the full context of each transaction.
Every deal in our database is represented as a high-dimensional vector encoding its complete profile — companies, asset characteristics, modality, indication, development phase, territory, deal economics, and strategic context. When you run a calculation, your inputs are similarly encoded and compared against every transaction using cosine similarity.
This means an "oral GLP-1 receptor agonist for obesity at Phase 2" query will surface deals like Zealand/Roche (petrelintide), Carmot/Roche (CT-388), and Structure/Roche (GSBR-1290) — even if the exact keywords don't overlap. The system finds deals that are structurally and strategically similar, not just categorically related.
Data Quality & Freshness
Stale data produces misleading benchmarks. Our automated ingestion pipeline processes new regulatory filings and deal announcements multiple times per day, ensuring benchmarks reflect the latest market activity.
Understanding Benchmark Ranges
Biopharma deal terms are not deterministic — they are the product of negotiation between parties with different leverage, information, and strategic objectives. Our benchmarks reflect this reality by providing ranges derived from the distribution of comparable transactions, not single-point predictions.
The ranges you see represent where similar deals have historically landed across different market conditions, competitive landscapes, and negotiation dynamics. They are designed to inform your deal strategy and provide data-driven anchor points for term sheet discussions — not to predict the exact outcome of any individual negotiation.
For definitive deal structuring, we recommend engaging qualified financial and legal advisors who can incorporate proprietary clinical data, specific IP considerations, and counterparty dynamics that quantitative models cannot fully capture.