Livaramed's patent-pending multi-provider adversarial AI placed the correct diagnosis in its top 3 on 274 of 300 New England Journal of Medicine Case Records — and at rank #1 on 58.3%, versus GPT-4's 38% top-1 on the same dataset.
Published medical AI diagnostic benchmarks on NEJM Case Records of the Massachusetts General Hospital, each as reported by its own study — metrics and protocols differ (see note below).
| System | Developer | Architecture | Cases | Protocol | Accuracy |
|---|---|---|---|---|---|
| CrucibleDx | Livaramed | Multi-provider adversarial (Opus + Gemini) | 300 | Single-pass, full case | 91.3% (top-3) |
| MAI-DxO | Microsoft / OpenAI o3 | Multi-agent ensemble | 304 | Interactive sequential | 85.5% (own metric) |
| AMIE | Google / DeepMind | Single model (Gemini-based) | 302 | Single-pass, HPI only | 59.1% (top-10) |
| GPT-4 | OpenAI (Kanjee 2023) | Single model | 70 | Single-pass, full case | 39% (top-1) |
| GPT-4 | OpenAI (Savage 2024) | Single model | 301 | Single-pass, full case | 38% (top-1) |
| Physicians | (MAI-DxO study) | Human experts | 56 | Interactive sequential | ~20% (top-1) |
Accuracy = correct diagnosis in top-3 differential (score ≥4/5). Savage et al. 2024 dataset (CC-BY-4.0) used for Livaramed and GPT-4 (Savage) benchmarks. MAI-DxO used 304 independently collected NEJM CPC cases with interactive sequential history-taking (AI asks questions); physician accuracy measured on 56 held-out test cases. AMIE used top-10 differential on HPI-only presentation. Livaramed and GPT-4 (Savage) use single-pass full-case protocol. Direct cross-study comparison has inherent limitations; presented for context. See Methodology below.
Each case scored on a 5-point Likert scale by an independent LLM judge. Scores of 4 or 5 indicate a correct diagnosis.
Every NEJM Case Record we analyzed, what the ground truth was, what our AI diagnosed, and how it scored. No cherry-picking — every case, every result.
| # | Case | Score | Result |
|---|
Our benchmark follows reproducible methods on a publicly available dataset.
Claude Opus generates initial differential diagnosis with 7-9 ranked hypotheses
Google Gemini independently challenges each hypothesis, probing for blind spots
Claude Opus integrates the critique, re-ranks diagnoses, produces final differential
Claude Sonnet scores the output against ground truth (1-5 Likert scale)
300 NEJM Case Records of the Massachusetts General Hospital from the Savage et al. 2024 dataset (CC-BY-4.0). Cases span 2015-2023, covering the full spectrum of diagnostic medicine.
Two AI providers (Anthropic Claude + Google Gemini) in a patent-pending adversarial arrangement. The critique step catches diagnostic anchoring bias and forces consideration of alternative hypotheses. Protected under U.S. Patent Application No. 64/071,743 (70 claims, 25 figures).
5-point Likert scale: 5 = exact match at rank #1, 4 = correct in top 2-3, 3 = correct category or lower rank, 2 = related but incorrect, 1 = completely wrong diagnostic direction.
Each case is presented once in full (history, exam, labs, imaging) with no interactive follow-up questions. This matches real clinical decision-making where the AI must reason from available information.
A protective mechanism ensures the adversarial critique step cannot improperly demote the primary model's top hypotheses — preventing diagnostic gaslighting between AI providers.
A separate model (Claude Sonnet) scores each case against the published ground truth. The judge has no access to the diagnostic pipeline's internal reasoning — only the final output and the answer.
Five years of misdiagnosis. Unnecessary heart surgery. Daily suffering. I built Livaramed to help the millions of others going through the same thing.