300-Case NEJM Benchmark

91.3% Top-3 Diagnostic Accuracy
on the Hardest Cases in Medicine

Livaramed's patent-pending multi-provider adversarial AI placed the correct diagnosis in its top 3 on 274 of 300 New England Journal of Medicine Case Records — and at rank #1 on 58.3%, versus GPT-4's 38% top-1 on the same dataset.

91.3%
Top-3 Accuracy
274
Correct Diagnoses
300
NEJM CPC Cases
175
Exact Match (#1)
U.S. Patent Application No. 64/071,743 · Multi-Provider Adversarial Diagnostic Architecture · Filed May 21, 2026

How We Compare

Published medical AI diagnostic benchmarks on NEJM Case Records of the Massachusetts General Hospital, each as reported by its own study — metrics and protocols differ (see note below).

System Developer Architecture Cases Protocol Accuracy
CrucibleDx Livaramed Multi-provider adversarial (Opus + Gemini) 300 Single-pass, full case 91.3% (top-3)
MAI-DxO Microsoft / OpenAI o3 Multi-agent ensemble 304 Interactive sequential 85.5% (own metric)
AMIE Google / DeepMind Single model (Gemini-based) 302 Single-pass, HPI only 59.1% (top-10)
GPT-4 OpenAI (Kanjee 2023) Single model 70 Single-pass, full case 39% (top-1)
GPT-4 OpenAI (Savage 2024) Single model 301 Single-pass, full case 38% (top-1)
Physicians (MAI-DxO study) Human experts 56 Interactive sequential ~20% (top-1)

Accuracy = correct diagnosis in top-3 differential (score ≥4/5). Savage et al. 2024 dataset (CC-BY-4.0) used for Livaramed and GPT-4 (Savage) benchmarks. MAI-DxO used 304 independently collected NEJM CPC cases with interactive sequential history-taking (AI asks questions); physician accuracy measured on 56 held-out test cases. AMIE used top-10 differential on HPI-only presentation. Livaramed and GPT-4 (Savage) use single-pass full-case protocol. Direct cross-study comparison has inherent limitations; presented for context. See Methodology below.

Patent-Pending Technology — U.S. Application No. 64/071,743
Livaramed's multi-provider adversarial diagnostic architecture is protected by a filed U.S. patent application covering 70 claims across the adversarial pipeline, dynamic questionnaire system, model fallback cascade, diagnostic gaslighting detection, and cost transparency mechanisms. Unlike single-model approaches used by other systems, Livaramed's architecture is structurally novel — requiring AI providers from different companies to independently challenge each other's reasoning before reaching a diagnosis.

How Our AI Performed Across All 300 Cases

Each case scored on a 5-point Likert scale by an independent LLM judge. Scores of 4 or 5 indicate a correct diagnosis.

175
Score 5
Exact match at #1
99
Score 4
Correct in top 2-3
14
Score 3
Correct category
9
Score 2
Related field
3
Score 1
Wrong direction
0
Score 0
No response

All 300 Case Results

Every NEJM Case Record we analyzed, what the ground truth was, what our AI diagnosed, and how it scored. No cherry-picking — every case, every result.

All 300
Correct 274
Missed 26
Score 5
Score 4
Showing 300 of 300 cases
# Case Ground Truth Our Diagnosis Score Result

How We Built and Tested This

Our benchmark follows reproducible methods on a publicly available dataset.

Diagnostic Pipeline

Step 1: Primary Analysis

Claude Opus generates initial differential diagnosis with 7-9 ranked hypotheses

Step 2: Adversarial Critique

Google Gemini independently challenges each hypothesis, probing for blind spots

Step 3: Synthesis

Claude Opus integrates the critique, re-ranks diagnoses, produces final differential

Step 4: Independent Judge

Claude Sonnet scores the output against ground truth (1-5 Likert scale)

Dataset

300 NEJM Case Records of the Massachusetts General Hospital from the Savage et al. 2024 dataset (CC-BY-4.0). Cases span 2015-2023, covering the full spectrum of diagnostic medicine.

Adversarial Architecture Patent Pending

Two AI providers (Anthropic Claude + Google Gemini) in a patent-pending adversarial arrangement. The critique step catches diagnostic anchoring bias and forces consideration of alternative hypotheses. Protected under U.S. Patent Application No. 64/071,743 (70 claims, 25 figures).

Scoring

5-point Likert scale: 5 = exact match at rank #1, 4 = correct in top 2-3, 3 = correct category or lower rank, 2 = related but incorrect, 1 = completely wrong diagnostic direction.

Single-Pass Protocol

Each case is presented once in full (history, exam, labs, imaging) with no interactive follow-up questions. This matches real clinical decision-making where the AI must reason from available information.

Demotion Guard

A protective mechanism ensures the adversarial critique step cannot improperly demote the primary model's top hypotheses — preventing diagnostic gaslighting between AI providers.

Independent Judging

A separate model (Claude Sonnet) scores each case against the published ground truth. The judge has no access to the diagnostic pipeline's internal reasoning — only the final output and the answer.

Built Because the Medical System Failed My Daughter

Five years of misdiagnosis. Unnecessary heart surgery. Daily suffering. I built Livaramed to help the millions of others going through the same thing.

Read Our Story The Science