# Additional file 1: Contradiction recall ladder data

**Table S1. Contradiction recall ladder across human and automated verification protocols.**

| Method | TP/N | Recall | Wilson 95% CI | Notes |
|---|---:|---:|---:|---|
| Human2 focused review (open-book) | 13/23 | 56.5% | [36.0%, 74.9%] | Focused known-contradiction review; not directly comparable to blind models |
| E1-glm-5.1 simple prompt | 5/23 | 21.7% | [9.6%, 42.2%] | Blind full-set LLM evaluation |
| GPT-5.5 simple verifier | 4/23 | 17.4% | [6.9%, 36.8%] | Blind full-set LLM evaluation |
| DeepSeek E1 simple prompt | 3/23 | 13.0% | [4.2%, 29.9%] | Blind full-set LLM evaluation |
| V2.1-glm-5.1 structured prompt | 3/23 | 13.0% | [4.2%, 29.9%] | Blind full-set LLM evaluation |
| DistilBERT-MNLI | 2/23 | 8.7% | [2.8%, 27.2%] | Blind full-set classifier baseline |
| PubMedBERT-MedNLI | 2/23 | 8.7% | [2.8%, 27.2%] | Blind full-set biomedical NLI baseline |
| DeepSeek V2.1 structured prompt | 1/23 | 4.3% | [0.5%, 20.8%] | Blind full-set LLM evaluation |
| Human2 sample30 first pass (blind) | 0/7 | 0.0% | — | Blind human first pass; denominator is 7 gold-contradicted cases in sample30 |
