# Supplementary Material S5: Contradiction Error Taxonomy and Truncation Confound Analysis

---

## S5.1 Error Type Definitions

Six error types were defined to classify why a gold-contradicted case was missed:

**Table S5-1. Contradiction error type taxonomy.**

| Code | Error Type | Definition |
|------|-----------|------------|
| T1 | Direct contradiction missed | Evidence contains explicit contradicteddictory statements visible in the input; model fails to recognize them |
| T2 | Truncation (input-side deficit) | Evidence text is truncated before the contradicteddictory signal appears; the contradiction signal is absent from the model's input |
| T3 | Polarity confusion | Model misreads the direction of evidence (e.g., evidence says "no effect," model interprets as "supports yes") |
| T4 | Hedged-as-insufficient | Evidence contains clear directional signals, but model absorbs them into "insufficient" due to hedging language or absence of definitive conclusions |
| T5 | Population/scope mismatch | Evidence concerns a different population, setting, or scope than the question, creating misalignment |
| T6 | Statistical non-significance misread | Evidence reports a non-significant result; model treats this as "no contradiction" rather than recognizing that a null result contradicteddicts an affirmative answer |

Each of the 23 gold-contradicteddicted items was classified into one primary error type per model configuration where the prediction was not "contradicteddicted."

---

## S5.2 Item-Level Table

**Table S5. 23 Gold-Contradicted Items: Per-Model Predictions, Error Types, and Truncation Status**

| ID | Truncated? | Gold | Human2 (focused) | GPT-5.5 | GPT-5.5 Error | DS E1 | DS E1 Error | DS V2.1 | DS V2.1 Error | Dominant Error |
|----|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|---|
| SB_0007 | Y | contradicted | insufficient | insufficient | T2 | insufficient | T2 | insufficient | T2 | T2 |
| SB_0008 | Y | contradicted | insufficient | insufficient | T2 | insufficient | T2 | insufficient | T2 | T2 |
| SB_0010 | Y* | contradicted | contradicteddicted | insufficient | T2† | insufficient | T2† | insufficient | T2† | T2† |
| SB_0011 | N | contradicted | insufficient | supported | T3 | insufficient | T4 | insufficient | T4 | T4/T3 |
| SB_0012 | N | contradicted | supported | supported | —‡ | supported | —‡ | supported | —‡ | —‡ |
| SB_0014 | Y* | contradicted | contradicteddicted | contradicteddicted | — | insufficient | T2† | insufficient | T2† | T2† |
| SB_0017 | N | contradicted | contradicteddicted | contradicteddicted | — | contradicteddicted | — | insufficient | T1 | T1 |
| SB_0027 | Y | contradicted | insufficient | insufficient | T2 | insufficient | T2 | insufficient | T2 | T2 |
| SB_0032 | N | contradicted | contradicteddicted | supported | T3 | insufficient | T4 | insufficient | T4 | T3/T4 |
| SB_0038 | Y | contradicted | insufficient | insufficient | T2 | insufficient | T2 | insufficient | T2 | T2 |
| SB_0040 | Y* | contradicted | contradicteddicted | supported | T4 | insufficient | T4 | insufficient | T4 | T4 |
| SB_0043 | N | contradicted | contradicteddicted | insufficient | T6 | insufficient | T6 | insufficient | T6 | T6 |
| SB_0044 | Y* | contradicted | contradicteddicted | insufficient | T2 | insufficient | T2 | insufficient | T2 | T2 |
| SB_0048 | N | contradicted | supported | insufficient | —‡ | supported | —‡ | insufficient | —‡ | —‡ |
| SB_0055 | Y | contradicted | insufficient | insufficient | T2 | insufficient | T2 | insufficient | T2 | T2 |
| SB_0082 | N | contradicted | supported | supported | —‡ | insufficient | T4 | supported | —‡ | —‡ |
| SB_0087 | N | contradicted | contradicteddicted | contradicteddicted | — | contradicteddicted | — | insufficient | T1 | T1 |
| SB_0093 | N | contradicted | contradicteddicted | insufficient | T1 | insufficient | T1 | insufficient | T1 | T1 |
| SB_0094 | N | contradicted | contradicteddicted | contradicteddicted | — | contradicteddicted | — | insufficient | T1 | T1 |
| SB_0099 | N | contradicted | supported | insufficient | —‡ | insufficient | T4 | supported | —‡ | —‡ |
| SB_0103 | N | contradicted | contradicteddicted | insufficient | T4 | insufficient | T4 | insufficient | T4 | T4 |
| SB_0106 | N | contradicted | contradicteddicted | insufficient | T4 | insufficient | T4 | insufficient | T4 | T4 |
| SB_0121 | N | contradicted | contradicteddicted | insufficient | T4 | insufficient | T4 | insufficient | T4 | T4 |

*Evidence truncated but partial contradicteddictory signal still visible in the available fragment.
—‡ These items were classified by Human2 focused review as supported (not contradicteddicted), suggesting the gold label may reflect a stricter contradiction standard that was not consistently applied. They are retained in the gold set but flagged as annotation-boundary cases.
†T2-truncation applies to DS E1 and DS V2.1 for SB_0010 and SB_0014: while the truncation cut point may differ from GPT-5.5's, the core issue is that the contradiction signal was degraded by input truncation.
T2 for SB_0044 under GPT-5.5: evidence truncated after the Chinese-language segment describing actual data usage — the contradiction signal was in the truncated portion.

---

## S5.3 Error Type Counts by Model

**Table S5b. Error Type Frequencies Across Configurations (23 gold-contradicteddicted items)**

| Error Type | GPT-5.5 | DS E1 | DS V2.1 | Description |
|:---|---:|---:|---:|:---|
| T1 (direct miss) | 1 | 1 | 4 | Contradiction present in input but not recognized |
| T2 (truncation) | 4 | 5 | 5 | Contradiction signal absent from truncated input |
| T3 (polarity confusion) | 2 | 0 | 0 | Evidence direction misread |
| T4 (hedged-as-insufficient) | 5 | 10 | 9 | Directional signal absorbed into insufficient |
| T5 (population/scope) | 0 | 0 | 0 | Not observed |
| T6 (non-significance) | 1 | 1 | 1 | Null result not recognized as contradiction |
| Correct (contradicteddicted) | 4 | 3 | 1 | — |
| Annotation-boundary§ | 6 | 3 | 3 | Human2 judged as not contradicteddicted |

§Annotation-boundary cases: items SB_0012, SB_0048, SB_0082, SB_0099 (and partially SB_0011, SB_0055) where Human2 focused review labeled them supported or insufficient, raising the possibility that the gold-contradicteddicted label may not reflect a straightforward contradiction. These cases are noted transparently; removing them from the denominator would not change the qualitative pattern.

---

## S5.4 Truncation Confound: Information Asymmetry vs. Reasoning Failure

Of the 23 gold-contradicteddicted items, **7 (30.4%)** had evidence truncated before the contradiction signal (T2 cases across at least one model: SB_0007, SB_0008, SB_0010, SB_0027, SB_0038, SB_0044, SB_0055). An additional **2 items (SB_0014, SB_0040)** had partial truncation where a fragment of the contradicteddictory signal was visible but degraded.

This means that **7–9 of 23 contradicteddicted cases (30–39%)** had degraded or absent contradiction signals in the model's input due to evidence truncation. For these cases, the model's failure to detect contradiction is not purely a reasoning failure — it is partly an *input-side deficit*: the information needed to identify the contradiction was never presented to the model.

The decomposition of contradiction miss rate is therefore:

- **Input-signal-absent misses (T2)**: 4–5 per model, reflecting a benchmark construction artifact (truncation)
- **Reasoning-failure misses (T1, T3, T4, T6)**: 10–14 per model, reflecting genuine difficulty in contradiction recognition

This distinction is methodologically important. A critic could argue: "Your gold contradiction labels were based on the annotator seeing the complete evidence, but your models only saw truncated fragments — of course they under-detected contradiction." The T2 analysis partially supports this critique: truncation accounts for approximately one-third of the missed contradictions. However, even when T2 cases are excluded from the denominator (i.e., considering only the ~14–16 cases with non-truncated contradicteddictory signals), the best model recall rises to at most 5/15 = 33.3% (GPT-5.5) — still substantially below acceptable levels for safety-critical applications. The contradiction-blindness finding is *attenuated* by truncation but not *explained away* by it.

---

## S5.5 Dominant Failure Direction: Contradiction Absorbed into Insufficient

Across all automated configurations, the dominant failure direction was **contradicteddicted → insufficient** rather than contradicteddicted → supported. Table S5c quantifies this asymmetry.

**Table S5c. Failure Direction: Where Do Missed Contradictions Go?**

| Configuration | Contra TP | Missed→insufficient | Missed→supported | Absorption Ratio (insuff./supp.) |
|:---|---:|---:|---:|:---:|
| GPT-5.5 | 4 | 14 | 5 | 2.8:1 |
| DeepSeek E1 | 3 | 17 | 3 | 5.7:1 |
| DeepSeek V2.1 | 1 | 19 | 3 | 6.3:1 |
| DistilBERT-MNLI | 2 | 19 | 2 | 9.5:1 |
| PubMedBERT-MedNLI | 2 | 19 | 2 | 9.5:1 |

This directional bias — contradicteddicted cases being absorbed into insufficient rather than flipped to supported — has an important safety interpretation. A verifier that misclassifies contradiction as "insufficient evidence" produces a *false negative for risk*: it tells the user that evidence is merely incomplete, when in fact the evidence actively contradicteddicts the answer. This is arguably more dangerous than misclassifying contradiction as supported, because "insufficient" may be perceived as a neutral or cautionary flag rather than an active error signal. The absorption pattern suggests that current verifiers lack the discrimination to distinguish "evidence does not support the answer" (insufficient) from "evidence actively refutes the answer" (contradicteddicted) — a distinction with material consequences in biomedical decision-support contexts.

---

## S5.6 Discussion: Implications for Benchmark Design

The error taxonomy reveals two improvement directions for future contradiction detection benchmarks. First, **evidence truncation must be controlled**: if the contradiction signal is systematically absent from shorter evidence windows, then recall measured on truncated evidence conflates retrieval quality with verification quality. Future benchmarks should ensure that contradicteddicted cases have the contradicteddictory portion of evidence present within the window provided to verifiers, or should report contradiction recall stratified by truncation status.

Second, the **insufficient/contradicteddicted boundary** is the central annotation challenge. Of the 23 gold-contradicteddicted cases, Human2's focused review reclassified 4 as supported and 6 as insufficient even under strict contradiction criteria, suggesting genuine ambiguity in the gold labels for some items. This is consistent with the known difficulty of contradiction annotation documented in SNLI [20], MNLI [21], and MedNLI [22], where contradiction consistently shows the lowest inter-annotator agreement. The moderate κ = 0.530 after calibration is not a weakness of this study's annotation protocol; it reflects an inherent property of contradiction annotation in biomedical evidence, where the absence/presence boundary is fundamentally fuzzy in truncated or domain-specialized contexts.

---

*All per-item classifications above are based on manual review of model outputs, evidence text, and Human2 annotations. The error type assignment involves judgment, particularly for boundary cases, and is provided as a qualitative analysis framework rather than a fully independent second-annotator adjudication.*
