================================================================================
 Supplementary code and data
 Paper: "Volume or Coupling? A Scale-Dependent Dissociation in Constraint
         Recovery of Language-Model Loops"
================================================================================

CONTENTS (code/)
  exp1_pilot_1p5B.py             Pilot disclosed in Sec. 2.4 (C vs S, n=10, window=5).
                                 Excluded from all analyses; archived for transparency.
  exp2_1p5B_coupled_vs_single.py Main 1.5B experiment: conditions C and S (Table 1).
  exp3_1p5B_context_matched.py   1.5B condition X + three-group comparison.
  exp4_1p5B_dose_response.py     1.5B condition D4 + dose-response summary.
  exp5_scale_3B_7B.py            3B and 7B tiers, all four conditions (Table 1).
                                 Run once per tier; set MODEL_NAME in the header.
  exp6_verify_7B_controls.py     Sec. 2.6 validity checks: no-perturbation control
                                 and per-agent replay of the opener-1 coupled run.
  exp7_replay_7B_trajectories.py Exact replays of all 8 recovered 7B coupled runs
                                 with STABLE/TRANSIENT classification (Sec. 3.3).
  make_figures.py                Generates fig_scale_rates.png and fig_7b_traj.png
                                 from embedded outcomes (no GPU needed).
  analysis_mcnemar.py            Reproduces every exact McNemar p-value in the
                                 paper from the embedded raw outcomes (no GPU).

CONTENTS (data/)
  raw_outcomes_v2.txt            Per-run recovery delays, all tiers + pilot +
                                 validity-check records.
  replay_output_7b.txt           Full per-exchange violation trajectories of the
                                 eight replayed 7B coupled runs, with the
                                 stability classification.

HOW TO RUN
  Environment: Google Colab. Install once per session:
      pip install -q -U transformers accelerate scipy
  GPU: T4 suffices for 1.5B; use an L4 for the 7B tier (fp16 needs ~15 GB VRAM;
  do not use 4-bit quantization, which would confound the scale comparison).
  Runtime: ~3 T4 GPU-hours (1.5B tier total) + ~5 L4 GPU-hours (3B+7B tiers).

DETERMINISM
  Greedy decoding + fixed openers: every run is a deterministic function of its
  opener, so all reported trajectories reproduce exactly (up to hardware-level
  numeric nondeterminism). exp7 asserts this by matching replayed delays
  against recorded delays (8/8).

NOTES
  Comments and console strings in the scripts are English translations of the
  scripts as run; all logic, parameters, prompts, and openers are unchanged.
  One kwarg was renamed for current library versions (torch_dtype -> dtype);
  behavior is identical.
