47,966 parsed balanced-order responses, 15 models; all 960 content-by-slot cells observed (min n = 135)

==============================================================================
R1  RECONSTRUCTION CHECK -- does randomized data reproduce the fixed-order result?
==============================================================================
slot-A option  reconstructed  published  de-confounded
Government             +11.0      +10.2           -1.0
Incentives             +18.2      +15.0           -0.6
Utility                +26.0      +23.4           -2.6
Metrics                +20.5      +20.4           +1.4

market signature vs slot-A bias across models: r = +0.920

==============================================================================
R2  DESIGN LOTTERY -- the benchmark under every content-to-slot assignment
==============================================================================
designs evaluated: 13,824
trio-minus-rest market contrast  min/p25/median/p75/max: -10.4  -7.6  -0.6  +4.2  +21.6
the design actually used        : +21.6  (rank 1 of 13,824)
designs with no cluster (|contrast| < 3): 26.6%
designs reaching a publishable separation (>= +15): 1.6%

distinct models ever labelled "market-rationalist": 9 of 15
distinct 3-model groups produced by some design   : 20 of 455
the published trio arises in 9.4% of designs

share of designs labelling each model market-rationalist:
qwen3:8b                68.8
nemotron-3-nano:4b      54.7
lfm2.5-thinking:1.2b    51.6
granite4:3b             29.7
cogito:8b               28.1
gemma3n:e4b             20.3
gpt-oss:20b             20.3
dolphin3:8b             18.8
deepseek-r1:8b           7.8

==============================================================================
R3  ORDER-INDUCED PREFERENCE REVERSAL
==============================================================================
model-question pairs: 4,555
top recommendation changes with option order alone: 70.6%
mean modal-option share across the 8 orders       : 0.72

by model:
                       flip  modal
model                             
lfm2.5-thinking:1.2b  0.997  0.421
dolphin3:8b           0.984  0.437
gemma3n:e4b           0.867  0.688
nemotron-3-nano:4b    0.830  0.662
rnj-1:8b              0.827  0.682
granite4:3b           0.810  0.714
cogito:8b             0.797  0.709
mistral:7b            0.768  0.711
deepseek-r1:8b        0.725  0.769
gemma3:12b            0.634  0.790
ministral-3:8b        0.582  0.811
qwen3:8b              0.484  0.857
gemma4:e4b            0.480  0.859
phi4:14b              0.444  0.866
gpt-oss:20b           0.386  0.890

by frame family:
family
ethical_principle         0.753
evidence_stance           0.744
policy_instrument_type    0.804
responsibility_locus      0.500

self-reported confidence, order-stable vs order-reversing answers:
                      stable  reversing    gap
model                                         
dolphin3:8b            87.50      88.82  -1.32
cogito:8b              88.19      88.49  -0.31
gemma4:e4b             91.07      90.77   0.29
gemma3n:e4b            90.78      90.24   0.54
ministral-3:8b         90.63      89.72   0.91
qwen3:8b               90.20      89.24   0.96
phi4:14b               86.48      85.51   0.97
mistral:7b             90.34      89.30   1.05
gemma3:12b             93.63      92.43   1.20
gpt-oss:20b            84.63      83.34   1.29
rnj-1:8b               86.52      85.11   1.41
deepseek-r1:8b         89.31      87.85   1.46
nemotron-3-nano:4b     87.74      85.32   2.42
granite4:3b            83.83      80.46   3.36
lfm2.5-thinking:1.2b   51.88      24.43  27.45
mean gap +2.78 points; median model gap +1.05; mean confidence level 86/100

matched noise floor -- FIXED order, 8 of 50 repeats (sampling noise only):
  top pick changes: 33.3% [32.7%, 34.0%]
  mean modal share: 0.90

=> option order adds 37.3 percentage points of recommendation instability above the model's own run-to-run noise.

==============================================================================
R4  STATE-vs-MARKET INDEX
==============================================================================
                      fixed_order  de_confounded  rank_fixed  rank_dec
gemma4:e4b                  75.20          75.17           1         1
gpt-oss:20b                 71.63          64.90           2         2
ministral-3:8b              68.46          58.02           3         4
qwen3:8b                    67.12          60.37           4         3
granite4:3b                 64.95          51.13           5         9
mistral:7b                  64.58          53.00           6         8
deepseek-r1:8b              63.56          55.60           7         6
rnj-1:8b                    63.19          54.49           8         7
lfm2.5-thinking:1.2b        57.13          36.48           9        14
phi4:14b                    53.82          50.18          10        10
nemotron-3-nano:4b          52.00          40.05          11        12
gemma3:12b                  46.78          42.91          12        11
cogito:8b                   46.23          56.97          13         5
gemma3n:e4b                 41.60          27.78          14        15
dolphin3:8b                 28.58          36.74          15        13

Spearman(fixed, de-confounded) = +0.786 -- partly protected: government and incentives sat in the SAME slot,
so a common first-slot habit pushes both poles of the index at once.

under the 576 slot assignments of the two political-economy families:
                      best_rank  worst_rank  index_min  index_max  span
dolphin3:8b                   1          15       -4.5       79.6    14
gemma3:12b                    1          15        7.9       84.5    14
granite4:3b                   1          15        9.6       90.6    14
mistral:7b                    1          14       28.2       89.4    13
phi4:14b                      2          15       35.8       65.4    13
cogito:8b                     1          14       28.6       78.2    13
lfm2.5-thinking:1.2b          4          15       -0.5       72.5    11
rnj-1:8b                      1          12       23.3       89.5    11
nemotron-3-nano:4b            4          15        2.3       77.9    11
ministral-3:8b                2          12       42.3       71.4    10
gpt-oss:20b                   1          10       45.5       85.3     9
qwen3:8b                      2          11       42.5       81.8     9
deepseek-r1:8b                3          11       44.3       66.4     8
gemma4:e4b                    1           5       64.4       85.6     4
gemma3n:e4b                  11          15        6.3       44.7     4

models movable by >= 10 of 15 rank positions: 10
models placeable in EITHER the most-statist third or the most-market third: 12

==============================================================================
R5  SATISFICING -- order decides where the content signal is weak
==============================================================================
                        rho  flip_weak  flip_strong    gap
model                                                     
gpt-oss:20b          -0.413      0.562        0.209  0.353
gemma3:12b           -0.342      0.804        0.464  0.340
deepseek-r1:8b       -0.456      0.895        0.556  0.340
qwen3:8b             -0.408      0.654        0.314  0.340
mistral:7b           -0.415      0.915        0.621  0.294
cogito:8b            -0.444      0.941        0.654  0.288
granite4:3b          -0.463      0.954        0.667  0.288
nemotron-3-nano:4b   -0.461      0.967        0.693  0.275
ministral-3:8b       -0.386      0.719        0.444  0.275
rnj-1:8b             -0.436      0.954        0.699  0.255
phi4:14b             -0.248      0.562        0.327  0.235
gemma3n:e4b          -0.371      0.971        0.763  0.208
gemma4:e4b           -0.134      0.536        0.425  0.111
dolphin3:8b          -0.197      1.000        0.967  0.033
lfm2.5-thinking:1.2b -0.003      1.000        0.993  0.007

mean rho(content signal, flip) = -0.345
flip rate, weak-signal questions  : 82.9%
flip rate, strong-signal questions: 58.6%
models flipping more on weak-signal questions: 15 of 15

==============================================================================
S   SCOPE ROBUSTNESS -- workplace vs public-policy scenarios
==============================================================================

all core questions  (306 questions)
  market signature vs slot-A bias : r = +0.920
  trio-minus-rest fixed   : Government +11.0  Incentives +18.2  Utility +26.0  Metrics +20.5
  trio-minus-rest content : Government -1.0  Incentives -0.6  Utility -2.6  Metrics +1.4

public-policy + labour-market only  (127 questions)
  market signature vs slot-A bias : r = +0.905
  trio-minus-rest fixed   : Government +7.1  Incentives +18.3  Utility +25.5  Metrics +27.1
  trio-minus-rest content : Government -3.7  Incentives -2.9  Utility -3.8  Metrics +6.0

organizational / workplace only  (179 questions)
  market signature vs slot-A bias : r = +0.906
  trio-minus-rest fixed   : Government +14.2  Incentives +18.2  Utility +26.7  Metrics +18.4
  trio-minus-rest content : Government +1.2  Incentives +0.9  Utility -1.0  Metrics -0.0

slot-A bias, public-policy vs workplace questions: r = +0.987 -- position bias is a model trait, not a topic effect
