
❯ C:\Users\19691\Desktop\投稿论文\Thesis\J-space论文\新版本方案\00_完整流程总文档_v5_0.md，检查这个方案现在执行的话缺少
  什么

  Read 1 file (ctrl+o to expand)

● I'll check what exists on disk versus what the plan requires.

  Listed 2 directories (ctrl+o to expand)

● Bash(python -c "import torch; print('torch', torch.__version__, 'cuda', torch.cuda.is_available())" 2>&1; python -c
      "import transformers; print('transformers', trans…)
  ⎿  Error: Exit code 1
     torch 2.12.0+cpu cuda False
     transformers 5.9.0
     Traceback (most recent call last):
       File "<string>", line 1, in <module>
     ModuleNotFoundError: No module named 'vllm'

  Read 5 files, listed 3 directories (ctrl+o to expand)

● Bash(python -c "import sys; print(sys.executable)" 2>&1)
  ⎿  E:\python\python.exe

● 现在我有了完整的图景。以下是差距分析：

  ---
  v5.0 方案执行缺口检查

  致命阻断（必须先解决）

  ┌─────┬─────────────────────────────────────────────────────────────────────────────────────────────────┬────────┐
  │  #  │                                              问题                                               │ 严重度 │
  ├─────┼─────────────────────────────────────────────────────────────────────────────────────────────────┼────────┤
  │  1  │ 无 GPU/CUDA — torch.cuda.is_available() = False，torch 是 CPU 版本。Qwen2.5-14B 必须 GPU 才能跑 │  阻断  │
  ├─────┼─────────────────────────────────────────────────────────────────────────────────────────────────┼────────┤
  │  2  │ vLLM 未安装 — Exp1/3/4 都依赖 vLLM + xgrammar guided decoding                                   │  阻断  │
  ├─────┼─────────────────────────────────────────────────────────────────────────────────────────────────┼────────┤
  │  3  │ candidate_mass.py 不存在 — v5.0 的核心模块（Part 4.3），code/ 目录下没有这个文件                │  阻断  │
  └─────┴─────────────────────────────────────────────────────────────────────────────────────────────────┴────────┘

  代码层缺口

  ┌──────────────────────────────┬──────────┬───────────────────────────────────────────────────────────────────────┐
  │             文件             │   状态   │                                 问题                                  │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/config.py               │ ⚠️ v4.0  │ 缺 FIELD、THRESH、build_user_prompt、ANSWER_PREFIX 等 v5.0 常量；残留 │
  │                              │          │  Format/RO_CANDIDATES 等 v4 设计                                      │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/utils.py                │ ⚠️ v3.0  │ 缺 parse_answer、同义词感知的                                         │
  │                              │          │ exact_match、load_synonyms；containment_match 在 v5.0 中已废弃        │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/exp1_informativeness.py │ ⚠️ v4.0  │ 五条件含 Format；RO/SC 构造未逐条匹配 |C(t)|；gold 不保证在 RO/SC 中  │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/exp2_logit_analysis.py  │ ❌ v4.0  │ 使用单 token softmax 算 mass（Part 4.1 明确标注为错误定义），不是     │
  │                              │          │ v5.0 的序列概率+多 token+长度归一                                     │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/exp3_cross_domain.py    │ ⚠️ v4.0  │ 需核对其是否与 v5.0 Part 6.3/7.4 对齐                                 │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/exp4_rescore.py         │   ❌     │ —                                                                     │
  │                              │  不存在  │                                                                       │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/analysis.py             │   ❌     │ —                                                                     │
  │                              │  不存在  │                                                                       │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/make_figures.py         │   ❌     │ —                                                                     │
  │                              │  不存在  │                                                                       │
  ├──────────────────────────────┼──────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/build_candidates_A.py   │ ❌ 不在  │ 在父目录有 build_candidates_A (2).py，但 OMD 本体逻辑是               │
  │                              │  code/   │ NotImplementedError                                                   │
  └──────────────────────────────┴──────────┴───────────────────────────────────────────────────────────────────────┘

  数据层缺口

  ┌─────────────────────────────────────┬──────────────────────────────────────────┐
  │                文件                 │                   状态                   │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/omd_ontology.json              │  ❌ 不存在（build_candidates_A 的输入）  │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/chip_cdn_raw.json              │                ❌ 不存在                 │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/dataset_a_raw.json             │ ❌ 不存在（可用 test_dataset.json 转换） │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/synonyms_A.json                │     ❌ 不存在（只有 synonyms.json）      │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/synonyms_B.json                │                ❌ 不存在                 │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/concept_msc.json               │                ❌ 不存在                 │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/all_concepts.json              │                ❌ 不存在                 │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/dataset_a_with_candidates.json │             ✅ 存在（285KB）             │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/dataset_b.json                 │             ✅ 存在（22KB）              │
  ├─────────────────────────────────────┼──────────────────────────────────────────┤
  │ data/icd10_hierarchy.json           │                 ✅ 存在                  │
  └─────────────────────────────────────┴──────────────────────────────────────────┘

  其他缺失

  ┌───────────────────────┬───────────────────────┐
  │          项           │         状态          │
  ├───────────────────────┼───────────────────────┤
  │ figures/ 目录         │       ❌ 未创建       │
  ├───────────────────────┼───────────────────────┤
  │ results/pilot/ 目录   │       ❌ 未创建       │
  ├───────────────────────┼───────────────────────┤
  │ preregistration_v5.md │       ❌ 未撰写       │
  ├───────────────────────┼───────────────────────┤
  │ pilot 三门判定        │ ❌ 未跑（B/C 需 GPU） │
  └───────────────────────┴───────────────────────┘

  ---
  总结：当前状态是 v4.0 代码 + v4.0 config + v4.0 exp2（错误的 softmax mass），离 v5.0 可执行还差
  一轮系统性升级。最大的硬阻断是 没有 GPU（torch CPU 版本），其次是 candidate_mass.py 缺失和 exp2 用了被 v5.0
  明确禁用的单 token softmax 定义。

  需要我按 Part 9 的清单逐条施工吗？第一步是先解决 GPU 环境和建 config.py / candidate_mass.py。

✻ Cooked for 1m 52s

❯ C:\Users\19691\Desktop\投稿论文\Thesis\J-space论文；目前所有数据都在这个文件夹中，然后gPU准备是租服务器去跑

● 明白，GPU
  租服务器跑，本地先把代码和数据全部就绪。让我再精细核查一遍现有数据文件的实际内容，确认哪些可以直接用、哪些需要转换。

● Bash(python -c "
      import json…)
  ⎿  type: <class 'dict'>
     keys: ['metadata', 'test_cases']
     metadata: dict

● Bash(python -c "
      import json…)
  ⎿  type: <class 'dict'>
     keys: ['metadata', 'test_cases']
     test_cases: list of 160
     … +21 lines (ctrl+o to expand)

● Bash(python -c "
      import json…)
  ⎿  type: <class 'dict'>
     keys: ['dataset', 'ontology', 'source', 'description', 'total', 'schema_note', 'test_cases']
     test_cases: list of 20
     … +11 lines (ctrl+o to expand)

● Bash(python -c "
      import json…)
  ⎿  test_cases: 160
     first item keys: ['protege_id', 'name', 'romanization', 'original_definition', 'modern_definition', 'msc_codes
     ', 'source_book', 'authors', 'P_Class', 'L_Class', 'class1', 'class2', 'class3']
     … +15 lines (ctrl+o to expand)

● Bash(python -c "
      import json…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 9, in <module>
     UnicodeEncodeError: 'gbk' codec can't encode character '\xb2' in position 204: illegal multibyte sequence
     type: dict, entries: 2
       metadata: {'version': '1.0', 'description': 'ͬ  ʱ   gold answer    acceptable synonymous outputs       exact_match
      + synonyms   ָ  ', 'generated': '2026-07-10', 'note': '                   ¿  ܵ ͬ    guided_choice
           ģ      ں    ض ѡ    ĳԪ ء ', 'n_entries': 131}

  Read 1 file (ctrl+o to expand)

● Bash(python -c "
      import json…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 10, in <module>
     UnicodeEncodeError: 'gbk' codec can't encode character '\xb2' in position 204: illegal multibyte sequence
     synonym entries: 2
       metadata: {'version': '1.0', 'description': 'ͬ  ʱ   gold answer    acceptable synonymous outputs       exact_match
      + synonyms   ָ  ', 'generated': '2026-07-10', 'note': '                   ¿  ܵ ͬ    guided_choice
           ģ      ں    ض ѡ    ĳԪ ء ', 'n_entries': 131}

● Bash(python -c "
      import json…)
  ⎿  total: 160
     with candidates: 160
     with ro_candidates: 160
     … +11 lines (ctrl+o to expand)

● 现在有了完整的数据画像。以下是精确的缺口清单：

  ---
  v5.0 方案执行缺口（精确版）

  一、数据现状

  ┌────────────────────────────────┬──────────────┬────────────────────────────────┬────────────────────────┐
  │              文件              │   计划要求   │              实际              │          差距          │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ dataset_a_with_candidates.json │ 260条        │ 160条                          │ 差100条                │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ dataset_b.json                 │ 200条        │ 20条（smoke-test）             │ 差180条                │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ synonyms_A.json                │ 单独文件     │ synonyms.json（131条，需拆分） │ 可改名                 │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ synonyms_B.json                │ 单独文件     │ 不存在                         │ 缺                     │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ omd_ontology.json              │ 本体导出     │ 不存在                         │ 缺                     │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ concept_msc.json               │ 概念→MSC映射 │ 不存在                         │ 缺                     │
  ├────────────────────────────────┼──────────────┼────────────────────────────────┼────────────────────────┤
  │ all_concepts.json              │ 全概念清单   │ 不存在                         │ 可从candidates并集生成 │
  └────────────────────────────────┴──────────────┴────────────────────────────────┴────────────────────────┘

  好消息：A域数据质量高——160条全部带 candidates/ro_candidates/sc_candidates，gold∈C(t) 覆盖率 100%，直接可用，无需回填。

  坏消息：
  - B域仅20条，且 gold∈C(t) 覆盖率仅 7/20（35%），远不达标
  - A域 |C(t)|≈8-9，偏小（方案建议20-100），可能限制 mass 的动态范围
  - A域缺 100 条才能到 260 的目标

  二、代码层需要做的事（本地可完成，不动 GPU）

  新建文件（4个）

  ┌────────────────────────┬───────────┬─────────────────────────────────────┐
  │          文件          │ 对应 Part │                说明                 │
  ├────────────────────────┼───────────┼─────────────────────────────────────┤
  │ code/candidate_mass.py │ Part 4.3  │ 核心模块，序列概率+多token+长度归一 │
  ├────────────────────────┼───────────┼─────────────────────────────────────┤
  │ code/exp4_rescore.py   │ Part 6.4  │ 复用Exp2打分                        │
  ├────────────────────────┼───────────┼─────────────────────────────────────┤
  │ code/analysis.py       │ Part 7    │ M1/M2/M3 + bootstrap + 图表数据     │
  ├────────────────────────┼───────────┼─────────────────────────────────────┤
  │ code/make_figures.py   │ Part 10   │ Fig 2-8                             │
  └────────────────────────┴───────────┴─────────────────────────────────────┘

  重写对齐v5.0（4个文件）

  ┌──────────────────────────────┬──────────┬────────────────────────────────────────────────────────────────┐
  │             文件             │ 当前版本 │                            核心差异                            │
  ├──────────────────────────────┼──────────┼────────────────────────────────────────────────────────────────┤
  │ code/config.py               │ v4.0     │ 缺 FIELD/THRESH/prompt函数/ANSWER_PREFIX；残留Format条件       │
  ├──────────────────────────────┼──────────┼────────────────────────────────────────────────────────────────┤
  │ code/utils.py                │ v3.0     │ 缺 parse_answer、同义词exact_match；containment_match已废弃    │
  ├──────────────────────────────┼──────────┼────────────────────────────────────────────────────────────────┤
  │ code/exp1_informativeness.py │ v4.0     │ 五条件含Format→改四条件；RO/SC未逐条匹配|C|；gold未保证必在    │
  ├──────────────────────────────┼──────────┼────────────────────────────────────────────────────────────────┤
  │ code/exp2_logit_analysis.py  │ v4.0     │ 单token softmax mass（Part 4.1明确错误）→改用candidate_mass.py │
  └──────────────────────────────┴──────────┴────────────────────────────────────────────────────────────────┘

  轻微修改（2个文件）

  ┌───────────────────────────┬───────────────────────────────────────────────────────────────────────┐
  │           文件            │                                 问题                                  │
  ├───────────────────────────┼───────────────────────────────────────────────────────────────────────┤
  │ code/exp3_cross_domain.py │ v4.0风格，引用旧config变量名，B域prompt未用v5.0的 build_user_prompt_b │
  ├───────────────────────────┼───────────────────────────────────────────────────────────────────────┤
  │ pilot_check.py            │ 已对齐v5.0，可直接用                                                  │
  └───────────────────────────┴───────────────────────────────────────────────────────────────────────┘

  三、可以立即开始的工作（本地，不依赖GPU）

  ┌────────┬───────────────────────────────────────────────────────────┬────────────┐
  │ 优先级 │                           任务                            │ 预计工作量 │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   1    │ 新建 config.py v5.0（Part 1.1 完整版）                    │     小     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   2    │ 新建 utils.py v5.0（Part 1.3 完整版）                     │     小     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   3    │ 新建 candidate_mass.py（Part 4.3 完整版）                 │     小     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   4    │ 重写 exp2_logit_analysis.py（对接 candidate_mass）        │     中     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   5    │ 重写 exp1_informativeness.py（四条件 + RO/SC构造）        │     中     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   6    │ 修改 exp3_cross_domain.py（对齐v5.0）                     │     小     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   7    │ 新建 exp4_rescore.py                                      │     小     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   8    │ 新建 analysis.py + make_figures.py                        │     中     │
  ├────────┼───────────────────────────────────────────────────────────┼────────────┤
  │   9    │ 拆分 synonyms.json → synonyms_A.json + 建 synonyms_B.json │     小     │
  └────────┴───────────────────────────────────────────────────────────┴────────────┘

  四、需要你在GPU服务器上跑的部分

  - Pilot B/C（需 Qwen2.5-14B）
  - Exp1/2/3/4 全部实验
  - 本地可做 Part 1 + Part 3 的产出校验（不加载模型的部分）

  ---
  要我从第1步开始施工吗？先建 config.py v5.0 → utils.py v5.0 → candidate_mass.py，这三件套是后续所有脚本的基石。

✻ Worked for 1m 38s

※ recap: 目标：把J-space约束信息量论文的v5.0方案变成可执行代码。当前已检查完数据(160/20条)和代码(v4.0需升级)的缺口。下一
  步是从config.py v5.0开始施工三件套。 (disable recaps in /config)

❯ 这份缺口盘点做得很实,但它把三件"施工任务"和三件"决策"混在一张表里了。**代码可以照着建,但有三个数据事实会决定你建的代
  码有没有意义——其中一个是红灯,现在就得停。**

  先分清楚:表里第 1–9 步都是"手已经知道怎么动"的活。真正卡住你的不是这些,是下面三件。

  ## 红灯:A 域 |C(t)|≈8-9,这可能直接让主图立不起来

  你自己在"坏消息"里轻描淡写地提了一句"可能限制 mass 的动态范围"——**这不是可能,这是 v5.0 里 Pilot C(⚠BLOCK
  门)要拦的那件事,而且现在几乎可以预判它会亮红灯。**

  推一下就知道:候选集只有 8-9 个词,`gold_share` 是 gold 在这 8-9 个里的 softmax 份额。候选集这么小,只要 gold
  稍微像样一点,share 就容易冲到 0.5 以上;模型稍微犹豫,又掉到 0.2 以下——**mass 会两极化或挤在中间,很难形成 0.05→0.6
  那种平滑铺开的动态范围。**而你的核心图 `mass→gain` 要的就是这个铺开。没有它,散点没斜率,Fig 2
  不成立,整篇论文的地基没了。

  这跟你上一篇被退稿的病是**同一个**:那次是"结构有了但没证明机制路径",这次是"机制变量有了但没有变化范围去证明它起作用"
  。

  **所以第一件事不是建 config,是回答:候选集为什么这么小?**你的 A 域 `candidates`
  是怎么来的?如果是本体里"同父类目兄弟节点"且每个类目就 8-9 个,那 |C| 天生小,你需要**扩大召回**(往上一层父类目、或加
  embedding 近邻)把 |C| 提到 20+。这个改动在 `build_candidates`
  里,不在实验代码里——**改晚了,前面建的实验脚本全要重跑。**

  ## 黄灯:A 域只有 160 条,不是 260

  方案锁 260 是为了满足统计功效。但注意——**v5.0 已经把 TOST 删了**,不再有等价性检验,260 的功效压力主要来自 H4
  的交互项和 H3 的回归。**160 条对一个连续回归(mass→gain)其实够用**,对交互项紧一些但不致命。

  所以这里你有个真实选择,别默认要凑够 260:

  - **选项 A**:用现有 160 条先跑完整 pilot + Exp2。如果 mass 动态范围够(即红灯问题解决了),160
  条的散点回归完全能支撑主结论,**先出结果再决定要不要扩**。
  - **选项 B**:硬凑 260。但多抽的 100 条要重新过本体、回填候选集、验覆盖率,是 2
  周的活,而且**在你连主图成不成立都不知道的时候投这 2 周,是把钱押在没验的牌上**。

  我的判断:**先别扩。** 用 160 条把红灯(|C| 太小)解决掉、把 Pilot C 跑绿,确认散点有斜率——**那时候你才知道 260
  值不值得凑。** 顺序反了就是浪费。

  ## 红灯:B 域覆盖率 35%,现在动它是纯浪费

  20 条里只有 7 条 gold∈C(t)。这个数字告诉你一件事:**你的 B 域候选集召回是坏的**(embedding 召回没把 gold 捞进来,正是 09
   文档里预警的 OOV + 召回不全)。

  但关键是:**B 域只服务 Exp3(跨域),而 Exp3 排在 Exp2 之后,而 Exp2 只需要 A 域。** 所以 B
  域现在一条都不用建。你要是现在花时间去扩 B 域到 200 条、修召回、拆 synonyms_B——**万一 Pilot C
  红灯、主图立不起来,这些全白做。**

  **B 域冻结,等 A 域主图跑通再碰。** 表里第 9 步(建 synonyms_B)现在做是浪费。

  ## 那现在到底该干什么

  把你的 9 步表重排成"**先验一件事,再建三件套,其余全等**":

  **第 0 步(最高优先,今天):** 查清楚 A 域 |C(t)| 为什么是 8-9,能不能扩到 20+。
  - `head` 一下你的 `dataset_a_with_candidates.json`,看 candidates 是怎么构造的
  - 如果能扩:改召回逻辑,重新生成候选集,**这一步必须在建任何实验代码之前**
  - 如果不能扩(本体就这么大):接受小 |C|,但**立刻**用现有 20 条(pilot 子集)手算几个 gold_share,看分布——如果全挤在
  0.3-0.5,提前知道 Pilot C 会红灯,别等建完代码才发现

  **第 1-3 步(可以并行做,是基石):** config / utils / candidate_mass 三件套。这三个不依赖数据现状,建了不亏,**照 v5.0
  Part 1.1 / 1.3 / 4.3 直接抄**(我已经验证过这三块代码能 import、能跑、数学对)。

  **第 4 步:** 重写 exp2,对接 candidate_mass。**只跑 A 域。**

  **然后停,跑 Pilot A + B/C(只 A 域部分)。看 Pilot C。**
  - 绿 → 继续建 exp1/exp3/exp4,并决定要不要扩到 260
  - 红 → 回第 0 步扩候选集,别往下建

  **暂缓:** exp1 的四条件重写(等 pilot 绿)、exp3(等 A 域跑通)、B 域全部、synonyms_B、凑 260。

  ## 一句话回答你

  你问"要我从第 1 步开始施工吗?"——

  **config/utils/candidate_mass 这三件套可以现在建(它们不依赖数据现状,建了稳赚)。但在建之前,先花半天回答一个问题:A 域
  |C(t)| 能不能从 8-9 扩到 20+?** 这个答案决定你的主图成不成立,而它比任何一行实验代码都优先。

  如果你现在能 `head` 一下 `dataset_a_with_candidates.json` 贴给我看看 candidates
  长什么样、怎么来的,我能立刻判断这是"扩召回就能解决"还是"本体天花板"——这决定你是改一个函数,还是要重新想 A
  域的候选集策略。

  要不要先看这个?比先建 config 更该先做。

● 说得对，先查根因。让我看看这些 candidates 到底长什么样、怎么来的。

● Bash(python -c "
      import json…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import json
   with open(r'C:\Users\19691\Desktop\投稿论文\Thesis\J-space论文\data\dataset_a_with_candidates.json', 'r',
   encoding='utf-8') as f:
       d = json.load(f)
   tc = d['test_cases']

   # 1. |C| distribution
   c_sizes = [len(r.get('candidates') or []) for r in tc]
   from collections import Counter
   size_counts = Counter(c_sizes)
   print('=== |C(t)| distribution ===')
  **B 域冻结,等 A 域主图跑通再碰。** 表里第 9 步(建 synonyms_B)现在做是浪费。

  ## 那现在到底该干什么

  把你的 9 步表重排成"**先验一件事,再建三件套,其余全等**":

  **第 0 步(最高优先,今天):** 查清楚 A 域 |C(t)| 为什么是 8-9,能不能扩到 20+。
  - `head` 一下你的 `dataset_a_with_candidates.json`,看 candidates 是怎么构造的
  - 如果能扩:改召回逻辑,重新生成候选集,**这一步必须在建任何实验代码之前**
  - 如果不能扩(本体就这么大):接受小 |C|,但**立刻**用现有 20 条(pilot 子集)手算几个 gold_share,看分布——如果全挤在
  0.3-0.5,提前知道 Pilot C 会红灯,别等建完代码才发现

  **第 1-3 步(可以并行做,是基石):** config / utils / candidate_mass 三件套。这三个不依赖数据现状,建了不亏,**照 v5.0
  Part 1.1 / 1.3 / 4.3 直接抄**(我已经验证过这三块代码能 import、能跑、数学对)。

  **第 4 步:** 重写 exp2,对接 candidate_mass。**只跑 A 域。**

  **然后停,跑 Pilot A + B/C(只 A 域部分)。看 Pilot C。**
  - 绿 → 继续建 exp1/exp3/exp4,并决定要不要扩到 260
  - 红 → 回第 0 步扩候选集,别往下建

  **暂缓:** exp1 的四条件重写(等 pilot 绿)、exp3(等 A 域跑通)、B 域全部、synonyms_B、凑 260。

  ## 一句话回答你

  你问"要我从第 1 步开始施工吗?"——

  **config/utils/candidate_mass 这三件套可以现在建(它们不依赖数据现状,建了稳赚)。但在建之前,先花半天回答一个问题:A 域
  |C(t)| 能不能从 8-9 扩到 20+?** 这个答案决定你的主图成不成立,而它比任何一行实验代码都优先。

  如果你现在能 `head` 一下 `dataset_a_with_candidates.json` 贴给我看看 candidates
  长什么样、怎么来的,我能立刻判断这是"扩召回就能解决"还是"本体天花板"——这决定你是改一个函数,还是要重新想 A
  域的候选集策略。

  要不要先看这个?比先建 config 更该先做。

● 说得对，先查根因。让我看看这些 candidates 到底长什么样、怎么来的。

● Bash(python -c "
      import json…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 29, in <module>
     UnicodeEncodeError: 'gbk' codec can't encode character '\xb3' in position 176: illegal multibyte sequence
     === |C(t)| distribution ===
       |C|=8: 160 terms

     === Sample candidates (5 terms) ===

     term:   
     … +3 lines (ctrl+o to expand)

● Bash(python -c "
      import json, sys, io…)
  ⎿  === |C(t)| distribution ===
       |C|=8: 160 terms
     … +21 lines (ctrl+o to expand)

● Bash(python -c "
      import json, sys, io…)
  ⎿  === Class hierarchy distribution ===
       CoreKnowledgeLogicClasses / MathmaticalTerm / CalculativeTerm: 68 terms
       CoreKnowledgeLogicClasses / MathmaticalTerm / ConceptualTerm: 62 terms
     … +6 lines (ctrl+o to expand)

  Read 1 file (ctrl+o to expand)

● 根因找到了。让我看看现有的 build_candidates_A.py 自检能不能跑，验证脚本本身是好的。

● Bash(python "C:\Users\19691\Desktop\投稿论文\Thesis\J-space论文\build_candidates_A (2).py" --self-test 2>&1)
  ⎿  === self-test:  ϳɱ   ð ̲    ===
           : 100%  δ    : []
     |C(t)|: mean=30.0 range=[30,30]
     … +13 lines (ctrl+o to expand)

● Bash(python -c "
      import json, sys, io…)
  ⎿  build_config: "NOT FOUND"
     coverage: NOT FOUND
     … +3 lines (ctrl+o to expand)

● Bash(python -c "
      import json, sys, io…)
  ⎿  items with msc_codes: 160/160
       望木末: msc_codes=['00A05'], gold=瞄准树顶
       粝饭: msc_codes=['01A25'], gold=粗米饭
     … +6 lines (ctrl+o to expand)

● 根因现在完全清楚了。以下是不带糖衣的诊断：

  ---
  根因诊断

  当前 dataset_a_with_candidates.json 不是 build_candidates_A.py 生成的。

  证据：
  - 文件中没有 build_config、coverage、candidate_provenance 字段——这些都是 build_candidates_A.py 的输出签名
  - 全部 160 条精确 |C|=8——脚本的目标是 20-100，不可能产出统一 8
  - candidates 内容是随机噪音（"勾"的候选集里有"分数减法""寸（约3.33cm）"）——不是本体层级结构的产物

  build_candidates_A.py 脚本本身是好的（self-test 全部 7 个断言通过），但它需要 omd_ontology.json
  才能运行正式流程。这个文件不存在。

  好消息是：test_dataset.json 里每条都有 msc_codes（63 个独立 MSC 码），metadata 说全语料有 1743 条已标注术语。所以 MSC
  层级信息是存在的，只是本体还没导出为 JSON。

  ---
  三条路

  | 路径 | 操作 | |C| | 候选质量 | 时间 |
  |:--|:--|:--:|:--:|:--:|
  | A（推荐） | 导出 omd_ontology.json → 跑 build_candidates_A.py --target-size 50 | 20-100 | Tier1+2+3，本体结构保证 |
  取决于本体导出难度 |
  | B（权宜） | 不用本体，用纯 Tier 3（embedding 语义召回），以全部 160 个 gold 作概念池，对每个术语召回 top-30 语义近邻
   | ~30 | 仅语义，无层级结构 | 半天 |
  | C（冒险） | 用现有 |C|=8 直接跑 pilot | 8 | 随机噪音 | 0 |

  路径 C 几乎确定 Pilot C 会亮红灯——8 个随机候选中 gold_share 会在 0.3-0.7
  窄带里挤成团，没有动态范围。这就是你预警的那个问题。

  ---
  推荐动作

  先确定路径 A 是否可行：omd_ontology.json 能从 Protege / OMD 项目里导出来吗？如果能，跑 build_candidates_A.py
  --target-size 50 --ontology data/omd_ontology.json --test-dataset
  data/test_dataset.json，一小时内就能拿到正确的候选集。

  如果不能，路径 B 是可行的 fallback：build_candidates_A.py 本身支持空本体退化——不给 ontology 时，Tier 1/2
  自动跳过，纯走 Tier 3 语义召回。唯一需要的是把 160 个 gold（或全语料 1743 个概念）喂进 concepts 列表。

  在我开始建 config/utils/candidate_mass 三件套之前——omd_ontology.json 能拿到吗？ 这个答案决定我接下来是走路径 A 还是
  B。

✻ Baked for 1m 56s

❯ C:\Users\19691\Desktop\投稿论文；C:\Users\19691\Desktop\论文与研究；你可以从这两个文件夹中检索并构建可用的数据集吗

  Listed 6 directories (ctrl+o to expand)

● 发现了两个 4MB+ 的 RDF 文件——这很可能是 OMD 本体导出。让我深入检查。

  Read 1 file, listed 2 directories (ctrl+o to expand)

● 找到了！827.rdf / 9.12.rdf 就是 OMD 本体。让我快速解析结构，确认能做候选集构造。

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import sys, io
   sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')

   # Quick parse of RDF to understand its structure
   import xml.etree.ElementTree as ET
  三条路

  | 路径 | 操作 | |C| | 候选质量 | 时间 |
  |:--|:--|:--:|:--:|:--:|
  | A（推荐） | 导出 omd_ontology.json → 跑 build_candidates_A.py --target-size 50 | 20-100 | Tier1+2+3，本体结构保证 |
  取决于本体导出难度 |
  | B（权宜） | 不用本体，用纯 Tier 3（embedding 语义召回），以全部 160 个 gold 作概念池，对每个术语召回 top-30 语义近邻
   | ~30 | 仅语义，无层级结构 | 半天 |
  | C（冒险） | 用现有 |C|=8 直接跑 pilot | 8 | 随机噪音 | 0 |

  路径 C 几乎确定 Pilot C 会亮红灯——8 个随机候选中 gold_share 会在 0.3-0.7
  窄带里挤成团，没有动态范围。这就是你预警的那个问题。

  ---
  推荐动作

  先确定路径 A 是否可行：omd_ontology.json 能从 Protege / OMD 项目里导出来吗？如果能，跑 build_candidates_A.py
  --target-size 50 --ontology data/omd_ontology.json --test-dataset
  data/test_dataset.json，一小时内就能拿到正确的候选集。

  如果不能，路径 B 是可行的 fallback：build_candidates_A.py 本身支持空本体退化——不给 ontology 时，Tier 1/2
  自动跳过，纯走 Tier 3 语义召回。唯一需要的是把 160 个 gold（或全语料 1743 个概念）喂进 concepts 列表。

  在我开始建 config/utils/candidate_mass 三件套之前——omd_ontology.json 能拿到吗？ 这个答案决定我接下来是走路径 A 还是
  B。

✻ Baked for 1m 56s

❯ C:\Users\19691\Desktop\投稿论文；C:\Users\19691\Desktop\论文与研究；你可以从这两个文件夹中检索并构建可用的数据集吗

  Listed 6 directories (ctrl+o to expand)

● 发现了两个 4MB+ 的 RDF 文件——这很可能是 OMD 本体导出。让我深入检查。

  Read 1 file, listed 2 directories (ctrl+o to expand)

● 找到了！827.rdf / 9.12.rdf 就是 OMD 本体。让我快速解析结构，确认能做候选集构造。

● Bash(python -c "
      import sys, io…)
  ⎿  === Top-level elements ===
       NamedIndividual: 7728
       Class: 612
     … +27 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import sys, io
   sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
   import xml.etree.ElementTree as ET

   tree = ET.parse(r'C:\Users\19691\Desktop\论文与研究\9.12.rdf')
   root = tree.getroot()

   ns = {'rdf': 'http://www.w3.org/1999/02/22-rdf-syntax-ns#',
         'rdfs': 'http://www.w3.org/2000/01/rdf-schema#',
         'owl': 'http://www.w3.org/2002/07/owl#',
         'ns0': 'http://ResearchPaper.com/ontology#',
         'ns2': 'http://ResearchPaper.com/ontology/Relation#',
         'ns1': 'http://ResearchPaper.com/ontology/dataproperty#'}

   # 1. Find all relations (ObjectPropertyAssertion or similar)
   from collections import Counter
   relations = Counter()
   # Look for elements that represent edges/relations
   for elem in root:
       tag = elem.tag.split('}')[-1] if '}' in elem.tag else elem.tag
       # Check for ns2: relations
       for child in elem:
           ctag = child.tag.split('}')[-1] if '}' in child.tag else child.tag
           # Some may be relation types
           pass

       # Also check for ObjectProperty
       if tag == 'ObjectProperty':
           for child in elem:
               if 'label' in child.tag:
                   relations['ObjectProperty'] += 1

   # Better: search by namespace
   | ~30 | 仅语义，无层级结构 | 半天 |
  | C（冒险） | 用现有 |C|=8 直接跑 pilot | 8 | 随机噪音 | 0 |

  路径 C 几乎确定 Pilot C 会亮红灯——8 个随机候选中 gold_share 会在 0.3-0.7
  窄带里挤成团，没有动态范围。这就是你预警的那个问题。

  ---
  推荐动作

  先确定路径 A 是否可行：omd_ontology.json 能从 Protege / OMD 项目里导出来吗？如果能，跑 build_candidates_A.py
  --target-size 50 --ontology data/omd_ontology.json --test-dataset
  data/test_dataset.json，一小时内就能拿到正确的候选集。

  如果不能，路径 B 是可行的 fallback：build_candidates_A.py 本身支持空本体退化——不给 ontology 时，Tier 1/2
  自动跳过，纯走 Tier 3 语义召回。唯一需要的是把 160 个 gold（或全语料 1743 个概念）喂进 concepts 列表。

  在我开始建 config/utils/candidate_mass 三件套之前——omd_ontology.json 能拿到吗？ 这个答案决定我接下来是走路径 A 还是
  B。

✻ Baked for 1m 56s

❯ C:\Users\19691\Desktop\投稿论文；C:\Users\19691\Desktop\论文与研究；你可以从这两个文件夹中检索并构建可用的数据集吗

  Listed 6 directories (ctrl+o to expand)

● 发现了两个 4MB+ 的 RDF 文件——这很可能是 OMD 本体导出。让我深入检查。

  Read 1 file, listed 2 directories (ctrl+o to expand)

● 找到了！827.rdf / 9.12.rdf 就是 OMD 本体。让我快速解析结构，确认能做候选集构造。

● Bash(python -c "
      import sys, io…)
  ⎿  === Top-level elements ===
       NamedIndividual: 7728
       Class: 612
     … +27 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === Relations (ns2) ===
       FormulaMappingTerm: 1356
       KnowledgeMappingTerm: 976
     … +30 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import sys, io
  推荐动作

  先确定路径 A 是否可行：omd_ontology.json 能从 Protege / OMD 项目里导出来吗？如果能，跑 build_candidates_A.py
  --target-size 50 --ontology data/omd_ontology.json --test-dataset
  data/test_dataset.json，一小时内就能拿到正确的候选集。

  如果不能，路径 B 是可行的 fallback：build_candidates_A.py 本身支持空本体退化——不给 ontology 时，Tier 1/2
  自动跳过，纯走 Tier 3 语义召回。唯一需要的是把 160 个 gold（或全语料 1743 个概念）喂进 concepts 列表。

  在我开始建 config/utils/candidate_mass 三件套之前——omd_ontology.json 能拿到吗？ 这个答案决定我接下来是走路径 A 还是
  B。

✻ Baked for 1m 56s

❯ C:\Users\19691\Desktop\投稿论文；C:\Users\19691\Desktop\论文与研究；你可以从这两个文件夹中检索并构建可用的数据集吗

  Listed 6 directories (ctrl+o to expand)

● 发现了两个 4MB+ 的 RDF 文件——这很可能是 OMD 本体导出。让我深入检查。

  Read 1 file, listed 2 directories (ctrl+o to expand)

● 找到了！827.rdf / 9.12.rdf 就是 OMD 本体。让我快速解析结构，确认能做候选集构造。

● Bash(python -c "
      import sys, io…)
  ⎿  === Top-level elements ===
       NamedIndividual: 7728
       Class: 612
     … +27 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === Relations (ns2) ===
       FormulaMappingTerm: 1356
       KnowledgeMappingTerm: 976
     … +30 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 19, in <module>
     NameError: name 'Counter' is not defined

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
  自动跳过，纯走 Tier 3 语义召回。唯一需要的是把 160 个 gold（或全语料 1743 个概念）喂进 concepts 列表。

  在我开始建 config/utils/candidate_mass 三件套之前——omd_ontology.json 能拿到吗？ 这个答案决定我接下来是走路径 A 还是
  B。

✻ Baked for 1m 56s

❯ C:\Users\19691\Desktop\投稿论文；C:\Users\19691\Desktop\论文与研究；你可以从这两个文件夹中检索并构建可用的数据集吗

  Listed 6 directories (ctrl+o to expand)

● 发现了两个 4MB+ 的 RDF 文件——这很可能是 OMD 本体导出。让我深入检查。

  Read 1 file, listed 2 directories (ctrl+o to expand)

● 找到了！827.rdf / 9.12.rdf 就是 OMD 本体。让我快速解析结构，确认能做候选集构造。

● Bash(python -c "
      import sys, io…)
  ⎿  === Top-level elements ===
       NamedIndividual: 7728
       Class: 612
     … +27 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === Relations (ns2) ===
       FormulaMappingTerm: 1356
       KnowledgeMappingTerm: 976
     … +30 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 19, in <module>
     NameError: name 'Counter' is not defined

● Bash(python -c "
      import sys, io…)
  ⎿  === Individual types ===
       TechnicalTerm: 570
       MathematicalKnowledgeComponent: 276
     … +244 lines (ctrl+o to expand)

● 数据源找到了。以下是完整诊断和方案。

  ---
  发现：9.12.rdf 包含了全部需要的数据

  827.rdf (4.15MB) / 9.12.rdf (4.22MB) 就是 OMD 本体完整导出：

  ┌────────────────────────────────────────────┬──────────────────┬─────────────────────────────────┐
  │                    实体                    │       数量       │              用途               │
  ├────────────────────────────────────────────┼──────────────────┼─────────────────────────────────┤
  │ TechnicalTerm（古籍术语）                  │       570        │ 与 test_dataset 的 160 条取交集 │
  ├────────────────────────────────────────────┼──────────────────┼─────────────────────────────────┤
  │ MathematicalKnowledgeComponent（现代概念） │       276        │ 候选集全集                      │
  ├────────────────────────────────────────────┼──────────────────┼─────────────────────────────────┤
  │ KnowledgeMappingTerm 边                    │       976        │ 术语→现代概念的黄金映射         │
  ├────────────────────────────────────────────┼──────────────────┼─────────────────────────────────┤
  │ hasPrimaryMSCClassification                │ 每条知识组件都有 │ Tier 1 MSC 同目召回             │
  ├────────────────────────────────────────────┼──────────────────┼─────────────────────────────────┤
  │ MathematicalFormula                        │       240        │ 附加语义关联                    │
  └────────────────────────────────────────────┴──────────────────┴─────────────────────────────────┘

  且 build_candidates_A.py 自带 load_ontology() 函数，接受 JSON 格式本体。只需把 RDF → JSON 转换一次即可驱动候选集构造。

  ---
  施工路径（精简为两步）

  第一步：RDF → JSON 本体转换（1-2小时）

  解析 9.12.rdf，提取 build_candidates_A.py 需要的四个字段：

  concepts        ← 所有 MathematicalKnowledgeComponent 的 name
  term_to_msc     ← TechnicalTerm → hasPrimaryMSCClassification
  msc_to_concepts ← MSC码 → 归属该码的知识组件列表
  edges           ← KnowledgeMappingTerm, EVOLUTION, SUBSUMES 等关系

  写一个 rdf_to_ontology_json.py，一次转换，产出 data/omd_ontology.json。

  第二步：跑 build_candidates_A.py

  python build_candidates_A.py \
      --test-dataset data/test_dataset.json \
      --ontology data/omd_ontology.json \
      --target-size 50 \
      --embed-model shibing624/text2vec-base-chinese \
      --verify-gold-blind

  预期产出：
  - data/dataset_a_with_candidates.json（|C(t)| ≈ 30–50，含 provenance 追溯）
  - data/synonyms_A.json（从 EQUIVALENT_TO 闭包自动生成）
  - data/candidates_A_audit.csv（逐条审计 + chance level）

  然后现有那版 |C|=8 的随机候选集直接废弃。

  ---
  回答核心问题

  |C(t)| 能不能从 8 扩到 20+？ 能。276 个现代概念 + MSC 层级 + 976 条映射边足以支撑 |C| 在 20–80 之间，Pilot C
  的动态范围问题有解。

  需要多久？ RDF→JSON 转换 1-2 小时 + 候选集构造 10 分钟（不加载 embedding 模型则用 n-gram 退化，更快但质量稍低；加
  embedding 需 GPU/网络下载模型）。

  ---
  要不要我现在写 rdf_to_ontology_json.py？这是整个管道的上游——这一步完成，build_candidates_A.py 就能跑出正确的候选集。

✻ Baked for 2m 0s

❯ 把所需要的文件打包给我，放在桌面上，新建文件夹我来处理

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import sys, io
   sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
   import xml.etree.ElementTree as ET
   from collections import Counter

   tree = ET.parse(r'C:\Users\19691\Desktop\论文与研究\9.12.rdf')
   root = tree.getroot()

   # Sample a few more TechnicalTerm to see if they have MSC directly
   count = 0
   for elem in root:
       tag = elem.tag.split('}')[-1] if '}' in elem.tag else elem.tag
       if tag != 'NamedIndividual':
           continue
       about = elem.get('{http://www.w3.org/1999/02/22-rdf-syntax-ns#}about', '')
       rdf_type = None
       props = {}
       for child in elem:
           ctag = child.tag.split('}')[-1] if '}' in child.tag else child.tag
           if ctag == 'type':
               rdf_type = child.get('{http://www.w3.org/1999/02/22-rdf-syntax-ns#}resource', '').split('#')[-1]
           else:
               val = (child.text or '').strip()
               res = child.get('{http://www.w3.org/1999/02/22-rdf-syntax-ns#}resource', '')
               if res:
                   val = res.split('#')[-1]
               props.setdefault(ctag, []).append(val)

       if rdf_type == 'TechnicalTerm':
  msc_to_concepts ← MSC码 → 归属该码的知识组件列表
  edges           ← KnowledgeMappingTerm, EVOLUTION, SUBSUMES 等关系

  写一个 rdf_to_ontology_json.py，一次转换，产出 data/omd_ontology.json。

  第二步：跑 build_candidates_A.py

  python build_candidates_A.py \
      --test-dataset data/test_dataset.json \
      --ontology data/omd_ontology.json \
      --target-size 50 \
      --embed-model shibing624/text2vec-base-chinese \
      --verify-gold-blind

  预期产出：
  - data/dataset_a_with_candidates.json（|C(t)| ≈ 30–50，含 provenance 追溯）
  - data/synonyms_A.json（从 EQUIVALENT_TO 闭包自动生成）
  - data/candidates_A_audit.csv（逐条审计 + chance level）

  然后现有那版 |C|=8 的随机候选集直接废弃。

  ---
  回答核心问题

  |C(t)| 能不能从 8 扩到 20+？ 能。276 个现代概念 + MSC 层级 + 976 条映射边足以支撑 |C| 在 20–80 之间，Pilot C
  的动态范围问题有解。

  需要多久？ RDF→JSON 转换 1-2 小时 + 候选集构造 10 分钟（不加载 embedding 模型则用 n-gram 退化，更快但质量稍低；加
  embedding 需 GPU/网络下载模型）。

  ---
  要不要我现在写 rdf_to_ontology_json.py？这是整个管道的上游——这一步完成，build_candidates_A.py 就能跑出正确的候选集。

✻ Baked for 2m 0s

❯ 把所需要的文件打包给我，放在桌面上，新建文件夹我来处理

● Bash(python -c "
      import sys, io…)
  ⎿  TERM with MSC: 角较度
       hasPrimaryMSCClassification: ['85A04']
     TERM with MSC: 连比例
     … +14 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import sys, io
   sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
   import xml.etree.ElementTree as ET
   from collections import defaultdict

   tree = ET.parse(r'C:\Users\19691\Desktop\论文与研究\9.12.rdf')
   root = tree.getroot()

   # Check: what are the ObjectProperty definitions?

  python build_candidates_A.py \
      --test-dataset data/test_dataset.json \
      --ontology data/omd_ontology.json \
      --target-size 50 \
      --embed-model shibing624/text2vec-base-chinese \
      --verify-gold-blind

  预期产出：
  - data/dataset_a_with_candidates.json（|C(t)| ≈ 30–50，含 provenance 追溯）
  - data/synonyms_A.json（从 EQUIVALENT_TO 闭包自动生成）
  - data/candidates_A_audit.csv（逐条审计 + chance level）

  然后现有那版 |C|=8 的随机候选集直接废弃。

  ---
  回答核心问题

  |C(t)| 能不能从 8 扩到 20+？ 能。276 个现代概念 + MSC 层级 + 976 条映射边足以支撑 |C| 在 20–80 之间，Pilot C
  的动态范围问题有解。

  需要多久？ RDF→JSON 转换 1-2 小时 + 候选集构造 10 分钟（不加载 embedding 模型则用 n-gram 退化，更快但质量稍低；加
  embedding 需 GPU/网络下载模型）。

  ---
  要不要我现在写 rdf_to_ontology_json.py？这是整个管道的上游——这一步完成，build_candidates_A.py 就能跑出正确的候选集。

✻ Baked for 2m 0s

❯ 把所需要的文件打包给我，放在桌面上，新建文件夹我来处理

● Bash(python -c "
      import sys, io…)
  ⎿  TERM with MSC: 角较度
       hasPrimaryMSCClassification: ['85A04']
     TERM with MSC: 连比例
     … +14 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === ObjectProperty definitions ===
       Contains:
       ContainsFormula:
     … +19 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   python -c "
   import sys, io
   sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
   import xml.etree.ElementTree as ET
   from collections import defaultdict

   tree = ET.parse(r'C:\Users\19691\Desktop\论文与研究\9.12.rdf')
   root = tree.getroot()

   # Check MSC property namespace

  预期产出：
  - data/dataset_a_with_candidates.json（|C(t)| ≈ 30–50，含 provenance 追溯）
  - data/synonyms_A.json（从 EQUIVALENT_TO 闭包自动生成）
  - data/candidates_A_audit.csv（逐条审计 + chance level）

  然后现有那版 |C|=8 的随机候选集直接废弃。

  ---
  回答核心问题

  |C(t)| 能不能从 8 扩到 20+？ 能。276 个现代概念 + MSC 层级 + 976 条映射边足以支撑 |C| 在 20–80 之间，Pilot C
  的动态范围问题有解。

  需要多久？ RDF→JSON 转换 1-2 小时 + 候选集构造 10 分钟（不加载 embedding 模型则用 n-gram 退化，更快但质量稍低；加
  embedding 需 GPU/网络下载模型）。

  ---
  要不要我现在写 rdf_to_ontology_json.py？这是整个管道的上游——这一步完成，build_candidates_A.py 就能跑出正确的候选集。

✻ Baked for 2m 0s

❯ 把所需要的文件打包给我，放在桌面上，新建文件夹我来处理

● Bash(python -c "
      import sys, io…)
  ⎿  TERM with MSC: 角较度
       hasPrimaryMSCClassification: ['85A04']
     TERM with MSC: 连比例
     … +14 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === ObjectProperty definitions ===
       Contains:
       ContainsFormula:
     … +19 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === MSC property samples ===
       hasPrimaryMSCClassification: 05A10
       hasRelatedMSCClassification: 05A17
     … +13 lines (ctrl+o to expand)

● 现在全部信息已就绪，开始打包。

● Bash(mkdir -p "C:/Users/19691/Desktop/J-space-候选集构建包" 2>&1)
  ⎿  Done

● Write(Desktop\J-space-候选集构建包\rdf_to_ontology_json.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 Desktop\J-space-候选集构建包\rdf_to_ontology_json.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
    1 #!/usr/bin/env python3
    2 # -*- coding: utf-8 -*-
    3 """
    4 rdf_to_ontology_json.py
    5 =======================
    6 将 OMD RDF/XML 本体（9.12.rdf）转换为 build_candidates_A.py 所需的 JSON 格式。
    7
    8 输入: 9.12.rdf（或 827.rdf）
    9 输出: omd_ontology.json
   10
   11 用法:
   12     python rdf_to_ontology_json.py --rdf 9.12.rdf --out omd_ontology.json
   13
   14 依赖: 仅 Python 标准库（xml.etree.ElementTree）
   15 """
   16
   17 import argparse
   18 import json
   19 import sys
   20 from collections import defaultdict
   21 from xml.etree import ElementTree as ET
   22
   23 # ========== 命名空间 ==========
   24 NS_RDF  = "http://www.w3.org/1999/02/22-rdf-syntax-ns#"
   25 NS_OWL  = "http://www.w3.org/2002/07/owl#"
   26 NS_BASE = "http://ResearchPaper.com/ontology#"
   27 NS_REL  = "http://ResearchPaper.com/ontology/Relation#"
   28 NS_MSC  = "http://www.co-ode.org/ontologies/ont.owl#"
   29 NS_ONT  = "http://www.co-ode.org/ontologies/ont.owl#"
   30
   31
   32 def tag_local(elem):
   33     return elem.tag.split("}")[-1] if "}" in elem.tag else elem.tag
   34
   35
   36 def ns_of(elem):
   37     return elem.tag.split("}")[0].strip("{") if "}" in elem.tag else ""
   38
   39
   40 def parse_rdf(path):
   41     """Parse RDF/XML, return dicts keyed by individual ID (the fragment after #)."""
   42     tree = ET.parse(path)
   43     root = tree.getroot()
   44
   45     # Collect all NamedIndividuals
   46     individuals = {}   # {fragment_id: {"type": str, "props": {prop_name: [values]}}}
   47
   48     for elem in root:
   49         if tag_local(elem) != "NamedIndividual":
   50             continue
   51         about = elem.get(f"{{{NS_RDF}}}about", "")
   52         frag = about.split("#")[-1] if "#" in about else about
   53         if not frag:
   54             continue
   55
   56         record = {"type": None, "props": defaultdict(list)}
   57         for child in elem:
   58             ctag = tag_local(child)
   59             child_ns = ns_of(child)
   60
   61             if ctag == "type":
   62                 res = child.get(f"{{{NS_RDF}}}resource", "")
   63                 record["type"] = res.split("#")[-1]
   64             elif child_ns == NS_REL:
   65                 # Relation edge: source = this individual, target = resource
   66                 dst = child.get(f"{{{NS_RDF}}}resource", "")
   67                 dst_frag = dst.split("#")[-1] if "#" in dst else dst
   68                 if dst_frag:
   69                     record["props"][ctag].append(dst_frag)
   70             elif child_ns in (NS_MSC, NS_ONT):
   71                 # MSC classification
   72                 dst = child.get(f"{{{NS_RDF}}}resource", "")
   73                 dst_frag = dst.split("#")[-1] if "#" in dst else dst
   74                 if dst_frag:
   75                     record["props"][ctag].append(dst_frag)
   76             else:
   77                 # Data property (name, description, etc.)
   78                 val = (child.text or "").strip()
   79                 if val:
   80                     record["props"][ctag].append(val)
   81
   82         individuals[frag] = record
   83
   84     return individuals
   85
   86
   87 def build_ontology(individuals):
   88     """Convert parsed RDF individuals into build_candidates_A ontology format."""
   89
   90     # ---- concepts: all MathematicalKnowledgeComponent names ----
   91     concepts = []
   92     for frag, rec in individuals.items():
   93         if rec["type"] == "MathematicalKnowledgeComponent":
   94             name = rec["props"].get("name", [frag])[0]
   95             concepts.append(name)
   96
   97     # Also include any concept-like individuals that have MSC classifications
   98     concept_names = set(concepts)
   99     for frag, rec in individuals.items():
  100         if rec["type"] in ("MathematicalKnowledgeComponent",):
  101             name = rec["props"].get("name", [frag])[0]
  102             concept_names.add(name)
  103
  104     concepts = sorted(concept_names)
  105
  106     # ---- msc_to_concepts: MSC code → list of concept names ----
  107     msc_to_concepts = defaultdict(set)
  108     for frag, rec in individuals.items():
  109         if rec["type"] != "MathematicalKnowledgeComponent":
  110             continue
  111         name = rec["props"].get("name", [frag])[0]
  112         for msc_list_key in ("hasPrimaryMSCClassification", "hasRelatedMSCClassification"):
  113             for msc in rec["props"].get(msc_list_key, []):
  114                 msc_to_concepts[msc].add(name)
  115
  116     msc_to_concepts = {k: sorted(v) for k, v in msc_to_concepts.items()}
  117
  118     # ---- term_to_msc: term name → MSC codes ----
  119     term_to_msc = defaultdict(set)
  120
  121     for frag, rec in individuals.items():
  122         if rec["type"] == "TechnicalTerm":
  123             term_name = rec["props"].get("name", [frag])[0]
  124             # Direct MSC on term
  125             for msc_list_key in ("hasPrimaryMSCClassification", "hasRelatedMSCClassification"):
  126                 for msc in rec["props"].get(msc_list_key, []):
  127                     term_to_msc[term_name].add(msc)
  128
  129     # Also derive term→MSC from KnowledgeMappingTerm links
  130     for frag, rec in individuals.items():
  131         if rec["type"] == "MathematicalKnowledgeComponent":
  132             msc_codes = set()
  133             for msc_list_key in ("hasPrimaryMSCClassification", "hasRelatedMSCClassification"):
  134                 msc_codes.update(rec["props"].get(msc_list_key, []))
  135             if not msc_codes:
  136                 continue
  137             # All terms this knowledge maps to
  138             for term_name in rec["props"].get("KnowledgeMappingTerm", []):
  139                 term_to_msc[term_name].update(msc_codes)
  140
  141     term_to_msc = {k: sorted(v) for k, v in term_to_msc.items()}
  142
  143     # ---- edges: (src_name, relation, dst_name) ----
  144     edges = []
  145
  146     # We need name lookups to convert frag IDs to display names
  147     frag_to_name = {}
  148     for frag, rec in individuals.items():
  149         name = rec["props"].get("name", [frag])[0]
  150         frag_to_name[frag] = name
  151
  152     for frag, rec in individuals.items():
  153         src_name = rec["props"].get("name", [frag])[0]
  154         for rel_name, targets in rec["props"].items():
  155             # Only keep relation-type properties (not data properties)
  156             if rel_name in ("name", "description", "original_text", "position",
  157                             "latex", "hasPrimaryMSCClassification",
  158                             "hasRelatedMSCClassification"):
  159                 continue
  160             for tgt_frag in targets:
  161                 tgt_name = frag_to_name.get(tgt_frag, tgt_frag)
  162                 edges.append([src_name, rel_name, tgt_name])
  163
  164     # ---- Build output ----
  165     onto = {
  166         "concepts": concepts,
  167         "term_to_msc": term_to_msc,
  168         "msc_to_concepts": msc_to_concepts,
  169         "edges": edges,
  170     }
  171
  172     return onto
  173
  174
  175 def main():
  176     ap = argparse.ArgumentParser(
  177         description="将 OMD RDF/XML 本体转换为 build_candidates_A.py 所需的 JSON 格式")
  178     ap.add_argument("--rdf", required=True, help="输入 RDF 文件路径（如 9.12.rdf）")
  179     ap.add_argument("--out", default="omd_ontology.json", help="输出 JSON 文件路径")
  180     args = ap.parse_args()
  181

  |C(t)| 能不能从 8 扩到 20+？ 能。276 个现代概念 + MSC 层级 + 976 条映射边足以支撑 |C| 在 20–80 之间，Pilot C
  的动态范围问题有解。

  需要多久？ RDF→JSON 转换 1-2 小时 + 候选集构造 10 分钟（不加载 embedding 模型则用 n-gram 退化，更快但质量稍低；加
  embedding 需 GPU/网络下载模型）。

  ---
  要不要我现在写 rdf_to_ontology_json.py？这是整个管道的上游——这一步完成，build_candidates_A.py 就能跑出正确的候选集。

✻ Baked for 2m 0s

❯ 把所需要的文件打包给我，放在桌面上，新建文件夹我来处理

● Bash(python -c "
      import sys, io…)
  ⎿  TERM with MSC: 角较度
       hasPrimaryMSCClassification: ['85A04']
     TERM with MSC: 连比例
     … +14 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === ObjectProperty definitions ===
       Contains:
       ContainsFormula:
     … +19 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === MSC property samples ===
       hasPrimaryMSCClassification: 05A10
       hasRelatedMSCClassification: 05A17
     … +13 lines (ctrl+o to expand)

● 现在全部信息已就绪，开始打包。

● Bash(mkdir -p "C:/Users/19691/Desktop/J-space-候选集构建包" 2>&1)
  ⎿  Done

● Write(Desktop\J-space-候选集构建包\rdf_to_ontology_json.py)
  ⎿  Wrote 223 lines to Desktop\J-space-候选集构建包\rdf_to_ontology_json.py
      1 #!/usr/bin/env python3
      2 # -*- coding: utf-8 -*-
      3 """
      4 rdf_to_ontology_json.py
      5 =======================
      6 将 OMD RDF/XML 本体（9.12.rdf）转换为 build_candidates_A.py 所需的 JSON 格式。
      7
      8 输入: 9.12.rdf（或 827.rdf）
      9 输出: omd_ontology.json
     … +213 lines (ctrl+o to expand)

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/build_candidates_A (2).py"
      "C:/Users/19691/Desktop/J-space-候选集构建包/build_candidates_A.py" 2>&1)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/data/test_dataset.json"
      "C:/Users/19691/Desktop/J-space-候选集构建包/test_dataset.json" 2>&1)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/论文与研究/9.12.rdf"
      "C:/Users/19691/Desktop/J-space-候选集构建包/omd_ontology_9.12.rdf" 2>&1)
  ⎿  Done

● Write(Desktop\J-space-候选集构建包\README_操作说明.md)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 Desktop\J-space-候选集构建包\README_操作说明.md
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
   1 # J-space 候选集构建包
   2
   3 本包用于从 OMD 本体 RDF 重新生成 Dataset A 的正确候选集 C(t)，将 |C(t)| 从当前的 8 扩到 20–50，解决 Pilot C（动
     态范围）的红灯风险。
   4
   5 ---
   6
   7 ## 包含文件
   8
   9 | 文件 | 说明 |
  10 |:--|:--|
  11 | `rdf_to_ontology_json.py` | **第1步**：RDF → JSON 本体转换 |
  12 | `omd_ontology_9.12.rdf` | OMD 本体 RDF/XML（4.2MB，最新版） |
  13 | `build_candidates_A.py` | **第2步**：gold-blind 候选集构造 |
  14 | `test_dataset.json` | 160 条古籍数学术语（带 MSC codes） |
  15
  16 ---
  17
  18 ## 操作步骤
  19
  20 ### 第1步：RDF → JSON 转换
  21
  22 ```bash
  23 python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf --out omd_ontology.json
  24 ```
  25
  26 预期输出：
  27 - concepts: ~276 个现代数学概念
  28 - term_to_msc: 数百条术语→MSC 映射
  29 - msc_to_concepts: 238 个 MSC 类目
  30 - edges: ~3000+ 条关系
  31
  32 ### 第2步：候选集构造（仅需标准库 + numpy）
  33
  34 ```bash
  35 # 基础版（字符 n-gram 语义召回，无外部依赖）
  36 python build_candidates_A.py \
  37     --test-dataset test_dataset.json \
  38     --ontology omd_ontology.json \
  39     --out dataset_a_with_candidates.json \
  40     --synonyms-out synonyms_A.json \
  41     --audit-out candidates_A_audit.csv \
  42     --target-size 50 \
  43     --hops 2 \
  44     --verify-gold-blind
  45 ```
  46
  47 如果有网络/GPU，加 embedding 后端提升语义召回质量：
  48 ```bash
  49 pip install sentence-transformers
  50 # 加上 --embed-model shibing624/text2vec-base-chinese
  51 ```
  52
  53 ### 第3步：触发阶梯放宽（如果覆盖率 < 100%）
       ContainsFormula:
     … +19 lines (ctrl+o to expand)

● Bash(python -c "
      import sys, io…)
  ⎿  === MSC property samples ===
       hasPrimaryMSCClassification: 05A10
       hasRelatedMSCClassification: 05A17
     … +13 lines (ctrl+o to expand)

● 现在全部信息已就绪，开始打包。

● Bash(mkdir -p "C:/Users/19691/Desktop/J-space-候选集构建包" 2>&1)
  ⎿  Done

● Write(Desktop\J-space-候选集构建包\rdf_to_ontology_json.py)
  ⎿  Wrote 223 lines to Desktop\J-space-候选集构建包\rdf_to_ontology_json.py
      1 #!/usr/bin/env python3
      2 # -*- coding: utf-8 -*-
      3 """
      4 rdf_to_ontology_json.py
      5 =======================
      6 将 OMD RDF/XML 本体（9.12.rdf）转换为 build_candidates_A.py 所需的 JSON 格式。
      7
      8 输入: 9.12.rdf（或 827.rdf）
      9 输出: omd_ontology.json
     … +213 lines (ctrl+o to expand)

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/build_candidates_A (2).py"
      "C:/Users/19691/Desktop/J-space-候选集构建包/build_candidates_A.py" 2>&1)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/data/test_dataset.json"
      "C:/Users/19691/Desktop/J-space-候选集构建包/test_dataset.json" 2>&1)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/论文与研究/9.12.rdf"
      "C:/Users/19691/Desktop/J-space-候选集构建包/omd_ontology_9.12.rdf" 2>&1)
  ⎿  Done

● Write(Desktop\J-space-候选集构建包\README_操作说明.md)
  ⎿  Wrote 95 lines to Desktop\J-space-候选集构建包\README_操作说明.md
       2 # J-space 候选集构建包
       3 本包用于从 OMD 本体 RDF 重新生成 Dataset A 的正确候选集 C(t)，将 |C(t)| 从当前的 8 扩到 20–50，解决 Pilo
         t C（动态范围）的红灯风险。
       4
       5 ---
       6
       7 ## 包含文件
       8
       9 | 文件 | 说明 |
      10 |:--|:--|
     … +86 lines (ctrl+o to expand)

● 现在验证转换脚本能在 RDF 上正常运行。

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf
      --out omd_ontology.json 2>&1)
  ⎿  Error: Exit code 1
          omd_ontology_9.12.rdf   
       ʵ      : 7704
       TechnicalTerm: 570
       MathematicalKnowledgeComponent: 271
       MathematicalFormula: 221

     :     ת
       concepts ( ִ      ): 271
       term_to_msc ӳ  : 292       
     … +8 lines (ctrl+o to expand)

● 转换逻辑完全正确（271 概念、292 术语-MSC、3444 边），只是 Windows GBK 终端吞了 ✓ 字符。修复打印并验证 JSON
  已成功写出。

  Read 1 file (ctrl+o to expand)

● Update(Desktop\J-space-候选集构建包\rdf_to_ontology_json.py)
  ⎿  Added 2 lines, removed 2 lines
      203      with open(args.out, "w", encoding="utf-8") as f:
      204          json.dump(onto, f, ensure_ascii=False, indent=2)
      205
      206 -    print(f"\n✓ 已输出: {args.out}")
      206 +    print(f"\n[OK] 已输出: {args.out}")
      207      print(f"  文件大小: {len(json.dumps(onto, ensure_ascii=False)) / 1024:.0f} KB")
      208
      209      # 自检
     ...
      216      if onto["term_to_msc"]:
      217          sample_terms = list(onto["term_to_msc"].keys())[:3]
      218          for t in sample_terms:
      219 -            print(f"  {t} → MSC: {onto['term_to_msc'][t][:3]}")
      219 +            print(f"  {t} -> MSC: {onto['term_to_msc'][t][:3]}")
      220
      221
      222  if __name__ == "__main__":

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf
      --out omd_ontology.json 2>&1)
  ⎿       omd_ontology_9.12.rdf   
       ʵ      : 7704
       TechnicalTerm: 570
     … +19 lines (ctrl+o to expand)

● 转换成功。现在验证整个管道能跑通。

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python -c "
      import json…)
  ⎿  concepts: 271
     term_to_msc: 292 terms
     msc_to_concepts: 219 MSC codes
     … +9 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python -c "
      import sys, io…)
  ⎿  Empty/nan concepts: 1
       ['nan']
     … +3 lines (ctrl+o to expand)

  Read 1 file (ctrl+o to expand)

● Update(rdf_to_ontology_json.py)
  ⎿  Added 4 lines, removed 1 line
      161                  tgt_name = frag_to_name.get(tgt_frag, tgt_frag)
      162                  edges.append([src_name, rel_name, tgt_name])
      163
      164 +    # ---- Filter bad concepts ----
      165 +    concepts = [c for c in concepts if c and c.strip() and c.lower() != 'nan']
      166 +
      167      # ---- Build output ----
      168      onto = {
      169          "concepts": concepts,
      170          "term_to_msc": term_to_msc,
      168 -        "msc_to_concepts": msc_to_concepts,
      171 +        "msc_to_concepts": {k: sorted(v) for k, v in msc_to_concepts.items()},
      172          "edges": edges,
      173      }
      174

● Update(rdf_to_ontology_json.py)
  ⎿  Added 2 lines
      180          description="将 OMD RDF/XML 本体转换为 build_candidates_A.py 所需的 JSON 格式")
      181      ap.add_argument("--rdf", required=True, help="输入 RDF 文件路径（如 9.12.rdf）")
      182      ap.add_argument("--out", default="omd_ontology.json", help="输出 JSON 文件路径")
      183 +    ap.add_argument("--test-dataset", default=None,
      184 +                    help="可选：test_dataset.json，将其 msc_codes 合并入 term_to_msc（仍是 gold-blind
          +，MSC 编码的是术语领域而非答案）")
      185      args = ap.parse_args()
      186
      187      print(f"解析 {args.rdf} …")

● Update(rdf_to_ontology_json.py)
  ⎿  Added 26 lines
      199
      200      onto = build_ontology(individuals)
      201
      202 +    # ---- Merge test_dataset MSC codes (optional) ----
      203 +    if args.test_dataset and args.test_dataset.strip():
      204 +        try:
      205 +            with open(args.test_dataset, "r", encoding="utf-8") as f:
      206 +                ds = json.load(f)
      207 +            cases = ds.get("test_cases", ds) if isinstance(ds, dict) else ds
      208 +            merged = 0
      209 +            for c in cases:
      210 +                name = c.get("name", "")
      211 +                msc_codes = c.get("msc_codes", [])
      212 +                if name and msc_codes:
      213 +                    clean = [m.split("#")[-1] for m in msc_codes]
      214 +                    existing = set(onto["term_to_msc"].get(name, []))
      215 +                    before = len(existing)
      216 +                    existing.update(clean)
      217 +                    if len(existing) > before:
      218 +                        onto["term_to_msc"][name] = sorted(existing)
      219 +                        merged += 1
      220 +                        # Also populate msc_to_concepts
      221 +                        for m in clean:
      222 +                            if m not in onto["msc_to_concepts"]:
      223 +                                onto["msc_to_concepts"][m] = []
      224 +            print(f"  从 test_dataset 合并 MSC: {merged} 个术语")
      225 +        except Exception as e:
      226 +            print(f"  [warn] 无法读取 test_dataset: {e}")
      227 +
      228      print(f"\n转换结果:")
      229      print(f"  concepts (现代概念池): {len(onto['concepts'])}")
      230      print(f"  term_to_msc 映射: {len(onto['term_to_msc'])} 个术语")

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf
      --out omd_ontology.json --test-dataset test_dataset.jso…)
  ⎿       omd_ontology_9.12.rdf   
       ʵ      : 7704
       TechnicalTerm: 570
     … +20 lines (ctrl+o to expand)

● 现在 141 个术语从 test_dataset 补了 MSC，term_to_msc 从 292→398。做最终的管道烟雾测试。

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python build_candidates_A.py --self-test 2>&1)
  ⎿  === self-test:  ϳɱ   ð ̲    ===
           : 100%  δ    : []
     |C(t)|: mean=30.0 range=[30,30]
     … +13 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python -c "
      # Quick smoke test: build candidates for first 3 terms only, verify output structure…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command
      209 +            for c in cases:
      210 +                name = c.get("name", "")
      211 +                msc_codes = c.get("msc_codes", [])
      212 +                if name and msc_codes:
      213 +                    clean = [m.split("#")[-1] for m in msc_codes]
      214 +                    existing = set(onto["term_to_msc"].get(name, []))
      215 +                    before = len(existing)
      216 +                    existing.update(clean)
      217 +                    if len(existing) > before:
      218 +                        onto["term_to_msc"][name] = sorted(existing)
      219 +                        merged += 1
      220 +                        # Also populate msc_to_concepts
      221 +                        for m in clean:
      222 +                            if m not in onto["msc_to_concepts"]:
      223 +                                onto["msc_to_concepts"][m] = []
      224 +            print(f"  从 test_dataset 合并 MSC: {merged} 个术语")
      225 +        except Exception as e:
      226 +            print(f"  [warn] 无法读取 test_dataset: {e}")
      227 +
      228      print(f"\n转换结果:")
      229      print(f"  concepts (现代概念池): {len(onto['concepts'])}")
      230      print(f"  term_to_msc 映射: {len(onto['term_to_msc'])} 个术语")

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf
      --out omd_ontology.json --test-dataset test_dataset.jso…)
  ⎿       omd_ontology_9.12.rdf   
       ʵ      : 7704
       TechnicalTerm: 570
     … +20 lines (ctrl+o to expand)

● 现在 141 个术语从 test_dataset 补了 MSC，term_to_msc 从 292→398。做最终的管道烟雾测试。

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python build_candidates_A.py --self-test 2>&1)
  ⎿  === self-test:  ϳɱ   ð ̲    ===
           : 100%  δ    : []
     |C(t)|: mean=30.0 range=[30,30]
     … +13 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python -c "
      # Quick smoke test: build candidates for first 3 terms only, verify output structure…)
  ⎿    ľĩ: |C|=50  provenance={'tier1_msc': 3, 'tier2_neighbor': 0, 'tier3_semantic': 47}
       gold=  ׼    
       candidates[:5]: [' ˷   ά  ѭ    ', '        ', '     ѧԭ  ', '45           ̶   ', '     ʷ      \u200b']
     … +11 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/J-space-候选集构建包" && python -c "
      import sys, io…)
  ⎿  Fields: ['protege_id', 'name', 'romanization', 'original_definition', 'modern_definition', 'msc_codes', 'sourc
     e_book', 'authors', 'P_Class', 'L_Class', 'class1', 'class2', 'class3']
     Has gold_concept: 0/160
     … +5 lines (ctrl+o to expand)

● 管道验证通过。现在清理生成物，最终确认包内容。

● Bash(rm "C:/Users/19691/Desktop/J-space-候选集构建包/omd_ontology.json" 2>&1 && echo "cleaned")
  ⎿  cleaned

  Listed 1 directory (ctrl+o to expand)

● 包在桌面 J-space-候选集构建包，5 个文件：

  J-space-候选集构建包/
  ├── rdf_to_ontology_json.py    ← 第1步：RDF→JSON 转换（仅标准库）
  ├── omd_ontology_9.12.rdf      ← 源数据：OMD 本体（270个现代概念/3444条边）
  ├── test_dataset.json          ← 源数据：160条古籍术语
  ├── build_candidates_A.py      ← 第2步：gold-blind 候选集构造
  └── README_操作说明.md          ← 操作说明

  你只需两行命令：

  # 第1步
  python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf --out omd_ontology.json --test-dataset test_dataset.json

  # 第2步
  python build_candidates_A.py --test-dataset test_dataset.json --ontology omd_ontology.json --target-size 50
  --verify-gold-blind

  验证结果：管道已烟雾测试通过——|C| 从 8→50+，provenance 分层清晰（Tier1 MSC / Tier3 semantic），gold-blind
  置换检验正常。

  有一个需要注意的细节在 README 里写了——部分术语的 modern_definition
  是长句描述而非概念名，可能导致覆盖率校验时需启用阶梯放宽或 --allow-gold-injection。这不是脚本 bug，是数据格式问题。

✻ Churned for 5m 29s

❯ C:\Users\19691\Desktop\投稿论文\Thesis\J-space论文\新版本方案\DATA1,检查一下这些数据

  Read 1 file, listed 1 directory (ctrl+o to expand)

● 已经有人跑过了。让我检查产出质量。

  Read 2 files (ctrl+o to expand)

● 数据质量超出预期。让我抽样验证实际内容。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1" && python -c "
      import sys, io…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1" && python -c "
   import sys, io
   sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
  Listed 1 directory (ctrl+o to expand)

● 包在桌面 J-space-候选集构建包，5 个文件：

  J-space-候选集构建包/
  ├── rdf_to_ontology_json.py    ← 第1步：RDF→JSON 转换（仅标准库）
  ├── omd_ontology_9.12.rdf      ← 源数据：OMD 本体（270个现代概念/3444条边）
  ├── test_dataset.json          ← 源数据：160条古籍术语
  ├── build_candidates_A.py      ← 第2步：gold-blind 候选集构造
  └── README_操作说明.md          ← 操作说明

  你只需两行命令：

  # 第1步
  python rdf_to_ontology_json.py --rdf omd_ontology_9.12.rdf --out omd_ontology.json --test-dataset test_dataset.json

  # 第2步
  python build_candidates_A.py --test-dataset test_dataset.json --ontology omd_ontology.json --target-size 50
  --verify-gold-blind

  验证结果：管道已烟雾测试通过——|C| 从 8→50+，provenance 分层清晰（Tier1 MSC / Tier3 semantic），gold-blind
  置换检验正常。

  有一个需要注意的细节在 README 里写了——部分术语的 modern_definition
  是长句描述而非概念名，可能导致覆盖率校验时需启用阶梯放宽或 --allow-gold-injection。这不是脚本 bug，是数据格式问题。

✻ Churned for 5m 29s

❯ C:\Users\19691\Desktop\投稿论文\Thesis\J-space论文\新版本方案\DATA1,检查一下这些数据

  Read 1 file, listed 1 directory (ctrl+o to expand)

● 已经有人跑过了。让我检查产出质量。

  Read 2 files (ctrl+o to expand)

● 数据质量超出预期。让我抽样验证实际内容。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1" && python -c "
      import sys, io…)
  ⎿  Total: 160

     === |C(t)| distribution ===
     … +38 lines (ctrl+o to expand)

● 数据质量扎实。以下是完整评估：

  ---
  DATA1 评估结果

  核心指标（全部绿灯）

  ┌─────────────────┬───────────┬───────────────────────────┬──────┐
  │      指标       │ v5.0 要求 │           实际            │ 判定 │
  ├─────────────────┼───────────┼───────────────────────────┼──────┤
  │ gold 覆盖率     │ 100%      │ 100%（156 结构 + 4 注入） │  ✅  │
  ├─────────────────┼───────────┼───────────────────────────┼──────┤
  │ |C(t)| 范围     │ 20–100    │ 50–101，均值 72.1         │  ✅  │
  ├─────────────────┼───────────┼───────────────────────────┼──────┤
  │ chance level    │ —         │ 1.54%（旧版 12.5%）       │  ✅  │
  ├─────────────────┼───────────┼───────────────────────────┼──────┤
  │ gold-blind 置换 │ 0 差异    │ 0/160                     │  ✅  │
  ├─────────────────┼───────────┼───────────────────────────┼──────┤
  │ 置换偶然命中    │ < 50%     │ 22.5%                     │  ✅  │
  ├─────────────────┼───────────┼───────────────────────────┼──────┤
  │ |C| vs 难度相关 │ < 0.30    │ 0.056                     │  ✅  │
  └─────────────────┴───────────┴───────────────────────────┴──────┘

  最关键的修复：Pilot C 动态范围

  |C| 从 8 → 50–101，gold_share 现在有充分的变化空间。旧版 |C|=8 时 chance level 12.5%（随机猜就有 1/8），新版
  1.54%。Pilot C 红灯已解除。

  |C(t)| 分布

  [50,60):  82 █████████████████████████  (主峰，Tier3 n-gram 补齐到目标50)
  [60,70):   5 █
  [80,90):  12 ███
  [90,101): 57 ██████████████            (次峰，Tier1 MSC 大桶)

  分布双峰，50 和 90+ 各一簇——这是 MSC 巨桶（如 51M04 初等几何含 166+ 概念）的自然结果，不是缺陷。

  候选质量（抽样）

  - "勾"→"直接映射"：候选全是几何概念（"直角边""半径""切线""椭圆轴"）→ 高质量
  - "移项"→"移项"：候选含"代入""四元一次方程组""方程行/列"→ 相关且区分度好
  - "佣价"→"雇工每日工资"：候选里混入"下等马""井""亩"等非数学词 → gold 本身不是数学概念导致

  需注意的4个问题

  ┌─────┬─────────────────────────────────────────┬───────────────────────┬─────────────────────────────────────────┐
  │  #  │                  问题                   │         影响          │                  处理                   │
  ├─────┼─────────────────────────────────────────┼───────────────────────┼─────────────────────────────────────────┤
  │  1  │ 4条注入项（邪之适出/袤/引葭赴岸/横之不  │ 论文需单独报告或从 HC │ GPU 上加 --embed-model                  │
  │     │ 出）                                    │  剔除                 │ 重跑大概率能纯结构捞回                  │
  ├─────┼─────────────────────────────────────────┼───────────────────────┼─────────────────────────────────────────┤
  │  2  │ synonyms_A.json 为空                    │ exact_match           │ 需手动补同义词表，或用论文中的同义匹配  │
  │     │                                         │ 无同义折叠            │ 规则                                    │
  ├─────┼─────────────────────────────────────────┼───────────────────────┼─────────────────────────────────────────┤
  │  3  │ Tier2 = 0（本体无边匹配                 │ 候选全来自 MSC +      │ 可接受；论文如实写 provenance           │
  │     │ NEIGHBOR_RELATIONS）                    │ n-gram                │                                         │
  ├─────┼─────────────────────────────────────────┼───────────────────────┼─────────────────────────────────────────┤
  │  4  │ 部分术语 gold                           │ 这些术语的 gain 预期  │ 分析时做敏感性检验：剔除这些条目后      │
  │     │ 本身非数学概念（如"粗米饭""下等马"）    │ 为零（不算真·约束）   │ mass→gain 是否依然成立                  │
  └─────┴─────────────────────────────────────────┴───────────────────────┴─────────────────────────────────────────┘

  结论

  可以直接用于实验。 用这个 dataset_a_with_candidates.json 替换掉旧的 |C|=8 版本，然后按顺序：pilot_check.py --stage A →
   GPU 上 --stage BC → Exp2 → Exp1。

✻ Sautéed for 1m 5s

❯ 可以用,但要区分"跑实验可用"和"写进论文站得住"两个标准——现在是前者已达标,后者还差三件事。

  **现在就能做的**:这份 `dataset_a_with_candidates.json` 的硬指标(覆盖率 100%、置换检验通过、chance level
  1.54%、规模-难度无相关)已经满足 pilot 的要求,你可以直接拿去跑 `pilot_check.py --stage A` 和 GPU 上的 stage BC,验证
  mass 动态范围。Pilot 阶段用它没有问题,这正是它的用途。

  **写进论文之前要补的三件事**:

  第一,**消灭或如实报告那 4 条注入**。最优解是在 GPU 机器上加 `--embed-model shibing624/text2vec-base-chinese`
  重跑一次(命令在交付说明里),如果 embedding 召回能把 4
  条纯结构捞回,就得到零注入的干净版本,论文里一句话都不用解释;捞不回就在 §Method 写明这 4 条被注入并在 HC
  分析中单独报告或剔除。审计表里 `gold_injected` 列已经标好,剔除很方便。

  第二,**方法部分要如实描述概念池和 MSC 链接的来源**。候选池 = 数据集标签空间(131)+
  代微积拾级独立标注的现代定义(440),这在分类任务里是标准做法(标签空间即候选全集),但要写出来;`msc_to_concepts`
  有一部分来自语料标注配对(provenance 文件可逐条溯源),这不是逐术语的 gold 泄漏——置换检验的 0/160 差异和 22.5%
  偶然命中率就是你的辩护证据,建议把"+77.5% 结构性增益"这个数直接写进论文。

  第三,**注意一个论文层面的表述风险**:目前 Tier 2(本体关系边)贡献为 0,因为 9_12 的边类型(KnowledgeMappingTerm
  等)不在邻域关系清单里,也没有 EQUIVALENT_TO。也就是说现在的"本体结构召回"实质上是 MSC
  分类结构,而非图谱关系结构。如果论文的核心主张是 OMD 图谱的关系结构带来增益,这个候选集支撑不了那个表述——要么在本体里补
   SUBSUMES/EVOLUTION/EQUIVALENT_TO 这类边再重建,要么把主张收窄为"MSC 分类结构 + 语义召回"。

  总结:Pilot 现在就跑,论文定稿前先做 embedding 重跑,这是成本最低、收益最大的一步。

● 清楚了。三个层次：

  - Pilot 用：现在就能跑，数据达标
  - 论文用：GPU embedding 重跑消注入 + 方法写清概念池来源 + 把"图谱关系"收窄为"MSC 分类结构"
  - 优先级：先跑 pilot 验 mass 动态范围，论文定稿前补 embedding 重跑

  现在继续建 config/utils/candidate_mass 三件套，为 pilot 和后续实验做准备？

✻ Cogitated for 7s

❯ 构建吧

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\config.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 ..\code\config.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
   1 # -*- coding: utf-8 -*-
   2 """全局配置 v5.0 —— 所有路径/字段/种子/阈值的唯一来源。禁止在子脚本另立常量。"""
   3 import os
   4
   5 # ========== 路径 ==========
   6 BASE_DIR    = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
   7 DATA_DIR    = os.path.join(BASE_DIR, "data")
   8 RESULTS_DIR = os.path.join(BASE_DIR, "results")
   9 FIG_DIR     = os.path.join(BASE_DIR, "figures")
  10 for d in (DATA_DIR, RESULTS_DIR, FIG_DIR):
  11     os.makedirs(d, exist_ok=True)
  12
  13 # 输入数据
  14 DATASET_A     = os.path.join(DATA_DIR, "dataset_a_with_candidates.json")
  15 DATASET_B     = os.path.join(DATA_DIR, "dataset_b.json")
  16 SYNONYMS_A    = os.path.join(DATA_DIR, "synonyms_A.json")
  17 SYNONYMS_B    = os.path.join(DATA_DIR, "synonyms_B.json")
  18 CONCEPT_MSC   = os.path.join(DATA_DIR, "concept_msc.json")
  19 ALL_CONCEPTS  = os.path.join(DATA_DIR, "all_concepts.json")
  20
  21 # 原始来源（回填用）
  22 OMD_ONTOLOGY_RAW = os.path.join(DATA_DIR, "omd_ontology.json")
  23 CHIP_CDN_RAW     = os.path.join(DATA_DIR, "chip_cdn_raw.json")
  24
  25 # ========== 模型 ==========
  26 MODEL_NAME   = "Qwen/Qwen2.5-14B-Instruct"
  27 DEVICE       = "cuda"
  28 DTYPE        = "float16"          # A100 40G；显存紧张改 "bfloat16" 或 4bit
  29 VLLM_BACKEND = "xgrammar"
  30 TEMPERATURE  = 0.0
  31 MAX_TOKENS   = 32                # 术语很短，避免生成跑飞
  32
  33 # ========== 实验参数 ==========
  34 N_REPEATS      = 3
  35 SEED_BASE      = 20260101
  36 BOOTSTRAP_SEED = 42
  37 BOOTSTRAP_N    = 10000
  38 CI_LEVEL       = 0.95
  39
  40 # 数据集规模（当前实际值；计划值 260/200 待扩充）
  41 N_DATASET_A = 160
  42 N_DATASET_B = 20
  43
  44 # ========== 字段名（内部标准名 → 真实 JSON 字段）==========
  45 # 若 JSON 字段不同，只改这里
  46 FIELD = {
  47     "id":          "protege_id",
  48     "term":        "name",
  49     "source":      "source_book",
  50     "definition":  "original_definition",
  51     "gold":        "modern_definition",
  52     "p_class":     "P_Class",
  53     "l_class":     "L_Class",
  54     "candidates":  "candidates",

❯ 可以用,但要区分"跑实验可用"和"写进论文站得住"两个标准——现在是前者已达标,后者还差三件事。

  **现在就能做的**:这份 `dataset_a_with_candidates.json` 的硬指标(覆盖率 100%、置换检验通过、chance level
  1.54%、规模-难度无相关)已经满足 pilot 的要求,你可以直接拿去跑 `pilot_check.py --stage A` 和 GPU 上的 stage BC,验证
  mass 动态范围。Pilot 阶段用它没有问题,这正是它的用途。

  **写进论文之前要补的三件事**:

  第一,**消灭或如实报告那 4 条注入**。最优解是在 GPU 机器上加 `--embed-model shibing624/text2vec-base-chinese`
  重跑一次(命令在交付说明里),如果 embedding 召回能把 4
  条纯结构捞回,就得到零注入的干净版本,论文里一句话都不用解释;捞不回就在 §Method 写明这 4 条被注入并在 HC
  分析中单独报告或剔除。审计表里 `gold_injected` 列已经标好,剔除很方便。

  第二,**方法部分要如实描述概念池和 MSC 链接的来源**。候选池 = 数据集标签空间(131)+
  代微积拾级独立标注的现代定义(440),这在分类任务里是标准做法(标签空间即候选全集),但要写出来;`msc_to_concepts`
  有一部分来自语料标注配对(provenance 文件可逐条溯源),这不是逐术语的 gold 泄漏——置换检验的 0/160 差异和 22.5%
  偶然命中率就是你的辩护证据,建议把"+77.5% 结构性增益"这个数直接写进论文。

  第三,**注意一个论文层面的表述风险**:目前 Tier 2(本体关系边)贡献为 0,因为 9_12 的边类型(KnowledgeMappingTerm
  等)不在邻域关系清单里,也没有 EQUIVALENT_TO。也就是说现在的"本体结构召回"实质上是 MSC
  分类结构,而非图谱关系结构。如果论文的核心主张是 OMD 图谱的关系结构带来增益,这个候选集支撑不了那个表述——要么在本体里补
   SUBSUMES/EVOLUTION/EQUIVALENT_TO 这类边再重建,要么把主张收窄为"MSC 分类结构 + 语义召回"。

  总结:Pilot 现在就跑,论文定稿前先做 embedding 重跑,这是成本最低、收益最大的一步。

● 清楚了。三个层次：

  - Pilot 用：现在就能跑，数据达标
  - 论文用：GPU embedding 重跑消注入 + 方法写清概念池来源 + 把"图谱关系"收窄为"MSC 分类结构"
  - 优先级：先跑 pilot 验 mass 动态范围，论文定稿前补 embedding 重跑

  现在继续建 config/utils/candidate_mass 三件套，为 pilot 和后续实验做准备？

✻ Cogitated for 7s

❯ 构建吧

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\config.py)
  ⎿  Wrote 91 lines to ..\code\config.py
       1 # -*- coding: utf-8 -*-
       2 """全局配置 v5.0 —— 所有路径/字段/种子/阈值的唯一来源。禁止在子脚本另立常量。"""
       3 import os
       4
       5 # ========== 路径 ==========
       6 BASE_DIR    = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
       7 DATA_DIR    = os.path.join(BASE_DIR, "data")
       8 RESULTS_DIR = os.path.join(BASE_DIR, "results")
       9 FIG_DIR     = os.path.join(BASE_DIR, "figures")
      10 for d in (DATA_DIR, RESULTS_DIR, FIG_DIR):
     … +81 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\utils.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 ..\code\utils.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
   1 # -*- coding: utf-8 -*-
   2 """共享工具 v5.0 —— 数据加载 / 匹配 / 解析。所有脚本 `from utils import *`。"""
   3 import json
   4 import os
   5 import csv
   6
   7 from config import FIELD
   8
   9
  10 def load_dataset(path):
  11     """读取数据集，把真实字段名映射到内部标准名。返回 list[dict]。"""
  12     if not os.path.exists(path):
  13         raise FileNotFoundError(f"数据集不存在: {path}")
  14     with open(path, encoding="utf-8") as f:
  15         raw = json.load(f)
  16     records = raw["test_cases"] if isinstance(raw, dict) and "test_cases" in raw else raw
  17     out = []
  18     for r in records:
  19         out.append({k: r.get(src) for k, src in FIELD.items()})
  20     return out
  21
  22
  23 def load_synonyms(path):
  24     """读取同义词表 {标准概念: [同义词列表]}。文件不存在返回空。"""
  25     if not os.path.exists(path):
  26         return {}
  27     with open(path, encoding="utf-8") as f:
  第二,**方法部分要如实描述概念池和 MSC 链接的来源**。候选池 = 数据集标签空间(131)+
  代微积拾级独立标注的现代定义(440),这在分类任务里是标准做法(标签空间即候选全集),但要写出来;`msc_to_concepts`
  有一部分来自语料标注配对(provenance 文件可逐条溯源),这不是逐术语的 gold 泄漏——置换检验的 0/160 差异和 22.5%
  偶然命中率就是你的辩护证据,建议把"+77.5% 结构性增益"这个数直接写进论文。

  第三,**注意一个论文层面的表述风险**:目前 Tier 2(本体关系边)贡献为 0,因为 9_12 的边类型(KnowledgeMappingTerm
  等)不在邻域关系清单里,也没有 EQUIVALENT_TO。也就是说现在的"本体结构召回"实质上是 MSC
  分类结构,而非图谱关系结构。如果论文的核心主张是 OMD 图谱的关系结构带来增益,这个候选集支撑不了那个表述——要么在本体里补
   SUBSUMES/EVOLUTION/EQUIVALENT_TO 这类边再重建,要么把主张收窄为"MSC 分类结构 + 语义召回"。

  总结:Pilot 现在就跑,论文定稿前先做 embedding 重跑,这是成本最低、收益最大的一步。

● 清楚了。三个层次：

  - Pilot 用：现在就能跑，数据达标
  - 论文用：GPU embedding 重跑消注入 + 方法写清概念池来源 + 把"图谱关系"收窄为"MSC 分类结构"
  - 优先级：先跑 pilot 验 mass 动态范围，论文定稿前补 embedding 重跑

  现在继续建 config/utils/candidate_mass 三件套，为 pilot 和后续实验做准备？

✻ Cogitated for 7s

❯ 构建吧

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\config.py)
  ⎿  Wrote 91 lines to ..\code\config.py
       1 # -*- coding: utf-8 -*-
       2 """全局配置 v5.0 —— 所有路径/字段/种子/阈值的唯一来源。禁止在子脚本另立常量。"""
       3 import os
       4
       5 # ========== 路径 ==========
       6 BASE_DIR    = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
       7 DATA_DIR    = os.path.join(BASE_DIR, "data")
       8 RESULTS_DIR = os.path.join(BASE_DIR, "results")
       9 FIG_DIR     = os.path.join(BASE_DIR, "figures")
      10 for d in (DATA_DIR, RESULTS_DIR, FIG_DIR):
     … +81 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\utils.py)
  ⎿  Wrote 64 lines to ..\code\utils.py
       1 # -*- coding: utf-8 -*-
       2 """共享工具 v5.0 —— 数据加载 / 匹配 / 解析。所有脚本 `from utils import *`。"""
       3 import json
       4 import os
       5 import csv
       6
       7 from config import FIELD
       8
       9
      10 def load_dataset(path):
     … +54 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\candidate_mass.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 ..\code\candidate_mass.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
   1 # -*- coding: utf-8 -*-
   2 """
   3 candidate_mass v5.0 —— 核心度量（序列概率 + 多 token + 长度归一）。
   4
   5 ❌ 禁止: candidate_mass = sum(softmax(logits))  （v4 错误定义，单 token / 受词表大小污染）
   6 ✅ 正确定义: gold_share = exp(score(gold)) / Σ_{c∈C(t)} exp(score(c))
   7    其中 score(c) = (1/|c|) · Σ_i log P(c_i | prompt, c_<i)  (per-token 平均 log 概率)
   8 """
   9 import math
  10 import torch
  11 import torch.nn.functional as F
  12 from config import DEVICE
  13
  14
  15 @torch.no_grad()
  16 def score_candidates(model, tok, prompt, candidates, length_norm=True, batch_size=16):
  17     """教师强制。返回 dict[c] = per-token 平均 log 概率（几何平均，消长度偏置）。"""
  18     prompt_ids = tok(prompt, return_tensors="pt").input_ids[0]
  19     out = {}
  20     for i in range(0, len(candidates), batch_size):
  21         chunk = candidates[i:i + batch_size]
  22         seqs, plens, clens = [], [], []
  23         for c in chunk:
  24             cid = tok(c, add_special_tokens=False, return_tensors="pt").input_ids[0]
  25             seqs.append(torch.cat([prompt_ids, cid]))
  26             plens.append(len(prompt_ids))
  27             clens.append(len(cid))
  28
  29         maxlen = max(len(s) for s in seqs)
  30         pad = tok.pad_token_id if tok.pad_token_id is not None else tok.eos_token_id
  31         batch = torch.full((len(seqs), maxlen), pad, dtype=torch.long)
  - Pilot 用：现在就能跑，数据达标
  - 论文用：GPU embedding 重跑消注入 + 方法写清概念池来源 + 把"图谱关系"收窄为"MSC 分类结构"
  - 优先级：先跑 pilot 验 mass 动态范围，论文定稿前补 embedding 重跑

  现在继续建 config/utils/candidate_mass 三件套，为 pilot 和后续实验做准备？

✻ Cogitated for 7s

❯ 构建吧

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\config.py)
  ⎿  Wrote 91 lines to ..\code\config.py
       1 # -*- coding: utf-8 -*-
       2 """全局配置 v5.0 —— 所有路径/字段/种子/阈值的唯一来源。禁止在子脚本另立常量。"""
       3 import os
       4
       5 # ========== 路径 ==========
       6 BASE_DIR    = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
       7 DATA_DIR    = os.path.join(BASE_DIR, "data")
       8 RESULTS_DIR = os.path.join(BASE_DIR, "results")
       9 FIG_DIR     = os.path.join(BASE_DIR, "figures")
      10 for d in (DATA_DIR, RESULTS_DIR, FIG_DIR):
     … +81 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\utils.py)
  ⎿  Wrote 64 lines to ..\code\utils.py
       1 # -*- coding: utf-8 -*-
       2 """共享工具 v5.0 —— 数据加载 / 匹配 / 解析。所有脚本 `from utils import *`。"""
       3 import json
       4 import os
       5 import csv
       6
       7 from config import FIELD
       8
       9
      10 def load_dataset(path):
     … +54 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\candidate_mass.py)
  ⎿  Wrote 68 lines to ..\code\candidate_mass.py
       1 # -*- coding: utf-8 -*-
       2 """
       3 candidate_mass v5.0 —— 核心度量（序列概率 + 多 token + 长度归一）。
       4
       5 ❌ 禁止: candidate_mass = sum(softmax(logits))  （v4 错误定义，单 token / 受词表大小污染）
       6 ✅ 正确定义: gold_share = exp(score(gold)) / Σ_{c∈C(t)} exp(score(c))
       7    其中 score(c) = (1/|c|) · Σ_i log P(c_i | prompt, c_<i)  (per-token 平均 log 概率)
       8 """
       9 import math
      10 import torch
     … +58 lines (ctrl+o to expand)

● Bash(mkdir -p "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/results" "C:/Users/19691/Desktop/投稿论文/The…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1/dataset_a_with_candidates.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_a_…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/data/dataset_b.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_b.json" 2>&1 && cp "C:/Us…)
  ⎿  copied

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
   import sys
   sys.path.insert(0, '.')
       3 import json
       4 import os
       5 import csv
       6
       7 from config import FIELD
       8
       9
      10 def load_dataset(path):
     … +54 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\candidate_mass.py)
  ⎿  Wrote 68 lines to ..\code\candidate_mass.py
       1 # -*- coding: utf-8 -*-
       2 """
       3 candidate_mass v5.0 —— 核心度量（序列概率 + 多 token + 长度归一）。
       4
       5 ❌ 禁止: candidate_mass = sum(softmax(logits))  （v4 错误定义，单 token / 受词表大小污染）
       6 ✅ 正确定义: gold_share = exp(score(gold)) / Σ_{c∈C(t)} exp(score(c))
       7    其中 score(c) = (1/|c|) · Σ_i log P(c_i | prompt, c_<i)  (per-token 平均 log 概率)
       8 """
       9 import math
      10 import torch
     … +58 lines (ctrl+o to expand)

● Bash(mkdir -p "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/results" "C:/Users/19691/Desktop/投稿论文/The…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1/dataset_a_with_candidates.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_a_…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/data/dataset_b.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_b.json" 2>&1 && cp "C:/Us…)
  ⎿  copied

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys…)
  ⎿  Error: Exit code 1
       File "<string>", line 30
         print(f'parse_answer("ӳ    :       \n    "): "{parse_answer("ӳ    :       \n    ", ANSWER_PREFIX)}"')
                                                                                                 ^
     SyntaxError: f-string expression part cannot include a backslash

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, '.')…)
  ⎿  Waiting…

     … +54 lines (ctrl+o to expand)

● Write(~\Desktop\投稿论文\Thesis\J-space论文\新版本方案\code\candidate_mass.py)
       4
       5 ❌ 禁止: candidate_mass = sum(softmax(logits))  （v4 错误定义，单 token / 受词表大小污染）
       6 ✅ 正确定义: gold_share = exp(score(gold)) / Σ_{c∈C(t)} exp(score(c))
       7    其中 score(c) = (1/|c|) · Σ_i log P(c_i | prompt, c_<i)  (per-token 平均 log 概率)
       8 """
       9 import math
      10 import torch
     … +58 lines (ctrl+o to expand)

● Bash(mkdir -p "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/results" "C:/Users/19691/Desktop/投稿论文/The…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1/dataset_a_with_candidates.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_a_…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/data/dataset_b.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_b.json" 2>&1 && cp "C:/Us…)
  ⎿  copied

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys…)
  ⎿  Error: Exit code 1
       File "<string>", line 30
         print(f'parse_answer("ӳ    :       \n    "): "{parse_answer("ӳ    :       \n    ", ANSWER_PREFIX)}"')
                                                                                                 ^
     SyntaxError: f-string expression part cannot include a backslash

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, '.')…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 3, in <module>
     ModuleNotFoundError: No module named 'code.config'; 'code' is not a package

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys…)
  ⎿  === config ===
     BASE_DIR: C:\Users\19691\Desktop\Ͷ      \Thesis\J-space    \ °汾    
     N_DATASET_A=160
     … +8 lines (ctrl+o to expand)

● 三件套全部通过导入测试。创建 __init__.py 解决 code 目录名冲突。

● Write(code\__init__.py)
  ⎿  Wrote 1 lines to code\__init__.py
      1 # code/__init__.py — v5.0 experiment package

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, 'code')…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
   import sys; sys.path.insert(0, 'code')
   from config import *
   from utils import *
● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/DATA1/dataset_a_with_candidates.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_a_…)
  ⎿  Done

● Bash(cp "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/data/dataset_b.json"
      "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/data/dataset_b.json" 2>&1 && cp "C:/Us…)
  ⎿  copied

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys…)
  ⎿  Error: Exit code 1
       File "<string>", line 30
         print(f'parse_answer("ӳ    :       \n    "): "{parse_answer("ӳ    :       \n    ", ANSWER_PREFIX)}"')
                                                                                                 ^
     SyntaxError: f-string expression part cannot include a backslash

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, '.')…)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "<string>", line 3, in <module>
     ModuleNotFoundError: No module named 'code.config'; 'code' is not a package

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys…)
  ⎿  === config ===
     BASE_DIR: C:\Users\19691\Desktop\Ͷ      \Thesis\J-space    \ °汾    
     N_DATASET_A=160
     … +8 lines (ctrl+o to expand)

● 三件套全部通过导入测试。创建 __init__.py 解决 code 目录名冲突。

● Write(code\__init__.py)
  ⎿  Wrote 1 lines to code\__init__.py
      1 # code/__init__.py — v5.0 experiment package

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, 'code')…)
  ⎿  Dataset A: 160 records
     Gold in C(t): 160/160
     |C|: min=50, max=101, mean=72.1
     … +14 lines (ctrl+o to expand)

● 三件套建成，全部验证通过：

  新版本方案/
  ├── code/
  │   ├── __init__.py
  │   ├── config.py          ← Part 1.1：FIELD/THRESH/prompt/路径
  │   ├── utils.py           ← Part 1.3：load_dataset/exact_match/parse_answer
  │   └── candidate_mass.py  ← Part 4.3：序列概率 mass/gold_share
  ├── data/
  │   ├── dataset_a_with_candidates.json  (|C|=50-101, 均值72)
  │   ├── dataset_b.json
  │   └── synonyms_A.json
  ├── results/
  └── figures/

  验证结果：160条加载、100% gold∈C(t)、candidate_mass_row 数学校验通过、THRESH 已锁定。

  下一步是按 Part 9.2 清单继续——pilot 三门（A 不用 GPU，BC 等租服务器）、还是先写 exp2/exp1 实验脚本？

✻ Baked for 2m 39s

❯ 我会选择**第三条路**，不是 Claude 提出的两个选项之一。

  > **先完成所有 CPU 能完成的 Pilot 和实验框架，再租 GPU。不要现在直接写完整的 Exp1/Exp2。**

  原因很简单：**GPU 是最贵的资源，而现在还有几个"生死门"没有验证。**

  ---

  # 我建议的执行顺序（v5.0）

  ## Phase 0（今天，CPU即可）★★★★★

  目标只有一个：

  **确认整个实验是否值得跑。**

  ### Pilot A（已经完成 70%）

  你已经验证了：

  ```
  Dataset A: 160
  gold ∈ C(t): 100%
  |C|=50~101
  ```

  很好。

  我建议再补三项统计：

  ### A1. candidate size 分布

  例如：

  ```
  mean
  std
  min
  max
  histogram
  ```

  以后论文 Figure 可以直接用。

  ---

  ### A2. MSC覆盖率

  统计：

  ```
  每个MSC

  term数量

  candidate数量

  平均|C|
  ```

  看看有没有极端类别。

  ---

  ### A3. synonym覆盖

  例如：

  ```
  gold

  ↓

  synonym数量
  ```

  看看有没有：

  ```
  0
  1
  10+
  ```

  这种极端。

  这些以后Discussion都会用。

  ---

  ## Phase 1（先把Exp2脚手架写好，不跑GPU）★★★★★

  我建议：

  **先写Exp2。**

  不是跑。

  而是：

  把整个pipeline写完。

  例如：

  ```
  exp2_collect_logits.py

  ↓

  mass

  ↓

  csv

  ↓

  regression.py

  ↓

  plot_mass_gain.py
  ```

  GPU部分可以空着。

  但是：

  整个接口必须固定。

  原因：

  candidate_mass是全论文核心。

  以后所有实验都依赖它。

  所以：

  **接口要最先冻结。**

  ---

  ## Phase 2（CPU模拟）★★★★★

  现在不要GPU。

  直接Mock。

  例如：

  candidate_mass.py

  输出：

  ```
  term

  mass

  gold_share

  candidate_size
  ```

  哪怕：

  mass

  先随机。

  你都可以先把：

  ```
  csv

  ↓

  regression

  ↓

  画图

  ↓

  Figure
  ```

  全部走通。

  以后GPU回来，

  只替换：

  ```
  mass
  ```

  即可。

  这叫：

  > Dry Run

  很多实验室都是这么干。

  ---

  # Phase 3（GPU开始）

  GPU第一件事：

  不是跑Exp1。

  而是：

  ## Pilot B

  采：

  ```
  Dataset A

  30
  ```

  *

  ```
  Dataset B

  30
  ```

  看看：

  ```
  mass
  ```

  有没有：

  ```
  common support
  ```

  如果：

  ```
  A

  0.08~0.15

  B

  0.52~0.68
  ```

  那：

  H4直接改。

  不用继续幻想。

  ---

  然后：

  ## Pilot C

  Dataset A：

  160条。

  看看：

  mass：

  是不是：

  ```
  0.04

  ↓

  0.60
  ```

  还是：

  ```
  0.31~0.35
  ```

  如果：

  没有动态范围。

  回归没意义。

  要先解决。

  ---

  # 真正GPU实验

  只有Pilot过了以后。

  才开始：

  ```
  Exp2
  ```

  完整采集。

  ---

  然后：

  ```
  Exp1
  ```

  验证。

  ---

  最后：

  ```
  Exp3
  ```

  跨域。

  ---

  # 为什么我不建议现在写Exp1？

  因为：

  Exp1现在已经不是论文中心。

  而：

  candidate_mass

  才是。

  所以：

  开发资源应该：

  ```
  70%

  Exp2

  20%

  Exp1

  10%

  Exp4
  ```

  不要反过来。

  ---

  # 我还建议现在加一个东西（非常值得）

  你的目录：

  ```
  code/

  config.py

  utils.py

  candidate_mass.py
  ```

  很好。

  我建议马上补：

  ```
  metrics.py
  ```

  里面统一：

  ```
  exact_match()

  candidate_mass()

  gold_share()

  entropy()

  candidate_size()
  ```

  以后：

  所有实验：

  ```
  from metrics import *
  ```

  论文以后改指标：

  只改这里。

  整个实验不会炸。

  这是大型实验非常重要的一步。

  ---

  # 我还建议再补一个目录

  例如：

  ```
  analysis/

  regression.py

  ablation.py

  plots.py
  ```

  以后：

  ```
  results/

  ↓

  analysis/

  ↓

  figures/
  ```

  完全分离。

  否则：

  后面：

  Notebook、

  Python、

  CSV、

  Figure

  会混一起。

  论文最后一个月最容易乱。

  ---

  # 我的最终建议（按优先级）

  我建议不要在 **"Pilot 三门 vs 写 Exp2/Exp1"** 二选一，而是按下面顺序推进：

  1. **完成 CPU 可做的 Pilot A**：补充候选集规模分布、MSC 覆盖、同义词覆盖等统计，把数据集画像做完整。
  2. **先把 Exp2 的完整框架写出来，但不跑 GPU**：固定 `candidate_mass` 的输入输出接口，搭好
  CSV、回归、绘图等全流程，哪怕先用模拟数据跑通。
  3. **GPU 到位后先跑 Pilot B 和 Pilot C**：验证 `candidate_mass` 的共同支撑（common
  support）和动态范围。这两个检查决定 H4 和回归分析是否站得住脚。
  4. **Pilot 通过后再正式跑 Exp2 → Exp1 → Exp3**。Exp2 是论文的核心机制实验，Exp1
  是验证性实验，因此开发和计算资源应优先保障 Exp2。

  **一句话总结：现在最值得投入的是"把 Exp2 的整个流水线（pipeline）搭好"，而不是急着写完整的 Exp1。** 这样 GPU
  一到，你几乎就是把真实的 `candidate_mass` 数据灌进去，就能开始正式分析，而不会因为脚本或分析流程反复返工。


● 收到。按这个顺序执行：Phase 0 Pilot A 补统计 → Phase 1 Exp2 脚手架 → metrics.py → analysis/ 目录结构 → dry run。

  先跑 Phase 0。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, 'code')…)
  ⎿  Waiting…

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Bash command

   cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
   import sys; sys.path.insert(0, 'code')
   from config import *
   from utils import *
   from collections import Counter, defaultdict
   import json

   ds = load_dataset(DATASET_A)
   print(f'n={len(ds)}')

   # === A1. |C| distribution ===
   sizes = [len(r['candidates']) for r in ds]
   import statistics
   print(f'\n=== A1: |C(t)| distribution ===')
   print(f'  mean={statistics.mean(sizes):.1f}  median={statistics.median(sizes):.1f}')
   print(f'  std={statistics.stdev(sizes):.1f}  min={min(sizes)}  max={max(sizes)}')


  否则：

  后面：

  Notebook、

  Python、

  CSV、

  Figure

  会混一起。

  论文最后一个月最容易乱。

  ---

  # 我的最终建议（按优先级）

  我建议不要在 **"Pilot 三门 vs 写 Exp2/Exp1"** 二选一，而是按下面顺序推进：

  1. **完成 CPU 可做的 Pilot A**：补充候选集规模分布、MSC 覆盖、同义词覆盖等统计，把数据集画像做完整。
  2. **先把 Exp2 的完整框架写出来，但不跑 GPU**：固定 `candidate_mass` 的输入输出接口，搭好
  CSV、回归、绘图等全流程，哪怕先用模拟数据跑通。
  3. **GPU 到位后先跑 Pilot B 和 Pilot C**：验证 `candidate_mass` 的共同支撑（common
  support）和动态范围。这两个检查决定 H4 和回归分析是否站得住脚。
  4. **Pilot 通过后再正式跑 Exp2 → Exp1 → Exp3**。Exp2 是论文的核心机制实验，Exp1
  是验证性实验，因此开发和计算资源应优先保障 Exp2。

  **一句话总结：现在最值得投入的是"把 Exp2 的整个流水线（pipeline）搭好"，而不是急着写完整的 Exp1。** 这样 GPU
  一到，你几乎就是把真实的 `candidate_mass` 数据灌进去，就能开始正式分析，而不会因为脚本或分析流程反复返工。


● 收到。按这个顺序执行：Phase 0 Pilot A 补统计 → Phase 1 Exp2 脚手架 → metrics.py → analysis/ 目录结构 → dry run。

  先跑 Phase 0。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, 'code')…)
  ⎿  n=160

     === A1: |C(t)| distribution ===
     … +30 lines (ctrl+o to expand)

● Phase 0 完成。关键发现：

  - A1: |C| 分布健康——mean=72, std=23, min=50, max=101
  - A2: 两个脏 MSC 码（:--- 和数值溢出），22 个单例 MSC。大桶 51M04 均值 |C|=100 是几何概念的主来源
  - A3: 同义词覆盖 = 0/131——这是个隐患，exact_match 会太严格。需要手动补或 GPU 后用 embedding 做软匹配

  继续 Phase 1/2——建 metrics.py、analysis/ 目录、Exp2 脚手架。

● Write(code\metrics.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 code\metrics.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
    1 # -*- coding: utf-8 -*-
    2 """统一指标 v5.0 —— 所有实验脚本 `from metrics import *`。论文改指标只改这里。"""
    3 import math
    4 import numpy as np
    5 from config import BOOTSTRAP_N, BOOTSTRAP_SEED, CI_LEVEL
    6
    7
    8 # ========== Core metrics ==========
    9
   10 def gold_share(scores, gold):
   11     """gold 在 C(t) 内的软最大份额 ∈ [0,1]。scores = {candidate: log_prob}。"""
   12     cs = list(scores.values())
   13     m = max(cs)
   14     lse = m + math.log(sum(math.exp(x - m) for x in cs))
   15     glp = scores.get(gold)
   16     return math.exp(glp - lse) if glp is not None else float("nan")
   17
   18
   19 def candidate_entropy(scores):
   20     """C(t) 内 softmax 分布的熵（nat）。高→模型不确定；低→模型确信。"""
   21     cs = list(scores.values())
   22     m = max(cs)
   23     probs = [math.exp(x - m) for x in cs]
   24     z = sum(probs)
   25     probs = [p / z for p in probs]
   26     return -sum(p * math.log(p) for p in probs if p > 0)
   27
   28
   29 def gold_rank(scores, gold):
   30     """gold 在 C(t) 内的排名（1 = 最高分）。"""
   31     ranked = sorted(scores, key=scores.get, reverse=True)
   32     return (ranked.index(gold) + 1) if gold in scores else -1
   33
   34
   35 # ========== Accuracy ==========
   36
   37 def constraint_gain(acc_hc, acc_uc):
   38     """约束增益 = HC − UC。行为层指标。"""
   39     return acc_hc - acc_uc
   40
   41
   42 # ========== Bootstrap (term-level) ==========
   43
   44 def bootstrap_ci(values, n=BOOTSTRAP_N, seed=BOOTSTRAP_SEED, ci=CI_LEVEL):
   45     """术语级重采样 bootstrap 均值 CI。values: array-like。"""
   46     arr = np.asarray(values, dtype=float)
   47     rng = np.random.RandomState(seed)
   48     means = [arr[rng.randint(0, len(arr), len(arr))].mean() for _ in range(n)]
   49     lo = (1 - ci) / 2 * 100
   50     hi = 100 - lo
   51     return float(arr.mean()), float(np.percentile(means, lo)), float(np.percentile(means, hi))
   52
   53
   54 def bootstrap_diff_paired(a, b, n=BOOTSTRAP_N, seed=BOOTSTRAP_SEED, ci=CI_LEVEL):
   55     """配对 bootstrap：对每对差值重采样。a,b: 等长 array-like。"""
   56     a, b = np.asarray(a, float), np.asarray(b, float)
   57     d = a - b
   58     rng = np.random.RandomState(seed)
   59     idx = rng.randint(0, len(d), size=(n, len(d)))
   60     boots = d[idx].mean(axis=1)
   61     lo = (1 - ci) / 2 * 100
   62     hi = 100 - lo
   63     return float(d.mean()), float(np.percentile(boots, lo)), float(np.percentile(boots, hi))
   64
   65

  论文最后一个月最容易乱。

  ---

  # 我的最终建议（按优先级）

  我建议不要在 **"Pilot 三门 vs 写 Exp2/Exp1"** 二选一，而是按下面顺序推进：

  1. **完成 CPU 可做的 Pilot A**：补充候选集规模分布、MSC 覆盖、同义词覆盖等统计，把数据集画像做完整。
  2. **先把 Exp2 的完整框架写出来，但不跑 GPU**：固定 `candidate_mass` 的输入输出接口，搭好
  CSV、回归、绘图等全流程，哪怕先用模拟数据跑通。
  3. **GPU 到位后先跑 Pilot B 和 Pilot C**：验证 `candidate_mass` 的共同支撑（common
  support）和动态范围。这两个检查决定 H4 和回归分析是否站得住脚。
  4. **Pilot 通过后再正式跑 Exp2 → Exp1 → Exp3**。Exp2 是论文的核心机制实验，Exp1
  是验证性实验，因此开发和计算资源应优先保障 Exp2。

  **一句话总结：现在最值得投入的是"把 Exp2 的整个流水线（pipeline）搭好"，而不是急着写完整的 Exp1。** 这样 GPU
  一到，你几乎就是把真实的 `candidate_mass` 数据灌进去，就能开始正式分析，而不会因为脚本或分析流程反复返工。


● 收到。按这个顺序执行：Phase 0 Pilot A 补统计 → Phase 1 Exp2 脚手架 → metrics.py → analysis/ 目录结构 → dry run。


● 收到。按这个顺序执行：Phase 0 Pilot A 补统计 → Phase 1 Exp2 脚手架 → metrics.py → analysis/ 目录结构 → dry run。

  先跑 Phase 0。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, 'code')…)
  ⎿  n=160

     === A1: |C(t)| distribution ===
     … +30 lines (ctrl+o to expand)

● Phase 0 完成。关键发现：

  - A1: |C| 分布健康——mean=72, std=23, min=50, max=101
  - A2: 两个脏 MSC 码（:--- 和数值溢出），22 个单例 MSC。大桶 51M04 均值 |C|=100 是几何概念的主来源
  - A3: 同义词覆盖 = 0/131——这是个隐患，exact_match 会太严格。需要手动补或 GPU 后用 embedding 做软匹配

  继续 Phase 1/2——建 metrics.py、analysis/ 目录、Exp2 脚手架。

● Write(code\metrics.py)
  ⎿  Wrote 107 lines to code\metrics.py
       1 # -*- coding: utf-8 -*-
       2 """统一指标 v5.0 —— 所有实验脚本 `from metrics import *`。论文改指标只改这里。"""
       3 import math
       4 import numpy as np
       5 from config import BOOTSTRAP_N, BOOTSTRAP_SEED, CI_LEVEL
       6
       7
       8 # ========== Core metrics ==========
       9
      10 def gold_share(scores, gold):
     … +97 lines (ctrl+o to expand)

● Bash(mkdir -p "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/analysis" 2>&1)
  ⎿  Done

● Write(analysis\__init__.py)
  ⎿  Wrote 1 lines to analysis\__init__.py
      1 # analysis/__init__.py — v5.0 analysis package

● Write(analysis\regression.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 analysis\regression.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
    1 # -*- coding: utf-8 -*-
    2 """
    3 回归分析 v5.0 —— tautology 三嵌套 + 跨域交互。
    4
    5 输入: merged/analysis_df.csv（每术语一行: gold_share, gain, UC_acc, HC_acc, dataset）
    6 输出: 回归表（M1/M2/M3 + H4交互项）
    7 """
    8 import os
    9 import sys
   10 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "code"))
   11
   12 import numpy as np
   13 import pandas as pd
   14 import statsmodels.formula.api as smf
   15
   16 from config import RESULTS_DIR
   17 from metrics import pearson_r_ci
   18
   19
   20 def run_tautology_regressions(df, dataset_label="A"):
   21     """
   22     M1: gain ~ mass
   23     M2: gain ~ mass + UC_acc       (控UC后mass系数是否存活)
   24     M3: HC_acc ~ mass              (换因变量，最干净)
   25     """
   26     sub = df[df["dataset"] == dataset_label].dropna(subset=["gold_share", "gain", "UC_acc", "HC_acc"])
   27     if len(sub) < 30:
   28         return {"error": f"n={len(sub)} insufficient"}
   29
   30     r, r_lo, r_hi = pearson_r_ci(sub["gold_share"], sub["gain"])
   31
   32     M1 = smf.ols("gain ~ gold_share", sub).fit()
   33     M2 = smf.ols("gain ~ gold_share + UC_acc", sub).fit()
   34     M3 = smf.ols("HC_acc ~ gold_share", sub).fit()
   35
   36     def extract(model, name):
   37         return {
   38             "model": name,
   39             "n": int(model.nobs),
   40             "r2": round(model.rsquared, 4),
   41             "r2_adj": round(model.rsquared_adj, 4),
   42             "mass_coef": round(model.params.get("gold_share", float("nan")), 4),
   43             "mass_p": round(model.pvalues.get("gold_share", float("nan")), 6),
   44             "mass_ci_low": round(model.conf_int().loc["gold_share", 0], 4)
   45                 if "gold_share" in model.conf_int().index else float("nan"),
   46             "mass_ci_high": round(model.conf_int().loc["gold_share", 1], 4)
   47                 if "gold_share" in model.conf_int().index else float("nan"),
   48             "f_pvalue": round(model.f_pvalue, 6),
   49         }
   50
   51     return {
   52         "pearson_r": round(r, 4),
   53         "pearson_ci": (round(r_lo, 4), round(r_hi, 4)),
   54         "M1": extract(M1, "M1: gain ~ mass"),
   55         "M2": extract(M2, "M2: gain ~ mass + UC_acc"),
   56         "M3": extract(M3, "M3: HC_acc ~ mass"),
   57         "m2_mass_survives": extract(M2, "")["mass_p"] < 0.05,
   58     }
   59
   60
   61 def run_cross_domain_interaction(df):
   62     """H4: gain ~ mass * dataset 交互项检验。"""
   63     sub = df.dropna(subset=["gold_share", "gain", "dataset"])
   64     if sub["dataset"].nunique() < 2:
   65         return {"error": "need both datasets"}
   66     m = smf.ols("gain ~ gold_share * C(dataset)", sub).fit()
   67     idx = [i for i in m.params.index if "gold_share" in i and "dataset" in i]
   68     inter_p = m.pvalues.get(idx[0], float("nan")) if idx else float("nan")
● 收到。按这个顺序执行：Phase 0 Pilot A 补统计 → Phase 1 Exp2 脚手架 → metrics.py → analysis/ 目录结构 → dry run。

  先跑 Phase 0。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python -c "
      import sys; sys.path.insert(0, 'code')…)
  ⎿  n=160

     === A1: |C(t)| distribution ===
     … +30 lines (ctrl+o to expand)

● Phase 0 完成。关键发现：

  - A1: |C| 分布健康——mean=72, std=23, min=50, max=101
  - A2: 两个脏 MSC 码（:--- 和数值溢出），22 个单例 MSC。大桶 51M04 均值 |C|=100 是几何概念的主来源
  - A3: 同义词覆盖 = 0/131——这是个隐患，exact_match 会太严格。需要手动补或 GPU 后用 embedding 做软匹配

  继续 Phase 1/2——建 metrics.py、analysis/ 目录、Exp2 脚手架。

● Write(code\metrics.py)
  ⎿  Wrote 107 lines to code\metrics.py
       1 # -*- coding: utf-8 -*-
       2 """统一指标 v5.0 —— 所有实验脚本 `from metrics import *`。论文改指标只改这里。"""
       3 import math
       4 import numpy as np
       5 from config import BOOTSTRAP_N, BOOTSTRAP_SEED, CI_LEVEL
       6
       7
       8 # ========== Core metrics ==========
       9
      10 def gold_share(scores, gold):
     … +97 lines (ctrl+o to expand)

● Bash(mkdir -p "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/analysis" 2>&1)
  ⎿  Done

● Write(analysis\__init__.py)
  ⎿  Wrote 1 lines to analysis\__init__.py
      1 # analysis/__init__.py — v5.0 analysis package

● Write(analysis\regression.py)
  ⎿  Wrote 110 lines to analysis\regression.py
       1 # -*- coding: utf-8 -*-
       2 """
       3 回归分析 v5.0 —— tautology 三嵌套 + 跨域交互。
       4
       5 输入: merged/analysis_df.csv（每术语一行: gold_share, gain, UC_acc, HC_acc, dataset）
       6 输出: 回归表（M1/M2/M3 + H4交互项）
       7 """
       8 import os
       9 import sys
      10 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "code"))
     … +100 lines (ctrl+o to expand)

● Write(code\exp2_logit_analysis.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 code\exp2_logit_analysis.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
    1 #!/usr/bin/env python3
    2 # -*- coding: utf-8 -*-
    3 """
    4 Exp2: candidate_mass 采集（★核心，先跑）
    5
    6 条件: 仅 UC
    7 模型: Qwen2.5-14B, HF transformers（要 logits，不是 vLLM）
    8 样本: Dataset A(160), 各 1 次（确定性）
    9 产出: results/exp2/mass_A.csv
   10
   11 用法:
   12   python exp2_logit_analysis.py --dry-run        # 模拟数据，验证全流程
   13   python exp2_logit_analysis.py --dataset A      # GPU 正式跑
   14 """
   15 import os, sys, json, argparse, math, random
   16 import numpy as np
   17 import pandas as pd
   18
   19 sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
   20 from config import *
   21 from utils import load_dataset, save_csv
   22 from candidate_mass import score_candidates, candidate_mass_row
   23
   24
   25 # ========== GPU section (marked) ==========
   26
   27 def load_model_hf():
   28     """[GPU] 加载 Qwen2.5-14B HF transformers。"""
   29     import torch
   30     from transformers import AutoModelForCausalLM, AutoTokenizer
   31
   32     dtype = {"float16": torch.float16, "bfloat16": torch.bfloat16}.get(DTYPE, torch.float16)
   33     tok = AutoTokenizer.from_pretrained(MODEL_NAME, trust_remote_code=True)
   34     model = AutoModelForCausalLM.from_pretrained(
   35         MODEL_NAME, torch_dtype=dtype, device_map=DEVICE, trust_remote_code=True)
   36     model.eval()
   37     return model, tok
   38
   39
   40 def compute_mass_for_dataset(records, model, tok, dataset_tag):
   41     """[GPU] 对每条术语计算 candidate_mass。"""
   42     rows = []
   43     for i, r in enumerate(records):
   44         cands = list(dict.fromkeys(r["candidates"]))  # 去重保序
   45         gold = r["gold"]
   46         if gold not in cands:
   47             cands.append(gold)
   48
   49         prompt = build_user_prompt(r["term"], r["source"], r["definition"])
   50         scores = score_candidates(model, tok, prompt, cands)
   51         cm = candidate_mass_row(scores, gold)
   52
   53         rows.append({
   54             "id": r["id"],
   55             "term": r["term"],
   56             "gold": gold,
   57             "dataset": dataset_tag,
   58             **cm,
   59         })
   60
   61         if (i + 1) % 20 == 0:
   62             print(f"  [{dataset_tag}] {i+1}/{len(records)}")
   63
   64     return rows
   65
   66
   67 # ========== Dry-run (CPU, no model) ==========
   68
   69 def generate_mock_mass(records, seed=42):
   70     """生成合理的模拟 candidate_mass 数据用于流程验证。"""
   71     rng = np.random.RandomState(seed)
   72     rows = []
   73     for r in records:
   74         n_cand = len(r["candidates"]) if r["candidates"] else 50
   75         # Simulate: larger candidate set → gold_share tends lower (more competition)
   76         # Add noise so we get a spread
   77         base = rng.beta(2, 3)  # beta distribution: skew toward low, tail to high
   78         gold_share_val = base * (1 + rng.normal(0, 0.05))
   79
   80         # gold_rank: roughly inverse of share
   81         rank = max(1, int((1 - gold_share_val) * n_cand) + rng.randint(-2, 2))
   82         rank = min(rank, n_cand)
   83
   84         # Mock candidate names for argmax
   85         mock_cands = [f"mock_{i:03d}" for i in range(n_cand)]
   86         rng.shuffle(mock_cands)
   87         argmax = mock_cands[0]
   88
   89         rows.append({
   90             "id": r["id"],
   91             "term": r["term"],
   92             "gold": r["gold"],
   93             "dataset": "A",
   94             "gold_share": round(gold_share_val, 6),
   95             "C_mass_lse": round(math.log(n_cand) + rng.normal(0, 0.3), 4),
   96             "gold_rank_in_C": rank,
   97             "argmax_in_C": argmax,
   98             "n_cand": n_cand,
   99         })
  100     return rows
  101
  102
  103 # ========== Main ==========
  104
  105 def main():
  - A2: 两个脏 MSC 码（:--- 和数值溢出），22 个单例 MSC。大桶 51M04 均值 |C|=100 是几何概念的主来源
  - A3: 同义词覆盖 = 0/131——这是个隐患，exact_match 会太严格。需要手动补或 GPU 后用 embedding 做软匹配

  继续 Phase 1/2——建 metrics.py、analysis/ 目录、Exp2 脚手架。

● Write(code\metrics.py)
  ⎿  Wrote 107 lines to code\metrics.py
       1 # -*- coding: utf-8 -*-
       2 """统一指标 v5.0 —— 所有实验脚本 `from metrics import *`。论文改指标只改这里。"""
       3 import math
       4 import numpy as np
       5 from config import BOOTSTRAP_N, BOOTSTRAP_SEED, CI_LEVEL
       6
       7
       8 # ========== Core metrics ==========
       9
      10 def gold_share(scores, gold):
     … +97 lines (ctrl+o to expand)

● Bash(mkdir -p "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案/analysis" 2>&1)
  ⎿  Done

● Write(analysis\__init__.py)
  ⎿  Wrote 1 lines to analysis\__init__.py
      1 # analysis/__init__.py — v5.0 analysis package

● Write(analysis\regression.py)
  ⎿  Wrote 110 lines to analysis\regression.py
       1 # -*- coding: utf-8 -*-
       2 """
       3 回归分析 v5.0 —— tautology 三嵌套 + 跨域交互。
       4
       5 输入: merged/analysis_df.csv（每术语一行: gold_share, gain, UC_acc, HC_acc, dataset）
       6 输出: 回归表（M1/M2/M3 + H4交互项）
       7 """
       8 import os
       9 import sys
      10 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "code"))
     … +100 lines (ctrl+o to expand)

● Write(code\exp2_logit_analysis.py)
  ⎿  Wrote 147 lines to code\exp2_logit_analysis.py
      1 #!/usr/bin/env python3
      2 # -*- coding: utf-8 -*-
      3 """
      4 Exp2: candidate_mass 采集（★核心，先跑）
      5
      6 条件: 仅 UC
      7 模型: Qwen2.5-14B, HF transformers（要 logits，不是 vLLM）
      8 样本: Dataset A(160), 各 1 次（确定性）
      9 产出: results/exp2/mass_A.csv
     … +137 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python code/exp2_logit_analysis.py
      --dry-run 2>&1)
  ⎿  Dataset A: 160 terms
     [dry-run] ʹ  ģ   candidate_mass     
     … +5 lines (ctrl+o to expand)

● Dry-run 通过——mock mass 分布 p10=0.12 到 p90=0.70，有足够动态范围做回归。现在验证 analysis pipeline。

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python analysis/regression.py 2>&1)
  ⎿  Error: Exit code 1
     Traceback (most recent call last):
       File "C:\Users\19691\Desktop\Ͷ      \Thesis\J-space    \ °汾    \analysis\regression.py", line 14, in <module>
         import statsmodels.formula.api as smf
     ModuleNotFoundError: No module named 'statsmodels'

● Bash(pip install statsmodels 2>&1 | tail -3)
  ⎿  [notice] A new release of pip is available: 24.0 -> 26.1.2
     [notice] To update, run: python.exe -m pip install --upgrade pip

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python analysis/regression.py 2>&1)
  ⎿  === regression dry-run ===

     Pearson r(mass, gain) = 0.0695
     … +8 lines (ctrl+o to expand)

● 回归管道通了。建 plots skeleton，然后做最终验证。

● Write(analysis\plots.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 analysis\plots.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
    1 # -*- coding: utf-8 -*-
    2 """
    3 绘图 v5.0 —— Fig 2–8。从 results/merged/ 读数据，写到 figures/。
    4 """
    5 import os
    6 import sys
    7 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "code"))
    8
    9 import numpy as np
   10 import pandas as pd
   11 import matplotlib
   12 matplotlib.use("Agg")
   13 import matplotlib.pyplot as plt
   14 from matplotlib import rcParams
   15 from config import FIG_DIR
   16
   17 # 中文字体（Windows: Microsoft YaHei; Linux/macOS 可能需要调整）
   18 try:
   19     rcParams["font.sans-serif"] = ["Microsoft YaHei", "SimHei", "DejaVu Sans"]
   20     rcParams["axes.unicode_minus"] = False
   21 except Exception:
   22     pass
   23
   24
   25 def fig2_mass_vs_gain(df, out=None):
   26     """Fig 2: mass → gain 散点 + LOESS + 回归线。主图。"""
   27     sub = df.dropna(subset=["gold_share", "gain"])
   28     fig, ax = plt.subplots(figsize=(6, 5))
   29     ax.scatter(sub["gold_share"], sub["gain"], alpha=0.5, s=30, edgecolors="none")
   30
   31     # LOESS smooth
   32     from statsmodels.nonparametric.smoothers_lowess import lowess
   33     smoothed = lowess(sub["gain"], sub["gold_share"], frac=0.3)
   34     ax.plot(smoothed[:, 0], smoothed[:, 1], "C1", lw=2, label="LOESS")
   35
   36     # Linear fit
   37     z = np.polyfit(sub["gold_share"], sub["gain"], 1)
   38     x_line = np.linspace(sub["gold_share"].min(), sub["gold_share"].max(), 100)
   39     ax.plot(x_line, np.polyval(z, x_line), "C3--", lw=1.5, label="Linear")
   40
   41     ax.set_xlabel("candidate_mass (gold_share)")
   42     ax.set_ylabel("constraint gain (HC − UC)")
   43     ax.legend()
   44     ax.set_title("Fig 2: Mass → Gain")
   45     fig.tight_layout()
   46
   47     out = out or os.path.join(FIG_DIR, "fig2_mass_vs_gain.png")
   48     os.makedirs(os.path.dirname(out), exist_ok=True)
   49     fig.savefig(out, dpi=150)
   50     plt.close(fig)
   51     print(f"  Fig2 -> {out}")
   52     return out
   53
   54
   55 def fig3_coefficient_forest(reg_results, out=None):
   56     """Fig 3: M1/M2/M3 mass 系数森林图。"""
   57     fig, ax = plt.subplots(figsize=(5, 3))
   58     models = []
   59     coefs, lows, highs = [], [], []
   60     for key in ["M1", "M2", "M3"]:
   61         m = reg_results.get(key, {})
   62         if m:
   63             models.append(m.get("model", key))
   64             coefs.append(m.get("mass_coef", 0))
   65             lows.append(m.get("mass_ci_low", 0))
   66             highs.append(m.get("mass_ci_high", 0))
   67
   68     y = range(len(models))
   69     ax.errorbar(coefs, y, xerr=[[abs(c - l) for c, l in zip(coefs, lows)],
   70                                   [abs(h - c) for c, h in zip(coefs, highs)]],
   71                 fmt="o", capsize=4)
   72     ax.axvline(0, color="gray", ls="--", lw=0.8)
   73     ax.set_yticks(y)
   74     ax.set_yticklabels(models)
   75     ax.set_xlabel("mass coefficient")
   76     ax.set_title("Fig 3: Tautology Check (M1/M2/M3)")
   77     fig.tight_layout()
   78
   79     out = out or os.path.join(FIG_DIR, "fig3_coefficient_forest.png")
   80     os.makedirs(os.path.dirname(out), exist_ok=True)
   81     fig.savefig(out, dpi=150)
   82     plt.close(fig)
   83     print(f"  Fig3 -> {out}")
   84     return out
   85
   86
● Write(code\exp2_logit_analysis.py)
  ⎿  Wrote 147 lines to code\exp2_logit_analysis.py
      1 #!/usr/bin/env python3
      2 # -*- coding: utf-8 -*-
      3 """
      4 Exp2: candidate_mass 采集（★核心，先跑）
      5
      6 条件: 仅 UC
      7 模型: Qwen2.5-14B, HF transformers（要 logits，不是 vLLM）
      8 样本: Dataset A(160), 各 1 次（确定性）
      9 产出: results/exp2/mass_A.csv
     … +137 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python code/exp2_logit_analysis.py
      --dry-run 2>&1)
  ⎿  Dataset A: 160 terms
     [dry-run] ʹ  ģ   candidate_mass     
     … +5 lines (ctrl+o to expand)

● Dry-run 通过——mock mass 分布 p10=0.12 到 p90=0.70，有足够动态范围做回归。现在验证 analysis pipeline。
     Traceback (most recent call last):
       File "C:\Users\19691\Desktop\Ͷ      \Thesis\J-space    \ °汾    \analysis\regression.py", line 14, in <module>
         import statsmodels.formula.api as smf
     ModuleNotFoundError: No module named 'statsmodels'

● Bash(pip install statsmodels 2>&1 | tail -3)
  ⎿  [notice] A new release of pip is available: 24.0 -> 26.1.2
     [notice] To update, run: python.exe -m pip install --upgrade pip

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python analysis/regression.py 2>&1)
  ⎿  === regression dry-run ===

     Pearson r(mass, gain) = 0.0695
     … +8 lines (ctrl+o to expand)

● 回归管道通了。建 plots skeleton，然后做最终验证。

● Write(analysis\plots.py)
  ⎿  Wrote 128 lines to analysis\plots.py
       1 # -*- coding: utf-8 -*-
       2 """
       3 绘图 v5.0 —— Fig 2–8。从 results/merged/ 读数据，写到 figures/。
       4 """
       5 import os
       6 import sys
       7 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "code"))
       8
       9 import numpy as np
      10 import pandas as pd
     … +119 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python analysis/plots.py 2>&1)
  ⎿  === plots dry-run ===
       Fig2 -> C:\Users\19691\Desktop\Ͷ      \Thesis\J-space    \ °汾    \figures\fig2_mass_vs_gain.png
     [OK] plots dry-run complete

● 全管道 dry-run 通过。最终验证——

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && find . -type f -not -path "./DATA1/*" -not
      -path "./__pycache__/*" -not -path "./code/__pycache__/*"…)
  ⎿  ./00_完整流程总文档_v5_0.md
     ./analysis/__init__.py
     ./analysis/plots.py
     … +14 lines (ctrl+o to expand)

● 建成。最终结构：

  新版本方案/
  ├── 00_完整流程总文档_v5_0.md       ← 施工蓝图
  ├── pilot_check.py                  ← 三门（已有）
  ├── code/
  │   ├── config.py                   ← Part 1.1：FIELD/THRESH/prompt/路径
  │   ├── utils.py                    ← Part 1.3：数据加载/匹配/解析
  │   ├── candidate_mass.py           ← Part 4.3：序列概率 mass（核心）
  │   ├── metrics.py                  ← 统一指标（bootstrap/cliffδ/pearson）
  │   └── exp2_logit_analysis.py      ← Exp2 脚手架（--dry-run 可用）
  ├── analysis/
  │   ├── regression.py               ← M1/M2/M3 + H4 交互项
  │   └── plots.py                    ← Fig 2–8 生成
  ├── data/
  │   └── dataset_a_with_candidates.json  (|C|=50-101, 160条)
  ├── results/
  │   └── exp2/mass_A.csv             ← dry-run 验证产物
  └── figures/
      └── fig2_mass_vs_gain.png       ← dry-run 验证产物

  已验证：exp2 --dry-run → mass_A.csv → regression.py → plots.py → fig2 全管道无断裂点。GPU 到位后去掉
  --dry-run，灌入真实 mass 数据即可。

✻ Worked for 4m 34s

❯ 我看了你前面几轮和 Claude 的讨论，以及现在三件套已经跑通的状态。

  **现在不要继续写实验脚本。**

  如果按论文研发流程，现在真正的优先级只有一个：

  > **先验证方案成立，再投入GPU。**

  你现在已经进入 **v5.0 的 Pilot 阶段**，不是 Exp1/Exp2 正式实验阶段。

  ---

  # 我建议的顺序（也是我认为风险最低的顺序）

  ## Step 1（今天完成）

  ## 做三个 Pilot Gate（Gatekeeper）

  这是前几轮所有讨论最后收敛出来的三个问题。

  ---

  ### Gate 1：Dataset A 是否有足够动态范围

  也就是看看：

  ```
  candidate_mass
  ```

  到底是不是都集中在一个小区间。

  例如最后统计：

  ```
  min
  max
  mean
  std

  Q1
  Median
  Q3
  ```

  再画

  ```
  Histogram
  ```

  如果最后发现

  ```
  0.43
  0.46
  0.45
  ```

  那论文直接结束。

  因为：

  ```
  没有动态范围
  ↓

  无法回归
  ↓

  核心假设失效
  ```

  这是第一优先。

  ---

  ### Gate 2：A/B 有没有 Common Support

  也就是：

  ```
  Dataset A mass

  Dataset B mass
  ```

  是不是有重叠。

  画：

  ```
  Density

  或者

  Violin plot
  ```

  如果

  ```
  A
  0.10~0.20

  B
  0.45~0.60
  ```

  完全不交。

  那么

  ```
  gain ~ mass + dataset
  ```

  后面的交互项全部没有意义。

  这是Claude最后一轮也强调的。

  必须先知道。

  ---

  ### Gate 3：Gold Share

  你已经写好了

  ```
  gold_share
  ```

  现在统计：

  ```
  gold_share distribution
  ```

  如果

  ```
  gold_share≈candidate_mass
  ```

  说明：

  模型基本只考虑gold。

  如果

  ```
  gold_share << candidate_mass
  ```

  说明：

  还有很多竞争候选。

  这是后面分析的重要依据。

  ---

  ## 只有三个 Gate 全绿

  才值得租GPU。

  ---

  # Step2（CPU）

  然后再写

  ```
  exp2_collect_mass.py
  ```

  不是正式跑。

  而是

  Pilot版。

  例如

  ```
  sample=30
  ```

  或者

  ```
  sample=50
  ```

  看看：

  整个candidate_mass流程有没有bug。

  包括：

  ```
  teacher forcing

  ↓

  sequence score

  ↓

  length normalization

  ↓

  logsumexp

  ↓

  gold_share
  ```

  全部走通。

  ---

  # Step3（CPU）

  这时候再写

  ```
  exp1.py
  ```

  注意：

  只是写。

  不要跑。

  因为：

  真正耗GPU的是：

  ```
  Qwen

  +

  vLLM

  +

  guided decoding
  ```

  脚本可以提前准备。

  ---

  # Step4（GPU）

  GPU开始以后：

  顺序建议改一下。

  不是以前那个：

  ```
  Exp1

  ↓

  Exp2
  ```

  而是：

  ```
  Exp2

  ↓

  Pilot统计

  ↓

  确认Gate

  ↓

  Exp1

  ↓

  Exp3
  ```

  原因很简单：

  如果Pilot发现

  ```
  mass没有动态范围
  ```

  那么：

  Exp1其实已经没有必要继续烧GPU。

  ---

  # 我还建议加一个 Gate（这是我自己的建议）

  Claude一直在讨论：

  ```
  candidate_mass
  ```

  但是没人真正检查：

  ## Candidate Size 会不会影响 mass

  因为你现在：

  ```
  |C|

  50

  ~

  101
  ```

  跨度其实不小。

  建议Pilot顺便画：

  ```
  candidate_size

  vs

  candidate_mass
  ```

  计算：

  ```
  Spearman

  Pearson
  ```

  如果最后发现：

  ```
  r=-0.7
  ```

  说明：

  mass主要是在测：

  ```
  候选数
  ```

  而不是：

  ```
  constraint informativeness
  ```

  这个必须尽早知道。

  ---

  # 所以我认为接下来不要直接开始正式实验，而是完成一个专门的 Pilot 分析脚本。

  建议写一个：

  ```
  pilot_analysis.py
  ```

  一次性输出所有关键诊断，包括：

  1. **Dataset A 动态范围**

     * mass 的均值、标准差、四分位数
     * 直方图

  2. **A/B Common Support**

     * 两域密度图（或小提琴图）
     * 重叠区间统计

  3. **Gold Share 分布**

     * gold_share 与 candidate_mass 的关系

  4. **|C| 对 mass 的影响**

     * 散点图
     * Pearson / Spearman 相关

  5. **Pilot Gate 判定**

     * Gate 1：PASS / FAIL（动态范围）
     * Gate 2：PASS / FAIL（共同支撑）
     * Gate 3：PASS / FAIL（gold_share）
     * Gate 4：PASS / FAIL（|C| 是否引入明显偏置）

  **这份报告会决定后续是否值得投入 GPU 跑正式实验。**如果四个 Gate 都通过，你就可以比较有把握地进入正式实验；如果有某个
   Gate 不通过，调整方案的成本远低于跑完整个实验后再返工。


● Write(code\pilot_analysis.py)

────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
 Create file
 code\pilot_analysis.py
╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
    1 #!/usr/bin/env python3
    2 # -*- coding: utf-8 -*-
    3 """
    4 pilot_analysis.py — 四门诊断（一次性输出，决定是否值得租GPU）
    5
    6 在 pilot_check.py 的 B/C 阶段之后运行。读 results/pilot/mass_{a,b}.npy，
    7 或任何包含 gold_share / n_cand / dataset 列的 CSV。
    8
    9 用法:
   10   python pilot_analysis.py --mass-a results/pilot/mass_a.npy        # pilot_check 产出
   11   python pilot_analysis.py --csv results/exp2/mass_A.csv            # Exp2 产出
   12   python pilot_analysis.py --csv results/exp2/mass_A.csv --csv-b results/exp2/mass_B.csv
   13   python pilot_analysis.py --dry-run                                # 模拟验证
   14 """
   15 import argparse, os, sys, math
   16 import numpy as np
   17 import pandas as pd
   18
   19 sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
   20 from config import THRESH, FIG_DIR, RESULTS_DIR
   21 from metrics import mass_distribution_stats
   22
   23 import matplotlib
   24 matplotlib.use("Agg")
   25 import matplotlib.pyplot as plt
   26
   27 try:
   28     plt.rcParams["font.sans-serif"] = ["Microsoft YaHei", "SimHei", "DejaVu Sans"]
   29     plt.rcParams["axes.unicode_minus"] = False
   30 except Exception:
   31     pass
   32
   33 OUT_DIR = os.path.join(RESULTS_DIR, "pilot")
   34 VERDICT = {True: "PASS", False: "FAIL", None: "SKIP"}
   35
   36
   37 # ========== Gate 1: Dynamic range ==========
   38
   39 def gate_dynamic_range(masses, label="A"):
   40     """A 域 mass 是否覆盖 low→high？Pilot C 的核心问题。"""
   41     stats = mass_distribution_stats(masses)
   42     spread_ok = stats["spread_p90_p10"] >= THRESH["dynrange_spread"]
   43     low_ok   = stats["frac_low"] >= THRESH["dynrange_low_frac"]
   44     high_ok  = stats["frac_high"] >= THRESH["dynrange_high_frac"]
   45
   46     if spread_ok and low_ok and high_ok:
   47         verdict, note = True, "mass 有充分动态范围 → 核心图有斜率可拟合"
   48     elif spread_ok and (low_ok or high_ok):
   49         verdict, note = None, "动态范围偏窄但可用 → limitation 声明分布偏斜"
   50     else:
   51         verdict, note = False, "【阻断】mass 挤在窄带 → 主图不成立，需重抽 A 域"
   52
   53     # Histogram
   54     fig, ax = plt.subplots(figsize=(6, 3.5))
   55     ax.hist(masses, bins=20, range=(0, 1), color="steelblue", edgecolor="white", alpha=0.85)
   56     for thresh_val, color, name in [
   57         (THRESH["dynrange_low_frac"], "C1", f"low<0.2 ({stats['frac_low']:.0%})"),
   58         (THRESH["dynrange_high_frac"], "C2", f"high>0.5 ({stats['frac_high']:.0%})")]:
   59         pass
   60     ax.axvline(0.2, color="C1", ls="--", lw=1)
   61     ax.axvline(0.5, color="C2", ls="--", lw=1)
   62     ax.set_xlabel("gold_share")
   63     ax.set_ylabel("count")
   64     ax.set_title(f"Gate 1: Dynamic Range ({label}) — spread={stats['spread_p90_p10']:.3f}  [{VERDICT[verdict]}]
      ")
   65     fig.tight_layout()
   66     os.makedirs(OUT_DIR, exist_ok=True)
   67     path = os.path.join(OUT_DIR, f"gate1_dynamic_range_{label}.png")
   68     fig.savefig(path, dpi=150)
   69     plt.close(fig)
   70
   71     return {**stats, "verdict": verdict, "note": note, "figure": path,
   72             "spread_ok": spread_ok, "low_ok": low_ok, "high_ok": high_ok}
   73
   74
   75 # ========== Gate 2: Common support ==========
   76
   77 def gate_common_support(mass_a, mass_b):
   78     """两域 mass 分布是否有重叠？决定 H4 能否用交互项。"""
   79     if mass_b is None or len(mass_b) == 0:
   80         return {"verdict": None, "note": "B 域数据缺失，H4 跨域主张本轮无法验证"}
   81
   82     iou = _hist_overlap_iou(mass_a, mass_b)
   83     a_lo, a_hi = np.percentile(mass_a, [10, 90])
   84     b_lo, b_hi = np.percentile(mass_b, [10, 90])
   85     overlap_width = max(0.0, min(a_hi, b_hi) - max(a_lo, b_lo))
   86
   87     if iou >= THRESH["common_support_iou"] and overlap_width > 0.05:
   88         verdict, note = True, "两域 mass 有共同支撑 → H4 可用交互项"
   89     else:
   90         verdict, note = None, ("两域 mass 不重叠 → H4 降级为描述性域差异，"
   91                                 "不写 cross-domain law")
   92
   93     # Density plot
   94     fig, axes = plt.subplots(1, 2, figsize=(10, 4))
   95     for ax, data, label, c in [(axes[0], mass_a, "A: Ancient Math", "steelblue"),
   96                                  (axes[1], mass_b, "B: Clinical", "C2")]:
   97         ax.hist(data, bins=20, range=(0, 1), color=c, edgecolor="white", alpha=0.85, density=True)
   98         ax.set_xlabel("gold_share"); ax.set_ylabel("density"); ax.set_title(label)
   99     axes[0].set_title(f"A: μ={mass_a.mean():.3f} p10={np.percentile(mass_a,10):.3f} p90={np.percentile(mass_a,9
      0):.3f}")
  100     axes[1].set_title(f"B: μ={mass_b.mean():.3f} p10={np.percentile(mass_b,10):.3f} p90={np.percentile(mass_b,9
      0):.3f}")
  101     fig.suptitle(f"Gate 2: Common Support — IoU={iou:.3f}, overlap={overlap_width:.3f}  [{VERDICT[verdict]}]")
  102     fig.tight_layout()
  103     path = os.path.join(OUT_DIR, "gate2_common_support.png")
  104     os.makedirs(OUT_DIR, exist_ok=True)
  105     fig.savefig(path, dpi=150)
  106     plt.close(fig)
  107
  108     return {"iou": round(iou, 4), "overlap_width": round(overlap_width, 4),
  109             "a_p10_p90": (float(a_lo), float(a_hi)),
  110             "b_p10_p90": (float(b_lo), float(b_hi)),
  111             "verdict": verdict, "note": note, "figure": path}
  112
  113
  114 def _hist_overlap_iou(a, b, bins=10, lo=0.0, hi=1.0):
  115     edges = np.linspace(lo, hi, bins + 1)
  116     ha, _ = np.histogram(a, bins=edges, density=True)
  117     hb, _ = np.histogram(b, bins=edges, density=True)
  118     inter = np.minimum(ha, hb).sum()
  119     union = np.maximum(ha, hb).sum()
  120     return inter / union if union > 0 else 0.0
  121
  122
  123 # ========== Gate 3: Gold share distribution ==========
  124
  125 def gate_gold_share(df):
  126     """gold_share 分布：模型是否只认 gold，还是有竞争？"""
  127     gs = df["gold_share"].dropna().values
  128     stats = {
  129         "n": len(gs), "mean": float(gs.mean()), "std": float(gs.std()),
  130         "p10": float(np.percentile(gs, 10)), "p50": float(np.median(gs)),
  131         "p90": float(np.percentile(gs, 90)),
  132         "frac_above_0_5": float((gs > 0.5).mean()),
  133         "frac_below_0_2": float((gs < 0.2).mean()),
  134     }
  135     if stats["frac_above_0_5"] > 0.8:
  136         verdict, note = True, "gold_share 偏高 → 模型在 C(t) 内明显偏好 gold，约束增益空间小"
  137     elif stats["frac_below_0_2"] > 0.5:
  138         verdict, note = True, "gold_share 偏低 → C(t) 内竞争激烈，约束有大量增益空间"
  139     else:
  140         verdict, note = True, "gold_share 分布均衡 → 有区分度"
  141
  142     # Scatter: gold_share vs n_cand
  143     fig, axes = plt.subplots(1, 2, figsize=(10, 4))
  144     axes[0].hist(gs, bins=20, range=(0, 1), color="steelblue", edgecolor="white", alpha=0.85)
  145     axes[0].axvline(0.5, color="gray", ls="--"); axes[0].set_xlabel("gold_share"); axes[0].set_ylabel("count")
  146     axes[0].set_title(f"gold_share dist (μ={stats['mean']:.3f})")
  147
  148     if "n_cand" in df.columns:
  149         axes[1].scatter(df["n_cand"], gs, alpha=0.4, s=20, edgecolors="none")
  150         axes[1].set_xlabel("|C(t)|"); axes[1].set_ylabel("gold_share")
  151         axes[1].set_title("gold_share vs |C(t)|")
  152     fig.suptitle(f"Gate 3: Gold Share  [{VERDICT[verdict]}]")
  153     fig.tight_layout()
  154     path = os.path.join(OUT_DIR, "gate3_gold_share.png")
  155     os.makedirs(OUT_DIR, exist_ok=True)
  156     fig.savefig(path, dpi=150)
  157     plt.close(fig)
  158
  159     return {**stats, "verdict": verdict, "note": note, "figure": path}
  160
  161
  162 # ========== Gate 4: |C| confounding ==========
  163
  164 def gate_candidate_size_bias(df):
  165     """|C| 是否系统地影响 gold_share？如果是，mass 在测候选规模而非 informativeness。"""
  166     sub = df.dropna(subset=["n_cand", "gold_share"])
  167     if len(sub) < 10:
  168         return {"verdict": None, "note": "样本不足"}
  169
  170     from scipy.stats import pearsonr, spearmanr
  171     pr, pp = pearsonr(sub["n_cand"], sub["gold_share"])
  172     sr, sp = spearmanr(sub["n_cand"], sub["gold_share"])
  173     abs_max = max(abs(pr), abs(sr))
  174
  175     if abs_max > 0.5:
  176         verdict, note = False, (f"|C| 与 mass 强相关 (r={abs_max:.3f}) → "
  177                                  "mass 主要在测候选规模。需固定 |C| 或在回归中控制。")
  178     elif abs_max > 0.3:
  179         verdict, note = None, (f"|C| 与 mass 中等相关 (r={abs_max:.3f}) → "
  180                                 "建议在回归中把 log|C| 作协变量")
  181     else:
  182         verdict, note = True, f"|C| 与 mass 弱相关 (r={abs_max:.3f}) → 无显著偏置"
  183
  184     fig, ax = plt.subplots(figsize=(5, 4))
  185     ax.scatter(sub["n_cand"], sub["gold_share"], alpha=0.4, s=20, edgecolors="none")
  186     z = np.polyfit(sub["n_cand"], sub["gold_share"], 1)
  187     xs = np.linspace(sub["n_cand"].min(), sub["n_cand"].max(), 50)
  188     ax.plot(xs, np.polyval(z, xs), "C1--", lw=1.5)
  189     ax.set_xlabel("|C(t)|"); ax.set_ylabel("gold_share")
  190     ax.set_title(f"Gate 4: |C| vs mass — r={pr:.3f} (Pearson), ρ={sr:.3f} (Spearman)  [{VERDICT[verdict]}]")
  191     fig.tight_layout()
  192     path = os.path.join(OUT_DIR, "gate4_candidate_size_bias.png")
  193     os.makedirs(OUT_DIR, exist_ok=True)
  194     fig.savefig(path, dpi=150)
  195     plt.close(fig)
  196
  197     return {"pearson_r": round(pr, 4), "pearson_p": round(pp, 6),
  198             "spearman_r": round(sr, 4), "spearman_p": round(sp, 6),
  199             "verdict": verdict, "note": note, "figure": path}
  200
  201
  202 # ========== Report ==========
  203
  204 def print_report(gates):
  205     print("\n" + "=" * 66)
  206     print("  PILOT GATE REPORT")
  207     print("=" * 66)
  208     for name, g in gates.items():
  209         v = g.get("verdict")
  210         flag = {True: "[PASS]", False: "[FAIL]", None: "[SKIP]"}.get(v, "[??]")
  211         print(f"\n  Gate {name}: {flag}")
  212         for k, val in g.items():
  213             if k in ("verdict", "note", "figure"):
  214                 continue
  215             if isinstance(val, float):
  216                 print(f"    {k}: {val:.4f}")
  217         print(f"    → {g.get('note', '')}")
  218         if g.get("figure"):
  219             print(f"    fig: {g['figure']}")
  220
  221     blocked = any(g.get("verdict") is False for g in gates.values())
  222     print("\n" + "=" * 66)
  223     if blocked:
  224         print("  [BLOCKED] 存在 FAIL 门 → 先修数据，再重跑本脚本。不要租 GPU。")
  225     else:
  226         degraded = [k for k, g in gates.items() if g.get("verdict") is None]
  227         if degraded:
  228             print(f"  [GO] 无阻断门。{degraded} 降级（见上），可租 GPU。")
  229         else:
  230             print("  [GO] 四门全绿。租 GPU，跑 Exp2。")
  231     print("=" * 66 + "\n")
  232
  233
  234 # ========== Main ==========
  235
  236 def main():
  237     ap = argparse.ArgumentParser(description="Pilot 四门诊断")
  238     ap.add_argument("--dry-run", action="store_true", help="模拟 mass 数据，验证全流程")
  239     ap.add_argument("--csv", default=None, help="mass CSV（含 gold_share, n_cand, dataset 列）")
  240     ap.add_argument("--csv-b", default=None, help="B 域 mass CSV")
  241     ap.add_argument("--mass-a", default=None, help="mass_a.npy（pilot_check 产出）")
  242     ap.add_argument("--mass-b", default=None, help="mass_b.npy（pilot_check 产出）")
  243     args = ap.parse_args()
  244
  245     os.makedirs(OUT_DIR, exist_ok=True)
  246
  247     # Load data
  248     if args.dry_run:
  249         print("[dry-run] 生成模拟 mass 数据")
  250         rng = np.random.RandomState(42)
  251         n_a, n_b = 160, 20
  252         # Scenario A: healthy dynamic range
  253         mass_a = np.clip(rng.beta(2, 3, n_a) + rng.normal(0, 0.03, n_a), 0.01, 0.99)
  254         n_cand_a = rng.randint(50, 101, n_a)
  255         # Scenario B: higher mass (knowledge-rich domain)
  256         mass_b = np.clip(rng.beta(4, 2, n_b) + rng.normal(0, 0.03, n_b), 0.01, 0.99)
  257         n_cand_b = rng.randint(20, 60, n_b)
  258         df_a = pd.DataFrame({"gold_share": mass_a, "n_cand": n_cand_a, "dataset": "A"})
  259         df_b = pd.DataFrame({"gold_share": mass_b, "n_cand": n_cand_b, "dataset": "B"})
  260         df = pd.concat([df_a, df_b], ignore_index=True)
  261     elif args.csv:
  262         df_a = pd.read_csv(args.csv)
  263         df_b = pd.read_csv(args.csv_b) if args.csv_b else None
  264         mass_a = df_a["gold_share"].dropna().values
  265         mass_b = df_b["gold_share"].dropna().values if df_b is not None else None
  266         if "n_cand" not in df_a.columns:
  267             df_a["n_cand"] = 50
  # 所以我认为接下来不要直接开始正式实验，而是完成一个专门的 Pilot 分析脚本。

  建议写一个：

  ```
  pilot_analysis.py
  ```

  一次性输出所有关键诊断，包括：

  1. **Dataset A 动态范围**

     * mass 的均值、标准差、四分位数
     * 直方图

  2. **A/B Common Support**

     * 两域密度图（或小提琴图）
     * 重叠区间统计

  3. **Gold Share 分布**

     * gold_share 与 candidate_mass 的关系

  4. **|C| 对 mass 的影响**

     * 散点图
     * Pearson / Spearman 相关

  5. **Pilot Gate 判定**

     * Gate 1：PASS / FAIL（动态范围）
     * Gate 2：PASS / FAIL（共同支撑）
     * Gate 3：PASS / FAIL（gold_share）
     * Gate 4：PASS / FAIL（|C| 是否引入明显偏置）

  **这份报告会决定后续是否值得投入 GPU 跑正式实验。**如果四个 Gate 都通过，你就可以比较有把握地进入正式实验；如果有某个
   Gate 不通过，调整方案的成本远低于跑完整个实验后再返工。


● Write(code\pilot_analysis.py)
  ⎿  Wrote 308 lines to code\pilot_analysis.py
       1 #!/usr/bin/env python3
       2 # -*- coding: utf-8 -*-
       3 """
       4 pilot_analysis.py — 四门诊断（一次性输出，决定是否值得租GPU）
       5
       6 在 pilot_check.py 的 B/C 阶段之后运行。读 results/pilot/mass_{a,b}.npy，
       7 或任何包含 gold_share / n_cand / dataset 列的 CSV。
       8
       9 用法:
      10   python pilot_analysis.py --mass-a results/pilot/mass_a.npy        # pilot_check 产出
     … +304 lines (ctrl+o to expand)

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python code/pilot_analysis.py --dry-run
      2>&1)
  ⎿  [dry-run]     ģ   mass     

     Gate 1: Dynamic Range (A)   
     … +43 lines (ctrl+o to expand)

● 四门全绿，dry-run 通过。验证产出文件——

  Listed 1 directory (ctrl+o to expand)

● 全部完成。pilot_analysis.py 已就绪：

  新版本方案/code/pilot_analysis.py   ← 四门诊断（一次性）
  新版本方案/results/pilot/
  ├── gate1_dynamic_range_A.png       ← 直方图 + spread/p10/p90
  ├── gate2_common_support.png        ← A/B 密度对比 + IoU
  ├── gate3_gold_share.png            ← gold_share 分布 + vs |C| 散点
  ├── gate4_candidate_size_bias.png   ← |C| 对 mass 的影响 + Pearson/Spearman
  └── gate_report.json                ← 结构化判定数据

  使用方式：
  - python code/pilot_analysis.py --dry-run → 模拟验证
  - python code/pilot_analysis.py --csv results/exp2/mass_A.csv → 真实数据
  - python code/pilot_analysis.py --mass-a results/pilot/mass_a.npy → pilot_check 产出

  GPU 到位后流程：pilot_check --stage BC → pilot_analysis --mass-a → 四门全绿 → exp2_logit_analysis → 正式实验。

✻ Baked for 1m 16s

❯ 我觉得你们现在已经到了一个**不要再改论文设计，而要开始验证论文设计**的阶段。

  从整个迭代过程来看，大概经历了：

  * **v3**：框架很大，但变量没有定义清楚。
  * **v4**：开始聚焦 `candidate_mass`，但暴露了 softmax、TOST、guided_choice 等问题。
  * **v5**：把这些问题基本都收敛了，并建立了 Pilot Gate。

  **现在的工程状态其实已经很成熟。**

  ---

  ## 我建议暂停增加任何新模块

  目前已经有：

  * ✅ config.py
  * ✅ utils.py
  * ✅ candidate_mass.py
  * ✅ pilot_check.py
  * ✅ pilot_analysis.py

  已经形成了完整的数据流：

  ```text
  Dataset
      │
      ▼
  candidate_mass.py
      │
      ▼
  pilot_check.py
      │
      ▼
  pilot_analysis.py
      │
      ▼
  Gate Report
      │
      ▼
  决定是否进入正式实验
  ```

  这是一个比较规范的科研 workflow。

  **如果继续增加模块，很容易又回到 v3 那种"设计越来越复杂"的状态。**

  ---

  # 我现在最建议做的一件事

  不是继续写代码。

  而是：

  > **真正跑一次 Pilot。**

  ---

  ## 为什么？

  因为目前所有 Gate 都只是

  ```text
  dry-run
  ```

  也就是：

  模拟数据。

  模拟数据只能说明：

  > **程序没问题。**

  不能说明：

  > **论文假设没问题。**

  真正需要知道的是：

  例如：

  ```
  Gate1
  PASS?
  ```

  到底是不是：

  ```
  spread = 0.18
  ```

  还是：

  ```
  spread = 0.03
  ```

  这是完全两回事。

  ---

  ## GPU到了以后第一件事

  不是：

  ```
  Exp1
  ```

  不是：

  ```
  Exp2
  ```

  而是：

  ```
  pilot_check.py
  ```

  建议：

  Dataset A

  随机30条

  Dataset B

  随机30条

  一共：

  ```
  60 samples
  ```

  就够了。

  ---

  然后：

  ```
  pilot_analysis.py
  ```

  生成四张图。

  ---

  ## 如果我是PI（导师）

  我甚至不会让学生继续写任何代码。

  我会要求：

  **先拿到这四张真实图。**

  包括：

  * Gate1 Histogram
  * Gate2 Density
  * Gate3 Gold Share
  * Gate4 Candidate Size

  然后我们一起看。

  很多论文就是死在：

  > 真数据和想象完全不同。

  ---

  # 我唯一还建议补充一个东西（很小）

  你的

  ```
  gate_report.json
  ```

  现在最好增加一个：

  ```json
  {
    "overall": "PASS",
    "recommendation": "Proceed to full experiments"
  }
  ```

  或者：

  ```json
  {
    "overall": "FAIL",
    "reason": [
        "Gate2 failed",
        "Gate4 failed"
    ]
  }
  ```

  这样以后：

  Exp2

  Exp3

  CI

  自动化实验

  都可以直接读取。

  不用人工判断。

  ---

  # 我觉得现在整个项目的成熟度

  如果按照科研开发阶段评价：

  | 阶段      | 状态      |
  | ------- | ------- |
  | 研究问题    | ✅ 已稳定   |
  | 理论假设    | ✅ 已稳定   |
  | 实验设计    | ✅ 已基本稳定 |
  | 指标定义    | ✅ 已稳定   |
  | 数据结构    | ✅ 已稳定   |
  | Pilot工具 | ✅ 已完成   |
  | GPU实验   | ⏳ 待开始   |
  | 统计分析    | ⏳ 待开始   |
  | 论文写作    | ⏳ 待开始   |

  我的建议只有一句话：

  > **停止继续设计，开始产生真实数据。**

  真正决定这篇论文能否成立的，不会是再讨论一次 `candidate_mass` 或 Exp1 的细节，而是第一批真实 Pilot
  数据是否支持你现在的假设。等 GPU 到位后，优先完成 30+30 条样本的 Pilot，跑完 `pilot_check` 和 `pilot_analysis`，再根
  据真实结果决定是否进入完整实验。这会比继续优化设计更有效，也能避免在未经验证的假设上投入大量计算资源。


● 收到。只补一个 overall 判定。

  Read 1 file (ctrl+o to expand)

● Update(code\pilot_analysis.py)
  ⎿  Added 17 lines, removed 7 lines
      219              print(f"    fig: {g['figure']}")
      220
      221      blocked = any(g.get("verdict") is False for g in gates.values())
      222 -    print("\n" + "=" * 66)
      222 +    degraded = [k for k, g in gates.items() if g.get("verdict") is None]
      223 +    fails   = [k for k, g in gates.items() if g.get("verdict") is False]
      224 +
      225      if blocked:
      224 -        print("  [BLOCKED] 存在 FAIL 门 → 先修数据，再重跑本脚本。不要租 GPU。")
      226 +        overall = "FAIL"
      227 +        recommendation = f"先修 {fails}，再重跑本脚本。不要租 GPU。"
      228 +    elif degraded:
      229 +        overall = "GO_DEGRADED"
      230 +        recommendation = f"无阻断门。{degraded} 降级（见上），可租 GPU，注意调整方案。"
      231      else:
      226 -        degraded = [k for k, g in gates.items() if g.get("verdict") is None]
      227 -        if degraded:
      228 -            print(f"  [GO] 无阻断门。{degraded} 降级（见上），可租 GPU。")
      229 -        else:
      230 -            print("  [GO] 四门全绿。租 GPU，跑 Exp2。")
      232 +        overall = "GO"
      233 +        recommendation = "四门全绿。租 GPU，跑 Exp2。"
      234 +
      235 +    summary = {"overall": overall, "recommendation": recommendation,
      236 +               "blocked_by": fails, "degraded": degraded}
      234 +
      235 +    summary = {"overall": overall, "recommendation": recommendation,
      236 +               "blocked_by": fails, "degraded": degraded}
      237 +
      238 +    print("\n" + "=" * 66)
      239 +    print(f"  [{overall}] {recommendation}")
      240      print("=" * 66 + "\n")
      241 +    return summary
      242
      243
      244  # ========== Main ==========

● Update(code\pilot_analysis.py)
  ⎿  Added 6 lines, removed 3 lines
      303      print("Gate 4: |C| Bias …")
      304      gates["4_cand_size_bias"] = gate_candidate_size_bias(df[df["dataset"] == "A"])
      305
      306 -    print_report(gates)
      306 +    summary = print_report(gates)
      307
      308      # Save gate data
      309      import json
      310      report_path = os.path.join(OUT_DIR, "gate_report.json")
      311 +    payload = {k: {kk: vv for kk, vv in v.items() if kk != "figure"}
      312 +               for k, v in gates.items()}
      313 +    payload["overall"] = summary["overall"]
      314 +    payload["recommendation"] = summary["recommendation"]
      315      with open(report_path, "w", encoding="utf-8") as f:
      312 -        json.dump({k: {kk: vv for kk, vv in v.items() if kk != "figure"}
      313 -                   for k, v in gates.items()}, f, ensure_ascii=False, indent=2)
      316 +        json.dump(payload, f, ensure_ascii=False, indent=2)
      317      print(f"Report saved: {report_path}")
      318
      319

● Bash(cd "C:/Users/19691/Desktop/投稿论文/Thesis/J-space论文/新版本方案" && python code/pilot_analysis.py --dry-run
      2>&1 | tail -15)
  ⎿      fig: C:\Users\19691\Desktop\Ͷ      \Thesis\J-space    \ °汾    \results\pilot\gate3_gold_share.png

       Gate 4_cand_size_bias: [PASS]
     … +13 lines (ctrl+o to expand)

● Bash(python -c "
      import json…)
  ⎿  overall: GO
     recommendation:     ȫ ̡    GPU     Exp2  

● 完成。gate_report.json 现在包含：

  {
    "overall": "GO",
    "recommendation": "四门全绿。租 GPU，跑 Exp2。",
    "1_dynamic_range": {...},
    "2_common_support": {...},
    ...
  }

  三个判定值：GO / GO_DEGRADED / FAIL。后续自动化实验脚本可以直接读这个 JSON 的 overall 字段决定是否继续。

  ---
  当前项目状态跟你列的表一致：研究问题/假设/设计/指标/数据结构已稳定，Pilot 工具链完整。不再加模块。GPU 到位后跑
  pilot_check --stage BC → pilot_analysis → 看四张真实图。

✻ Sautéed for 44s

※ recap: 论文 v5.0 工具链已全部建成并通过 dry-run 验证。四个 Pilot Gate 脚本就绪，等 GPU 到位后跑第一批 60 条真实 mass
  数据，看核心假设是否成立。 (disable recaps in /config)