Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261001-233756-search-t1-abc-r0-B-population

节点 n22

实现类型内「表达-斜率对齐」(future-ness)第三抽样轴并做增殖残差化;实测 de_recovery 完全不动、cell_state/covariation 显著下降,按预定判定回退 β_future=0,发布 node18 参数(γ=-1.6、slope_min=0.1)为默认。

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261001-233756-search-t1-abc-r0-B-population
父节点n16
子节点—
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 54.74(+0.2) · proxy 57.11(+0.3) · proxy2 57.11(+0.3) · X3 50.00(+0.0) · 3 次复测均分 54.66
审查通过 1 未发现问题:run.py 仅通过 view_io 的 load_manifest/read_stage/panel_genes 读取 --data 视图内数据;唯一的直接文件读取是 run.py:58 的 view_dir/prior/reactome/gene_sets.gmt,在 manifest['prior'] 清单内,且默认 lam=0 不触发;无绝对路径、..、网络访问或打分器路径。; 2 未发现问题:数值常量仅为超参数(run.py:45-54 的 BETA/GAMMA/SLOPE_MIN 等);PROLIF(run.py:80)是通用细胞周期标记基因列表,非目标阶段统计量…
用时?从运行开始到结束(或到现在)的挂钟时间。12 分
程序版本1a3a4ad0436707ef2c24ca8a63a7cf19c7b568d4 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 1a3a4ad043:solution/METHOD.md

实现类型内「表达-斜率对齐」(future-ness)第三抽样轴并做增殖残差化;实测 de_recovery 完全不动、cell_state/covariation 显著下降,按预定判定回退 β_future=0,发布 node18 参数(γ=-1.6、slope_min=0.1)为默认。

方法

基底 = node16 代码(= node7 逻辑),默认参数改为 node18 已验证的 γ=-1.6、slope_min=0.1(树上最佳 rank3 54.66):copy_last 最新官方阶段 + 两级增殖重加权抽样(β_type=-4、β_cell=-1,Efraimidis–Spirakis 无放回)+ 类型内「增殖得分–基因表达」OLS 斜率逐细胞表达调整 x_adj = x + γ·slope_g·(tmean−p_i),clip≥0;外部/退化视图(X3)跳过表达调整。

本节点新机制(PLAN:future-ness 抽样轴,最终默认关闭):

  • 每类型 c 对 |slope|>slope_min 的基因算 F_i = Σ_g slope_g·(x_ig−mean_cg)/σ_cg(σ 下限 0.05;参与基因 <50 时 F=0;类型 <20 细胞跳过)。
  • 残差化:诊断实测 F 与类型内增殖偏离高度共线(Pearson r=0.57–0.95,PLAN 风险 1 命中),故 F 先对 prolif 偏离做正交投影取残差再标准化(残差化后 r≈0),避免与 β_cell 轴冗余。
  • 第三级权重 w_future = max(1+β_future·Fstd_i, 1e-6),w = w_type·w_cell·w_future,仍用 E-S 无放回抽样;只改选哪些细胞,不改表达值。
  • 外部/退化视图(X3)同样跳过 future-ness。

查分记录(proxy A 半 seed0 共 4 次 + seed1 1 次 + X3 1 次 = 6/20)

配置板分de_reccovcell_statedirection
对照 β_f=0(γ=-1.6,sm=0.1,=node18)57.5952.4853.0263.1559.69
β_f=+1(残差化)49.9150.0044.7251.7651.75
β_f=−1(残差化)51.2852.4846.4946.9259.16
β_f=+0.3(残差化)56.6252.4851.9661.0559.17
对照 seed157.3152.4852.4862.2460.10
X3 默认参数50.00————

判定(PLAN 预定规则:β_f=−1 时 de_recovery<52.5 即判机制失败):

  • 机制被证伪:de_score 在 β_f=−1/+0.3 下与对照完全相同(0.0909),de_recovery 对 future-ness 抽样轴零响应;β_f=+1 反而把 de_score 打到 0。同时所有 β_f≠0 都伤 cell_state(最多 −16)与 covariation(最多 −8),弱到 β_f=+0.3 仍净降 0.97。
  • 结论:de_recovery 的 de_score 疑似粗粒量(对照/β−1 同为 0.0909),在「不改表达值、只改抽样」的机制族内不动;node18 ANALYSIS 的「对抽样不敏感」判断在第三条抽样轴上再次成立。
  • 未测非残差化原始 F 的查分(本地诊断 r>0.7 即按 PLAN 切残差化,省额度)。

发布

默认 = β_future=0、γ=-1.6、slope_min=0.1、β_type=-4、β_cell=-1(λ 通路平滑代码保留但关闭,node16 已证伪)。相对父节点 16(γ=-0.8、sm=0.05)的预期收益即 node18 的已验证增益(板分 54.54→54.74,rank3 54.48→54.66,3 种子 A 半一致)。β_f=0 时代码路径与改动前逐字节等价(final_proxy 与首轮 ctrl md5 相同验证)。

验证过

  • proxy / proxy2 / X3 三视图跑通 + vec-check 全 ok;proxy2 输出与 proxy md5 相同(Qiu 外部输入照旧忽略)。
  • X3 默认查分 = 50.0(退化+外部保护生效,与父一致)。
  • 对照双种子 A 半 57.59 / 57.31,de_recovery 稳定 52.48。
  • 运行 ~5 s/视图,峰值内存 <2 GB;确定性仅 np.random.default_rng(seed)。

没验证 / 风险

  • final 视图(E8.5+E9.5→E10.5)未实测(无该视图);退化保护逻辑与父相同。
  • future-ness 只做了线性残差化;非线性去共线(如按增殖分层内排序)未测——但鉴于 de_score 对抽样零响应,预期同败。
  • γ=-1.6/sm=0.1 的 B 半增益已由 node18 三种子验证,本节点未重复多种子裁决。

外部知识来源

  • PROLIF 基因列表为通用 cell-cycle 标记(Mki67、Top2a、Pcna、Cdk1、Ccna2、Ccnb1、Aurkb、Bub1、Cenpf、Nusc1),非阶段特异知识;发布路径不读任何 prior 文件。

调研员的计划

名称表达-斜率对齐加权抽样:按类型内未来方向得分重加权选细胞以提 de_recovery
动机de_recovery 是全树最弱且最顽固的分组(node7/16/18 均 51.69),node18 ANALYSIS 明确称其'对抽样和表达扰动完全不敏感'。但该'不敏感'仅针对增殖水平轴(β_cell)和小幅表达修饰;尚未有节点尝试用斜率方向本身作为抽样轴。node7 的 de_recovery +1.03(50.65→51.69)来自表达调整,说明斜率信息对该组有微弱信号。node13 证明乘性表达修改可提 de_recovery 至 52.48 但崩 cell_state(59.64→41.94),因此需要一种不改表达值、仅改选哪些细胞的机制。本方案用类型内'表达-斜率对齐得分'(future-ness score)作为第三级抽样权重,选出表达已沿斜率方向偏移的真实细胞,从而在不修改任何表达值的前提下增强群体 DE 信号。
做法基底 = node7 代码(= node16 发布),起始参数取 node18 已验证的 γ=-1.6、slope_min=0.1(若 Engineer 能直接取 node18 run.py 则以其为基底)。新增步骤:

1. 在现有逐类型 OLS 斜率计算后(已有代码),对每个细胞 i(类型 c)计算 future-ness 得分:F_i = Σ_{g: |slope_g|>slope_min} slope_g · (x_ig − mean_cg) / σ_cg,其中 σ_cg 为类型 c 内基因 g 的标准差(防高表达基因主导)。仅用 |slope|>slope_min 的基因(与表达调整共用阈值),类型内细胞数 < MIN_TYPE_CELLS 时 F_i=0。

2. 第三级抽样权重:w_future = max(1 + β_future · (F_i − mean_F_c), 1e-6),其中 mean_F_c 为类型 c 内 F 均值。最终权重 w = w_type · w_cell · w_future,仍用 Efraimidis–Spirakis 无放回抽样。

3. 参数:β_future 初值 -1.0(负号 = 选 F 高的细胞,即表达已沿斜率正方向偏移的细胞)。搜索范围 β_future ∈ {-0.5, -1.0, -1.5, -2.0}。γ 保持 -1.6,slope_min 保持 0.1。

4. 表达调整(γ·slope·Δprolif)保留不变——future-ness 只影响选哪些细胞,不改表达值。

5. 单输入阶段(proxy):全部计算在单快照内完成(类型内均值、标准差、斜率),无需第二时间点,天然适用。proxy2:同样只用 E8.5 官方阶段计算,Qiu 外部输入照旧忽略。final(E8.5+E9.5→E10.5):若 E9.5 池退化则跳过(与父逻辑相同保护)。X3:跳过表达调整 + 跳过 future-ness 加权(is_external 检查),退回纯 copy_last。

6. 快速筛选流程:先跑 proxy A-half seed 0,仅看 de_recovery 和 covariation;若 de_recovery ≥ 53.5(+1.8)且 covariation ≥ 51(未降 ≥2),再跑 seed 1 确认;若 de_recovery ≥ 54(+2.3,超噪声)且无组降 ≥2,跑 seed 2/3 共 4 种子裁决。总查分 ≤ 8 次。

7. 若 β_future 全范围无效(所有点 de_recovery < 53 或任一组降 ≥2),回退 β_future=0(等价父逻辑),零风险收尾。
风险1) future-ness 得分与增殖偏离高度共线(低增殖细胞可能恰好是斜率方向偏移的细胞),导致 β_future 与 β_cell 冗余、无额外增益——Engineer 应在第一次查分后计算 F_i 与 prolif_i 的类型内 Pearson r,若 |r|>0.8 则机制退化,应尝试用残差 F_i − proj(F|prolif) 替代;2) 斜率估计噪声大(类型内细胞少时 σ_cg 不稳定),future-ness 变成噪声抽样——已有 MIN_TYPE_CELLS=20 保护,但可额外要求参与求和的基因数 ≥50 否则 F_i=0;3) 选出'未来方向'细胞可能使细胞类型组成偏移,伤 cell_state——若 cell_state 降 ≥2 立即停止;4) de_recovery 可能确实对抽样不敏感(如 node18 ANALYSIS 所言),此时所有 β_future 值均无效——预判标准:β_future=-1 时 de_recovery < 52.5 即可判定机制失败,不必扫完全网格。

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 6edcfbbbdb。改动的文件:solution/METHOD.md +32 −31、solution/run.py +122 −76

diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 8dda540..9d99793 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,47 +1,48 @@-扫描了 Reactome 通路约束斜率平滑(λ)与 |slope| 阈值 / γ 网格:全部低于对照或在噪声内,按预定判定保留 node7 原逻辑,λ=0 默认输出与父逐字节等价。+实现类型内「表达-斜率对齐」(future-ness)第三抽样轴并做增殖残差化;实测 de_recovery 完全不动、cell_state/covariation 显著下降,按预定判定回退 β_future=0,发布 node18 参数(γ=-1.6、slope_min=0.1)为默认。  ## 方法 -基底 = 节点 7(完整保留,默认参数不变):copy_last 最新官方阶段 + 两级增殖重加权抽样(β_type=-4、β_cell=-1,Efraimidis–Spirakis)+ 类型内「增殖得分–基因表达」OLS 斜率的逐细胞表达调整 x_adj = x + γ·slope_g·(tmean_prolif − prolif_i),clip≥0;外部来源或抽样退化视图(X3)跳过表达调整。+基底 = node16 代码(= node7 逻辑),默认参数改为 node18 已验证的 γ=-1.6、slope_min=0.1(树上最佳 rank3 54.66):copy_last 最新官方阶段 + 两级增殖重加权抽样(β_type=-4、β_cell=-1,Efraimidis–Spirakis 无放回)+ 类型内「增殖得分–基因表达」OLS 斜率逐细胞表达调整 x_adj = x + γ·slope_g·(tmean−p_i),clip≥0;外部/退化视图(X3)跳过表达调整。 -本节点新增(默认关闭):-- `--lam λ`:读视图 `prior/reactome/gene_sets.gmt`(1843 个通路,通用通路注释,非阶段特异测量),取与面板交叠 ≥10 基因的集合,行归一化成稀疏矩阵 M;对每类型的阈值化斜率做子空间平滑 slope' = λ·(含 g 的通路载荷 L_P 的均值) + (1−λ)·slope_g,未被任何集合覆盖的基因保留原斜率。λ=0 即父逻辑(已验证输出 md5 与父逐字节一致)。-- `--slope-min` / `--top-frac`:斜率筛选阈值变体(绝对阈值 / 类型内相对分位)。+本节点新机制(PLAN:future-ness 抽样轴,最终默认关闭):+- 每类型 c 对 |slope|>slope_min 的基因算 F_i = Σ_g slope_g·(x_ig−mean_cg)/σ_cg(σ 下限 0.05;参与基因 <50 时 F=0;类型 <20 细胞跳过)。+- **残差化**:诊断实测 F 与类型内增殖偏离高度共线(Pearson r=0.57–0.95,PLAN 风险 1 命中),故 F 先对 prolif 偏离做正交投影取残差再标准化(残差化后 r≈0),避免与 β_cell 轴冗余。+- 第三级权重 w_future = max(1+β_future·Fstd_i, 1e-6),w = w_type·w_cell·w_future,仍用 E-S 无放回抽样;只改选哪些细胞,不改表达值。+- 外部/退化视图(X3)同样跳过 future-ness。 -## 查分记录(proxy A 半,seed 0,共 9 次)+## 查分记录(proxy A 半 seed0 共 4 次 + seed1 1 次 + X3 1 次 = 6/20) -| 配置 | 板分 | cov | cell_state |-|---|---|---|---|-| λ=0(对照,=父) | 57.20(父 METHOD 记录,输出 md5 相同) | 53.56 | 61.45 |-| λ=0.5, γ=-0.8 | 56.68 | 53.32 | 60.30 |-| λ=1.0, γ=-0.8 | 56.02 | 52.60 | 59.03 |-| λ=0.5, γ=-1.6 | 56.78 | 50.87 | 61.81 |-| λ=1.0, γ=-1.6 | 55.54 | 49.79 | 59.44 |-| top_frac=0.25, γ=-0.8 | 56.91 | 51.68 | 61.49 |-| slope_min=0.1, γ=-0.8 | 57.21 | 53.96 | 61.26 |-| slope_min=0.1, γ=-1.2 | 57.45(seed1: 57.17 vs 父 seed1 56.89) | 53.50 | 62.37 |-| slope_min=0.2, γ=-0.8 | 57.09 | 54.27 | 60.68 |+| 配置 | 板分 | de_rec | cov | cell_state | direction |+|---|---|---|---|---|---|+| 对照 β_f=0(γ=-1.6,sm=0.1,=node18) | **57.59** | 52.48 | 53.02 | 63.15 | 59.69 |+| β_f=+1(残差化) | 49.91 | 50.00 | 44.72 | 51.76 | 51.75 |+| β_f=−1(残差化) | 51.28 | 52.48 | 46.49 | 46.92 | 59.16 |+| β_f=+0.3(残差化) | 56.62 | 52.48 | 51.96 | 61.05 | 59.17 |+| 对照 seed1 | 57.31 | 52.48 | 52.48 | 62.24 | 60.10 |+| X3 默认参数 | 50.00 | — | — | — | — | -判定(PLAN 预定规则:任一组相对对照 +≥2 且无组 −≥2 才采纳):-- **通路平滑机制被证伪**:λ 越大分越低且 covariation 单调恶化(λ=1,γ=-1.6 时 cov 49.8)。通路平均把斜率幅度收缩(等效减小 |γ|),而 node7 的增益恰来自放大类型内增殖梯度;相干化没能补偿幅度损失。-- slope_min=0.1+γ=-1.2 双种子一致 +0.25~0.28,方向为正但远低于噪声阈 2,且属父节点已警告的「噪声内细扫 γ」,为控 A 半过拟合风险**不采纳**。-- 最终发布 = 父逻辑(λ=0、sm=0.05、γ=-0.8、β_type=-4、β_cell=-1),新代码路径默认全部关闭。+判定(PLAN 预定规则:β_f=−1 时 de_recovery<52.5 即判机制失败):+- **机制被证伪**:de_score 在 β_f=−1/+0.3 下与对照完全相同(0.0909),de_recovery 对 future-ness 抽样轴零响应;β_f=+1 反而把 de_score 打到 0。同时所有 β_f≠0 都伤 cell_state(最多 −16)与 covariation(最多 −8),弱到 β_f=+0.3 仍净降 0.97。+- 结论:de_recovery 的 de_score 疑似粗粒量(对照/β−1 同为 0.0909),在「不改表达值、只改抽样」的机制族内不动;node18 ANALYSIS 的「对抽样不敏感」判断在第三条抽样轴上再次成立。+- 未测非残差化原始 F 的查分(本地诊断 r>0.7 即按 PLAN 切残差化,省额度)。++## 发布++默认 = β_future=0、γ=-1.6、slope_min=0.1、β_type=-4、β_cell=-1(λ 通路平滑代码保留但关闭,node16 已证伪)。相对父节点 16(γ=-0.8、sm=0.05)的预期收益即 node18 的已验证增益(板分 54.54→54.74,rank3 54.48→54.66,3 种子 A 半一致)。β_f=0 时代码路径与改动前逐字节等价(final_proxy 与首轮 ctrl md5 相同验证)。  ## 验证过 -- λ=0 默认输出与父节点逐字节相同(md5);proxy / proxy2 / X3 三视图跑通 + vec-check ok;proxy2 输出与 proxy md5 相同(外部 Qiu E9.0 照旧忽略)。-- X3 默认参数查分 = 50.0(保护生效)。-- 运行 ~5 s / 视图,峰值内存 <2 GB,远低于限额。-- 确定性:仅 np.random.default_rng(seed)。+- proxy / proxy2 / X3 三视图跑通 + vec-check 全 ok;proxy2 输出与 proxy md5 相同(Qiu 外部输入照旧忽略)。+- X3 默认查分 = 50.0(退化+外部保护生效,与父一致)。+- 对照双种子 A 半 57.59 / 57.31,de_recovery 稳定 52.48。+- 运行 ~5 s/视图,峰值内存 <2 GB;确定性仅 np.random.default_rng(seed)。  ## 没验证 / 风险 -- 发布配置与父完全相同,B 半分数预期 ≈ 父(54.54 板分),无新增收益——本节点的贡献是**否证**通路约束方向并封住后续节点重复尝试。-- λ 混合只对 Reactome 做过;GO / MSigDB hallmark / CollecTRI 基因集未测(λ 机制已证伪,预计同理)。-- slope_min=0.1+γ=-1.2 的 +0.27(2 seeds)若在 B 半重现则是小增益,留给后续节点用更多种子裁决。-- final 视图(E8.5+E9.5→E10.5)未实测:若 E9.5 池 ≤ 目标细胞数,退化保护使表达调整跳过,退回 copy_last(安全无增益)。+- final 视图(E8.5+E9.5→E10.5)未实测(无该视图);退化保护逻辑与父相同。+- future-ness 只做了线性残差化;非线性去共线(如按增殖分层内排序)未测——但鉴于 de_score 对抽样零响应,预期同败。+- γ=-1.6/sm=0.1 的 B 半增益已由 node18 三种子验证,本节点未重复多种子裁决。  ## 外部知识来源 -- prior/reactome/gene_sets.gmt(视图内挂载的通用通路注释),仅用于已关闭的 λ 平滑实验;发布路径不读任何先验文件。-- PROLIF 基因列表为通用 cell-cycle 标记(同父节点),非阶段特异知识。+- PROLIF 基因列表为通用 cell-cycle 标记(Mki67、Top2a、Pcna、Cdk1、Ccna2、Ccnb1、Aurkb、Bub1、Cenpf、Nusc1),非阶段特异知识;发布路径不读任何 prior 文件。diff --git a/solution/run.py b/solution/run.pyindex 1e55337..add3c00 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,21 +1,25 @@ #!/usr/bin/env python3 """copy_last + two-level proliferation reweighting + proliferation-gradient-expression extrapolation.--Base (node 4/6): output a deterministic weighted subsample of the latest-official input stage. Type-level weight uses the type-mean proliferation-score (beta_type=-4, tuned); cell-level weight uses the within-type-deviation (beta_cell=-1, node 6).--New (this node): per cell type, regress each gene's expression on the cell's-proliferation score (OLS slope over ALL source cells of that type). Each-selected cell's expression is nudged along that slope toward its type's mean-proliferation level: x_adj[i,g] = x[i,g] + gamma * slope_g * (tmean - p_i),-applied only where |slope_g| > SLOPE_MIN, clipped to >= 0. Types with-< MIN_TYPE_CELLS cells are left untouched. gamma=0 reproduces the base.-Proliferation markers are generic cell-cycle genes (not stage-specific-knowledge); everything is computed within the input snapshot, so the method-works with a single input stage.+expression extrapolation + future-ness sampling axis.++Base (node 7/18): deterministic weighted subsample of the latest official+input stage. Type-level weight uses type-mean proliferation (beta_type=-4);+cell-level weight uses within-type deviation (beta_cell=-1). Expression of+each selected cell is nudged along per-type proliferation-expression OLS+slopes: x_adj = x + gamma*slope_g*(tmean - p_i), |slope_g|>slope_min,+clip>=0 (node 18: gamma=-1.6, slope_min=0.1).++New (this node): third sampling axis "future-ness". Per type c, for genes+passing the slope threshold, F_i = sum_g slope_g*(x_ig-mean_cg)/sigma_cg,+standardized within type, then w_future = max(1 + beta_future*Fstd_i, 1e-6).+beta_future>0 up-weights cells whose expression already leans along the+proliferation-gradient direction; beta_future=0 reproduces the base exactly.+Expression values are NOT modified by this axis (only which cells are kept).++External-only / degenerate views (e.g. test X3) skip both the expression+adjustment and the future-ness axis (pure copy_last fallback), as in parent.+Proliferation markers are generic cell-cycle genes; everything is computed+within the input snapshot, so a single input stage suffices. """  from __future__ import annotations@@ -38,20 +42,19 @@ from src.task1_temporal.view_io import (     write_prediction, ) -BETA_TYPE = -4.0   # tuned on T1 proxy A-half (node 4)-BETA_CELL = -1.0   # within-type weight (node 6)-GAMMA = -0.8       # expression-adjustment strength (negative tuned on proxy A-half)-SLOPE_MIN = 0.05   # only adjust genes with |slope| above this+BETA_TYPE = -4.0    # tuned on T1 proxy A-half (node 4)+BETA_CELL = -1.0    # within-type proliferation weight (node 6)+GAMMA = -1.6        # expression-adjustment strength (node 18, 3 seeds A-half)+SLOPE_MIN = 0.1     # |slope| threshold (node 18)+BETA_FUTURE = 0.0   # future-ness sampling axis strength (0 = node18 logic) MIN_TYPE_CELLS = 20-LAM = 0.0          # pathway-space smoothing weight for slopes (0 = node7 logic)+MIN_SLOPE_GENES = 50   # types with fewer thresholded slope genes get F=0+SIGMA_FLOOR = 0.05     # floor for within-type gene std in F normalization+LAM = 0.0           # pathway smoothing weight (falsified by node 16; keep 0) SET_MIN_OVERLAP = 10   def load_pathway_matrix(view_dir: str, genes) -> sparse.csr_matrix | None:-    """Sparse (n_sets x n_genes) row-normalized membership matrix from-    prior/reactome/gene_sets.gmt; None if unavailable. Sets with fewer than-    SET_MIN_OVERLAP panel genes are dropped. Generic pathway annotations-    (Reactome), not stage-specific measurements."""     path = os.path.join(view_dir, "prior", "reactome", "gene_sets.gmt")     if not os.path.exists(path):         return None@@ -73,6 +76,7 @@ def load_pathway_matrix(view_dir: str, genes) -> sparse.csr_matrix | None:         (data, (rows, cols)), shape=(max(rows) + 1, len(genes))     ) + PROLIF = ["Mki67", "Top2a", "Pcna", "Cdk1", "Ccna2", "Ccnb1",           "Aurkb", "Bub1", "Cenpf", "Nusc1"] @@ -102,10 +106,12 @@ def main() -> None:     parser.add_argument("--gamma", type=float, default=None)     parser.add_argument("--beta-cell", type=float, default=None)     parser.add_argument("--beta-type", type=float, default=None)+    parser.add_argument("--beta-future", type=float, default=None)     parser.add_argument("--lam", type=float, default=None)     parser.add_argument("--slope-min", type=float, default=None)-    parser.add_argument("--top-frac", type=float, default=None,-                        help="if set, keep only top fraction of |slope| genes per type")+    parser.add_argument("--top-frac", type=float, default=None)+    parser.add_argument("--debug-stats", default=None,+                        help="write F/prolif correlation diagnostics here")     args = parser.parse_args()     slope_min = SLOPE_MIN if args.slope_min is None else args.slope_min     top_frac = args.top_frac@@ -113,6 +119,7 @@ def main() -> None:     gamma = GAMMA if args.gamma is None else args.gamma     beta_cell = BETA_CELL if args.beta_cell is None else args.beta_cell     beta_type = BETA_TYPE if args.beta_type is None else args.beta_type+    beta_future = BETA_FUTURE if args.beta_future is None else args.beta_future      manifest = load_manifest(args.data)     genes = panel_genes(args.data, manifest)@@ -129,7 +136,14 @@ def main() -> None:         Xc = Xc.toarray() if sparse.issparse(Xc) else np.asarray(Xc)         score = Xc.mean(axis=1).astype(np.float64) +    X = last.X+    degenerate = n_target >= last.n_obs  # no real subsampling (e.g. test X3)+    adjust_ok = (score is not None and not is_external(entry) and not degenerate)+     inv = None+    type_score = None+    S = None+    fstd = None  # per-cell standardized future-ness (0 where unavailable)     if score is not None:         lab = labels_of(last)         uniq, inv = np.unique(lab, return_inverse=True)@@ -139,63 +153,95 @@ def main() -> None:         type_score = np.zeros(n_types, dtype=np.float64)         np.add.at(type_score, inv, score)         type_score /= counts-        ts = type_score[inv]-        mean_s = np.average(type_score, weights=counts)-        w_type = np.maximum(1.0 + beta_type * (ts - mean_s), 1e-6)-        w_cell = np.maximum(1.0 + beta_cell * (score - ts), 1e-6) if beta_cell != 0 else np.ones_like(w_type)-        rows = weighted_sample(w_type * w_cell, n_target, rng)-    else:-        rows = sample_rows(last.n_obs, n_target, rng) -    X = last.X-    Xsel = X[rows]-    # Guard: only adjust expression on official stages. In external-only views-    # (e.g. test question X3: different platform/technology, author cell-type-    # labels) the within-type proliferation regression proved harmful-    # (X3 A-half 50.0 -> 43.2), so we copy the cells untouched there.-    degenerate = n_target >= last.n_obs  # no real subsampling (e.g. test X3)-    if gamma != 0.0 and score is not None and not is_external(entry) and not degenerate:-        n_types = int(inv.max()) + 1-        S = np.zeros((n_types, X.shape[1]), dtype=np.float32)-        for t in range(n_types):-            idx = np.flatnonzero(inv == t)-            if len(idx) < MIN_TYPE_CELLS:-                continue-            pc = score[idx] - score[idx].mean()-            denom = float(pc @ pc)-            if denom <= 1e-8:-                continue-            slope = np.asarray(X[idx].T @ pc, dtype=np.float64).ravel() / denom-            if top_frac is not None:-                thr = np.quantile(np.abs(slope), 1.0 - top_frac)-                slope[np.abs(slope) <= thr] = 0.0-            else:-                slope[np.abs(slope) <= slope_min] = 0.0-            S[t] = slope.astype(np.float32)-        if lam != 0.0 and np.any(S):+        if adjust_ok or beta_future != 0.0:+            # Per-type OLS slopes of expression on proliferation score.+            n_types = int(inv.max()) + 1+            S = np.zeros((n_types, X.shape[1]), dtype=np.float32)+            fstd = np.zeros(last.n_obs, dtype=np.float64)+            dbg = []+            for t in range(n_types):+                idx = np.flatnonzero(inv == t)+                if len(idx) < MIN_TYPE_CELLS:+                    continue+                Xt = X[idx]+                pc = score[idx] - score[idx].mean()+                denom = float(pc @ pc)+                if denom <= 1e-8:+                    continue+                slope = np.asarray(Xt.T @ pc, dtype=np.float64).ravel() / denom+                if top_frac is not None:+                    thr = np.quantile(np.abs(slope), 1.0 - top_frac)+                    slope[np.abs(slope) <= thr] = 0.0+                else:+                    slope[np.abs(slope) <= slope_min] = 0.0+                S[t] = slope.astype(np.float32)+                if beta_future != 0.0:+                    nz = np.flatnonzero(slope != 0.0)+                    if len(nz) >= MIN_SLOPE_GENES:+                        Xn = Xt[:, nz]+                        mu = np.asarray(Xn.mean(axis=0)).ravel()+                        sq = np.asarray(Xn.multiply(Xn).mean(axis=0)).ravel()+                        sig = np.sqrt(np.maximum(sq - mu * mu, 0.0))+                        sig = np.maximum(sig, SIGMA_FLOOR)+                        v = slope[nz] / sig+                        f = np.asarray(Xn @ v, dtype=np.float64).ravel()+                        f -= f.mean()+                        # residualize against within-type proliferation+                        # deviation: raw F is collinear with prolif (r~0.6-0.95)+                        # and would be redundant with beta_cell axis.+                        pcc = pc - pc.mean()+                        d = float(pcc @ pcc)+                        if d > 1e-9:+                            f -= (float(f @ pcc) / d) * pcc+                        sd = f.std()+                        f = f / sd if sd > 1e-9 else np.zeros_like(f)+                        fstd[idx] = f+                        if args.debug_stats:+                            fs = f.std()+                            r = float(np.corrcoef(f, pc)[0, 1]) if fs > 0 else 0.0+                            dbg.append((int(t), len(idx), len(nz), r))+            if args.debug_stats and dbg:+                with open(args.debug_stats, "w") as fh:+                    for t, nc, ng, r in dbg:+                        fh.write(f"{t}\t{nc}\t{ng}\t{r:.4f}\n")++        if lam != 0.0 and S is not None and np.any(S):             M = load_pathway_matrix(args.data, genes)             if M is not None:-                # L_P per set (row-normalized mean), then back-project to genes:-                # slope' = lam * mean(L_P over sets containing g) + (1-lam)*slope-                L = M @ S.T                      # (n_sets, n_types)+                L = M @ S.T                 cnt = np.asarray(M.sum(axis=1)).ravel()                 L[cnt < 1e-9] = 0.0-                G = M.T @ L                      # (n_genes, n_types) sums of L_P+                G = M.T @ L                 nsets_g = np.asarray(M.T.sum(axis=1)).ravel()                 G[nsets_g > 0] /= nsets_g[nsets_g > 0, None]                 S = ((1.0 - lam) * S + lam * G.T).astype(np.float32)-        if np.any(S):-            c = (-gamma) * (score[rows] - type_score[inv[rows]])-            trow = inv[rows]-            parts = []-            chunk = 1024-            for a in range(0, len(rows), chunk):-                b = min(a + chunk, len(rows))-                D = np.asarray(Xsel[a:b].todense(), dtype=np.float32)-                D += (c[a:b, None] * S[trow[a:b]]).astype(np.float32)-                np.clip(D, 0.0, None, out=D)-                parts.append(sparse.csr_matrix(D))-            Xsel = sparse.vstack(parts).tocsr()++    if score is not None:+        ts = type_score[inv]+        mean_s = np.average(type_score, weights=counts)+        w_type = np.maximum(1.0 + beta_type * (ts - mean_s), 1e-6)+        w_cell = np.maximum(1.0 + beta_cell * (score - ts), 1e-6) if beta_cell != 0 else np.ones_like(w_type)+        w = w_type * w_cell+        if beta_future != 0.0 and fstd is not None and adjust_ok:+            w = w * np.maximum(1.0 + beta_future * fstd, 1e-6)+        rows = weighted_sample(w, n_target, rng)+    else:+        rows = sample_rows(last.n_obs, n_target, rng)++    Xsel = X[rows]+    if adjust_ok and gamma != 0.0 and S is not None and np.any(S):+        c = (-gamma) * (score[rows] - type_score[inv[rows]])+        trow = inv[rows]+        parts = []+        chunk = 1024+        for a in range(0, len(rows), chunk):+            b = min(a + chunk, len(rows))+            D = np.asarray(Xsel[a:b].todense(), dtype=np.float32)+            D += (c[a:b, None] * S[trow[a:b]]).astype(np.float32)+            np.clip(D, 0.0, None, out=D)+            parts.append(sparse.csr_matrix(D))+        Xsel = sparse.vstack(parts).tocsr()      write_prediction(Xsel, genes, args.out, seed=args.seed) 

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k041Within-stage pseudotime and graph toolkit offline: scanpy DPT/PAGA/Leiden, Palantir, CellRank 210.1186/s13059-019-1663-x (PAGA); 10.1038/s41587-019-0068-4 (Palantir); 10.1038/s41592-024-02303-9 (CellRank 2)
k031Offline OT toolkit in the sandbox: moscot TemporalProblem, wot OTModel, POT, geomloss10.1038/s41586-024-08453-2 (moscot); 10.1016/j.cell.2019.01.006 (Waddington-OT)
k018Damped per-type shift: shrinkage alpha on the observed deltanotes/plan/cards/T1.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么实现了 PLAN 的 future-ness 第三抽样轴(类型内表达-斜率对齐得分 F_i,含对增殖偏离的正交残差化,w = w_type·w_cell·w_future),本地实测证伪后回退 β_future=0;真正发布的改动是把默认参数从父节点的 γ=-0.8/slope_min=0.05 升级为 node18 已验证的 γ=-1.6/slope_min=0.1,并重构 run.py 使 adjust_ok 判定(非外部、非退化)同时门控表达调整与 future-ness。
各组分数的变化X3:零变化:50.00 → 50.00,退化+外部保护按预期生效。
board:54.54 → 54.74(Δ=+0.20),proxy/proxy2 各 +0.31,全部在 T1 约 2 分噪声内;该微量正向与 node18 三种子 A 半结论一致,但不能据此宣称 B 半有效。
cell_state:在噪声内偏正:56.59 → 57.60(Δ=+1.01,T1 噪声约 2 分),不能称为有效增益。
covariation:在噪声内偏负:52.88 → 52.42(Δ=-0.46)。
de_recovery:零变化:51.6867 → 51.6867(Δ=0.00),已是 node7/16/18/22 四个节点的同一数值,PLAN 的目标分组完全没被撬动。
direction:在噪声内:56.25 → 56.23(Δ=-0.02)。
假设是否成立否
经验
  1. 条件:de_recovery 目标组;做法:用类型内表达-斜率对齐得分做第三级抽样权重(不改任何表达值);结果:de_score 在 β_future=0/-1/+0.3 下完全相同(0.0909),榜上 de_recovery 精确 Δ=0.00——纯抽样轴对该组零响应,node18 的"de_recovery 对抽样不敏感"判断在第三条独立抽样轴上再次成立。
  2. PLAN 风险 1 命中且可提前廉价验证:原始 F_i 与类型内增殖偏离 Pearson r=0.57–0.95,新抽样轴与已有 β_cell 轴高度冗余;用 --debug-stats 类诊断在查分前算相关性,能在不花榜额度的情况下发现轴冗余。
  3. 残差化能消除共线(r≈0)但救不了机制:正交化后 β_future≠0 仍伤 cell_state(本地最多 -16)和 covariation(本地最多 -8),说明"沿任意新方向重加权选细胞"本身就会扭曲类型组成与基因协方差结构,与共线性无关。
  4. 预注册的硬判据(β_future=-1 时 de_recovery<52.5 即判机制失败)把本节点查分压到 6/20 次,避免了扫完 {-0.5,-1.0,-1.5,-2.0} 全网格;给新机制写单点否决判据是省额度的有效做法。
  5. 本地 A 半分组数值不能当榜分证据:Engineer 本地 A 半 de_recovery=52.48,榜上为 51.6867;A 半只适合做方向筛选,任何"本地 +X"结论必须回到变化量表核对。
  6. 参数升级(γ=-1.6、slope_min=0.1)在 B 半只兑现 +0.20,远小于 node18 在 A 半三种子看到的相对增益——A 半调出的超参迁移到 B 半会大幅衰减,属过拟合 A 半的典型表现。
下一步建议
  1. 针对 de_recovery:停止在抽样族(β_type/β_cell/future-ness 及其变体)上继续投入,四节点恒为 51.6867 说明该组只吃表达值修改;改试温和乘性表达调整(如按类型内 DE 方向对 |slope|>0.1 基因乘 1±δ,δ∈{0.05,0.1}),并预注册硬否决:covariation 或 cell_state 任一组降 ≥2 立即回退(node13 曾因乘性修改把 cell_state 从 59.64 打到 41.94)。
  2. 针对 de_recovery:先做指标粒度诊断而非再调参——de_score 在多个配置下重复取 0.0909,疑似粗粒/量化量;用本地脚本直接算该指标的分子分母,确认其最小可分辨步长,避免为不可动的 0.001 级差异烧额度。
  3. 针对整体榜分:γ=-1.6/slope_min=0.1 的 +0.20 在噪声内,若后续节点要以 node22 为父,应先在 B 半补 3–4 种子确认(对照 γ=-0.8/sm=0.05),否则整棵树会在一个 A 半过拟合的参数点上继续细扫。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。4
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 4 次:bash 3、write 1
用时1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 14,520 · 输出 1,509 · 思考 933
任务(第一行)分析节点 n22 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
ANALYSIS.json written
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/analyst.jsonl 13 KB
/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/analyst.stderr

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数16
工具调用共 18 次:bash 13、read 2、write 2、edit 1
用时8 分
token 数输入 43,549 · 输出 7,590 · 思考 7,238
任务(第一行)按计划修改程序:表达-斜率对齐加权抽样:按类型内未来方向得分重加权选细胞以提 de_recovery
最后的回答(摘录)
**做了什么**:实现 PLAN 的 future-ness 第三抽样轴(类型内表达-斜率对齐得分,含增殖残差化——诊断发现原始 F 与增殖偏离共线 r=0.57–0.95);实测机制被证伪后按预定判定回退 β_future=0,发布默认参数升级为 node18 已验证的 γ=-1.6、slope_min=0.1(树上最佳 rank3 54.66)。

**查分**(6/20):对照(=node18 配置)A 半 seed0/1 = 57.59/57.31;β_f=+1 → 49.91、β_f=−1 → 51.28、β_f=+0.3 → 56.62,de_score 在所有非零 β_f 下与对照相同(0.0909)或更差,de_recovery 对纯抽样机制零响应,且 cell_state/covariation 显著受损。X3 = 50.0(保护生效)。proxy/proxy2/X3 三视图跑通、vec-check 全 ok、输出确定性验证(md5)。

**下一步最值得试**:de_recovery 的 de_score 粗粒且只对表达值修改(node7 的 γ 项 +1.03)有响应、对任何抽样轴(node18 β_cell 扫描、本节点 future-ness)无响应——建议放弃抽样族,转向温和的乘性表达修改并守住 covariation ≥ −2 硬否决;另可对 γ=-1.6/sm=0.1 在 B 半补多种子确认增益稳定性。
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/engineer.jsonl 86 KB
/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/engineer.stderr

调研员

角色调研员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数1
工具调用共 0 次:—
用时2 分
token 数输入 14,428 · 输出 1,187 · 思考 4,416
任务(第一行)为节点 n22(改进,父节点 n16)写改进计划 PLAN.json
最后的回答(摘录)
(β_cell)和小幅表达修饰;尚未有节点尝试用斜率方向本身作为抽样轴。node7 的 de_recovery +1.03(50.65→51.69)来自表达调整,说明斜率信息对该组有微弱信号。node13 证明乘性表达修改可提 de_recovery 至 52.48 但崩 cell_state(59.64→41.94),因此需要一种不改表达值、仅改选哪些细胞的机制。本方案用类型内'表达-斜率对齐得分'(future-ness score)作为第三级抽样权重,选出表达已沿斜率方向偏移的真实细胞,从而在不修改任何表达值的前提下增强群体 DE 信号。",
  "approach": "基底 = node7 代码(= node16 发布),起始参数取 node18 已验证的 γ=-1.6、slope_min=0.1(若 Engineer 能直接取 node18 run.py 则以其为基底)。新增步骤:\n\n1. 在现有逐类型 OLS 斜率计算后(已有代码),对每个细胞 i(类型 c)计算 future-ness 得分:F_i = Σ_{g: |slope_g|>slope_min} slope_g · (x_ig − mean_cg) / σ_cg,其中 σ_cg 为类型 c 内基因 g 的标准差(防高表达基因主导)。仅用 |slope|>slope_min 的基因(与表达调整共用阈值),类型内细胞数 < MIN_TYPE_CELLS 时 F_i=0。\n\n2. 第三级抽样权重:w_future = max(1 + β_future · (F_i − mean_F_c), 1e-6),其中 mean_F_c 为类型 c 内 F 均值。最终权重 w = w_type · w_cell · w_future,仍用 Efraimidis–Spirakis 无放回抽样。\n\n3. 参数:β_future 初值 -1.0(负号 = 选 F 高的细胞,即表达已沿斜率正方向偏移的细胞)。搜索范围 β_future ∈ {-0.5, -1.0, -1.5, -2.0}。γ 保持 -1.6,slope_min 保持 0.1。\n\n4. 表达调整(γ·slope·Δprolif)保留不变——future-ness 只影响选哪些细胞,不改表达值。\n\n5. 单输入阶段(proxy):全部计算在单快照内完成(类型内均值、标准差、斜率),无需第二时间点,天然适用。proxy2:同样只用 E8.5 官方阶段计算,Qiu 外部输入照旧忽略。final(E8.5+E9.5→E10.5):若 E9.5 池退化则跳过(与父逻辑相同保护)。X3:跳过表达调整 + 跳过 future-ness 加权(is_external 检查),退回纯 copy_last。\n\n6. 快速筛选流程:先跑 proxy A-half seed 0,仅看 de_recovery 和 covariation;若 de_recovery ≥ 53.5(+1.8)且 covariation ≥ 51(未降 ≥2),再跑 seed 1 确认;若 de_recovery ≥ 54(+2.3,超噪声)且无组降 ≥2,跑 seed 2/3 共 4 种子裁决。总查分 ≤ 8 次。\n\n7. 若 β_future 全范围无效(所有点 de_recovery < 53 或任一组降 ≥2),回退 β_future=0(等价父逻辑),零风险收尾。",
  "expected_groups": ["de_recovery"],
  "risks": "1) future-ness 得分与增殖偏离高度共线(低增殖细胞可能恰好是斜率方向偏移的细胞),导致 β_future 与 β_cell 冗余、无额外增益——Engineer 应在第一次查分后计算 F_i 与 prolif_i 的类型内 Pearson r,若 |r|>0.8 则机制退化,应尝试用残差 F_i − proj(F|prolif) 替代;2) 斜率估计噪声大(类型内细胞少时 σ_cg 不稳定),future-ness 变成噪声抽样——已有 MIN_TYPE_CELLS=20 保护,但可额外要求参与求和的基因数 ≥50 否则 F_i=0;3) 选出'未来方向'细胞可能使细胞类型组成偏移,伤 cell_state——若 cell_state 降 ≥2 立即停止;4) de_recovery 可能确实对抽样不敏感(如 node18 ANALYSIS 所言),此时所有 β_future 值均无效——预判标准:β_future=-1 时 de_recovery < 52.5 即可判定机制失败,不必扫完全网格。",
  "sources": []
}
```
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/researcher.jsonl 5 KB
/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/researcher.stderr

审查员

角色审查员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数4
工具调用共 5 次:bash 2、read 2、write 1
用时不到 1 分
token 数输入 9,339 · 输出 897 · 思考 807
任务(第一行)审查节点 n22 的程序是否越界读取、写死目标、钻评分器漏洞
最后的回答(摘录)
REVIEW.json written
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/reviewer.jsonl 69 KB
/home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/22/reviewer.stderr