总览 · ← 返回运行 20261001-233756-search-t1-abc-r0-B-population
节点 n29
copy_last 两级增殖重加权(β_type=-4、β_cell=-1)+ 增殖-表达斜率放大(γ=-1.6、|slope|>0.1);新增的类型内离散度第三级抽样权重实测净负(de_recovery +0.52 / covariation −1.63),默认关闭。
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261001-233756-search-t1-abc-r0-B-population |
|---|---|
| 父节点 | n28 |
| 子节点 | — |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 改进 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 54.74(+0.0) · proxy 57.11(+0.0) · proxy2 57.11(+0.0) · X3 50.00(+0.0) · 3 次复测均分 54.66 |
| 审查 | 通过 1 越界读取:未发现问题——run.py 只通过 src.task1_temporal.view_io 的 load_manifest/read_stage/panel_genes 等接口读 --data 视图(run.py:45-55, 199-202),无绝对路径、'..'、/mnt、/home、打分器路径,无联网。; 2 硬编码目标统计量:未发现问题——代码中的常量仅为规则权重与超参(BETA_TYPE=-4、BETA_CELL=-1、GAMMA=-1.6、SLOPE_MIN=0.1、MIN_TYPE_CELLS=20、DISP_V=1500,run.py:57-67),无写死的类型比… |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 18 分 |
| 程序版本 | 40df6520a7663c096de42f9a97b7f572713fa81f (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git 40df6520a7:solution/METHOD.md
copy_last 两级增殖重加权(β_type=-4、β_cell=-1)+ 增殖-表达斜率放大(γ=-1.6、|slope|>0.1);新增的类型内离散度第三级抽样权重实测净负(de_recovery +0.52 / covariation −1.63),默认关闭。
方法(发布配置)
基底(继承 node 4/6/7/18/28,已验证):输出 = 最新「官方」输入阶段的 Efraimidis–Spirakis 加权无放回抽样,权重 = w_type × w_cell:
- w_type = max(1e-6, 1 + β_type·(prolif_type − mean_prolif)),β_type = -4(node 4);
- w_cell = max(1e-6, 1 + β_cell·(prolif_cell − prolif_type)),β_cell = -1(node 6)。
prolif 得分 = 通用 cell-cycle 基因(Mki67, Top2a, Pcna, Cdk1, Ccna2, Ccnb1, Aurkb, Bub1, Cenpf, Nusc1;来源:通用细胞周期标记注释,非阶段特异)log 表达逐细胞均值,全部在输入快照内现场计算,单输入阶段可用。
表达调整(node 7 机制、node 18 参数):每类型(≥20 细胞)内对全部源细胞做 OLS slope_g = cov(x_g, prolif)/var(prolif),保留 |slope_g| > 0.1;被选细胞 x_adj = x − γ·slope_g·(prolif_i − type_mean),γ = -1.6,clip ≥0,1024 行分块。保护开关:所选阶段为外部来源、或抽样退化(n_target ≥ 池大小,如 X3)时跳过表达调整与 w_disp。
本节点新机制:类型内表达离散度第三级权重(PLAN,已实现,默认关闭)
每类型(≥MIN_TYPE_CELLS=20)内按类型内方差取 top-V(V=1500)高变基因,对每个源细胞算 d_i = mean_g ((x_ig − μ_tg)/σ_tg)²(偏离类型质心的杠杆度),跨细胞标准化成 z_disp,w = w_type·w_cell·max(1e-6, 1 + β_disp·z_disp)(另有 --disp-mode exp 的 exp(β·z) 变体)。对称绝对值权重 → 类型比例基本不变。实现用 one-hot 稀疏矩乘一次算完所有类型的均值/平方均值(O(nnz)),全池 ~3 s。
proxy A 半 seed0 实测(对照 = 父配置 β_disp=0 → 57.59)
| β_disp | board | cell_state | covariation | de_recovery | direction | variogram | de_score |
|---|---|---|---|---|---|---|---|
| 0(对照/发布) | 57.59 | 63.15 | 53.02 | 52.48 | 59.69 | 0.000943 | 0.0909 |
| +0.5 | 57.29 | 62.80 | 51.39 | 53.00 | 59.67 | 0.000999 | 0.1091 |
| +1.0 | 54.63 | 58.40 | 45.97 | 53.00 | 58.68 | 0.001215 | 0.1091 |
| +1.5 | 53.39 | 55.82 | 44.55 | 53.00 | 57.92 | 0.001280 | 0.1091 |
| −1.0 | 50.61 | 51.55 | 40.91 | 51.96 | 55.89 | 0.001468 | 0.0727 |
判定:按 PLAN 预置判据(最高格 ≤ 对照 + 2 即放弃)放弃该轴,β_disp=0。
γ 精扫(补 A 半峰值定位):γ=−1.2 → 57.45(cov 53.50 / cell_state 62.37)、−1.4 → 57.54、−1.6 → 57.59、−1.8 → 57.48(cell_state 63.40 但 de_recovery 51.96)。−1.6 仍是峰值,且与 node28 的 −2.0/−2.4 更差一致,γ 族已收敛。
结论 / 供后续节点复用
- de_recovery 是离散计数:de_score = k/55(0.0727=4、0.0909=5、0.1091=6),组分 51.96 / 52.48 / 53.00 各差 0.52。它确实只对「抽样是否包含表达极端/高杠杆细胞」敏感(β_disp>0 就能把 k 从 5 抬到 6,且 β=0.5/1.0/1.5 完全饱和),对表达级扰动(γ、slope 阈值、PC 投影)免疫——证实了全树 10 个节点的观察并给出了机制。
- 但撬动它的代价远大于收益:k 每 +1 只值 +0.52 组分(≈ +0.13 board),而选中极端细胞使 variogram 变差、covariation 掉 1.6~8.1、cell_state 掉 0.35~7.3。covariation 与 cell_state 沿「样本离散度」轴都在 β_disp=0 附近取极大(两个方向都变差),说明 E8.5 快照的类型内协方差结构已经贴近 E9.5 真值,没有可挤的余量。
- γ 与 β_disp 是同一条 trade-off 曲线的两端:γ 更负 → cell_state↑ / covariation↓ / de_recovery↓;β_disp 更正 → de_recovery↑ / covariation↓↓。board 峰值就在父配置(γ=−1.6, sm=0.1, β=−4/−1, β_disp=0)。
- 未试但由 (2) 推出可能有微弱正收益的方向:把「抽样组成」与「表达离散度」解耦——用 β_disp≈1 选细胞、再把这些细胞的类型内残差按 ρ<1 收缩回去(ρ∈{0.9,0.95}),试图保住 de_recovery 的 k=6 同时把 variogram 拉回对照值。上限 ≈ +0.1 board(< T1 噪声 2 分),且 node15 已测过纯收缩会伤 covariation,故本节点未花查分额度。
验证过
- 三视图跑通:proxy(5118 细胞)、proxy2(外部 Qiu E9.0 被
pick_stage忽略,输出与 proxy 逐元素相同,max|ΔX|=0)、X3(退化保护生效,输出 = copy_last);vec-check三个视图均 ok。 - 确定性:只用
np.random.default_rng(seed);β_disp=0 时新分支完全不执行,输出与父 node28 的代码路径一致。 - 资源:proxy 全量 ~5 s、峰值内存 1.6 GB(β_disp≠0 时 2.4 GB),远低于 limits(28 GB / 30 min)。
- 查分用量:9/20(对照 1 + β_disp 四格 4 + γ 三点 3 + 无额外确认查分)。
没验证 / 风险
- 发布配置与 node18/21/23/28 相同,A 半 57.59 复现父记录(逐值一致),因此本节点不产生新分数,只产生负面结果与机制解释(B 半预期 ≈ 54.74)。
- β_disp 网格只在 proxy A 半 seed0 上测;因四格全部低于对照且幅度远超噪声(−0.3 到 −7.0),未做第二 seed 复核。
- V(top-V 基因数)未扫(PLAN 规定只在 β_disp 有信号时扫),β_disp 无信号故 V=1500 单点。
- final 视图(E8.5+E9.5→E10.5)未测;若 E9.5 池 ≤ 目标细胞数会触发退化保护,退回 copy_last。
调研员的计划
| 名称 | 类型内表达离散度第三级抽样权重(covariation+de_recovery 定向) |
|---|---|
| 动机 | 父节点28四组里 de_recovery=51.69 最弱,且在全树至少 9 个节点(7/12/13/15/18/21/22/23/24/26)中恒为 51.69、对一切表达级扰动(γ、slope、top-N、PC投影)完全免疫——说明 de_recovery 只能靠抽样组成/选细胞来撬动。次弱是 covariation=52.42,随 γ 增大还小幅下滑(52.88→52.42)。现有两级权重只用增殖(β_type=-4、β_cell=-1),类型内选哪些细胞完全由增殖随机性决定,导致下抽样丢失类型内协方差结构与表达极端细胞。已试过的分化轴全部无效/有害(node19 Reactome 轴 β_diff=0 打平、node20 DPT/n_genes 轴两向均降分、node26 top-N 斜率 wash),故不重复任何分化/成熟度轴。本方案用对称(绝对值)的类型内离散度做第三级抽样权重:既保均值(不偏组成、低风险伤 cell_state)、又保类型内协方差与极端态(同时定向 covariation 与 de_recovery),是与全树已试机制不同的新族。 |
| 做法 | 机制:在 β_type、β_cell 之外加第三级权重 w_disp。每类型(≥MIN_TYPE_CELLS=20 细胞)内:1) 取该类型内 top-V 高变基因(按类型内方差选,V=1500 初值,范围 800–3000),避免看家基因噪声;2) 对每个源细胞算类型内残差 z 分数平方均值 d_i = mean_g z_g(i)^2(即该细胞偏离类型质心的幅度/杠杆度),跨细胞标准化为 z_disp;3) w_disp_i = max(1e-6, 1 + β_disp·z_disp_i);总权重 w = w_type·w_cell·w_disp,沿用全局 Efraimidis–Spirakis 无放回抽 n_target。对称绝对值使两尾同权→类型比例基本不变(保 cell_state/direction 均值),但偏向保留高杠杆/表达极端细胞,提升下抽样样本的类型内协方差保真度(covariation)并富集已部分激活 E9.5 DE 程序的稀有亚态(de_recovery)。关键参数初值与扫描:β_disp ∈ {+0.5, +1.0, +1.5} 为主,另加 {-1.0} 一点确认符号方向(预期 + 向更优);V ∈ {1500} 单点起步,仅当 β_disp 有信号再扫 {800,3000}。保护开关沿用父逻辑:外部来源阶段或抽样退化(n_target≥池大小,如 X3)时 β_disp=0 跳过。单输入阶段退路:本机制只用单快照的类型内统计,天然适用于 proxy(唯一官方输入阶段);final(E8.5+E9.5 两阶段)取最新官方阶段同样可用,无第二时间点依赖。vec-score 快速筛选:第一步只在 proxy A-half seed0 跑 β_disp 四格(含对照 β_disp=0),用 5 次查分内定夺;若最高格相对对照 ≤ +2(T1 噪声)则立即判定 de_recovery 不可由抽样撬动、回退发布父配置,不再追加查分(吸取 node28 用 14/20 确认零效应的教训);若 >+2 且四组中 covariation 或 de_recovery 同向改善,用第 2 个 seed 复核该格后再定发布参数。预期节点分:若成功主要来自 covariation/de_recovery 各组 +1~3,板分目标 >54.74。 |
| 风险 | 1) de_recovery 可能根本上不可动:若 E9.5 的 DE 程序对应 E8.5 池中根本不存在的细胞态,任何抽样都救不回——Engineer 判据:β_disp 四格在 A-half 全部 ≤ 对照+2 即放弃并回退父配置,勿耗查分。2) 高杠杆细胞富集可能轻微拉偏均值伤 cell_state:用对称绝对值权重+小 β_disp 缓解,并监控 cell_state 若掉 >2 分即减小 β_disp 或回退。3) top-V 高变基因选择引入新自由度,V 过大会混入噪声:V 固定 1500 起步、仅在 β_disp 有信号时才扫。4) 提升幅度可能 < T1 噪声(~2分):务必用第 2 个 seed 复核,差距 <2 不算进步。5) 过早烧查分:第一步四格定夺,禁止在多组上重复确认零效应。 |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:父节点版本 ddc49f6e3a。改动的文件:solution/METHOD.md +32 −26、solution/run.py +106 −23
diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 5732847..143df7d 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,44 +1,50 @@-# copy_last 两级增殖重加权(β_type=-4、β_cell=-1)+ 增殖-表达斜率逐细胞放大(γ=-1.6、|slope|>0.1);PLAN 的确定性类型配额与类型自适应 γ 实测无效/有害,均默认关闭。+# copy_last 两级增殖重加权(β_type=-4、β_cell=-1)+ 增殖-表达斜率放大(γ=-1.6、|slope|>0.1);新增的类型内离散度第三级抽样权重实测净负(de_recovery +0.52 / covariation −1.63),默认关闭。 -## 方法+## 方法(发布配置) -基底(继承 node 4/6/7,已验证):输出 = 最新「官方」输入阶段的 Efraimidis–Spirakis 加权无放回抽样,权重 = w_type × w_cell:+基底(继承 node 4/6/7/18/28,已验证):输出 = 最新「官方」输入阶段的 Efraimidis–Spirakis 加权无放回抽样,权重 = w_type × w_cell: - w_type = max(1e-6, 1 + β_type·(prolif_type − mean_prolif)),β_type = **-4**(node 4); - w_cell = max(1e-6, 1 + β_cell·(prolif_cell − prolif_type)),β_cell = **-1**(node 6)。 -prolif 得分 = PROLIF 通用 cell-cycle 基因(Mki67, Top2a, Pcna, Cdk1, Ccna2, Ccnb1, Aurkb, Bub1, Cenpf, Nusc1,非阶段特异知识,来源:通用细胞周期标记注释)log 表达逐细胞均值,全部在输入快照内计算,单输入阶段可用。+prolif 得分 = 通用 cell-cycle 基因(Mki67, Top2a, Pcna, Cdk1, Ccna2, Ccnb1, Aurkb, Bub1, Cenpf, Nusc1;来源:通用细胞周期标记注释,非阶段特异)log 表达逐细胞均值,全部在输入快照内现场计算,单输入阶段可用。 -表达调整(继承 node 7 机制、node 18 参数):每类型(≥20 细胞)内对全部源细胞做 OLS slope_g = cov(x_g, prolif)/var(prolif),保留 |slope_g| > **0.1**;被选细胞 x_adj = x + γ·slope_g·(prolif_i − type_mean),γ = **-1.6**(放大类型内增殖-表达耦合),clip ≥0,1024 行分块。保护开关:所选阶段为外部来源、或抽样退化(n_target ≥ 池大小,如 X3)时跳过表达调整。+表达调整(node 7 机制、node 18 参数):每类型(≥20 细胞)内对全部源细胞做 OLS slope_g = cov(x_g, prolif)/var(prolif),保留 |slope_g| > **0.1**;被选细胞 x_adj = x − γ·slope_g·(prolif_i − type_mean),γ = **-1.6**,clip ≥0,1024 行分块。保护开关:所选阶段为外部来源、或抽样退化(n_target ≥ 池大小,如 X3)时跳过表达调整与 w_disp。 -## 本节点实测(PLAN 三机制,proxy A 半 seed0,对照 = γ-1.6/sm0.1 随机抽样 = 57.59)+## 本节点新机制:类型内表达离散度第三级权重(PLAN,已实现,默认关闭) -| 机制 | 结果 | 判定 |-|---|---|---|-| 确定性类型配额(最大余数法 + 类型内 E-S) | **55.60**(covariation 53.0→49.1、cell_state 63.2→58.7、direction +0.5) | **有害,回退随机 E-S**(PLAN 风险 1 兑现:全局 E-S 竞争的配额"噪声"本身有益) |-| 类型自适应 γ:γ_t = γ·clip((std_t/median_std)^α, 0.25, 4) | α=0.5/1.0/2.0 → 57.62/57.65/57.56(α=2 的 de_recovery 52.5→52.0) | 全部在 ±0.06 噪声内,无效,α=0 |-| 小类型斜率收缩 slope·n_t/(n_t+k)(n_t<50) | k=20/50 → 57.59(与对照逐字节相同:proxy 池无 20≤n_t<50 的类型) | 无效果,k=0 |+每类型(≥MIN_TYPE_CELLS=20)内按**类型内方差**取 top-V(V=1500)高变基因,对每个源细胞算 d_i = mean_g ((x_ig − μ_tg)/σ_tg)²(偏离类型质心的杠杆度),跨细胞标准化成 z_disp,w = w_type·w_cell·max(1e-6, 1 + β_disp·z_disp)(另有 `--disp-mode exp` 的 exp(β·z) 变体)。对称绝对值权重 → 类型比例基本不变。实现用 one-hot 稀疏矩乘一次算完所有类型的均值/平方均值(O(nnz)),全池 ~3 s。 -附加扫描:γ=-2.0 → 57.47、γ=-2.4 → 57.31(更负无益,-1.6 已是 A 半峰值区);sm=0.15 → 57.60(与 0.1 打平);β_cell=-2 → 55.28(大幅变差,维持 -1)。+### proxy A 半 seed0 实测(对照 = 父配置 β_disp=0 → **57.59**) -## 发布配置与查分记录(A 半)+| β_disp | board | cell_state | covariation | de_recovery | direction | variogram | de_score |+|---|---|---|---|---|---|---|---|+| 0(对照/发布) | **57.59** | 63.15 | 53.02 | 52.48 | 59.69 | 0.000943 | 0.0909 |+| +0.5 | 57.29 | 62.80 | 51.39 | **53.00** | 59.67 | 0.000999 | 0.1091 |+| +1.0 | 54.63 | 58.40 | 45.97 | 53.00 | 58.68 | 0.001215 | 0.1091 |+| +1.5 | 53.39 | 55.82 | 44.55 | 53.00 | 57.92 | 0.001280 | 0.1091 |+| −1.0 | 50.61 | 51.55 | 40.91 | 51.96 | 55.89 | 0.001468 | 0.0727 | -发布默认 = 对照配置(quota=0、α=0、k=0、γ=-1.6、sm=0.1、β=-4/-1),默认参数输出与查分文件逐字节一致(已核)。+判定:按 PLAN 预置判据(最高格 ≤ 对照 + 2 即放弃)**放弃该轴**,β_disp=0。 -- proxy seed0:**57.59**(cell_state 63.15 / covariation 53.02 / de_recovery 52.48 / direction 59.69)-- proxy seed1:57.31(一致,父配置 γ=-0.8/sm=0.05 的 A 半为 57.20/56.89)-- proxy2 seed0:57.59(外部 Qiu E9.0 被 pick_stage 忽略,输出同 proxy)-- X3 seed0:50.0(退化保护生效,= copy_last)-- 预期节点分 ≈ (57.6+57.6+50)/3 ≈ 55.1 A 半(父 54.54 板分 / node18 同配置 54.74 板分)+γ 精扫(补 A 半峰值定位):γ=−1.2 → 57.45(cov 53.50 / cell_state 62.37)、−1.4 → 57.54、−1.6 → **57.59**、−1.8 → 57.48(cell_state 63.40 但 de_recovery 51.96)。−1.6 仍是峰值,且与 node28 的 −2.0/−2.4 更差一致,γ 族已收敛。++## 结论 / 供后续节点复用++1. **de_recovery 是离散计数**:de_score = k/55(0.0727=4、0.0909=5、0.1091=6),组分 51.96 / 52.48 / 53.00 各差 0.52。它确实**只**对「抽样是否包含表达极端/高杠杆细胞」敏感(β_disp>0 就能把 k 从 5 抬到 6,且 β=0.5/1.0/1.5 完全饱和),对表达级扰动(γ、slope 阈值、PC 投影)免疫——证实了全树 10 个节点的观察并给出了机制。+2. **但撬动它的代价远大于收益**:k 每 +1 只值 +0.52 组分(≈ +0.13 board),而选中极端细胞使 variogram 变差、covariation 掉 1.6~8.1、cell_state 掉 0.35~7.3。covariation 与 cell_state 沿「样本离散度」轴都在 β_disp=0 附近取极大(两个方向都变差),说明 E8.5 快照的类型内协方差结构已经贴近 E9.5 真值,**没有可挤的余量**。+3. γ 与 β_disp 是同一条 trade-off 曲线的两端:γ 更负 → cell_state↑ / covariation↓ / de_recovery↓;β_disp 更正 → de_recovery↑ / covariation↓↓。board 峰值就在父配置(γ=−1.6, sm=0.1, β=−4/−1, β_disp=0)。+4. 未试但由 (2) 推出可能有微弱正收益的方向:把「抽样组成」与「表达离散度」解耦——用 β_disp≈1 选细胞、再把这些细胞的类型内残差按 ρ<1 收缩回去(ρ∈{0.9,0.95}),试图保住 de_recovery 的 k=6 同时把 variogram 拉回对照值。上限 ≈ +0.1 board(< T1 噪声 2 分),且 node15 已测过纯收缩会伤 covariation,故本节点未花查分额度。 ## 验证过 -- 三视图(proxy/proxy2/X3)跑通、vec-check ok、运行 ~5 s、内存远低于限额。-- 确定性:仅 np.random.default_rng(seed)。-- 配额分配器单独验证过总和 = n_target(在 quota=1 分支,虽默认关闭)。+- 三视图跑通:proxy(5118 细胞)、proxy2(外部 Qiu E9.0 被 `pick_stage` 忽略,输出与 proxy 逐元素相同,max|ΔX|=0)、X3(退化保护生效,输出 = copy_last);`vec-check` 三个视图均 ok。+- 确定性:只用 `np.random.default_rng(seed)`;β_disp=0 时新分支完全不执行,输出与父 node28 的代码路径一致。+- 资源:proxy 全量 ~5 s、峰值内存 1.6 GB(β_disp≠0 时 2.4 GB),远低于 limits(28 GB / 30 min)。+- 查分用量:9/20(对照 1 + β_disp 四格 4 + γ 三点 3 + 无额外确认查分)。 ## 没验证 / 风险 -- 发布配置与 node 18/21/23 机制相同(γ=-1.6、sm=0.1),A 半 57.59 vs 其记录的 A 半未见更高点;对父 node 7 的提升(+0.39 A 半)在 T1 噪声(~2 分)内,但方向与 node18 的 B 半结果(54.74 > 54.54)一致。-- 确定性配额在 **B 半**或其他视图是否同样有害未验证(A 半差距 -2.0,接近噪声边缘但四组指标同向恶化,判定可信度较高)。-- final 视图(E8.5+E9.5→E10.5)未测:若 E9.5 池 ≤ 目标细胞数则触发退化保护,退回 copy_last(安全无增益)。-- 类型自适应 γ 的无效可能因 proxy 池各类型 prolif 离散度差异小(median 归一后 ratio≈1);在类型异质性更大的视图上未测。+- 发布配置与 node18/21/23/28 相同,A 半 57.59 复现父记录(逐值一致),因此本节点**不产生新分数**,只产生负面结果与机制解释(B 半预期 ≈ 54.74)。+- β_disp 网格只在 proxy A 半 seed0 上测;因四格全部低于对照且幅度远超噪声(−0.3 到 −7.0),未做第二 seed 复核。+- V(top-V 基因数)未扫(PLAN 规定只在 β_disp 有信号时扫),β_disp 无信号故 V=1500 单点。+- final 视图(E8.5+E9.5→E10.5)未测;若 E9.5 池 ≤ 目标细胞数会触发退化保护,退回 copy_last。diff --git a/solution/run.py b/solution/run.pyindex 6ac89a6..8c90e6c 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,30 +1,38 @@ #!/usr/bin/env python3-"""copy_last + two-level proliferation reweighting + type-adaptive-proliferation-gradient expression extrapolation.+"""copy_last + two-level proliferation reweighting + within-type expression+dispersion as a third sampling weight + type-adaptive proliferation-gradient+expression extrapolation. -Base (node 4/6/18): output a weighted subsample of the latest official input-stage. Type-level weight uses the type-mean proliferation score+Base (node 4/6/18/28): output a weighted subsample of the latest official+input stage. Type-level weight uses the type-mean proliferation score (beta_type=-4); cell-level weight uses the within-type deviation (beta_cell=-1). Expression adjustment: per cell type, regress each gene on the cell's proliferation score (OLS over all source cells of that type) and shift each selected cell along the slope (gamma=-1.6, |slope|>0.1). -This node tested the PLAN's three mechanisms on the proxy A-half and kept-the winner; the shipped default is the strongest slope configuration-(gamma=-1.6, |slope|>0.1, node 18 parameters) with the parent's global-Efraimidis-Spirakis sampling:-1. Deterministic type quotas (QUOTA flag): REJECTED - proxy A-half 55.60 vs- 57.59 for random E-S (covariation -3.9, cell_state -4.4). The E-S quota- "noise" is beneficial; code kept, default off.-2. Type-adaptive gamma, gamma_t = gamma*clip((std_t/median_std)^alpha,- 0.25, 4) (ALPHA flag): NO EFFECT - alpha in {0.5,1,2} within +-0.06 of- alpha=0 (57.59); alpha=2 slightly hurts de_recovery. Default 0.-3. Small-type slope shrinkage slope *= n_t/(n_t+k) for n_t<50 (KSHRINK- flag): NO EFFECT - output byte-identical for k in {20,50}. Default 0.-Also scanned: gamma=-2.0/-2.4 worse (57.47/57.31), |slope|>0.15 ties-(57.60), beta_cell=-2 much worse (55.28).-Guards unchanged: skip expression adjustment on external stages or when-sampling degenerates (n_target >= pool size, e.g. test X3).+This node adds the PLAN's third-level sampling weight w_disp: within each+cell type (>= MIN_TYPE_CELLS cells) take the top-V most variable genes+(within-type variance), compute each cell's mean squared residual z-score+d_i over those genes (how far the cell sits from its type centroid),+standardise d across cells into z_disp, and use+w_disp_i = max(1e-6, 1 + beta_disp * z_disp_i) (mode "lin") or+exp(beta_disp * z_disp_i) (mode "exp"). The weight is symmetric in |z| so+type proportions are essentially preserved while the subsample keeps more+high-leverage / extreme cells, aiming at covariation and de_recovery.++Measured on the proxy A-half (seed 0, control = beta_disp 0 -> 57.59):+beta_disp +0.5 -> 57.29, +1.0 -> 54.63, +1.5 -> 53.39, -1.0 -> 50.61.+REJECTED by the PLAN's own criterion; default beta_disp=0 keeps the parent+(node 28) output byte-identical. Mechanistic finding: de_recovery is a+discrete count (de_score = k/55) that only reacts to which cells are+sampled - beta_disp>0 lifts k from 5 to 6 and saturates - but every gain+(+0.52 group points) costs far more covariation (-1.6 to -8.1) and+cell_state (-0.35 to -7.3); both directions of the dispersion axis move the+variogram away from the control, i.e. the E8.5 snapshot's within-type+covariance already sits near the E9.5 truth. Also re-scanned gamma:+-1.2 -> 57.45, -1.4 -> 57.54, -1.6 -> 57.59 (peak), -1.8 -> 57.48.+Guards: w_disp and the expression adjustment are skipped on external stages+or when sampling degenerates (n_target >= pool size, e.g. test X3). """ from __future__ import annotations@@ -54,6 +62,9 @@ MIN_TYPE_CELLS = 20 QUOTA = 0 # 1: deterministic type quotas; 0: parent's global E-S ALPHA = 0.0 # type-adaptive gamma exponent (0 = global gamma) KSHRINK = 0.0 # small-type slope shrinkage constant (0 = off)+BETA_DISP = 0.0 # within-type dispersion third-level weight (0 = off)+DISP_V = 1500 # top-V within-type variable genes for the dispersion score+DISP_MODE = "lin" # "lin": max(1e-6, 1+b*z) ; "exp": exp(b*z) PROLIF = ["Mki67", "Top2a", "Pcna", "Cdk1", "Ccna2", "Ccnb1", "Aurkb", "Bub1", "Cenpf", "Nusc1"]@@ -103,6 +114,61 @@ def quota_allocate(mass: np.ndarray, counts: np.ndarray, n_target: int) -> np.nd return alloc +def dispersion_z(X, inv: np.ndarray, n_types: int, counts: np.ndarray,+ v: int, min_cells: int) -> np.ndarray:+ """Per-cell within-type dispersion z-score (0 for cells in small types).++ d_i = mean over the top-v within-type variable genes of+ ((x_ig - mean_tg) / std_tg)^2+ """+ n_cells, n_genes = X.shape+ Xcsr = X.tocsr()+ onehot = sparse.csr_matrix(+ (np.ones(n_cells, dtype=np.float64), (np.arange(n_cells), inv)),+ shape=(n_cells, n_types),+ )+ sums = np.asarray((onehot.T @ Xcsr).todense(), dtype=np.float64) # types x genes+ sq = Xcsr.copy()+ sq.data = sq.data.astype(np.float64) ** 2+ sumsq = np.asarray((onehot.T @ sq).todense(), dtype=np.float64)+ cnt = np.maximum(counts.astype(np.float64), 1.0)[:, None]+ mean = sums / cnt+ var = np.maximum(sumsq / cnt - mean ** 2, 0.0)++ d = np.zeros(n_cells, dtype=np.float64)+ v = int(min(max(v, 1), n_genes))+ for t in range(n_types):+ idx = np.flatnonzero(inv == t)+ if len(idx) < min_cells:+ continue+ vt = var[t]+ top = np.argpartition(-vt, v - 1)[:v]+ sd = np.sqrt(vt[top])+ keep = sd > 1e-8+ top = top[keep]+ sd = sd[keep]+ if len(top) == 0:+ continue+ mu = mean[t, top]+ inv_sd = 1.0 / sd+ chunk = 1024+ for a in range(0, len(idx), chunk):+ sub = idx[a:a + chunk]+ D = np.asarray(Xcsr[sub][:, top].todense(), dtype=np.float64)+ Z = (D - mu) * inv_sd+ d[sub] = np.einsum("ij,ij->i", Z, Z) / len(top)+ valid = d > 0+ if valid.sum() < 10:+ return np.zeros(n_cells, dtype=np.float64)+ m = d[valid].mean()+ s = d[valid].std()+ if not np.isfinite(s) or s <= 1e-12:+ return np.zeros(n_cells, dtype=np.float64)+ z = (d - m) / s+ z[~valid] = 0.0+ return z++ def main() -> None: parser = argparse.ArgumentParser() parser.add_argument("--data", required=True)@@ -115,6 +181,9 @@ def main() -> None: parser.add_argument("--quota", type=int, default=None) parser.add_argument("--alpha", type=float, default=None) parser.add_argument("--kshrink", type=float, default=None)+ parser.add_argument("--beta-disp", type=float, default=None)+ parser.add_argument("--disp-v", type=int, default=None)+ parser.add_argument("--disp-mode", default=None) args = parser.parse_args() gamma = GAMMA if args.gamma is None else args.gamma slope_min = SLOPE_MIN if args.slope_min is None else args.slope_min@@ -123,6 +192,9 @@ def main() -> None: use_quota = QUOTA if args.quota is None else args.quota alpha = ALPHA if args.alpha is None else args.alpha kshrink = KSHRINK if args.kshrink is None else args.kshrink+ beta_disp = BETA_DISP if args.beta_disp is None else args.beta_disp+ disp_v = DISP_V if args.disp_v is None else args.disp_v+ disp_mode = DISP_MODE if args.disp_mode is None else args.disp_mode manifest = load_manifest(args.data) genes = panel_genes(args.data, manifest)@@ -141,6 +213,7 @@ def main() -> None: inv = None type_score = None+ degenerate = n_target >= last.n_obs if score is not None: lab = labels_of(last) uniq, inv = np.unique(lab, return_inverse=True)@@ -156,18 +229,29 @@ def main() -> None: w_type_t = np.maximum(1.0 + beta_type * (type_score - mean_s), 1e-6) w_cell = np.maximum(1.0 + beta_cell * (score - ts), 1e-6) if beta_cell != 0 else np.ones_like(w_type) + if beta_disp != 0.0 and not is_external(entry) and not degenerate:+ z_disp = dispersion_z(last.X, inv, n_types, counts, disp_v, MIN_TYPE_CELLS)+ if disp_mode == "exp":+ w_disp = np.clip(np.exp(np.clip(beta_disp * z_disp, -20.0, 20.0)), 1e-6, 1e6)+ else:+ w_disp = np.maximum(1.0 + beta_disp * z_disp, 1e-6)+ else:+ w_disp = None+ if use_quota and n_target < last.n_obs: alloc = quota_allocate(counts.astype(np.float64) * w_type_t, counts, n_target) parts = [] for t in range(n_types): idx = np.flatnonzero(inv == t)- sel = weighted_sample(w_cell[idx], int(alloc[t]), rng)+ wc = w_cell[idx] if w_disp is None else w_cell[idx] * w_disp[idx]+ sel = weighted_sample(wc, int(alloc[t]), rng) parts.append(idx[sel]) rows = np.sort(np.concatenate(parts)) if parts else np.zeros(0, dtype=np.int64) elif use_quota: rows = sample_rows(last.n_obs, n_target, rng) else:- rows = weighted_sample(w_type * w_cell, n_target, rng)+ w_all = w_type * w_cell if w_disp is None else w_type * w_cell * w_disp+ rows = weighted_sample(w_all, n_target, rng) rows = np.sort(rows) else: rows = sample_rows(last.n_obs, n_target, rng)@@ -175,7 +259,6 @@ def main() -> None: X = last.X Xsel = X[rows] # Guard: only adjust expression on official, non-degenerate views (X3).- degenerate = n_target >= last.n_obs if gamma != 0.0 and score is not None and not is_external(entry) and not degenerate: n_types = int(inv.max()) + 1 counts = np.bincount(inv, minlength=n_types).astype(np.float64)
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
用到的知识库条目
| 编号 | 标题 | 出处 |
|---|---|---|
| k031 | Offline OT toolkit in the sandbox: moscot TemporalProblem, wot OTModel, POT, geomloss | 10.1038/s41586-024-08453-2 (moscot); 10.1016/j.cell.2019.01.006 (Waddington-OT) |
| k041 | Within-stage pseudotime and graph toolkit offline: scanpy DPT/PAGA/Leiden, Palantir, CellRank 2 | 10.1186/s13059-019-1663-x (PAGA); 10.1038/s41587-019-0068-4 (Palantir); 10.1038/s41592-024-02303-9 (CellRank 2) |
| k018 | Damped per-type shift: shrinkage alpha on the observed delta | notes/plan/cards/T1.md |
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
| 改了什么 | 实现了 PLAN 的类型内离散度第三级抽样权重 w_disp(每类型 top-1500 高变基因的残差 z² 杠杆度,对称、lin/exp 两模式),并补扫 γ∈{-1.2,-1.4,-1.8};proxy A 半实测全部劣于对照后按预置判据放弃,发布默认 β_disp=0,输出与父 node28 逐元素相同,故榜分 54.74 不变(+0.00)。 |
|---|---|
| 各组分数的变化 | cell_state:噪声内(+0.00,57.60) covariation:噪声内(+0.00,52.42);A 半实验显示 β_disp≠0 会使其掉 1.6~8.1 de_recovery:噪声内(+0.00,51.69);但 A 半内部数据显示 β_disp>0 可把 de_score 的 k 从 5 抬到 6(组分 +0.52),代价过大未发布 direction:噪声内(+0.00,56.23) |
| 假设是否成立 | 否 |
| 经验 |
|
| 下一步建议 |
|
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。
分析员
| 角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。 | 分析员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数?大模型一共回复了几轮(每轮可以调用多个工具)。 | 3 |
| 工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。 | 共 2 次:write 1、bash 1 |
| 用时 | 不到 1 分 |
| token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。 | 输入 11,796 · 输出 1,163 · 思考 104 |
| 任务(第一行) | 分析节点 n29 的打分结果,写 ANALYSIS.json |
| 最后的回答(摘录) | ANALYSIS.json written |
| 原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。 | /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/analyst.jsonl 7 KB /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/analyst.stderr |
工程师
| 角色 | 工程师 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 15 |
| 工具调用 | 共 18 次:bash 12、read 2、write 2、edit 2 |
| 用时 | 11 分 |
| token 数 | 输入 48,519 · 输出 10,127 · 思考 10,159 |
| 任务(第一行) | 按计划修改程序:类型内表达离散度第三级抽样权重(covariation+de_recovery 定向) |
| 最后的回答(摘录) | **做了什么**:在父(node28)基底上实现 PLAN 的第三级抽样权重 w_disp(类型内 top-1500 高变基因的残差 z² 杠杆度,对称、E-S 无放回,one-hot 稀疏矩乘 O(nnz)),扫 β_disp ∈ {+0.5,+1.0,+1.5,−1.0},另补 γ 精扫 {−1.2,−1.4,−1.8}。按预置判据判负,发布配置 β_disp=0(默认参数输出与已查分的 57.59 文件逐元素相同,三视图 vec-check ok,~5 s / 1.6 GB)。查分用 9/20。
**结果(proxy A 半 seed0)**:对照 57.59;β_disp=+0.5 → 57.29、+1.0 → 54.63、+1.5 → 53.39、−1.0 → 50.61;γ=−1.6 仍是峰值。关键发现:de_recovery 是离散计数(de_score=k/55,k=4/5/6 → 51.96/52.48/53.00),只对「抽到哪些细胞」敏感、β_disp>0 即把 k 从 5 抬到 6 并饱和,但每 +0.52 组分要付 covariation −1.6~−8.1、cell_state −0.35~−7.3;variogram 在 β_disp=0 两侧都变差,说明 E8.5 快照的类型内协方差已贴近 E9.5 真值,没有余量。
**下一步最值得试**:解耦「抽样组成」与「表达离散度」——用 β_disp≈1 选细胞、再把选中细胞的类型内残差按 ρ∈{0.9,0.95} 收缩,试图保住 k=6 同时把 variogram 拉回对照(上限 ≈ +0.1 board,低于噪声,性价比低)。更实际的是换族:γ 与 β_disp 是同一条 trade-off 曲线两端,board 已在 (γ=−1.6, sm=0.1, β=−4/−1) 收敛,继续在「单快照抽样+斜率放大」族内调参收益为零。 |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/engineer.jsonl 102 KB /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/engineer.stderr |
调研员
| 角色 | 调研员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 1 |
| 工具调用 | 共 0 次:— |
| 用时 | 6 分 |
| token 数 | 输入 16,108 · 输出 6,357 · 思考 8,192 |
| 任务(第一行) | 为节点 n29(改进,父节点 n28)写改进计划 PLAN.json |
| 最后的回答(摘录) | json
{"name": "类型内表达离散度第三级抽样权重(covariation+de_recovery 定向)",
"motivation": "父节点28四组里 de_recovery=51.69 最弱,且在全树至少 9 个节点(7/12/13/15/18/21/22/23/24/26)中恒为 51.69、对一切表达级扰动(γ、slope、top-N、PC投影)完全免疫——说明 de_recovery 只能靠抽样组成/选细胞来撬动。次弱是 covariation=52.42,随 γ 增大还小幅下滑(52.88→52.42)。现有两级权重只用增殖(β_type=-4、β_cell=-1),类型内选哪些细胞完全由增殖随机性决定,导致下抽样丢失类型内协方差结构与表达极端细胞。已试过的分化轴全部无效/有害(node19 Reactome 轴 β_diff=0 打平、node20 DPT/n_genes 轴两向均降分、node26 top-N 斜率 wash),故不重复任何分化/成熟度轴。本方案用对称(绝对值)的类型内离散度做第三级抽样权重:既保均值(不偏组成、低风险伤 cell_state)、又保类型内协方差与极端态(同时定向 covariation 与 de_recovery),是与全树已试机制不同的新族。",
"approach": "机制:在 β_type、β_cell 之外加第三级权重 w_disp。每类型(≥MIN_TYPE_CELLS=20 细胞)内:1) 取该类型内 top-V 高变基因(按类型内方差选,V=1500 初值,范围 800–3000),避免看家基因噪声;2) 对每个源细胞算类型内残差 z 分数平方均值 d_i = mean_g z_g(i)^2(即该细胞偏离类型质心的幅度/杠杆度),跨细胞标准化为 z_disp;3) w_disp_i = max(1e-6, 1 + β_disp·z_disp_i);总权重 w = w_type·w_cell·w_disp,沿用全局 Efraimidis–Spirakis 无放回抽 n_target。对称绝对值使两尾同权→类型比例基本不变(保 cell_state/direction 均值),但偏向保留高杠杆/表达极端细胞,提升下抽样样本的类型内协方差保真度(covariation)并富集已部分激活 E9.5 DE 程序的稀有亚态(de_recovery)。关键参数初值与扫描:β_disp ∈ {+0.5, +1.0, +1.5} 为主,另加 {-1.0} 一点确认符号方向(预期 + 向更优);V ∈ {1500} 单点起步,仅当 β_disp 有信号再扫 {800,3000}。保护开关沿用父逻辑:外部来源阶段或抽样退化(n_target≥池大小,如 X3)时 β_disp=0 跳过。单输入阶段退路:本机制只用单快照的类型内统计,天然适用于 proxy(唯一官方输入阶段);final(E8.5+E9.5 两阶段)取最新官方阶段同样可用,无第二时间点依赖。vec-score 快速筛选:第一步只在 proxy A-half seed0 跑 β_disp 四格(含对照 β_disp=0),用 5 次查分内定夺;若最高格相对对照 ≤ +2(T1 噪声)则立即判定 de_recovery 不可由抽样撬动、回退发布父配置,不再追加查分(吸取 node28 用 14/20 确认零效应的教训);若 >+2 且四组中 covariation 或 de_recovery 同向改善,用第 2 个 seed 复核该格后再定发布参数。预期节点分:若成功主要来自 covariation/de_recovery 各组 +1~3,板分目标 >54.74。",
"expected_groups": ["covariation", "de_recovery"],
"risks": "1) de_recovery 可能根本上不可动:若 E9.5 的 DE 程序对应 E8.5 池中根本不存在的细胞态,任何抽样都救不回——Engineer 判据:β_disp 四格在 A-half 全部 ≤ 对照+2 即放弃并回退父配置,勿耗查分。2) 高杠杆细胞富集可能轻微拉偏均值伤 cell_state:用对称绝对值权重+小 β_disp 缓解,并监控 cell_state 若掉 >2 分即减小 β_disp 或回退。3) top-V 高变基因选择引入新自由度,V 过大会混入噪声:V 固定 1500 起步、仅在 β_disp 有信号时才扫。4) 提升幅度可能 < T1 噪声(~2分):务必用第 2 个 seed 复核,差距 <2 不算进步。5) 过早烧查分:第一步四格定夺,禁止在多组上重复确认零效应。"}
``` |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/researcher.jsonl 26 KB /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/researcher.stderr |
审查员
| 角色 | 审查员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 6 |
| 工具调用 | 共 9 次:bash 4、read 4、write 1 |
| 用时 | 1 分 |
| token 数 | 输入 21,186 · 输出 1,221 · 思考 1,261 |
| 任务(第一行) | 审查节点 n29 的程序是否越界读取、写死目标、钻评分器漏洞 |
| 最后的回答(摘录) | REVIEW.json written |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/reviewer.jsonl 93 KB /home/spark-longxinyang/vec/runs/formal/20261001-233756-search-t1-abc-r0-B-population/nodes/29/reviewer.stderr |