Virtual Embryo Challenge更新于 10-04 03:53(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-171955-search-t2-embryo-interp-chain-12h

节点 n53 终选程序?按该运行锁定的规则最终选出的程序;可能是候选节点,也可能由护栏回退到基线。在终选来历上

QREM:均匀 15-NN 平滑(β=0.35)后按秩回填原始非零值,精确恢复逐基因边际分布与伪批量,保住 nbhd 增益、消除 mmd_u 收缩损失;PLAN 的杂质门控自适应 β 经 6 配置证否。

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-171955-search-t2-embryo-interp-chain-12h
父节点n51
子节点n54、n58
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 66.48(+0.1) · proxy 66.48(+0.1) · 3 次复测均分 66.03
审查通过 检查1(越界读取):未发现问题——文件访问仅限 view 内路径:load_manifest/read_stage/panel_genes 走 --data 视图(run.py:3355-3365),prior 读取仅 view/prior/{reactome,go,msigdb,tf_regulons}(run.py:1020、1204-1211),均在 view_manifest prior 清单内;无绝对路径、..、/mnt、/home、data/raw、打分器路径,无联网(无 requests/urlopen),不读目标阶段文件(bracket 由 interp_bracket 从 …
用时?从运行开始到结束(或到现在)的挂钟时间。25 分
程序版本5ed9840601e413c786cdb18804c93e626ddd6ea1 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 5ed9840601:solution/METHOD.md

QREM:均匀 15-NN 平滑(β=0.35)后按秩回填原始非零值,精确恢复逐基因边际分布与伪批量,保住 nbhd 增益、消除 mmd_u 收缩损失;PLAN 的杂质门控自适应 β 经 6 配置证否。

节点 53:QREM(rank-quantile marginal restoration after neighborhood smoothing)

  • family_id: T2EI-01(表达 / cell_state + local_spatial)
  • 父节点:51(NBSMOOTH,均匀 β=0.35 的 15-NN 邻域平滑,A 半复现 65.70)
  • 提交机制:QREM;PLAN 指定机制(ADPNBSMOOTH 杂质门控自适应 β)已按 PLAN 实现并在 A 半证否(见下), 按任务书要求改交针对同一弱项(nbhd 增益与 mmd_u 损失的权衡)的备选机制。

方法

在父节点 NBSMOOTH 步骤(所有表达改动之后、坐标定型处;k=15 最近邻、排除自身、nnz-only)内部:

  1. 与父节点相同地计算平滑值 x' = (1−β)x + β·mean_15NN(x)(β=0.35 均匀,nnz-only:零位不动,支持集不变; 由 β<1 且 x'>0 当 x>0,支持集精确保持)。
  2. QREM(本节点新机制):对每个基因 g,把其非零位置上的平滑值按秩替换为平滑前该基因非零值的 同秩次序统计量(stable argsort 两次取秩 → 索引排序后的原值数组)。逐基因的非零值多重集因此逐位恢复 为平滑前的边际分布:池化伪批量、每基因均值/方差/分位数精确不变(不再需要 pb-neutral 的 g 复原,代码里 QREM 开启时跳过该步),而平滑造成的秩结构(哪些细胞对该基因偏高/偏低——现在与空间邻域相干)被保留。
  3. 坐标完全不动 → shape_scale 三项与父逐位相同;伪批量精确复原 → de_score / de_direction 逐位冻结 (A 半实测四组中 expression_change 62.23、shape_scale 78.23 在所有配置间逐位不变)。

原理:父节点 ANALYSIS 表明 nbhd 收益随 β 饱和(mean β≥0.27 后 nbhd raw 0.0451→0.0448 仅微降),而 mmd_u 损失来自逐细胞向邻域均值收缩导致的分布展宽丢失。QREM 把"空间相干"(秩)与"边际展宽"(值)解耦: 允许平滑改变细胞间的相对高低关系(nbhd 读取的邻域平均因此更相干),同时把每个基因的边际分布精确复位 (mmd_u/variogram 读取的单基因展宽不损失)。

关键参数(全部沿用父节点已验证值,QREM 无新自由参数):k=15、β=0.35、nnz-only、排除自身。 QREM 的 β 扫描(0.35/0.5/0.7/1.0)显示 0.35 最优:β≥0.7 时秩结构被过度平滑(邻域均值秩趋同), 回填后细胞值与自身局部环境失配,nbhd/mmd_u 同时崩坏(65.47 / 64.73)。

单输入回退:b=None → copy_last 提前返回(未触及);NBSMOOTH/QREM 块要求 ia.size 且 ib.size,单侧时整体跳过。 --ablate mechanism:关闭 QREM 与(默认已关的)自适应 β,保留父的均匀 NBSMOOTH β=0.35,逐位还原父节点 51 (已验证 seed 0 与 seed 3 两个种子 sha256 相同)。

PLAN 机制(ADPNBSMOOTH)的实现与证否

按 PLAN 实现:输出细胞的侧标签 = mix_indices 顺序(前 ia.size 个来自 a 侧,其余 b 侧;断言长度匹配), 最终坐标上 cKDTree 取 15-NN(排除自身),impurity_i = 1 − 同侧邻居占比,β_i = clip(β_max·impurity_i, β_min, β_max)。 诊断(PLAN mechanism_evidence):mean impurity = 0.181(高于 PLAN 风险 2 的 0.15 下限,机制有空间)、 中位 0.067;β_max=0.7 时 β_mean=0.127、p50=0.047、p90=0.373,2215/5000 个细胞(44.3%)β>0.05(机制生效于非平凡子集)。 但 A 半查分(seed 0,全部与同 seed 均匀 β=0.35 基线配对):

配置榜分nbhd rawmmd_u raw对照
均匀 β=0.35(=父 51,A 半基线)65.6960.045070.01034—
无平滑(=节点 49)65.4750.047410.01008−0.22
自适应 β_max=0.765.5510.046070.01038−0.15
自适应 β_max=1.165.4730.045420.01098−0.22
自适应 β_max=1.665.1390.045340.01240−0.56
自适应 β_max=0.7 + β_min=0.165.6100.045730.01033−0.09
自适应 β_max=0.9 + β_min=0.265.6510.045060.01049−0.05

结论(6 配置全劣于父):把平滑集中到混合邻域比摊平更伤 mmd_u(高 β 细胞变成分布中的"平均化"离群点), 而 nbhd 只随总平滑量饱和——PLAN 的"纯邻域过平滑导致 mmd_u 损失"假设不成立,β 异质性本身就是 mmd_u 的损失源。 β_min=0.2+β_max=0.9 已把 nbhd 追平均匀基线(0.04506 vs 0.04507)但 mmd_u 更差(0.01049 vs 0.01034)→ 被严格支配。

顺带按父建议 #1 完成均匀 k×β 细调(A 半 seed 0 配对):k10β0.35=65.693、k10β0.40=65.694、k20β0.35=65.687、 k20β0.40=65.692、k15β0.40=65.701、k15β0.45=65.692——全部落在 ±0.015(纯噪声),确认均匀族在 (k15, β0.35) 已饱和, k×β 方向无剩余空间(也说明必须换机制而非调参)。

查分记录(共 20 次,全部 T2:embryo:val_interp,A 半)

见上两表(7 + 6 + 无平滑基线 1 + 均匀基线 1),另 QREM β 扫描 4 次与种子验证 3 次:

配置seed榜分nbhd rawmmd_u rawvariogram raw
QREM β=0.35(提交)065.7990.044440.010300.007099
QREM β=0.5065.7800.044160.010500.007100
QREM β=0.7065.4720.045450.011000.007101
QREM β=1.0064.7330.048060.012610.007091
QREM β=0.35(提交)165.4390.044840.009990.007317
均匀 β=0.35(父)165.3320.045490.010060.007320
QREM β=0.35(提交)265.3190.044500.010440.007030

配对差(QREM − 均匀,同 seed 同锚点):s0 +0.103、s1 +0.107,两种子方向一致,且三个受影响指标 nbhd / mmd_u / variogram 的 raw 在两个种子上全部同向改善(6/6);expression_change 与 shape_scale 逐位冻结。 按父节点经验(A 半配对均值 +0.196 → 榜上 +0.24),预期正式收益 ~+0.1,量级小但机制风险低 (QREM 数学上不可能改变支持集、伪批量、逐基因边际;唯一可动的是 nbhd/mmd_u/variogram 读取的联合结构,实测三赢)。

验证过 / 未验证

  • 验证:seed 0 默认输出 == QREM β=0.35 筛选输出(sha256);--ablate mechanism 在 seed 0 与 seed 3 均逐位 还原均匀 β=0.35(父 51 输出);同 seed 重跑逐位确定;伪装视图(+1 天平移、manifest 键序打乱、 文件重命名 input_0/1.h5ad、换路径、重排版)输出与真实视图逐位相同(sha256 be25f360…); vec-check 通过;运行 ~3 s / ~0.6 GB,纯 CPU(EXECUTION.json {"gpu": false});QREM 循环 498 基因全部命中。
  • 未验证:B 半与真实括号(E6.75+E8.0 之外的括号宽度/共有类型数组合)上的收益——QREM 不含任何依赖视图、 绝对时间或阶段名的逻辑,参数全部现场从输入计算,预期可迁移,但没有实测;seed 2 的配对差(额度用尽, 仅单臂 65.319,nbhd 0.04450 与前两种子模式一致)。

知识来源

无生物学先验被写入程序:机制纯粹是统计/几何的(邻域平滑 + 秩-分位数边际恢复),只使用视图内数据。 未使用 prior/、external/(本视图 external 为空)、任何保留阶段/基因型的测量值或文献数值。

下一步建议

  1. QREM 与 β 的联合再扫(β∈{0.4, 0.45} + QREM):s0 上 β0.5 的 nbhd 更好(0.04416)但 mmd_u 略差,中间点未测。
  2. QREM 的秩源替换:用"邻域均值的秩"回填"原值"已隐含在 β=1 情形(崩),但 β=0.35–0.5 区间秩-值失配最小; 可试按类型分层的秩回填(型内边际恢复),检查 variogram。
  3. mmd_u 的剩余差距(nosmooth 0.01008 是上界参照)只能靠支持集/检出率通道(DETR 族)继续,不要再搜再居中/幅度族。

调研员的计划

名称ADPNBSMOOTH: heterogeneity-gated adaptive β for NBSMOOTH
动机Node 51 NBSMOOTH (β=0.35 uniform) gained nbhd 3/3 (+0.33 pts, local_spatial +1.30) but lost mmd_u −0.08 pts / cell_state −0.35 because uniform β over-averages already-coherent neighborhoods. Parent ANALYSIS suggestion #2: selective smoothing only where neighborhood expression variance is high. mmd_u raw 0.00982 (skill 0.648) and variogram 0.00706 (skill 0.564) are the two weakest above-floor metrics; recovering the mmd_u loss while keeping nbhd gain targets cell_state (60.61, weakest group).
做法Replace uniform β=0.35 with per-cell adaptive β_i = β_max × impurity_i, where impurity_i = 1 − (fraction of cell i's k-NN that share the same bracket-side origin). Implementation: (1) After mix_converge and all expression edits, before the existing NBSMOOTH step, compute for each output cell the fraction of its 15 nearest neighbors (cKDTree on final coords, exclude self) that originated from the same side (a or b). This side label is already tracked internally. (2) β_i = β_max × clip(1 − same_frac, 0, 1). Pure neighborhoods (same_frac=1) → β_i=0 (no smoothing, preserving mmd_u); maximally mixed (same_frac≈0.5) → β_i=β_max×0.5. (3) Apply x_i ← (1−β_i)x_i + β_i·mean_NN(x_j) nnz-only as before. (4) Pooled pseudobulk restore g=clip(pb_before/pb_after, 0.5, 2) unchanged → DE frozen. Coords untouched → shape frozen. Key parameters: β_max initial 0.70 (so effective β at 50% mixing = 0.35, matching current optimum); search β_max ∈ {0.5, 0.7, 0.9, 1.1}. k stays 15. Single-input fallback: b=None → copy_last early return, NBSMOOTH gate skips (ia.size and ib.size check already present). vec-score screening: run seed 0 A-half, compare mmd_u raw (target: ≤0.00954, i.e. recover parent-49 level) …
风险1) Side-origin labels may not survive all pipeline transforms (check array lengths match at NBSMOOTH point; Engineer should assert len(side_labels)==n_cells before use). 2) If most neighborhoods are already side-pure after stratified draw (high same_frac everywhere), adaptive β collapses to near-zero everywhere → output ≈ parent 49, no gain. Detect early: print mean impurity; if <0.15, the mechanism has no room and Engineer should fall back to k×β fine-tune (k∈{10,20}, β∈{0.35,0.40}) as backup. 3) Overfitting to A-half noise: require 3-seed paired direction consistency before concluding.

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 01713da477。改动的文件:solution/METHOD.md +95 −72、solution/README.md +13 −7、solution/run.py +159 −26

diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 9a994a6..2590edb 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,78 +1,101 @@-NBSMOOTH:把非零表达向 15-NN 邻域均值轻度平滑(β=0.35,nnz-only、池化伪批量中和、坐标不动)提升邻域相干;PLAN 的加性 VPRECENT 与 James-Stein 收缩均证否。+QREM:均匀 15-NN 平滑(β=0.35)后按秩回填原始非零值,精确恢复逐基因边际分布与伪批量,保住 nbhd 增益、消除 mmd_u 收缩损失;PLAN 的杂质门控自适应 β 经 6 配置证否。++# 节点 53:QREM(rank-quantile marginal restoration after neighborhood smoothing)++- family_id: T2EI-01(表达 / cell_state + local_spatial)+- 父节点:51(NBSMOOTH,均匀 β=0.35 的 15-NN 邻域平滑,A 半复现 65.70)+- 提交机制:**QREM**;PLAN 指定机制(ADPNBSMOOTH 杂质门控自适应 β)已按 PLAN 实现并在 A 半证否(见下),+  按任务书要求改交针对同一弱项(nbhd 增益与 mmd_u 损失的权衡)的备选机制。++## 方法++在父节点 NBSMOOTH 步骤(所有表达改动之后、坐标定型处;k=15 最近邻、排除自身、nnz-only)内部:++1. 与父节点相同地计算平滑值 `x' = (1−β)x + β·mean_15NN(x)`(β=0.35 均匀,nnz-only:零位不动,支持集不变;+   由 β<1 且 x'>0 当 x>0,支持集精确保持)。+2. **QREM(本节点新机制)**:对每个基因 g,把其非零位置上的平滑值按**秩**替换为平滑前该基因非零值的+   **同秩次序统计量**(stable argsort 两次取秩 → 索引排序后的原值数组)。逐基因的非零值多重集因此逐位恢复+   为平滑前的边际分布:池化伪批量、每基因均值/方差/分位数精确不变(不再需要 pb-neutral 的 g 复原,代码里+   QREM 开启时跳过该步),而平滑造成的**秩结构**(哪些细胞对该基因偏高/偏低——现在与空间邻域相干)被保留。+3. 坐标完全不动 → shape_scale 三项与父逐位相同;伪批量精确复原 → de_score / de_direction 逐位冻结+   (A 半实测四组中 expression_change 62.23、shape_scale 78.23 在所有配置间逐位不变)。++原理:父节点 ANALYSIS 表明 nbhd 收益随 β 饱和(mean β≥0.27 后 nbhd raw 0.0451→0.0448 仅微降),而 mmd_u+损失来自逐细胞向邻域均值收缩导致的分布展宽丢失。QREM 把"空间相干"(秩)与"边际展宽"(值)解耦:+允许平滑改变细胞间的相对高低关系(nbhd 读取的邻域平均因此更相干),同时把每个基因的边际分布精确复位+(mmd_u/variogram 读取的单基因展宽不损失)。++关键参数(全部沿用父节点已验证值,QREM 无新自由参数):k=15、β=0.35、nnz-only、排除自身。+QREM 的 β 扫描(0.35/0.5/0.7/1.0)显示 0.35 最优:β≥0.7 时秩结构被过度平滑(邻域均值秩趋同),+回填后细胞值与自身局部环境失配,nbhd/mmd_u 同时崩坏(65.47 / 64.73)。++单输入回退:b=None → copy_last 提前返回(未触及);NBSMOOTH/QREM 块要求 ia.size 且 ib.size,单侧时整体跳过。+`--ablate mechanism`:关闭 QREM 与(默认已关的)自适应 β,保留父的均匀 NBSMOOTH β=0.35,**逐位还原父节点 51**+(已验证 seed 0 与 seed 3 两个种子 sha256 相同)。++## PLAN 机制(ADPNBSMOOTH)的实现与证否++按 PLAN 实现:输出细胞的侧标签 = mix_indices 顺序(前 ia.size 个来自 a 侧,其余 b 侧;断言长度匹配),+最终坐标上 cKDTree 取 15-NN(排除自身),impurity_i = 1 − 同侧邻居占比,β_i = clip(β_max·impurity_i, β_min, β_max)。+诊断(PLAN mechanism_evidence):mean impurity = **0.181**(高于 PLAN 风险 2 的 0.15 下限,机制有空间)、+中位 0.067;β_max=0.7 时 β_mean=0.127、p50=0.047、p90=0.373,**2215/5000 个细胞(44.3%)β>0.05**(机制生效于非平凡子集)。+但 A 半查分(seed 0,全部与同 seed 均匀 β=0.35 基线配对):++| 配置 | 榜分 | nbhd raw | mmd_u raw | 对照 |+|---|---:|---:|---:|---|+| 均匀 β=0.35(=父 51,A 半基线) | 65.696 | 0.04507 | 0.01034 | — |+| 无平滑(=节点 49) | 65.475 | 0.04741 | 0.01008 | −0.22 |+| 自适应 β_max=0.7 | 65.551 | 0.04607 | 0.01038 | −0.15 |+| 自适应 β_max=1.1 | 65.473 | 0.04542 | 0.01098 | −0.22 |+| 自适应 β_max=1.6 | 65.139 | 0.04534 | 0.01240 | −0.56 |+| 自适应 β_max=0.7 + β_min=0.1 | 65.610 | 0.04573 | 0.01033 | −0.09 |+| 自适应 β_max=0.9 + β_min=0.2 | 65.651 | 0.04506 | 0.01049 | −0.05 |++结论(6 配置全劣于父):**把平滑集中到混合邻域比摊平更伤 mmd_u**(高 β 细胞变成分布中的"平均化"离群点),+而 nbhd 只随总平滑量饱和——PLAN 的"纯邻域过平滑导致 mmd_u 损失"假设不成立,β 异质性本身就是 mmd_u 的损失源。+β_min=0.2+β_max=0.9 已把 nbhd 追平均匀基线(0.04506 vs 0.04507)但 mmd_u 更差(0.01049 vs 0.01034)→ 被严格支配。++顺带按父建议 #1 完成均匀 k×β 细调(A 半 seed 0 配对):k10β0.35=65.693、k10β0.40=65.694、k20β0.35=65.687、+k20β0.40=65.692、k15β0.40=65.701、k15β0.45=65.692——全部落在 ±0.015(纯噪声),确认均匀族在 (k15, β0.35) 已饱和,+k×β 方向无剩余空间(也说明必须换机制而非调参)。++## 查分记录(共 20 次,全部 T2:embryo:val_interp,A 半)++见上两表(7 + 6 + 无平滑基线 1 + 均匀基线 1),另 QREM β 扫描 4 次与种子验证 3 次:++| 配置 | seed | 榜分 | nbhd raw | mmd_u raw | variogram raw |+|---|---|---:|---:|---:|---:|+| QREM β=0.35(提交) | 0 | **65.799** | 0.04444 | 0.01030 | 0.007099 |+| QREM β=0.5 | 0 | 65.780 | 0.04416 | 0.01050 | 0.007100 |+| QREM β=0.7 | 0 | 65.472 | 0.04545 | 0.01100 | 0.007101 |+| QREM β=1.0 | 0 | 64.733 | 0.04806 | 0.01261 | 0.007091 |+| QREM β=0.35(提交) | 1 | **65.439** | 0.04484 | 0.00999 | 0.007317 |+| 均匀 β=0.35(父) | 1 | 65.332 | 0.04549 | 0.01006 | 0.007320 |+| QREM β=0.35(提交) | 2 | 65.319 | 0.04450 | 0.01044 | 0.007030 |++配对差(QREM − 均匀,同 seed 同锚点):s0 **+0.103**、s1 **+0.107**,两种子方向一致,且三个受影响指标+nbhd / mmd_u / variogram 的 raw 在两个种子上全部同向改善(6/6);expression_change 与 shape_scale 逐位冻结。+按父节点经验(A 半配对均值 +0.196 → 榜上 +0.24),预期正式收益 ~+0.1,量级小但机制风险低+(QREM 数学上不可能改变支持集、伪批量、逐基因边际;唯一可动的是 nbhd/mmd_u/variogram 读取的联合结构,实测三赢)。 -## 本节点做了什么(improve,父节点 49)--父节点 49(WIDENP,榜分 66.10 / 3 种子 65.66)四组最弱是 **cell_state**(A 半 59.6),其中 **variogram**(skill≈0.55)是地板以上最弱指标。PLAN 指定机制 **VPRECENT**:把 WITHINP 的乘性再居中因子 f_g 换成加性、"保方差" 的逐基因位移 δ_g=μ_full−μ_drawn(只作用于非零值、≤0 截 0)。--**PLAN 机制被证否**,改交父节点 ANALYSIS 建议 #2 的备选机制 **NBSMOOTH**(针对 local_spatial/nbhd,权重 25 最高)。--## PLAN 机制 VPRECENT(加性再居中)——证否--实现于 `WITHINP_MODE="add"`(`T2_WITHINP_MODE=add` 复现)。δ_g=μ_full−μ_drawn,可加下限 δ_g≥c·μ_drawn,nnz-only、x+δ≤0→0,之后仍走原池化伪批量复原 g=clip(pb0/pb1,0.5,2)。A 半 seed 0,父锚点 **65.4745**(variogram 0.007101、mmd_u 0.01008、nbhd 0.04741、de_dir 0.3784):--| 解码 | 榜分 | variogram | mmd_u | nbhd | de_dir |-|---|---|---|---|---|---|-| δ-floor c=−0.3 | 65.4029 | 0.007131 | 0.01010 | 0.04772 | 0.3761 |-| δ-floor c=−0.5(PLAN 初值) | 65.3944 | 0.007124 | 0.01012 | 0.04774 | 0.3755 |-| δ-floor c=−0.8(=无限制,−0.8 未绑定) | 65.3928 | 0.007128 | 0.01013 | 0.04772 | 0.3756 |-| 1/p_nz 精确再居中(c=−0.5) | 未查分:δ 放大到 25、支持集损失 25%、g 饱和到 2.0,本地即判灾难 | — | — | — | — |--四档全部劣于父节点,且 **variogram(目标指标)不降反升**(0.00710→0.00713),mmd_u、nbhd 同向变差。de_score 逐位冻结在 0.3214(pb-neutral 生效),de_dir 的 ±0.003 是 g 复原 float32 舍入在近零 dp 基因上的秩抖动,非真实信号。--**证否原因(机制证据)**:PLAN 假设"加性只移均值、保方差"不成立——实测 `max|var_after/var_before−1| = 109`(不是 ≈0)。因为位移是 **nnz-only + 0 截断**:零保持不变、负值截 0,这不是对分布的常数平移,方差照样被扭曲,和它想替换的乘性 f 同类。进一步,代数上"保方差的乘性再居中"恰好退化成加性位移,所以整个再居中族(乘性/加性)已探尽,父节点的乘性 f(含其 f² 方差缩放)反而是族内最优。支持集损失(非零→零)平均 0.5%、最大 7.2%,在 2% 阈内,不是主因;主因是加性位移本身扭曲了 nnz 幅值分布。--## 备选机制 1:James-Stein 收缩(父建议 #1)——证否--实现于 `WITHINP_MODE="js"`(`T2_JS_N0/T2_JS_P/T2_JS_WMIN`)。把硬 clip 的 f 换成按可靠度收缩 f→1:`f=clip(r^(W·w_n),0.6,1.6)`。两种可靠度:--- **尺寸可靠度**(JS_P=0,w_n=n/(n+n0)):n0=15→65.4666、n0=30→65.4603;-- **逐基因 SNR 可靠度**(JS_P=2,w=snr²/(snr²+n0),snr=|log r|/CV):n0=0.001→65.4376、n0=0.003→65.4351。--四档全部劣于父(65.4745),且 variogram 随收缩单调变差(0.00710→0.00712–0.00713)。`JS_N0=0` 逐位还原父节点(sha256 一致),证明是干净推广。结论:把 f 往 1 收=削弱再居中,单调变差,印证父的"全量乘性再居中 + 硬 clip [0.6,1.6]"是族内最优。--## 备选机制 2(提交):NBSMOOTH 邻域相干平滑(父建议 #2)--实现于 `NBSMOOTH_ENABLE`(默认开),位于 mix_converge 末尾、所有表达改动之后、坐标已定型处:--- 对最终输出坐标建 cKDTree,每个细胞取 **k=15** 最近邻(默认排除自身,`T2_NBSMOOTH_SELF=0`);-- `x_i ← (1−β)·x_i + β·mean_{j∈15NN(i)} x_j`,**nnz-only**(零保持零,不做零→非零注入,遵节点 37 教训),负值截 0;-- **池化伪批量中和**:平滑后按基因 `g=clip(pb_before/pb_after,0.5,2)` 整体复原,dp 逐位还原 → de_score/de_direction 结构性冻结;-- **坐标不动** → shape_scale 三项逐位不变;-- β=0.35(默认)。单输入视图 b=None 走 copy_last 早退,NBSMOOTH 被 `ia.size and ib.size` 门跳过,自动 no-op。--**为什么与 PSEUDOSTEP 不同**:PSEUDOSTEP 把细胞向 **型均值** 收缩(与位置无关),解耦了表达-位置配对,故恶化 nbhd(父节点已证否)。NBSMOOTH 把细胞向 **空间邻域均值** 收缩,正是 neighborhood_mmd 读取的"表达-位置相干"同方向:混合云把同型的 a 侧/b 侧细胞放在对齐坐标上,其邻域内表达散布大于真实单一中间期,轻度 β 平滑把邻域均值分布收紧向真值。--### 提交配置的机制证据(3 种子配对,A 半)--| seed | 父 49 | NBSMOOTH β=0.35 | Δ | nbhd(父→nb) | local_spatial(父→nb) | mmd_u(父→nb) | variogram(父→nb) |-|---|---|---|---|---|---|---|---|-| 0 | 65.4745 | 65.6959 | **+0.221** | 0.04741→0.04507 | 61.81→63.02 | 0.01008→0.01034 | 0.007101→0.007117 |-| 1 | 65.1269 | 65.3320 | **+0.205** | 0.04777→0.04549 | 61.62→62.80 | 0.00976→0.01006 | 0.007318→0.007320 |-| 2 | 65.0317 | 65.1917 | **+0.160** | 0.04744→0.04514 | 61.79→62.98 | 0.01010→0.01057 | 0.007032→0.007043 |--3 种子均值 +0.196,**方向全正**。**nbhd 3/3 改善**(0.0474→0.0451,权重 25 + 结构门),local_spatial 3/3 约 +1.2;代价是 mmd_u 3/3 微升、cell_state 微降约 −0.3…−0.5。de_score/de_direction 逐位冻结(pb-neutral),shape_scale 逐位不变(坐标不动)——与 PLAN 的 expected_groups(cell_state、local_spatial)中 local_spatial 命中,cell_state 未命中(反被小幅牺牲,因 nbhd 权重远高于 mmd_u/variogram 的净损失)。--β 扫描(A 半 s0):0.1→65.560、0.2→65.633、0.3→65.681、0.35→65.696、0.4→65.701、0.5→65.670。峰值 0.35–0.4,0.5 掉头(cell_state 损失盖过 nbhd 收益)。取 **β=0.35**(近峰、对 cell_state 更保守,B 半更稳健)。--### A 半→B 半迁移预期--净增益 +0.20(A 半)虽在 T2 ~1 分噪声内,但按 §5 应以配对差方向一致性判定——3 种子全正、nbhd 3/3 改善,是真实机制增益而非噪声。nbhd 改善源自空间相干这一通用性质,应迁移到 B 半(方法卡:插值两榜本地↔官网接近)。父 B 半 66.10,预期 NBSMOOTH B 半 ≈ 66.3。--## 关键参数--- `T2_WITHINP_MODE=mul`(默认,父 49 乘性再居中,clip [0.6,1.6]、g [0.5,2.0])-- `T2_NBSMOOTH=1`(默认开)、`T2_NBSMOOTH_BETA=0.35`、`T2_NBSMOOTH_K=15`、`T2_NBSMOOTH_SELF=0`、`T2_NBSMOOTH_NNZ=1`、`T2_NBSMOOTH_PBNEUTRAL=1`-- VPRECENT/JS 作为被证否机制留在代码后(`T2_WITHINP_MODE=add`/`js`),默认不启用。--## --ablate mechanism(机制关闭对照)+## 验证过 / 未验证 -`--ablate mechanism` 设 `NBSMOOTH_ENABLE=False`、`WITHINP_MODE="mul"`,其余管线(WIDENP clip、SIDEFRIM、NBHDCOH、DETR 等)不动 → **逐位还原父节点 49**(sha256 = dabb16d3…,已验证 seed 0)。默认输出(NBSMOOTH 开)sha256 = e2a8efff…,与 ablate 不同 → mechanism_active=yes。+- 验证:seed 0 默认输出 == QREM β=0.35 筛选输出(sha256);`--ablate mechanism` 在 seed 0 与 seed 3 均逐位+  还原均匀 β=0.35(父 51 输出);同 seed 重跑逐位确定;**伪装视图**(+1 天平移、manifest 键序打乱、+  文件重命名 input_0/1.h5ad、换路径、重排版)输出与真实视图逐位相同(sha256 be25f360…);+  vec-check 通过;运行 ~3 s / ~0.6 GB,纯 CPU(EXECUTION.json {"gpu": false});QREM 循环 498 基因全部命中。+- 未验证:B 半与真实括号(E6.75+E8.0 之外的括号宽度/共有类型数组合)上的收益——QREM 不含任何依赖视图、+  绝对时间或阶段名的逻辑,参数全部现场从输入计算,预期可迁移,但没有实测;seed 2 的配对差(额度用尽,+  仅单臂 65.319,nbhd 0.04450 与前两种子模式一致)。 -## 验证过 / 未验证+## 知识来源 -- **验证过**:VPRECENT 4 解码 + JS 4 配置全劣于父(证否);NBSMOOTH 3 种子配对全正、nbhd 3/3 改善;default/ablate 的 sha256;vec-check 通过;单输入 no-op;无 RNG(确定);无绝对时间/视图路径依赖(视图无关)。查分 15 次(额度内)。-- **未验证**:B 半(官网)实际分——只有 A 半;未在真实 final 视图(可能不同 bracket/细胞数)上跑,但机制数据驱动、单输入自动退化;β 更细网格(0.35–0.4 之间)与 k≠15 未扫(k=15 锁定为 nbhd 度量尺度)。+无生物学先验被写入程序:机制纯粹是统计/几何的(邻域平滑 + 秩-分位数边际恢复),只使用视图内数据。+未使用 prior/、external/(本视图 external 为空)、任何保留阶段/基因型的测量值或文献数值。 -## 生物学知识来源+## 下一步建议 -NBSMOOTH 不含硬编码的保留阶段测量值:k-NN 邻域、β、坐标全部由程序从 view 的输入表达与坐标现场计算。用到的通用机制知识仅为"空间邻近细胞表达相干"(组织内局部表达连续性,属细胞状态/空间转录组的通用性质,非某保留阶段的定量测量)。无外部数据、无保留阶段/基因型数据。+1. QREM 与 β 的联合再扫(β∈{0.4, 0.45} + QREM):s0 上 β0.5 的 nbhd 更好(0.04416)但 mmd_u 略差,中间点未测。+2. QREM 的秩源替换:用"邻域均值的秩"回填"原值"已隐含在 β=1 情形(崩),但 β=0.35–0.5 区间秩-值失配最小;+   可试按类型分层的秩回填(型内边际恢复),检查 variogram。+3. mmd_u 的剩余差距(nosmooth 0.01008 是上界参照)只能靠支持集/检出率通道(DETR 族)继续,不要再搜再居中/幅度族。diff --git a/solution/README.md b/solution/README.mdindex 485b3f9..9f6e756 100644--- a/solution/README.md+++ b/solution/README.md@@ -1,11 +1,17 @@-# mix + 表达收敛管线 + DETR / DETR_EXT / NBHDCOH(_SHARED) / WITHINP + SIDEFRIM + WIDENP + NBSMOOTH(T2:embryo:val_interp)+# mix + 表达收敛管线 + DETR / DETR_EXT / NBHDCOH(_SHARED) / WITHINP + SIDEFRIM + WIDENP + NBSMOOTH + QREM(T2:embryo:val_interp) -> **本节点(51)提交机制 = NBSMOOTH**(详见 METHOD.md):在父节点 49(WIDENP,乘性再居中)之上,-> 对最终非零表达向 15-NN 邻域均值做轻度平滑(β=0.35,nnz-only、池化伪批量中和、坐标不动),-> 提升 neighborhood_mmd 读取的表达-位置相干。A 半 3 种子配对 vs 父 +0.221/+0.205/+0.160(全正,-> nbhd 3/3 改善、local_spatial +1.2)。`--ablate mechanism`(或 `T2_NBSMOOTH=0`)逐位还原父节点 49。-> PLAN 指定的加性 VPRECENT 与备选 James-Stein 收缩均经多档查分证否(见 METHOD.md),代码保留在-> `T2_WITHINP_MODE=add/js` 后默认不启用。以下为父节点 49 及上游管线的历史说明。+> **本节点(53)提交机制 = QREM**(详见 METHOD.md):在父节点 51(NBSMOOTH,均匀 β=0.35 的 15-NN+> 邻域平滑)之上,对平滑后的每个基因按**秩**回填平滑前的非零值(rank-quantile 边际恢复):逐基因边际+> 分布、池化伪批量、支持集精确复位(DE/shape 组逐位冻结),只保留平滑带来的空间相干**秩结构**。+> A 半配对 vs 父:s0 +0.103、s1 +0.107,nbhd/mmd_u/variogram 三项 raw 6/6 同向改善。+> `--ablate mechanism`(或 `T2_NBSMOOTH_QREM=0`)逐位还原父节点 51。PLAN 指定的杂质门控自适应 β+> (ADPNBSMOOTH)经 6 配置证否(β 异质性本身伤 mmd_u),代码保留在 `T2_NBSMOOTH_ADAPTIVE=0` 默认关;+> 父建议的均匀 k×β 细调(6 配置)全部落在噪声内,确认均匀族饱和。以下为父节点 51/49 及上游管线的历史说明。++> **节点 51 机制 = NBSMOOTH**:对最终非零表达向 15-NN 邻域均值做轻度平滑(β=0.35,nnz-only、+> 池化伪批量中和、坐标不动),提升 neighborhood_mmd 读取的表达-位置相干。A 半 3 种子配对 vs 节点 49+> +0.221/+0.205/+0.160(全正,nbhd 3/3 改善、local_spatial +1.2)。`T2_NBSMOOTH=0` 还原节点 49。+> 该节点 PLAN 的加性 VPRECENT 与 James-Stein 收缩均证否,代码保留在 `T2_WITHINP_MODE=add/js` 默认不启用。  父节点 47 管线原样保留(mix 混抽 + α=5 类型级收敛位移 + 软阈值基因权重 + 相关扩散 η=1/asym=0.85 + λ=6 投影加权 + β=0.2 型内配对收缩 + aniso 坐标整形 damp=1.25 + SIDEFRIMdiff --git a/solution/run.py b/solution/run.pyindex 3d79dbf..58bd124 100644--- a/solution/run.py+++ b/solution/run.py@@ -8,6 +8,37 @@ exp(log r_a + SCALE_DAMP·t·Δlog r), and draws cells stratified by type: round(t·n) from the later stage, the rest from the earlier one. Coordinates travel with the cells. n is log-linear in t, clipped to the board range. +This node (53, QREM, family T2EI-01): the PLAN mechanism (ADPNBSMOOTH: replace+the uniform β=0.35 with a per-cell β_i = β_max·impurity_i, impurity_i = 1 −+same-side fraction of cell i's 15-NN) was implemented and FALSIFIED on the+A-half over 6 paired configs (β_max 0.7/1.1/1.6, floors β_min 0.1/0.2; parent+anchor 65.696): best 65.651, all BELOW the uniform parent. Diagnostics: mean+impurity 0.181 (mechanism has room), 44.3% of cells β>0.05, yet CONCENTRATING+smoothing on mixed neighbourhoods hurts mmd_u MORE than spreading it (high-β+cells become distribution outliers: β_max=1.6 → mmd_u 0.01240 vs uniform+0.01034) while nbhd only tracks the TOTAL smoothing mass and saturates for+mean β ≥ 0.27 (β_max=0.9+floor 0.2 matches uniform nbhd 0.04506 vs 0.04507 at+worse mmd_u → strictly dominated). A uniform k×β fine grid (k∈{10,15,20} ×+β∈{0.35,0.40,0.45}, 6 configs) all landed within ±0.015 (noise), confirming the+uniform family is saturated at (k15, β0.35). The SUBMITTED alternative is QREM+(rank-quantile marginal restoration): after the SAME uniform β=0.35 nnz-only+15-NN smoothing, each gene's nonzero values are replaced cell-by-cell by the+PRE-smoothing values of the same within-gene rank (stable double-argsort ranks+index the sorted original multiset). Support, per-gene marginals and hence the+pooled pseudobulk are restored EXACTLY (de_score/de_direction bit-frozen, the+pb-neutral g step becomes a no-op and is skipped), while the smoothed, spatially+coherent RANK structure is kept — decoupling the nbhd gain (coherence) from the+mmd_u loss (per-cell contraction toward local means shrinks the spread). β sweep+with QREM: 0.35 → 65.799 / 0.5 → 65.780 / 0.7 → 65.472 / 1.0 → 64.733 (β≥0.7+over-smooths the ranks: neighbourhood-mean ranks collapse and restored values+mismatch the local environment). Paired A-half vs parent (same seed, same+anchors): s0 +0.103, s1 +0.107, with nbhd/mmd_u/variogram raw improving 6/6+(0.04444/0.04484 vs 0.04507/0.04549 nbhd; 0.01030/0.00999 vs 0.01034/0.01006+mmd_u). shape_scale and expression_change bit-identical (coords untouched, pb+exact). --ablate mechanism (or T2_NBSMOOTH_QREM=0) disables QREM and reproduces+parent node 51 bit-for-bit (sha256-verified at seeds 0 and 3). The falsified+adaptive β stays behind T2_NBSMOOTH_ADAPTIVE=0 (default off). See METHOD.md.+ This node (51, NBSMOOTH, family T2EI-01): the PLAN mechanism (VPRECENT: replace WITHINP's multiplicative recentering factor f_g with an ADDITIVE, supposedly variance-preserving per-gene shift δ_g=μ_full−μ_drawn applied to nnz entries and@@ -769,6 +800,46 @@ NBSMOOTH_NNZ = os.environ.get("T2_NBSMOOTH_NNZ", "1") == "1" NBSMOOTH_PBNEUTRAL = os.environ.get("T2_NBSMOOTH_PBNEUTRAL", "1") == "1" NBSMOOTH_G_CLIP = (float(os.environ.get("T2_NBSMOOTH_GLO", "0.5")),                    float(os.environ.get("T2_NBSMOOTH_GHI", "2.0")))+# ADPNBSMOOTH (this node 53, family T2EI-01, PLAN mechanism): replace the+# uniform β with a per-cell β_i gated by the local side-mixing impurity of the+# cell's k-NN neighborhood: impurity_i = 1 − (fraction of cell i's k nearest+# spatial neighbours, self excluded, that share cell i's bracket-side origin),+# β_i = clip(β_max · impurity_i, β_min, β_max). Rationale: node 51's uniform+# β=0.35 gained nbhd (+0.33) but lost mmd_u (−0.08) / cell_state (−0.35)+# because it over-averages cells whose neighbourhoods are already side-pure+# (expression-coherent); the mmd_u loss comes from shrinking the state+# distribution of those cells. Gating keeps the smoothing where the mixed+# cloud actually has excess within-neighbourhood scatter (a/b interfaces) and+# leaves pure neighbourhoods untouched (distribution shape preserved).+# T2_NBSMOOTH_ADAPTIVE=0 restores the uniform node-51 behaviour (β=0.35);+# --ablate mechanism disables NBSMOOTH entirely (reproduces node 49).+# ADPNBSMOOTH (PLAN mechanism of this node 53) was FALSIFIED on the A-half:+# impurity-gated per-cell β (mean impurity 0.18, β_max 0.7/1.1/1.6, floor+# β_min 0.1/0.2 — 5 configs) never reached the uniform-β=0.35 parent (best+# 65.65 vs 65.70): concentrating smoothing on mixed neighbourhoods hurts+# mmd_u MORE than spreading it (high-β cells become distribution outliers),+# while nbhd saturates for mean β ≥ 0.27. Default off; kept for --ablate+# diagnostics and reproduction of the screening runs.+# The SUBMITTED mechanism is QREM (below): rank-quantile marginal restoration+# after the same uniform neighbourhood smoothing.+NBSMOOTH_ADAPTIVE = os.environ.get("T2_NBSMOOTH_ADAPTIVE", "0") == "1"+NBSMOOTH_BETA_MAX = float(os.environ.get("T2_NBSMOOTH_BETA_MAX", "0.7"))+NBSMOOTH_BMIN = float(os.environ.get("T2_NBSMOOTH_BMIN", "0.0"))+# QREM (alternative mechanism of this node 53, after the impurity-gated adaptive+# β was falsified on the A-half): rank-quantile marginal restoration after the+# neighborhood smoothing. The nbhd gain of smoothing saturates with β while the+# mmd_u loss comes from the per-cell contraction toward the local mean (spread+# shrinkage). QREM decouples the two: after computing the smoothed values x',+# each gene's nonzero values are replaced, cell by cell, by the PRE-smoothing+# values of the same within-gene rank (order statistics of x'_g index the sorted+# multiset of x_g). Support is invariant (nnz-only smoothing), so the per-gene+# marginal — hence the pooled pseudobulk exactly, and the univariate spread — is+# restored bit-for-bit, while the smoothed RANK structure (which cells are high/+# low, now spatially coherent) is kept. This allows much stronger smoothing+# (β up to 1.0) without the distribution-shrinkage cost measured by mmd_u.+# Deterministic (stable sorts only), CPU-cheap. T2_NBSMOOTH_QREM=0 disables and+# (with T2_NBSMOOTH_ADAPTIVE=0) reproduces parent node 51 bit-for-bit.+NBSMOOTH_QREM = os.environ.get("T2_NBSMOOTH_QREM", "1") == "1" # AMPSHRINK (mechanism of this node 39, family T2EI-01): DETR repaired the # zero/nonzero SUPPORT channel; the nnz AMPLITUDE channel is still frozen at # the source stage (a-side cells keep the E_a nnz-value distribution, b-side@@ -3140,13 +3211,23 @@ def mix_converge(stage_a, stage_b, t: float, params: dict, alpha: float, view: s     # coherence smoothing of the final nnz expression toward each cell's k-NN     # neighborhood average (k=15, the neighborhood_mmd scale). nnz-only +     # pb-neutral + coords fixed. See the config block for the full rationale.-    nbsmooth_info = {"nbsmooth_enable": bool(NBSMOOTH_ENABLE and NBSMOOTH_BETA > 0.0),+    nbsmooth_info = {"nbsmooth_enable": bool(NBSMOOTH_ENABLE+                                             and (NBSMOOTH_BETA > 0.0+                                                  or (NBSMOOTH_ADAPTIVE and NBSMOOTH_BETA_MAX > 0.0))),                      "nbsmooth_beta": NBSMOOTH_BETA, "nbsmooth_k": NBSMOOTH_K,                      "nbsmooth_self": bool(NBSMOOTH_SELF), "nbsmooth_nnz": bool(NBSMOOTH_NNZ),                      "nbsmooth_pbneutral": bool(NBSMOOTH_PBNEUTRAL),                      "nbsmooth_move_rel_mean": None, "nbsmooth_pb_maxshift": None,-                     "nbsmooth_g_min": None, "nbsmooth_g_max": None}-    if (NBSMOOTH_ENABLE and NBSMOOTH_BETA > 0.0 and ia.size and ib.size+                     "nbsmooth_g_min": None, "nbsmooth_g_max": None,+                     "nbsmooth_adaptive": bool(NBSMOOTH_ADAPTIVE),+                     "nbsmooth_beta_max": NBSMOOTH_BETA_MAX, "nbsmooth_bmin": NBSMOOTH_BMIN,+                     "nbsmooth_imp_mean": None, "nbsmooth_imp_p50": None,+                     "nbsmooth_beta_mean": None, "nbsmooth_beta_p50": None,+                     "nbsmooth_beta_p90": None, "nbsmooth_beta_maxreal": None,+                     "nbsmooth_n_beta_gt005": None, "nbsmooth_frac_beta_gt005": None,+                     "nbsmooth_qrem": bool(NBSMOOTH_QREM), "nbsmooth_qrem_genes": None}+    if (NBSMOOTH_ENABLE and (NBSMOOTH_BETA > 0.0 or (NBSMOOTH_ADAPTIVE and NBSMOOTH_BETA_MAX > 0.0))+            and ia.size and ib.size             and expr.shape[0] >= NBSMOOTH_K + 2 and coords.shape[0] == expr.shape[0]):         from scipy.spatial import cKDTree         pb_before_nb = expr.mean(axis=0).astype(np.float64)@@ -3156,17 +3237,61 @@ def mix_converge(stage_a, stage_b, t: float, params: dict, alpha: float, view: s         if not NBSMOOTH_SELF:             idx = idx[:, 1:]  # drop self (col 0)         idx = np.asarray(idx, dtype=np.int64)+        n_cells = expr.shape[0]+        if NBSMOOTH_ADAPTIVE:+            # side origin: first ia.size output cells come from bracket side a,+            # the remaining ib.size cells from side b (mix_indices ordering).+            assert n_cells == ia.size + ib.size, "side labels must match output cells"+            side = np.zeros(n_cells, dtype=bool)+            side[ia.size:] = True+            same_nb = side[idx] == side[:, None]           # (n, k)+            same_frac = same_nb.mean(axis=1)+            impurity = np.clip(1.0 - same_frac, 0.0, 1.0)+            beta = np.clip(NBSMOOTH_BETA_MAX * impurity, NBSMOOTH_BMIN, NBSMOOTH_BETA_MAX)+            beta = beta.astype(np.float64)+            nbsmooth_info.update(+                nbsmooth_imp_mean=float(impurity.mean()),+                nbsmooth_imp_p50=float(np.median(impurity)),+                nbsmooth_beta_mean=float(beta.mean()),+                nbsmooth_beta_p50=float(np.median(beta)),+                nbsmooth_beta_p90=float(np.percentile(beta, 90)),+                nbsmooth_beta_maxreal=float(beta.max()),+                nbsmooth_n_beta_gt005=int((beta > 0.05).sum()),+                nbsmooth_frac_beta_gt005=float((beta > 0.05).mean()),+            )+        else:+            beta = np.full(n_cells, float(NBSMOOTH_BETA), dtype=np.float64)         nb_mean = expr[idx].mean(axis=1).astype(np.float64)  # (n, G)         cur = expr.astype(np.float64)-        new = (1.0 - NBSMOOTH_BETA) * cur + NBSMOOTH_BETA * nb_mean+        new = cur + beta[:, None] * (nb_mean - cur)         if NBSMOOTH_NNZ:             new = np.where(cur > 0, new, cur)  # preserve support (zeros stay)         new = np.clip(new, 0.0, None)+        qrem_done = False+        if NBSMOOTH_QREM:+            # rank-quantile marginal restoration: per gene, the sorted multiset+            # of pre-smoothing nonzero values is reassigned to the post-smoothing+            # ranks (stable sorts → deterministic). Per-gene marginals (hence the+            # pooled pseudobulk and univariate spreads) are restored exactly.+            out_q = new.copy()+            n_qrem_genes = 0+            for gq in range(new.shape[1]):+                mq = cur[:, gq] > 0+                ngq = int(mq.sum())+                if ngq < 2:+                    continue+                ranks = np.argsort(np.argsort(new[mq, gq], kind="stable"), kind="stable")+                out_q[mq, gq] = np.sort(cur[mq, gq], kind="stable")[ranks]+                n_qrem_genes += 1+            new = out_q+            qrem_done = True+            nbsmooth_info["nbsmooth_qrem"] = True+            nbsmooth_info["nbsmooth_qrem_genes"] = n_qrem_genes         move = np.abs(new - cur)         denom = float(np.abs(cur).sum()) + 1e-12         nbsmooth_info["nbsmooth_move_rel_mean"] = float(move.sum() / denom)         expr = new.astype(np.float32)-        if NBSMOOTH_PBNEUTRAL:+        if NBSMOOTH_PBNEUTRAL and not qrem_done:             pb_after_nb = expr.mean(axis=0).astype(np.float64)             nbsmooth_info["nbsmooth_pb_maxshift"] = float(np.abs(pb_after_nb - pb_before_nb).max())             gnb = np.ones_like(pb_before_nb)@@ -3196,29 +3321,31 @@ def main() -> None:     parser.add_argument("--out", required=True)     parser.add_argument("--seed", type=int, default=0)     parser.add_argument("--ablate", default=None,-                        help="mechanism-off control: 'mechanism' (or any name) reverts "-                             "VPRECENT (this node's additive variance-preserving recentering) "-                             "to the parent's multiplicative WITHINP f — an exact no-op that "-                             "reproduces parent node 49 bit-for-bit")+                        help="mechanism-off control: 'mechanism' (or any name) disables "+                             "QREM (this node's rank-quantile marginal restoration) and the "+                             "(default-off) adaptive β, leaving the parent's uniform NBSMOOTH "+                             "β=0.35 — an exact no-op that reproduces parent node 51 bit-for-bit")     args = parser.parse_args()     global RESID_GAMMA, SPATRESID_THETA, DETR_FLIP_ON, AMPS_ENABLE, DETR_EXT, NBHDGATE, NBHDCOH     global NBHDCOH_SHARED, PROGCOH, WITHINP, ANISO2_GAMMA, SIDEFRIM, PSEUDOSTEP     global WITHINP_CLIP, WITHINP_G_CLIP, WITHINP_MODE, VPRE_FLOOR, VPRE_PNZ-    global NBSMOOTH_ENABLE, NBSMOOTH_BETA, JS_N0+    global NBSMOOTH_ENABLE, NBSMOOTH_BETA, JS_N0, NBSMOOTH_QREM, NBSMOOTH_ADAPTIVE     if args.ablate:-        # Only THIS node's submitted mechanism is switched off (contract G39.7);-        # WIDENP's clips (node 49), SIDEFRIM (node 47), NBHDCOH, DETR and-        # everything upstream stay active. The submitted mechanism is whichever-        # of VPRECENT (additive) / JS (James-Stein) / NBSMOOTH (spatial smoothing)-        # is default-on; reverting WITHINP_MODE to the parent's multiplicative f-        # (node 49's widened clips kept), disabling NBSMOOTH, and zeroing JS's-        # shrinkage is an exact no-op relative to node 49, so the ablated output-        # reproduces parent node 49 bit-for-bit. PSEUDOSTEP (the falsified-        # node-49 PLAN mechanism) is default-off and also forced off here;-        # PROGCOH/ANISO2 are default-off.+        # Only THIS node's submitted mechanism is switched off (contract G39.7).+        # The submitted mechanism is QREM (rank-quantile marginal restoration+        # after the uniform neighbourhood smoothing); the PLAN's impurity-gated+        # adaptive β was falsified and is default-off. Disabling QREM and+        # adaptive β leaves the parent's uniform NBSMOOTH (β=0.35, k=15,+        # nnz-only, pb-neutral) untouched, so the ablated output reproduces+        # parent node 51 bit-for-bit. WIDENP's clips (node 49), SIDEFRIM+        # (node 47), NBHDCOH, DETR and everything upstream stay active.+        # PSEUDOSTEP (the falsified node-49 PLAN mechanism) is default-off and+        # also forced off here; PROGCOH/ANISO2 are default-off.         WITHINP_MODE = "mul"-        NBSMOOTH_ENABLE = False-        NBSMOOTH_BETA = 0.0+        NBSMOOTH_QREM = False+        NBSMOOTH_ADAPTIVE = False+        NBSMOOTH_ENABLE = True+        NBSMOOTH_BETA = 0.35         PSEUDOSTEP = False         PROGCOH = False         ANISO2_GAMMA = 0.0@@ -3372,10 +3499,16 @@ def main() -> None:                                         "vpre_shrink_frac_max",                                         "js_n0", "js_p", "js_wmin", "js_w_mean",                                         "js_w_min", "js_w_max",-                                        "nbsmooth_enable", "nbsmooth_beta", "nbsmooth_k",-                                        "nbsmooth_self", "nbsmooth_nnz", "nbsmooth_pbneutral",-                                        "nbsmooth_move_rel_mean", "nbsmooth_pb_maxshift",-                                        "nbsmooth_g_min", "nbsmooth_g_max",+                                         "nbsmooth_enable", "nbsmooth_beta", "nbsmooth_k",+                                         "nbsmooth_self", "nbsmooth_nnz", "nbsmooth_pbneutral",+                                         "nbsmooth_move_rel_mean", "nbsmooth_pb_maxshift",+                                         "nbsmooth_g_min", "nbsmooth_g_max",+                                         "nbsmooth_adaptive", "nbsmooth_beta_max", "nbsmooth_bmin",+                                         "nbsmooth_imp_mean", "nbsmooth_imp_p50",+                                         "nbsmooth_beta_mean", "nbsmooth_beta_p50",+                                         "nbsmooth_beta_p90", "nbsmooth_beta_maxreal",+                                         "nbsmooth_n_beta_gt005", "nbsmooth_frac_beta_gt005",+                                         "nbsmooth_qrem", "nbsmooth_qrem_genes",                                         "pseudostep_enable", "pseudostep_lambda", "pseudostep_tau",                                         "pseudostep_nnz_only", "pseudostep_pbneutral",                                         "pseudostep_n_types", "pseudostep_n_cells",

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k007Interval staging and held-out-window filtering of external datanotes/official/来件/virtualembryo.ai/rules.md
k024World-model evaluation dimensions for state-transition predictorsnotes/competition/07_biomedical_world_models.md
k026Canonicalise predicted 3D coordinates before submissionnotes/pitfalls/04_scorer_invariance.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么PLAN 的杂质门控自适应 β(ADPNBSMOOTH)按 PLAN 实现后在 A 半 6 配置全部劣于父被证否;实际提交备选机制 QREM:在父节点均匀 β=0.35 的 15-NN nnz 平滑后,逐基因按秩回填平滑前的非零值,精确恢复支持集/逐基因边际/池化伪批量(DE 与 shape 逐位冻结),只保留平滑后的空间相干秩结构;坐标不动,无新自由参数。
各组分数的变化cell_state:+0.10(60.61→60.71),在噪声内:mmd_u raw 0.00982→0.00975(skill 0.648→0.650,+0.02 分)、variogram raw 0.00706→0.007049(+0.005 分)
expression_change:不变:de_score raw 0.3448、de_direction raw 0.397,两项得分 +0.00(QREM 精确复原伪批量,逐位冻结,符合设计)
local_spatial:+0.48(64.42→64.90),主要动项:neighborhood_mmd raw 0.04309→0.0422(skill 0.644→0.649,+0.12 分);榜分净变 +0.15 仍在 T2 ~1 分噪声内,但消融对照显示该 +0.15 全部来自 QREM(关闭后 −0.15),且 A 半 2 种子配对方向一致(+0.103/+0.107)、nbhd/mmd_u/variogram raw 6/6 同向改善
shape_scale:不变:三项 raw 与得分逐位相同(+0.00),坐标未动,符合设计
family_idT2EI-01
假设是否成立否
经验
  1. 在分层混抽后的插值云上,把邻域平滑集中到 a/b 混合邻域(按 side-impurity 门控 β_i)比均匀摊平更伤 mmd_u:高 β 细胞变成分布中的平均化离群点,6 配置 A 半配对全劣于均匀 β=0.35(最好 65.65 vs 65.70;β_max=1.6 时 mmd_u raw 0.01034→0.01240),β 异质性本身就是 mmd_u 的损失源。
  2. neighborhood_mmd 的收益只随总平滑量走且在 mean β≥0.27 后饱和;均匀 k×β 网格(k∈{10,15,20}×β∈{0.35,0.40,0.45})全部落在 ±0.015 纯噪声内,确认均匀平滑族在 (k15, β0.35) 已饱和,继续调参无空间,必须换机制。
  3. 平滑后做逐基因秩-分位数回填(QREM)可把空间相干(秩)与边际展宽(值)解耦:支持集/伪批量/逐基因边际精确复位使 DE 与 shape 逐位冻结,同时 nbhd/mmd_u/variogram 三项 raw 同向小幅改善——是'平滑收益保留、分布收缩损失消除'的干净构造。
  4. QREM 下 β 不能加大:β≥0.7 时秩结构被过度平滑(邻域均值秩趋同),回填值与局部环境失配,nbhd 与 mmd_u 同时崩坏(β=1.0 → 64.73);β=0.35–0.5 是秩-值失配最小的区间。
mechanism_active是
下一步建议
  1. QREM×β 联合细扫 β∈{0.40, 0.45}(cell_state + local_spatial):A 半 β=0.5 的 nbhd raw 更好(0.04416 vs 0.04444)但 mmd_u 略差,中间点未测,可能还有 ~0.05–0.1 的配对增益。
  2. 按细胞类型分层的秩回填(型内边际恢复,local_spatial/cell_state):QREM 目前只恢复逐基因边际,型内恢复可能进一步改善 variogram(skill 0.565,仍是地板以上最弱项之一)。
  3. mmd_u 的剩余差距(无平滑参照 raw ~0.01008)只能靠支持集/检出率通道(DETR 族)继续(cell_state);再居中/幅度族已探尽(乘性/加性/JS 收缩/自适应 β 全部证否),不要再搜。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。3
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 3 次:bash 1、read 1、write 1
用时1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 24,235 · 输出 1,405 · 思考 1,786
任务(第一行)分析节点 n53 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
ANALYSIS.json written
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/analyst.jsonl 17 KB
/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/analyst.stderr

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数45
工具调用共 47 次:bash 26、edit 13、read 7、write 1
用时21 分
token 数输入 112,674 · 输出 16,297 · 思考 21,457
任务(第一行)按计划修改程序:ADPNBSMOOTH: heterogeneity-gated adaptive β for NBSMOOTH
最后的回答(摘录)
All checks pass. Final state of `solution/`: default = uniform β=0.35 NBSMOOTH + QREM (bit-identical to screened best and to disguised-view run), `--ablate mechanism` reproduces parent node 51 bit-for-bit, vec-check ok, ~3 s / 0.6 GB CPU.

**做了什么**:实现 PLAN 的杂质门控自适应 β(ADPNBSMOOTH)——诊断显示 mean impurity 0.181、44% 细胞 β>0.05(机制确实生效),但 6 个配置(β_max 0.7/1.1/1.6、floor 0.1/0.2)A 半配对全劣于父(最好 65.65 vs 65.70):β 集中到混合邻域比摊平更伤 mmd_u,nbhd 只随总平滑量饱和;顺带把父建议的均匀 k×β 网格(6 配置)证实在噪声内饱和。改交备选机制 **QREM**:同样的均匀 β=0.35 15-NN 平滑后,逐基因按秩回填平滑前非零值,精确恢复边际分布/伪批量/支持集(DE 与 shape 逐位冻结),只保留空间相干的秩结构。

**查分**(20/20 用尽):QREM β=0.35 A 半 seed0 **65.799**(父 65.696),nbhd 0.04444 vs 0.04507、mmd_u 0.01030 vs 0.01034、variogram 同向改善;seed1 配对 65.439 vs 65.332(+0.107),三指标 6/6 同向;β 扫描 0.35 最优(0.5→65.78、0.7→65.47、1.0→64.73)。

**下一步最值得试**:QREM×β∈{0.4,0.45} 中间点;型内分层秩回填(查 variogram);mmd_u 剩余差距只在支持集/检出率通道(DETR 族),不要再搜再居中/幅度族。
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/engineer.jsonl 335 KB
/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/engineer.stderr

调研员

角色调研员 alibaba-token-plan-cn/qwen3.8-max
调用次数2 次
轮数2
工具调用共 0 次:—
用时3 分
token 数输入 22,456 · 输出 2,684 · 思考 4,082
任务(第一行)为节点 n53(改进,父节点 n51)写改进计划 PLAN.json
最后的回答(摘录)
e β_max or add a floor β_min=0.1. If both improve, confirm with seeds 1,2 paired.",
  "expected_groups": ["cell_state", "local_spatial"],
  "risks": "1) Side-origin labels may not survive all pipeline transforms (check array lengths match at NBSMOOTH point; Engineer should assert len(side_labels)==n_cells before use). 2) If most neighborhoods are already side-pure after stratified draw (high same_frac everywhere), adaptive β collapses to near-zero everywhere → output ≈ parent 49, no gain. Detect early: print mean impurity; if <0.15, the mechanism has no room and Engineer should fall back to k×β fine-tune (k∈{10,20}, β∈{0.35,0.40}) as backup. 3) Overfitting to A-half noise: require 3-seed paired direction consistency before concluding.",
  "family_id": "T2EI-01",
  "mechanism": "Per-cell smoothing strength gated by local side-mixing impurity: cells in mixed a/b neighborhoods are smoothed (restoring spatial coherence), cells in pure neighborhoods are left untouched (preserving distribution shape for mmd_u).",
  "vs_constant_shift": "Not a per-gene or per-type constant offset. β varies per cell based on its spatial neighborhood composition; the operation is a local weighted average with spatially-varying strength, not a uniform additive/multiplicative transform.",
  "mechanism_evidence": "Engineer should report: (a) distribution of β_i (histogram: expect bimodal — near 0 for pure, ~0.3-0.35 for mixed); (b) mmd_u raw compared to node 51 (0.00982) and node 49 (0.00954) — target ≤0.0096; (c) nbhd raw compared to node 51 (0.04309) — target ≤0.044; (d) cell_state group score vs 60.61; (e) number of cells with β_i > 0.05 (confirms mechanism is active on a non-trivial subset).",
  "mechanism_off_control": "--ablate mechanism sets T2_NBSMOOTH=0, reproducing parent node 49 bit-for-bit (sha256 verified). Secondary control: T2_NBSMOOTH_ADAPTIVE=0 with β_max=0.35 uniform reproduces node 51. If adaptive output == node 51 output, mechanism not running (impurity all ~1 or all ~0)."
}
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/researcher.jsonl 5 KB
/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/researcher.1.jsonl 5 KB
/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/researcher.stderr

审查员

角色审查员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数14
工具调用共 21 次:bash 10、read 8、todowrite 2、write 1
用时3 分
token 数输入 76,018 · 输出 2,577 · 思考 2,377
任务(第一行)审查节点 n53 的程序是否越界读取、写死目标、钻评分器漏洞
最后的回答(摘录)
REVIEW.json written
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/reviewer.jsonl 375 KB
/home/spark-longxinyang/vec/runs/formal/20261003-171955-search-t2-embryo-interp-chain-12h/nodes/53/reviewer.stderr