Virtual Embryo Challenge更新于 10-03 20:28(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261002-204523-search-t2-heart-interp-g24q

节点 n5 终选程序?按该运行锁定的规则最终选出的程序;可能是候选节点,也可能由护栏回退到基线。在终选来历上

T2HI-05:把 mix 的全局 RMS 缩放换成按细胞类型插值型内 RMS 的缩放(提交态只作用于较晚端点),仅改型内空间离散度,表达不动。

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261002-204523-search-t2-heart-interp-g24q
父节点n4
子节点—
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 59.84(+0.3) · proxy 59.84(+0.3) · 3 次复测均分 59.59
审查通过 检查1(越界读取):未发现问题——run.py 不直接 open 任何文件,读取全部经 manifest 驱动的 load_manifest/read_stage/panel_genes;import 来自 src.task2_spatial.*(框架),非 src/common/evaluation;无绝对路径、.. 、/mnt、/home、data/raw、downloads,也无联网。; 检查2(硬编码目标统计量):未发现问题——细胞类型来自 np.unique(labels)(run.py L78-80),rms_radius/_within_rms/log_interp 等全部现场…
用时?从运行开始到结束(或到现在)的挂钟时间。13 分
程序版本38c51d1baf7fc975f5ccd3d9cedb1a2cbc1cd225 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 38c51d1baf:solution/METHOD.md

T2HI-05:把 mix 的全局 RMS 缩放换成按细胞类型插值型内 RMS 的缩放(提交态只作用于较晚端点),仅改型内空间离散度,表达不动。

方法族与实现(family_id: T2HI-05)

父节点(node 4)提交态 = seed mix(interpolate,align=procrustes,scale_damp=1.0)。本节点在其 align_pair 之后、mix 之前,把对两端点云的单一全局 scale_to_rms(target_rms) 替换为按类型缩放:

  1. 对 aligned_a / aligned_b 按 celltype 分组,算每型质心 c_k 与型内 RMS r_k(细胞到自身质心的均方根距离)。
  2. 两端共有且两侧均 ≥ MIN_CELLS(默认 20)的类型:target_k = log_interp(r_a_k, r_b_k, t, scale_damp),新坐标 = c_k + (target_k / r_k)·(coord − c_k)(质心不动,只改型内离散度)。其余类型沿用全局缩放。
  3. 之后完全按父流程 mix:mix_indices 组成插值、n log-linear 夹到 [min,max]、_jitter 去重叠、最终整体 scale_to_rms(target_rms)。表达 = 真实端点细胞原样携带(未改动)。
  4. 单输入退路(b=None)与父节点一致(分层抽到 max_cells 原样输出)。

开关与变体(环境变量,提交默认 PERTYPE=1、SIDE=b、GAMMA=1.0、MINCELLS=20):T2HI05_SIDE ∈ {a,b,both} 选择对哪端点做按类型缩放;T2HI05_GAMMA<1 时按类型结果与全局结果线性混合;T2HI05_PERTYPE=0 直接调用父节点 interpolate(机制关闭对照)。

对照结果(vec-score,proxy,A 半;注意实验表里父的 59.50 是 B 半,A 半实测如下)

配置seed 0seed 1shape_scale (s0/s1)local_spatial (s0/s1)
机制关(=父 mix,逐位一致)59.0759.2153.07 / 52.7653.46 / 53.52
SIDE=b(提交态)59.2659.5253.46 / 53.7253.83 / 53.76
SIDE=both, γ=159.24—53.54 / —53.66 / —
SIDE=both, γ=0.559.24—53.69 / —53.50 / —
SIDE=a58.96—52.92 / —53.16 / —

关闭态对照:T2HI05_PERTYPE=0 与父节点 run.py 在 proxy seed 0 和 seed 1 上 .X 与 obsm["spatial_3D"] 均逐位相等(本地 numpy 验证),四组分与父完全相同。

机制生效的证据(stderr diag)

  • 缩放因子跨型明显离散(seed 0,proxy 5 个共有类型全部 ≥457 细胞,均超 MIN_CELLS):V-CM s_a=1.384 / s_b=0.614(两端型内 RMS 106 vs 240,差异最大,触发 >2× 离散警告)、Peri s_a=1.073 / s_b=0.900、NCC 0.994/1.009、aPHM 0.978/1.033、pPHM 1.009/0.987。若机制退化为全局缩放,所有 s_k 应等于 g_a=0.978 / g_b=1.034——实际不等,且提交态只把 b 侧(较晚端点,含 E9.5 的 V-CM 大离散云)收回到自身插值目标。
  • 只改坐标:.X 与关闭态逐位相等(本地验证),故 expression_change / cell_state 的表达类指标(de_score、mmd_u)完全不变;实际变化集中在 shape_scale(d2_shape 0.0393→0.0368、occupancy_dice 0.834→0.837)与 local_spatial(neighborhood_mmd 0.0834→0.0822),两组在两个 seed 上同向改善,与 PLAN 预期组一致。
  • 净效应:+0.19(seed 0)/ +0.30(seed 1),A 半。小于噪声(~0.5),但方向在两个 seed、三个变体(both/γ0.5/b)上一致,且 SIDE=a 变差说明效应来自坐标缩放而非随机涨落的概率较高。

提交决定

提交 SIDE=b(机制开启)。A 半两 seed 均优于关闭态对照;B 半正式分预期 ≈59.5–59.8,改善幅度在噪声边缘,风险为中性。查分共用 8 次(含 1 次关闭态基线、1 次 seed1 对照)。

验证过 / 未验证

  • 验证:proxy seed 0/1 上关闭态与父逐位一致;SIDE=b 两 seed 榜分与四组分;sideA/sideB/both/γ0.5 变体 seed 0;vec-check 通过;运行 ~2.5s、内存 <1GB(远低于 30min/28GB)。
  • 未验证:final 视图(真实括号 E8.25↔E8.75,31 个共有类型、t=0.5——按类型缩放会作用于更多类型,V-CM 式的极端型间尺度差是否存在未知;SIDE=b 的选择基于 proxy,可能在 final 上不是最优);seed 2;MIN_CELLS=10/50(proxy 上所有共有类型 ≥457 细胞,改阈值输出不变,未浪费查分)。
  • 生物学知识来源:无外部知识;仅用视图内数据(标签、坐标、时间差)与方法卡中的通用说明(z 为切片轴、心脏尺度非单调)。未使用任何保留阶段/保留基因型信息。
  • 视图无关性:输出只依赖数据与时间差(t、Δlog RMS 均平移不变),不读路径/榜名/绝对时间;--seed 经 np.random.default_rng 使用,确定。

调研员的计划

名称per-type RMS scaling (T2HI-05): 按类型插值尺度替代全局单一 target_rms
动机父节点(=node 2 mix)四组中 shape_scale 53.37、local_spatial 54.03 最弱。node 3(质心位移)和 node 4(配对插值)分别改了坐标位置和表达,均因伤害其它组而净负。两者都未触碰每型尺度:当前管线把两阶段全部细胞缩放到同一个全局 target_rms(log_interp(rms_a_global, rms_b_global, t)),忽略了各型自身尺度趋势可能不同(心脏尺度非单调,T2HI-05 方向)。按类型分别插值 RMS 只改变型内离散度,不移动质心、不合成细胞、不改表达,因此不会重蹈 node 3/4 的覆辙。
做法在父节点 run.py 的 align_pair 之后、mix 之前,将全局 scale_to_rms 替换为按类型缩放:
1. 对 aligned_a、aligned_b 分别按标签分组,计算每型质心 c_k 和型内 RMS r_k(细胞到自身质心的均方根距离)。
2. 对两端共有且两侧均 ≥ MIN_CELLS(默认 20)的类型:target_rms_k = log_interp(r_a_k, r_b_k, t, damp);缩放因子 s_k = target_rms_k / r_k;新坐标 = c_k + s_k·(coord − c_k)。对仅在一端出现的类型或细胞数 < MIN_CELLS 的类型,仍用全局 target_rms 缩放(与父节点一致)。
3. 缩放后按父节点原有流程做 mix(组成插值、n 的 log-linear 插值、jitter 去重叠)。
4. 关键参数:MIN_CELLS 初值 20,搜索范围 [10, 50];damp 沿用 1.0。
5. 单输入阶段退路(b=None):无配对阶段,直接输出 a 的细胞,与父节点一致。
6. vec-score 快筛:① 先跑 T2HI05_PERTYPE=0(关闭),确认与父输出逐位一致(1 次);② 开启、MIN_CELLS=20、seed 0 查 proxy(1 次);③ 若 ≥59.5 再查 seed 1 确认(1 次);④ 若 <59.5 试 MIN_CELLS=10(1 次)。共 ≤6 次,留余量。
风险1) 小类型(<20 细胞)的型内 RMS 估计噪声大,反而引入随机缩放——用 MIN_CELLS 阈值兜底,Engineer 应输出每型 cell count 和 r_a_k、r_b_k、target_rms_k 的 diag 表,检查是否有型的 target 偏离全局值过远(>2×)。2) 全局 RMS 差异本身很小(354→335),型间差异可能也小,净提升 <1 分噪声——用 seed 0+1 双查确认,若两次均 <59.5 则放弃。3) 缩放改变型内坐标可能轻微影响 local_spatial 的邻居结构——若 local_spatial 降 >0.5 而 shape_scale 升,需权衡;若两者均无变化则机制可能无效。

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 414f874ab9。改动的文件:solution/METHOD.md +27 −28、solution/run.py +97 −282

diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 0109467..fdcad64 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,42 +1,41 @@-实现 T2HI-01 类型内配对插值并对照:五个变体在代理上均不优于父 mix,提交为机制关闭(=父节点 mix,逐位一致)。+T2HI-05:把 mix 的全局 RMS 缩放换成按细胞类型插值型内 RMS 的缩放(提交态只作用于较晚端点),仅改型内空间离散度,表达不动。 -## 方法族与实现+## 方法族与实现(family_id: T2HI-05) -PLAN 指定 T2HI-01:对每个共有细胞类型,在表达 PCA(30 维,全局拟合两阶段)空间用 KDTree 最近邻把 E8.25_late 与 E9.5 的细胞配对(数量不等时对少的一侧有放回采样补齐),生成 expr=(1−t)·a+t·b、coords=(1−t)·ca+t·cb(z 同插值)的合成中间态细胞;组成按 (1−t)·frac_a+t·frac_b,n 沿用 log-linear 夹到 [min,max];Procrustes 对齐 + log-linear RMS 缩放与父节点相同。单输入退路(b=None)与父节点一致。+父节点(node 4)提交态 = seed mix(`interpolate`,align=procrustes,scale_damp=1.0)。本节点在其 `align_pair` 之后、mix 之前,把对两端点云的单一全局 `scale_to_rms(target_rms)` 替换为按类型缩放: -实现了三个变体(`run.py` 内,环境变量开关,提交默认关闭):+1. 对 aligned_a / aligned_b 按 `celltype` 分组,算每型质心 c_k 与型内 RMS r_k(细胞到自身质心的均方根距离)。+2. 两端共有且两侧均 ≥ MIN_CELLS(默认 20)的类型:`target_k = log_interp(r_a_k, r_b_k, t, scale_damp)`,新坐标 = `c_k + (target_k / r_k)·(coord − c_k)`(质心不动,只改型内离散度)。其余类型沿用全局缩放。+3. 之后完全按父流程 mix:`mix_indices` 组成插值、n log-linear 夹到 [min,max]、`_jitter` 去重叠、最终整体 `scale_to_rms(target_rms)`。表达 = 真实端点细胞原样携带(未改动)。+4. 单输入退路(b=None)与父节点一致(分层抽到 max_cells 原样输出)。 -1. `pair`(PAIR_FRAC=1.0/0.5):合成插值池(可选与真实端点细胞混合成池)。-2. `hybrid_bernoulli`(提交默认 MODE):父 mix 的真实细胞与坐标不变,每个细胞的表达与其类型内 NN 伙伴做逐基因 Bernoulli 混合(a 侧以概率 t 取伙伴值),保留单细胞稀疏结构。-3. `hybrid_convex`:同上但表达做凸组合 (1−w)·self+w·partner。+开关与变体(环境变量,提交默认 PERTYPE=1、SIDE=b、GAMMA=1.0、MINCELLS=20):`T2HI05_SIDE` ∈ {a,b,both} 选择对哪端点做按类型缩放;`T2HI05_GAMMA`<1 时按类型结果与全局结果线性混合;`T2HI05_PERTYPE=0` 直接调用父节点 `interpolate`(机制关闭对照)。 -## 对照结果(vec-score,proxy,seed 0,A 半)+## 对照结果(vec-score,proxy,A 半;注意实验表里父的 59.50 是 B 半,A 半实测如下) -| 变体 | 榜分 | cell_state | expression_change | local_spatial | shape_scale |-|---|---:|---:|---:|---:|---:|-| 父 mix(=机制关闭,逐位一致) | 59.50 | 66.70 | 63.90 | 54.03 | 53.37 |-| pair 全合成 | 58.50 | 63.63 | 66.28 | 53.42 | 50.66 |-| pair 半合成 | 58.85 | 64.52 | 66.25 | 53.27 | 51.36 |-| hybrid_bernoulli | 59.31 | 66.43 | 64.70 | 53.04 | 53.07 |-| hybrid_bernoulli 仅 a 侧 | 58.97 | 65.97 | 64.04 | 52.79 | 53.07 |-| hybrid_convex | 59.06 | 65.61 | 64.69 | 52.88 | 53.07 |+| 配置 | seed 0 | seed 1 | shape_scale (s0/s1) | local_spatial (s0/s1) |+|---|---:|---:|---:|---:|+| 机制关(=父 mix,逐位一致) | 59.07 | 59.21 | 53.07 / 52.76 | 53.46 / 53.52 |+| **SIDE=b(提交态)** | **59.26** | **59.52** | **53.46 / 53.72** | **53.83 / 53.76** |+| SIDE=both, γ=1 | 59.24 | — | 53.54 / — | 53.66 / — |+| SIDE=both, γ=0.5 | 59.24 | — | 53.69 / — | 53.50 / — |+| SIDE=a | 58.96 | — | 52.92 / — | 53.16 / — | -机制关闭对照:`T2HI_INTERPOLATE=0` 时输出与父节点 run.py(node 2 的 mix)在 proxy seed 0 上 `.X` 与 `obsm` 逐位相等(本地 cmp 验证),四组分应与父节点完全相同。+关闭态对照:`T2HI05_PERTYPE=0` 与父节点 run.py 在 proxy seed 0 和 seed 1 上 `.X` 与 `obsm["spatial_3D"]` 均逐位相等(本地 numpy 验证),四组分与父完全相同。 -## 机制生效的证据(pair 变体,stderr diag)+## 机制生效的证据(stderr diag) -- 输出 RMS 346.3,介于 rms_a=354.1 与 rms_b=335.0 对应的 target_rms(log-linear,scale_damp=1)。-- 每个共有类型内插值坐标总方差远小于端点混合方差(NCC 2.96e4 vs 9.85e4;Peri 3.88e4 vs 8.33e4;V-CM 1.38e4 vs 3.65e4;aPHM 1.73e4 vs 6.54e4;pPHM 1.28e4 vs 7.48e4)→ 配对确实产生了更紧凑的中间态邻域,机制按设计改变了这 5 个共有类型的全部细胞(proxy 上共有类型仅 NCC/Peri/V-CM/aPHM/pPHM)。-- 配对距离中位数:NCC 19.2、aPHM 17.4、pPHM 18.9 超过 2× 型内 a 侧中位距离(~7.4–8.7),已按 PLAN 风险 1 输出警告——proxy 两端相隔 1.25 天且标签词汇不同,跨阶段表达配对质量差;Peri 11.7、V-CM 10.8 未触发警告。-- 四组分变化方向:expression_change 升(+2.4 全合成 / +0.8 hybrid),cell_state 降(凸组合过平滑伤分布真实性;Bernoulli 保稀疏基本恢复 66.4),shape_scale 降(合成坐标填在两云之间,d2_shape 变差),local_spatial 持平或略降。净效应为负。+- 缩放因子跨型明显离散(seed 0,proxy 5 个共有类型全部 ≥457 细胞,均超 MIN_CELLS):V-CM s_a=1.384 / s_b=0.614(两端型内 RMS 106 vs 240,差异最大,触发 >2× 离散警告)、Peri s_a=1.073 / s_b=0.900、NCC 0.994/1.009、aPHM 0.978/1.033、pPHM 1.009/0.987。若机制退化为全局缩放,所有 s_k 应等于 g_a=0.978 / g_b=1.034——实际不等,且提交态只把 b 侧(较晚端点,含 E9.5 的 V-CM 大离散云)收回到自身插值目标。+- 只改坐标:`.X` 与关闭态逐位相等(本地验证),故 expression_change / cell_state 的表达类指标(de_score、mmd_u)完全不变;实际变化集中在 shape_scale(d2_shape 0.0393→0.0368、occupancy_dice 0.834→0.837)与 local_spatial(neighborhood_mmd 0.0834→0.0822),两组在两个 seed 上同向改善,与 PLAN 预期组一致。+- 净效应:+0.19(seed 0)/ +0.30(seed 1),A 半。小于噪声(~0.5),但方向在两个 seed、三个变体(both/γ0.5/b)上一致,且 SIDE=a 变差说明效应来自坐标缩放而非随机涨落的概率较高。 -## 结论与提交+## 提交决定 -与方法卡一致(mix 家族在该榜是平台,OT 类更差):类型内配对插值把 expression_change 的提升被 cell_state/shape_scale/local_spatial 的损失抵消,5 个变体全部 ≤ 父节点。提交为机制关闭状态(`T2HI_INTERPOLATE` 默认 "0"),逐位等于父 mix(proxy seed 0 本地验证 `.X` 与 `obsm` 逐位相等;预期 proxy 59.50,rank3 ≈59.36)。vec-score 共用掉 6 次额度。+提交 SIDE=b(机制开启)。A 半两 seed 均优于关闭态对照;B 半正式分预期 ≈59.5–59.8,改善幅度在噪声边缘,风险为中性。查分共用 8 次(含 1 次关闭态基线、1 次 seed1 对照)。  ## 验证过 / 未验证 -- 验证:proxy seed 0 上 5 个变体的 vec-score;关闭态与父程序输出逐位一致;vec-check 通过;运行时间(关闭态 ~3s,机制开 ~60s,均远低于 30min/28GB 限制)。-- 未验证:final 视图(31 个共有类型、t=0.5、两端更近,配对质量应好于 proxy——若后续节点重试 T2HI-01,final 上结论可能不同,但本节点无法查证);seed 1/2 上机制开的复跑(差距为负,不值得额度);坐标加权配对(表达+坐标联合特征)未试。-- 生物学知识来源:无外部知识;仅使用视图内数据(表达、标签、时间差)与任务书/方法卡中的通用说明(z 为切片轴、心脏尺度非单调)。未使用任何保留阶段/保留基因型信息。-- 视图无关性:程序只读 manifest 数据与时间差,不读绝对时间/路径/榜名;伪装视图(时间平移)下输出不变(t、Δlog RMS、类型比例均平移不变)。+- 验证:proxy seed 0/1 上关闭态与父逐位一致;SIDE=b 两 seed 榜分与四组分;sideA/sideB/both/γ0.5 变体 seed 0;vec-check 通过;运行 ~2.5s、内存 <1GB(远低于 30min/28GB)。+- 未验证:final 视图(真实括号 E8.25↔E8.75,31 个共有类型、t=0.5——按类型缩放会作用于更多类型,V-CM 式的极端型间尺度差是否存在未知;SIDE=b 的选择基于 proxy,可能在 final 上不是最优);seed 2;MIN_CELLS=10/50(proxy 上所有共有类型 ≥457 细胞,改阈值输出不变,未浪费查分)。+- 生物学知识来源:无外部知识;仅用视图内数据(标签、坐标、时间差)与方法卡中的通用说明(z 为切片轴、心脏尺度非单调)。未使用任何保留阶段/保留基因型信息。+- 视图无关性:输出只依赖数据与时间差(t、Δlog RMS 均平移不变),不读路径/榜名/绝对时间;`--seed` 经 `np.random.default_rng` 使用,确定。diff --git a/solution/run.py b/solution/run.pyindex 11ae220..f71b2a8 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,26 +1,29 @@ #!/usr/bin/env python3-"""mix + within-type paired interpolation (T2HI-01) for T2 heart interpolation.--Bracket the target with the nearest inputs before/after, align both clouds in-one frame (procrustes: xy Kabsch on shared-type centroids, z kept as slice-axis, sign fixed), rescale both to the log-linear RMS. Then, instead of only-mixing whole endpoint cells:--* for every cell type present in BOTH stages, pair a-cells with b-cells by-  nearest neighbour in expression-PCA space (30 dims, fitted on both stages)-  within the type, drawing with replacement on the smaller side to equalise-  counts, and synthesise cells-      expr   = (1-t) * expr_a + t * expr_b-      coords = (1-t) * ca_pair + t * cb_pair   (z interpolated too)-* cells of types present in only one stage stay real (expression carried,-  coordinates from that stage's aligned+scaled cloud);-* the output composition matches the parent mix expectation: per-type-  fraction (1-t)*frac_a + t*frac_b, total n log-linear in t and clipped to-  the board range.--Set INTERPOLATE=False (or env T2HI_INTERPOLATE=0) to fall back bit-for-bit to-the parent seed mix path (same params, same rng usage) as a mechanism-off-control.+"""mix + per-type RMS scaling (T2HI-05) for T2 heart interpolation.++Parent pipeline: bracket the target with the nearest inputs before/after,+align both clouds (procrustes: xy Kabsch on shared-type centroids, z kept),+rescale both to a single global log-linear target RMS, then mix whole real+cells (composition and n log-interpolated in t).++T2HI-05 change: replace the single global ``scale_to_rms`` on the aligned+endpoint clouds with a per-type rescale. For every cell type present in BOTH+endpoints with >= MIN_CELLS cells on each side, compute the type centroid c_k+and within-type RMS r_k (root-mean-square distance of cells to their own+centroid) on both sides, interpolate log-linearly+``target_k = log_interp(r_a_k, r_b_k, t, scale_damp)`` and rescale that type's+cells around their own centroid: ``c_k + (target_k / r_k) * (coord - c_k)``.+Types present on one side only, or too small, keep the parent's global+scaling. Centroids are never moved, no cells are synthesised and expression is+carried unchanged, so only within-type spatial dispersion is affected. The+final mixed cloud is still rescaled to the global target RMS as in the parent.++Default variant (T2HI05_SIDE=b): per-type scaling is applied to the later+bracket endpoint only; on the proxy this scored best and most consistently+across seeds. Set env ``T2HI05_PERTYPE=0`` to fall back bit-for-bit to the+parent mix path (mechanism-off control). ``T2HI05_MINCELLS`` sets MIN_CELLS+(default 20); ``T2HI05_GAMMA`` blends per-type toward global scaling+(default 1.0 = pure per-type); ``T2HI05_SIDE`` in {a,b,both} (default b). """  from __future__ import annotations@@ -33,8 +36,8 @@ import sys import numpy as np  from src.task2_spatial.frame import align_pair, log_interp, rms_radius, scale_to_rms-from src.task2_spatial.methods import interpolate-from src.task2_spatial.sample import interp_count, take+from src.task2_spatial.methods import _jitter, _limits, interpolate+from src.task2_spatial.sample import mix_indices, take from src.task2_spatial.transport import as_dense from src.task2_spatial.view_io import (     board_params,@@ -46,204 +49,67 @@ from src.task2_spatial.view_io import ( )  PARAMS = {"align": "procrustes", "scale_damp": 1.0}-INTERPOLATE = os.environ.get("T2HI_INTERPOLATE", "0") == "1"-PAIR_FRAC = float(os.environ.get("T2HI_PAIR_FRAC", "1.0"))-MODE = os.environ.get("T2HI_MODE", "hybrid_bernoulli")-BLEND = os.environ.get("T2HI_BLEND", "bernoulli")-SIDE = os.environ.get("T2HI_SIDE", "both")-PCA_DIM = 30-PCA_FIT_CELLS = 20000---def _jitter(coords: np.ndarray, rng: np.random.Generator) -> np.ndarray:-    if len(coords) < 2:-        return coords-    rounded = np.round(coords, 5)-    _, inv, counts = np.unique(rounded, axis=0, return_inverse=True, return_counts=True)-    if counts.max() <= 1:-        return coords-    rms = rms_radius(coords) + 1e-8-    noise = rng.normal(0.0, 1e-4 * rms, size=coords.shape)-    out = coords.copy()-    dup = counts[inv] > 1-    out[dup] = out[dup] + noise[dup]-    return out---def _largest_remainder(fracs: dict, n: int, caps: dict) -> dict:-    types = sorted(fracs)-    raw = np.array([fracs[t] * n for t in types], dtype=np.float64)-    alloc = np.floor(raw).astype(int)-    caps_arr = np.array([caps[t] for t in types], dtype=int)-    alloc = np.minimum(alloc, caps_arr)-    rem = int(n - alloc.sum())-    order = np.argsort(-(raw - np.floor(raw)))-    i = 0-    guard = 0-    while rem > 0 and guard < 10 * len(types) + 100:-        j = order[i % len(order)]-        if alloc[j] < caps_arr[j]:-            alloc[j] += 1-            rem -= 1-        i += 1-        guard += 1-        if i > 0 and i % len(order) == 0 and all(alloc >= caps_arr):-            break-    return {t: int(alloc[k]) for k, t in enumerate(types)}---def pair_interpolate(stage_a, stage_b, t: float, params: dict):-    t = float(t)-    damp = float(params.get("scale_damp", 1.0))-    align = str(params.get("align", "procrustes"))-    rng = np.random.default_rng(int(params.get("seed", 0)))-    from scipy.spatial import cKDTree-    from sklearn.decomposition import PCA--    aligned_a, aligned_b, info = align_pair(-        stage_a.coords, stage_b.coords, stage_a.labels, stage_b.labels, align-    )-    rms_a = rms_radius(stage_a.coords)-    rms_b = rms_radius(stage_b.coords)-    target_rms = log_interp(rms_a, rms_b, t, damp)-    ca = scale_to_rms(aligned_a, target_rms)-    cb = scale_to_rms(aligned_b, target_rms)-    lo = int(params["min_cells"])-    hi = int(params["max_cells"])-    n = interp_count(stage_a.n, stage_b.n, t, lo, hi, 1.0)--    la = np.asarray(stage_a.labels).astype(str)-    lb = np.asarray(stage_b.labels).astype(str)-    types_a = {t_: np.flatnonzero(la == t_) for t_ in np.unique(la)}-    types_b = {t_: np.flatnonzero(lb == t_) for t_ in np.unique(lb)}-    shared = sorted(set(types_a) & set(types_b))-    only_a = sorted(set(types_a) - set(types_b))-    only_b = sorted(set(types_b) - set(types_a))--    # global expression PCA for pairing features-    sub_a = rng.choice(stage_a.n, size=min(PCA_FIT_CELLS, stage_a.n), replace=False)-    sub_b = rng.choice(stage_b.n, size=min(PCA_FIT_CELLS, stage_b.n), replace=False)-    fit_mat = np.vstack([as_dense(stage_a.X, sub_a), as_dense(stage_b.X, sub_b)]).astype(np.float32)-    dim = min(PCA_DIM, fit_mat.shape[1], fit_mat.shape[0] - 1)-    pca = PCA(n_components=max(dim, 2), svd_solver="full")-    pca.fit(fit_mat)-    del fit_mat-    Za = pca.transform(as_dense(stage_a.X).astype(np.float32))-    Zb = pca.transform(as_dense(stage_b.X).astype(np.float32))--    Xa_full = None-    Xb_full = None-    if shared:-        Xa_full = as_dense(stage_a.X)-        Xb_full = as_dense(stage_b.X)--    expr_parts = []-    coord_parts = []-    pool_type = []-    pool_sizes = {}-    diag = {"pair_dist": {}, "var_interp": {}, "var_mix": {}}--    for typ in shared + only_a + only_b:-        ia = types_a.get(typ)-        ib = types_b.get(typ)-        if ia is not None and ib is not None:-            m = int(max(ia.size, ib.size))-            sa = ia if ia.size == m else rng.choice(ia, m, replace=True)-            sb = ib if ib.size == m else rng.choice(ib, m, replace=True)-            tree = cKDTree(Zb[sb])-            dist, partner = tree.query(Za[sa], k=1)-            # pairing-quality diagnostics (PLAN risk 1)-            within, _ = cKDTree(Za[ia]).query(Za[ia], k=2)-            med_pair = float(np.median(dist))-            med_within = float(np.median(within[:, 1]))-            diag["pair_dist"][typ] = {"pair": med_pair, "within_a": med_within,-                                      "warn": bool(med_pair > 2.0 * med_within)}-            pa = sa-            pb = sb[partner]-            xe_s = np.clip((1.0 - t) * Xa_full[pa] + t * Xb_full[pb], 0.0, None).astype(np.float32)-            xc_s = (1.0 - t) * ca[pa] + t * cb[pb]-            mix_c = np.vstack([ca[pa], cb[pb]])-            diag["var_interp"][typ] = float(xc_s.var(axis=0).sum())-            diag["var_mix"][typ] = float(mix_c.var(axis=0).sum())-            if PAIR_FRAC >= 1.0:-                xe, xc = xe_s, xc_s-            else:-                k = int(round(PAIR_FRAC * m))-                keep_s = np.zeros(m, dtype=bool)-                keep_s[rng.choice(m, k, replace=False)] = True-                slots = np.flatnonzero(~keep_s)-                n_real = slots.size-                n_ra = int(round((1.0 - t) * n_real))-                ra = slots[rng.choice(n_real, n_ra, replace=False)] if n_ra else np.array([], int)-                rb = slots[rng.choice(n_real, n_real - n_ra, replace=False)] if n_real - n_ra else np.array([], int)-                xe = np.vstack([xe_s[keep_s], Xa_full[pa[ra]], Xb_full[pb[rb]]]).astype(np.float32)-                xc = np.vstack([xc_s[keep_s], ca[pa[ra]], cb[pb[rb]]])-        elif ia is not None:-            xe = Xa_full[ia].astype(np.float32) if Xa_full is not None else as_dense(stage_a.X, ia)-            xc = ca[ia]-        else:-            xe = Xb_full[ib].astype(np.float32) if Xb_full is not None else as_dense(stage_b.X, ib)-            xc = cb[ib]-        expr_parts.append(xe)-        coord_parts.append(np.asarray(xc, dtype=np.float64))-        pool_type.append(np.full(xe.shape[0], len(pool_sizes), dtype=int))-        pool_sizes[len(pool_sizes)] = typ--    expr_pool = np.clip(np.vstack(expr_parts), 0.0, None).astype(np.float32)-    coord_pool = np.vstack(coord_parts)-    pool_type = np.concatenate(pool_type)--    # composition = (1-t)*frac_a + t*frac_b (matches parent mix expectation)-    na_tot = float(stage_a.n)-    nb_tot = float(stage_b.n)-    all_types = shared + only_a + only_b-    fr = {}-    caps = {}-    for k, typ in enumerate(all_types):-        fa = len(types_a.get(typ, [])) / na_tot-        fb = len(types_b.get(typ, [])) / nb_tot-        fr[typ] = (1.0 - t) * fa + t * fb-        caps[typ] = int((pool_type == k).sum())-    s = sum(fr.values())-    fr = {k_: v / s for k_, v in fr.items()}-    counts = _largest_remainder(fr, n, caps)--    picks = []-    for k, typ in enumerate(all_types):-        c = counts[typ]-        if c <= 0:+PERTYPE = os.environ.get("T2HI05_PERTYPE", "1") == "1"+MIN_CELLS = int(os.environ.get("T2HI05_MINCELLS", "20"))+GAMMA = float(os.environ.get("T2HI05_GAMMA", "1.0"))+SIDE = os.environ.get("T2HI05_SIDE", "b")+++def _within_rms(x: np.ndarray) -> tuple[np.ndarray, float]:+    c = x.mean(axis=0)+    d = x - c+    return c, float(np.sqrt((d * d).sum(axis=1).mean()))+++def per_type_rescale(A, B, labels_a, labels_b, t, damp, target_rms, min_cells):+    """Global scale by default; shared, large-enough types scale around their+    own centroid to their own log-interpolated within-type RMS."""+    A = np.asarray(A, dtype=np.float64)+    B = np.asarray(B, dtype=np.float64)+    ra = rms_radius(A)+    rb = rms_radius(B)+    ga = (target_rms / ra) if (ra > 1e-8 and target_rms > 0) else 1.0+    gb = (target_rms / rb) if (rb > 1e-8 and target_rms > 0) else 1.0+    out_a = A * ga+    out_b = B * gb++    la = np.asarray(labels_a).astype(str)+    lb = np.asarray(labels_b).astype(str)+    idx_a = {typ: np.flatnonzero(la == typ) for typ in np.unique(la)}+    idx_b = {typ: np.flatnonzero(lb == typ) for typ in np.unique(lb)}+    shared = sorted(set(idx_a) & set(idx_b))++    diag = {"global": {"scale_a": ga, "scale_b": gb, "target_rms": target_rms}, "types": {}}+    for typ in shared:+        ia, ib = idx_a[typ], idx_b[typ]+        if min(ia.size, ib.size) < int(min_cells):+            diag["types"][typ] = {"n_a": int(ia.size), "n_b": int(ib.size), "skipped_small": True}             continue-        idx = np.flatnonzero(pool_type == k)-        if c <= idx.size:-            picks.append(rng.choice(idx, c, replace=False))-        else:-            picks.append(rng.choice(idx, c, replace=True))-    sel = np.sort(np.concatenate(picks)) if picks else np.array([], dtype=int)--    expr = expr_pool[sel]-    coords = _jitter(coord_pool[sel], rng)-    coords = scale_to_rms(coords, target_rms)-    info.update(-        t=t, n=int(expr.shape[0]), rms_a=rms_a, rms_b=rms_b,-        target_rms=target_rms, out_rms=rms_radius(coords),-        scale_damp=damp, align=align, n_shared_types=len(shared),-        method="mix_pair_interp", diag=diag,-    )-    return expr, coords.astype(np.float32), info---def hybrid_bernoulli(stage_a, stage_b, t: float, params: dict):-    """Parent mix selection (real cells + real coords), expression blended with-    the within-type nearest-neighbour partner from the other stage by a-    per-gene Bernoulli draw (keeps single-cell sparsity, shifts levels to t)."""+        ca_k, r_a_k = _within_rms(A[ia])+        cb_k, r_b_k = _within_rms(B[ib])+        target_k = log_interp(r_a_k, r_b_k, t, damp)+        s_a = (target_k / r_a_k) if r_a_k > 1e-8 else 1.0+        s_b = (target_k / r_b_k) if r_b_k > 1e-8 else 1.0+        if SIDE in ("both", "a"):+            pa = ca_k + s_a * (A[ia] - ca_k)+            out_a[ia] = GAMMA * pa + (1.0 - GAMMA) * out_a[ia]+        if SIDE in ("both", "b"):+            pb = cb_k + s_b * (B[ib] - cb_k)+            out_b[ib] = GAMMA * pb + (1.0 - GAMMA) * out_b[ib]+        diag["types"][typ] = {+            "n_a": int(ia.size), "n_b": int(ib.size),+            "r_a": r_a_k, "r_b": r_b_k, "target": target_k,+            "s_a": s_a, "s_b": s_b,+            "warn_dev": bool(max(s_a, s_b) > 2.0 * min(s_a, s_b)),+        }+    return out_a, out_b, diag+++def mix_pertype(stage_a, stage_b, t: float, params: dict):     t = float(t)     damp = float(params.get("scale_damp", 1.0))     align = str(params.get("align", "procrustes"))     rng = np.random.default_rng(int(params.get("seed", 0)))-    from scipy.spatial import cKDTree-    from sklearn.decomposition import PCA-    from src.task2_spatial.sample import mix_indices-    from src.task2_spatial.methods import _jitter as _jit      aligned_a, aligned_b, info = align_pair(         stage_a.coords, stage_b.coords, stage_a.labels, stage_b.labels, align@@ -251,73 +117,24 @@ def hybrid_bernoulli(stage_a, stage_b, t: float, params: dict):     rms_a = rms_radius(stage_a.coords)     rms_b = rms_radius(stage_b.coords)     target_rms = log_interp(rms_a, rms_b, t, damp)-    ca = scale_to_rms(aligned_a, target_rms)-    cb = scale_to_rms(aligned_b, target_rms)-    lo = int(params["min_cells"])-    hi = int(params["max_cells"])-    n = interp_count(stage_a.n, stage_b.n, t, lo, hi, 1.0)+    ca, cb, diag = per_type_rescale(+        aligned_a, aligned_b, stage_a.labels, stage_b.labels, t, damp, target_rms, MIN_CELLS+    )+    n = _limits(params, stage_a.n, stage_b.n, t, "interp")     ia, ib = mix_indices(stage_a.labels, stage_b.labels, t, n, rng) -    Xa = as_dense(stage_a.X)-    Xb = as_dense(stage_b.X)-    la = np.asarray(stage_a.labels).astype(str)-    lb = np.asarray(stage_b.labels).astype(str)--    sub_a = rng.choice(stage_a.n, size=min(PCA_FIT_CELLS, stage_a.n), replace=False)-    sub_b = rng.choice(stage_b.n, size=min(PCA_FIT_CELLS, stage_b.n), replace=False)-    fit_mat = np.vstack([Xa[sub_a], Xb[sub_b]])-    dim = min(PCA_DIM, fit_mat.shape[1], fit_mat.shape[0] - 1)-    pca = PCA(n_components=max(dim, 2), svd_solver="full")-    pca.fit(fit_mat)-    del fit_mat-    Za = pca.transform(Xa)-    Zb = pca.transform(Xb)--    expr = np.vstack([Xa[ia], Xb[ib]]).astype(np.float64)-    other = np.empty_like(expr)-    has_partner = np.zeros(expr.shape[0], dtype=bool)-    weights = np.empty(expr.shape[0], dtype=np.float64)-    weights[: ia.size] = t          # a-side cells take b value with prob t-    weights[ia.size :] = 1.0 - t    # b-side cells take a value with prob 1-t-    if SIDE == "a":-        weights[ia.size :] = 0.0-    elif SIDE == "b":-        weights[: ia.size] = 0.0-    diag = {"pair_dist": {}}-    for side, sel, Zself, Zoth, Xoth, labs_self, labs_oth, base in (-        ("a", ia, Za, Zb, Xb, la, lb, 0),-        ("b", ib, Zb, Za, Xa, lb, la, ia.size),-    ):-        if sel.size == 0:-            continue-        for typ in np.unique(labs_self[sel]):-            rows = np.flatnonzero(labs_self[sel] == typ)-            self_all = np.flatnonzero(labs_self == typ)-            oth_all = np.flatnonzero(labs_oth == typ)-            if oth_all.size == 0:-                continue-            tree = cKDTree(Zoth[oth_all])-            dist, part = tree.query(Zself[sel[rows]], k=1)-            diag["pair_dist"][f"{side}:{typ}"] = float(np.median(dist))-            gi = base + rows-            other[gi] = Xoth[oth_all[part]]-            has_partner[gi] = True-    if BLEND == "convex":-        w = weights[:, None]-        expr = np.where(has_partner[:, None], (1.0 - w) * expr + w * other, expr)-    else:-        u = rng.random(size=expr.shape)-        take_other = has_partner[:, None] & (u < weights[:, None])-        expr = np.where(take_other, other, expr)-    expr = np.clip(expr, 0.0, None).astype(np.float32)--    coords = _jit(np.vstack([ca[ia], cb[ib]]), rng)+    expr = np.clip(+        np.vstack([as_dense(stage_a.X, ia), as_dense(stage_b.X, ib)]), 0.0, None+    ).astype(np.float32)+    coords = _jitter(np.vstack([ca[ia], cb[ib]]), rng)     coords = scale_to_rms(coords, target_rms)     info.update(         t=t, n=int(expr.shape[0]), rms_a=rms_a, rms_b=rms_b,         target_rms=target_rms, out_rms=rms_radius(coords),-        scale_damp=damp, align=align, method="hybrid_bernoulli",-        n_paired=int(has_partner.sum()), diag=diag,+        scale_damp=damp, align=align, method="mix_pertype",+        n_shared_types=len(set(np.asarray(stage_a.labels).astype(str)) &+                           set(np.asarray(stage_b.labels).astype(str))),+        min_cells_pertype=MIN_CELLS, diag=diag,     )     return expr, coords.astype(np.float32), info @@ -341,10 +158,8 @@ def main() -> None:     stage_a = read_stage(args.data, a, genes)     stage_b = read_stage(args.data, b, genes)     params = board_params(manifest, "mix", PARAMS, args.seed)-    if INTERPOLATE and MODE == "hybrid_bernoulli":-        expr, coords, info = hybrid_bernoulli(stage_a, stage_b, t, params)-    elif INTERPOLATE:-        expr, coords, info = pair_interpolate(stage_a, stage_b, t, params)+    if PERTYPE:+        expr, coords, info = mix_pertype(stage_a, stage_b, t, params)     else:         expr, coords, info = interpolate(stage_a, stage_b, t, params)     keep = {k: info.get(k) for k in ("t", "n", "rms_a", "rms_b", "out_rms", "n_shared_types", "z_dot", "z_flipped", "align", "method")}

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k007Interval staging and held-out-window filtering of external datanotes/official/来件/virtualembryo.ai/rules.md
k026Canonicalise predicted 3D coordinates before submissionnotes/pitfalls/04_scorer_invariance.md
k023Time-split validation with a held-out intermediate or next time pointnotes/handover/02_知识学习路线.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么在父 mix 的 align_pair 之后、混合之前,把全局单一 scale_to_rms 替换为按细胞类型缩放(每型质心不动、型内 RMS 做 log 插值后缩放),只对两端共有且两侧 >=20 细胞的类型生效;提交态 SIDE=b(仅缩放较晚端点),表达完全不动。
各组分数的变化cell_state:不变(+0.00,66.70)——符合设计
expression_change:不变(+0.00,63.90)——符合设计,表达未触碰
local_spatial:噪声内:+0.38(54.03→54.41)
shape_scale:变好方向但接近噪声:+0.99(53.37→54.36),T2 噪声约 1 分,在噪声边缘
family_idT2HI-05
假设是否成立unclear
经验
  1. 只改坐标不动表达的机制(按类型缩放型内 RMS)能保证 expression_change/cell_state 两组严格不变(+0.00),是零风险改动的正确形态,避免 node 3/4 改一处伤多组的净负结局。
  2. 榜分 +0.34 小于 T2 噪声(~1 分),即使方向与 proxy 双 seed 一致(+0.19/+0.30)也不能宣称有效;shape_scale +0.99 也只是在噪声边缘。
  3. 机制关闭对照(T2HI05_PERTYPE=0)与父节点逐位一致且开启后输出确实不同(s_k 跨型离散,如 V-CM s_b=0.614 vs 全局 g_b=1.034),证明机制生效但收益量级太小:全局 RMS 差异本来就小(354→335),可改的空间有限。
  4. SIDE=b 优于 SIDE=a/both,提示收益可能主要来自收缩较晚端点中个别离散过大的类型(E9.5 的 V-CM),而非普遍的按类型尺度演化——机制的大部分作用可能集中于极少数极端类型。
  5. 在 proxy 上所有共有类型 >=457 细胞时,MIN_CELLS 在 [10,50] 内不改变输出,此类阈值参数不必消耗查分额度去扫。
mechanism_active是
下一步建议
  1. 针对 shape_scale:只对 |s_k - g| 超阈值的极端类型(如 V-CM)做按类型缩放、其余保持全局,验证收益是否全部来自极端型(proxy 上先 diag 再查分,双 seed 确认)。
  2. 针对 shape_scale/local_spatial:在 SIDE=b 基础上加大收缩力度,如 target_k 用 log_interp 加 damp<1 或对 b 侧极端型直接缩到全局 target 的 γ 倍,看 +0.99 能否放大到超出噪声。
  3. final 括号有 31 个共有类型且两端更近,先跑 diag 检查是否存在 V-CM 式极端型间尺度差再决定是否沿用 SIDE=b;若 final 上极端型不明显,此方向收益预期 <1 分,应转向其他组(expression_change 63.90 仍是短板)。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。3
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 3 次:bash 1、read 1、write 1
用时不到 1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 14,361 · 输出 992 · 思考 322
任务(第一行)分析节点 n5 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
ANALYSIS.json written
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/analyst.jsonl 11 KB
/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/analyst.stderr

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数24
工具调用共 26 次:bash 20、edit 3、write 2、read 1
用时10 分
token 数输入 56,116 · 输出 9,741 · 思考 8,975
任务(第一行)按计划修改程序:per-type RMS scaling (T2HI-05): 按类型插值尺度替代全局单一 target_rms
最后的回答(摘录)
**做了什么**:实现 PLAN 的 T2HI-05——在父 mix 的对齐之后,把全局单一 target_rms 缩放替换为按细胞类型插值型内 RMS 的缩放(质心不动、表达不动),并试了 SIDE∈{a,b,both}、γ∈{0.5,1} 变体;提交态为机制开启、SIDE=b(只缩放较晚端点)。

**查分结果**(proxy A 半,8 次额度):关闭态与父逐位一致,基线 59.07(s0)/ 59.21(s1);SIDE=b 59.26 / 59.52(+0.19 / +0.30),shape_scale +0.4~+1.0、local_spatial +0.2~+0.4,两个 seed 同向;SIDE=a 反而 −0.11。vec-check 通过,运行 ~2.5s。

**下一步最值得试**:改善幅度在噪声边缘且 SIDE=b 优于 both,提示收益主要来自收缩 E9.5 端 V-CM 的过大离散(s_b=0.61)——可在 final 括号(31 个共有类型)验证是否仍成立;另一个方向是只对 |s_k−g| 超阈值的极端类型做按类型缩放(其余保持全局),或对型内 RMS 用稳健估计(去质心后的中位距离)以降低小类型噪声。
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/engineer.jsonl 149 KB
/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/engineer.stderr

调研员

角色调研员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数1
工具调用共 0 次:—
用时2 分
token 数输入 9,886 · 输出 1,197 · 思考 5,659
任务(第一行)为节点 n5(改进,父节点 n4)写改进计划 PLAN.json
最后的回答(摘录)
al_spatial 54.03 最弱。node 3(质心位移)和 node 4(配对插值)分别改了坐标位置和表达,均因伤害其它组而净负。两者都未触碰每型尺度:当前管线把两阶段全部细胞缩放到同一个全局 target_rms(log_interp(rms_a_global, rms_b_global, t)),忽略了各型自身尺度趋势可能不同(心脏尺度非单调,T2HI-05 方向)。按类型分别插值 RMS 只改变型内离散度,不移动质心、不合成细胞、不改表达,因此不会重蹈 node 3/4 的覆辙。",
  "approach": "在父节点 run.py 的 align_pair 之后、mix 之前,将全局 scale_to_rms 替换为按类型缩放:\n1. 对 aligned_a、aligned_b 分别按标签分组,计算每型质心 c_k 和型内 RMS r_k(细胞到自身质心的均方根距离)。\n2. 对两端共有且两侧均 ≥ MIN_CELLS(默认 20)的类型:target_rms_k = log_interp(r_a_k, r_b_k, t, damp);缩放因子 s_k = target_rms_k / r_k;新坐标 = c_k + s_k·(coord − c_k)。对仅在一端出现的类型或细胞数 < MIN_CELLS 的类型,仍用全局 target_rms 缩放(与父节点一致)。\n3. 缩放后按父节点原有流程做 mix(组成插值、n 的 log-linear 插值、jitter 去重叠)。\n4. 关键参数:MIN_CELLS 初值 20,搜索范围 [10, 50];damp 沿用 1.0。\n5. 单输入阶段退路(b=None):无配对阶段,直接输出 a 的细胞,与父节点一致。\n6. vec-score 快筛:① 先跑 T2HI05_PERTYPE=0(关闭),确认与父输出逐位一致(1 次);② 开启、MIN_CELLS=20、seed 0 查 proxy(1 次);③ 若 ≥59.5 再查 seed 1 确认(1 次);④ 若 <59.5 试 MIN_CELLS=10(1 次)。共 ≤6 次,留余量。",
  "expected_groups": ["shape_scale", "local_spatial"],
  "risks": "1) 小类型(<20 细胞)的型内 RMS 估计噪声大,反而引入随机缩放——用 MIN_CELLS 阈值兜底,Engineer 应输出每型 cell count 和 r_a_k、r_b_k、target_rms_k 的 diag 表,检查是否有型的 target 偏离全局值过远(>2×)。2) 全局 RMS 差异本身很小(354→335),型间差异可能也小,净提升 <1 分噪声——用 seed 0+1 双查确认,若两次均 <59.5 则放弃。3) 缩放改变型内坐标可能轻微影响 local_spatial 的邻居结构——若 local_spatial 降 >0.5 而 shape_scale 升,需权衡;若两者均无变化则机制可能无效。",
  "family_id": "T2HI-05",
  "mechanism": "将全局单一 target_rms 缩放替换为按细胞类型分别插值型内 RMS 再缩放,使每型的空间离散度按自身两端尺度趋势独立演化,而非被全局均值抹平。",
  "vs_constant_shift": "常数位移对所有细胞施加同一缩放因子,不区分类型;本方案按类型计算各自的尺度轨迹并分别缩放,型间缩放因子不同,且缩放中心是各型自身质心而非全局原点,因此改变的是型内离散度的相对比例,而非整体平移。",
  "mechanism_evidence": "Engineer 在 diag 中输出:(a) 每型的 r_a_k、r_b_k、target_rms_k 和缩放因子 s_k,验证 s_k 跨型存在离散(若所有 s_k 相同则机制退化为全局缩放);(b) 输出点云中每型的实际 RMS 与 target_rms_k 的吻合度;(c) 四组分各自的变化方向,重点看 shape_scale 和 local_spatial 是否同向改善。",
  "mechanism_off_control": "环境变量 T2HI05_PERTYPE=0 时,所有类型统一使用全局 target_rms(即父节点原有逻辑),输出应与父节点 mix 在 proxy seed 0 上逐位一致(.X 和 obsm 逐元素相等)。预期差别:关闭态四组分与父节点完全相同(59.50 / 66.70 / 63.90 / 54.03 / 53.37)。",
  "sources": []
}
```
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/researcher.jsonl 5 KB
/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/researcher.stderr

审查员

角色审查员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数10
工具调用共 13 次:read 6、bash 5、grep 1、write 1
用时2 分
token 数输入 23,861 · 输出 2,251 · 思考 5,099
任务(第一行)审查节点 n5 的程序是否越界读取、写死目标、钻评分器漏洞
最后的回答(摘录)
REVIEW.json written
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/reviewer.jsonl 90 KB
/home/spark-longxinyang/vec/runs/formal/20261002-204523-search-t2-heart-interp-g24q/nodes/5/reviewer.stderr