总览 · ← 返回运行 20261003-043412-search-t2-heart-interp-g24-D-s1
节点 n7 在终选来历上
mix 组成不变,把每类型内的 i.i.d. 随机抽样换成空间连续簇抽样(k-NN 贪心区域生长,k=25、每型约 10 簇),输出继承源阶段局部邻域结构;A 半 62.4,local_spatial 53.5→64.6。
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261003-043412-search-t2-heart-interp-g24-D-s1 |
|---|---|
| 父节点 | n5 |
| 子节点 | n8 |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 改进 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 62.61(+3.1) · proxy 62.61(+3.1) · 3 次复测均分 61.92 |
| 审查 | 通过 检查1(越界读取):未发现问题。run.py 只经 src.task2_spatial.view_io 从 --data 视图目录读取(load_manifest/panel_genes/read_stage),cluster_sample.py 仅用 numpy/scipy,无绝对路径、'..'、/mnt、data/raw、评分器路径,无联网代码(grep 全文无 requests/urllib/http)。; 检查2(硬编码目标统计量):未发现问题。代码中唯一的常量是方法超参 PARAMS={align, scale_damp} 和 CLUSTER_PARAMS={k=25, n_clu… |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 13 分 |
| 程序版本 | 6739569794efe89014a2b87d8b1dcc3f4b1521ce (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git 6739569794:solution/METHOD.md
mix 组成不变,把每类型内的 i.i.d. 随机抽样换成空间连续簇抽样(k-NN 贪心区域生长,k=25、每型约 10 簇),输出继承源阶段局部邻域结构;A 半 62.4,local_spatial 53.5→64.6。
方法
- 路径与父节点 mix(node 2)一致:procrustes 对齐(z 保持切片轴)、log 线性 RMS 缩放(scale_damp=1)、
_limits定细胞数、round(t*n)从后侧阶段取、每类型配额用与stratified_choice完全相同的确定性分配算术(组成不变)。父节点的 OT 位移已删除(node 5 证明单调有害)。 - 机制(
CLUSTER_SAMPLING=True,solution/cluster_sample.py):在每个类型内,用scipy.spatial.cKDTree在对齐坐标上建 k-NN 图,贪心区域生长成簇——随机选种子,用最小堆反复弹出距当前簇最近的未选邻居加入,直到簇大小达到 ceil(配额/n_clusters);重复开新簇直到填满配额。只改变"选哪些细胞",表达和坐标仍来自真实细胞。 - 参数:
k=25, n_clusters=10(代理上搜索所得)。单输入阶段(b 为 None)走原有单阶段回退,不受影响。 - 确定性:只用视图数据 +
np.random.default_rng(seed);seed 0 重跑输出逐位相同(本地验证)。无绝对时间、无视图路径依赖(只用 t 与时间差)。
机制生效证据(proxy,E8.25_late+E9.5→E8.75,t=0.4,seed 0)
- 输出细胞的平均 5-NN 距离:随机 mix 21.1 → 簇抽样 14.2(−33%),局部邻域确实更连续。
- 关闭对照(
CLUSTER_SAMPLING=False,同一程序走库interpolate,即父节点 mix 路径):A 半 59.07,local_spatial 53.46;开启:62.37,local_spatial 64.61。四组分变化符合 PLAN 预期——local_spatial +11.2、neighborhood_mmd 0.083→0.053,expression_change +0.8、cell_state +0.9(同一批真实细胞、组成相同,仅集合不同带来的小幅波动),shape_scale +0.5(occupancy_dice 0.834→0.828 略降,被 d2_shape 改善抵消)。
查分(vec-score A 半,proxy;对照=父 mix 路径 59.07)
| 配置 (k, n_clusters, cluster_frac) | skill | expr | state | shape | local |
|---|---|---|---|---|---|
| (25, 10, 1.0) seed0(提交) | 62.37 | 64.6 | 66.8 | 53.5 | 64.6 |
| (25, 10) seed1 / seed2 | 61.66 / 62.31 | ||||
| (15, 10) | 62.25 | 64.8 | 66.6 | 53.5 | 64.1 |
| (8, 15) | 62.06 | 64.1 | 66.5 | 53.8 | 63.9 |
| (10, 10) | 62.00 | 64.3 | 66.9 | 52.9 | 63.9 |
| (15, 20) | 61.90 | 63.4 | 66.5 | 54.4 | 63.3 |
| (25, 7) | 61.72 | 65.0 | 66.8 | 49.8 | 65.3 |
| (15, 5) | 61.80 | 64.1 | 66.2 | 52.3 | 64.7 |
| (25, 15) | 61.41 | 64.1 | 66.2 | 52.1 | 63.4 |
| (25, 10, frac 0.85) | 62.05 | 64.4 | 66.4 | 54.0 | 63.4 |
| (25, 10, frac 0.7) | 60.87 | 63.9 | 65.8 | 52.6 | 61.2 |
| (15, 5) / (10, 5) / (40, 10) | 61.80 / 62.00* / 62.37* |
*k=40 与 k=25 输出逐位相同(生长顺序对 k 饱和);(10,5) 未单独查分,(10,10)=62.00。
- 趋势:n_clusters=10 附近是峰(5 与 15 都更低);k≥25 饱和;部分随机回填(cluster_frac<1)单调稀释 local_spatial 收益,不取。
- 提交配置 3 种子 A 半均值 62.11,比对照 59.07 高 +3.0,超出 T2 约 1 分噪声。
未验证 / 风险
- 只在代理括号(E8.25_late↔E9.5,5 个共有类型)上验证;final 括号(E8.25+E8.75,31 个共有类型、t=0.5)上细胞更多、类型更细,簇大小=配额/10 的绝对值会变,但机制不依赖跨阶段对齐(只用单侧坐标建图),预期方向一致。
- occupancy_dice 略降(簇间留空洞);若 final 上 shape_scale 受损明显,可增大 n_clusters 或用 frac 0.85 折中(代理上 frac 0.85 = 62.05,shape 54.0 最高)。
- 知识来源:无外部生物知识;仅用视图内坐标与标签。
调研员的计划
| 名称 | 空间连续簇抽样替代随机抽样(T2HI-04 保邻域) |
|---|---|
| 动机 | 父节点 5 的 OT 位移已证实单调有害(γ→0 收敛回父分),四组分全在噪声内。最弱组 local_spatial 53.85、shape_scale 53.62。当前 mix 的分层抽样对每个类型做独立随机抽取,打散了源阶段的空间邻域结构,导致输出中同类型细胞散布、局部密度不一致,是 local_spatial 偏低的结构原因。节点 4 已证明对齐改进可提升 shape_scale(+1.3),但抽样策略尚未被改进。 |
| 做法 | 步骤:1) 设 OT_ENABLED=False,回退到已验证的 mix 路径(bit-identical)。2) 新增 solution/cluster_sample.py:对每个类型,用 scipy.spatial.cKDTree 在已对齐坐标上建 k-NN(k=15);用贪心扩展法生成空间连续簇——随机选种子细胞,迭代加入当前簇边界的最近未选邻居,直到簇大小达到该类型需抽样数的 1/5(即约 5 个簇覆盖配额);从这些簇中取全部细胞填满配额(若簇总大小超过配额则截断最远边界细胞)。3) 替换 run.py 中 mix_indices 的调用为新的 cluster_mix_indices,保持每类型抽样总数不变(组成不变),仅改变选哪些细胞。4) 单输入阶段(b is None)走原有单阶段回退,不受影响。5) 参数:k=15(搜索 10-25),簇目标大小=配额/5(搜索 /3 到 /8)。先用 proxy 查 1 次(A 半),若 local_spatial ≥55 且 cell_state 不降,再查第 2 次确认。预计实现+跑分 <20 min。 |
| 风险 | 1) 若评分器对局部邻域的度量不敏感于抽样方式(例如只看全局密度),改善可能为零——Engineer 第一次查分即可发现(local_spatial 无变化)。2) 簇太小(/8)退化为随机抽样,簇太大(/3)丧失多样性——用参数搜索覆盖。3) 代理括号仅 5 个共有类型、跨度大,空间结构本身可能不可靠;但本机制不依赖跨阶段对齐质量(只用单侧坐标建图),风险低于 OT。4) 耗时增加(KDTree 建树),但细胞数 <5000 时 <1s,在时限内。 |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:父节点版本 488f9376ae。改动的文件:solution/EXECUTION.json +1 −0、solution/METHOD.md +32 −23、solution/cluster_sample.py +120 −0、solution/ot_interp.py +0 −144、solution/run.py +27 −26
diff --git a/solution/EXECUTION.json b/solution/EXECUTION.jsonnew file mode 100644index 0000000..6d8012e--- /dev/null+++ b/solution/EXECUTION.json@@ -0,0 +1 @@+{"gpu": false}\ No newline at end of filediff --git a/solution/METHOD.md b/solution/METHOD.mdindex d4f73a6..5e88a3e 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,32 +1,41 @@-分型 Sinkhorn OT 重心投影插值(T2HI-02):mix 组成不变,共有类型细胞沿表达匹配的传输路径按 gamma·t 向对侧重心位移坐标;机制在代理上单调伤 shape/cell_state,最终取最温和配置 coords+gamma0.25,仍略低于父节点 mix。+mix 组成不变,把每类型内的 i.i.d. 随机抽样换成空间连续簇抽样(k-NN 贪心区域生长,k=25、每型约 10 簇),输出继承源阶段局部邻域结构;A 半 62.4,local_spatial 53.5→64.6。 ## 方法 -- 加载、procrustes 对齐(z 保持切片轴)、log 线性 RMS 缩放、细胞数与分层抽样组成全部与父节点 2(mix)逐位一致(同一 rng 序列,`mix_indices`)。-- 机制(`OT_ENABLED=True`,`solution/ot_interp.py`):对每个在两侧括号阶段都出现、且两侧各 ≥5 个细胞的类型,用表达上的平方欧氏代价跑熵正则 Sinkhorn(eps=0.05×median(C),60 次迭代,均匀边际,每侧上限 2500 细胞,其余细胞按表达最近邻借用子集的耦合行)。每个被抽中的输出细胞向其对侧耦合行的重心投影位移:coords_out=(1−gamma·t)·own+gamma·t·bary(side a;side b 对称)。-- 单侧独有类型原样通过(与父节点一致)。无括号(b 为 None)时走父节点的单阶段回退。-- 最终配置 `OT_PARAMS={eps_frac:0.05, iters:60, cap:2500, min_per_type:5, mode:"coords", gamma:0.25}`:只位移坐标,表达保留真实细胞(mode="both" 会把 cell_state 从 66.7 拉到 63-64,表达混合是主要伤害源)。+- 路径与父节点 mix(node 2)一致:procrustes 对齐(z 保持切片轴)、log 线性 RMS 缩放(scale_damp=1)、`_limits` 定细胞数、`round(t*n)` 从后侧阶段取、每类型配额用与 `stratified_choice` 完全相同的确定性分配算术(组成不变)。父节点的 OT 位移已删除(node 5 证明单调有害)。+- 机制(`CLUSTER_SAMPLING=True`,`solution/cluster_sample.py`):在每个类型内,用 `scipy.spatial.cKDTree` 在对齐坐标上建 k-NN 图,贪心区域生长成簇——随机选种子,用最小堆反复弹出距当前簇最近的未选邻居加入,直到簇大小达到 ceil(配额/n_clusters);重复开新簇直到填满配额。只改变"选哪些细胞",表达和坐标仍来自真实细胞。+- 参数:`k=25, n_clusters=10`(代理上搜索所得)。单输入阶段(b 为 None)走原有单阶段回退,不受影响。+- 确定性:只用视图数据 + `np.random.default_rng(seed)`;seed 0 重跑输出逐位相同(本地验证)。无绝对时间、无视图路径依赖(只用 t 与时间差)。 ## 机制生效证据(proxy,E8.25_late+E9.5→E8.75,t=0.4,seed 0) -- 3553 个细胞(5 个共有类型 × 两侧)被位移;99.7% 的细胞位移向量与其类型平均位移之差 >10% 位移 RMS —— 是逐细胞位移,不是常数平移。-- 对照(`OT_ENABLED=False`,走父节点 `interpolate` 原路径):输出 X 与坐标与父节点程序逐位相同(本地 diff 验证),共享加载/对齐路径无 bug。+- 输出细胞的平均 5-NN 距离:随机 mix 21.1 → 簇抽样 14.2(−33%),局部邻域确实更连续。+- 关闭对照(`CLUSTER_SAMPLING=False`,同一程序走库 `interpolate`,即父节点 mix 路径):A 半 59.07,local_spatial 53.46;开启:62.37,local_spatial 64.61。四组分变化符合 PLAN 预期——local_spatial +11.2、neighborhood_mmd 0.083→0.053,expression_change +0.8、cell_state +0.9(同一批真实细胞、组成相同,仅集合不同带来的小幅波动),shape_scale +0.5(occupancy_dice 0.834→0.828 略降,被 d2_shape 改善抵消)。 -## 查分(vec-score A 半,proxy;父节点 A 半 59.12)+## 查分(vec-score A 半,proxy;对照=父 mix 路径 59.07) -| 配置 | skill | expr | state | shape | local |+| 配置 (k, n_clusters, cluster_frac) | skill | expr | state | shape | local | |---|---:|---:|---:|---:|---:|-| both, gamma=1, eps=.05 | 56.38 | 63.9 | 63.1 | 47.3 | 51.3 |-| coords, gamma=1 | 57.55 | 63.8 | 65.9 | 47.3 | 53.2 |-| both, gamma=0.5 | 57.57 | 64.2 | 64.1 | 50.7 | 51.3 |-| both, gamma=0.25 | 58.49 | 64.0 | 64.8 | 52.5 | 52.6 |-| **coords, gamma=0.25(提交)** | **58.90** | 63.8 | 65.9 | 52.5 | 53.3 |-| coords, gamma=0.25, eps=.01 | 58.86 | 63.8 | 65.9 | 52.5 | 53.2 |--## 结论与未验证--- 机制在此代理括号上单调有害:任何 gamma>0 的位移都拉低 occupancy_dice / shape_scale 和 cell_state,gamma→0 单调收敛回父节点分。原因推测:E8.25↔E9.5 跨度大、只有 5 个共有类型参与对齐,procrustes 残差大,重心位移把细胞拉进错误占位区。-- 提交版本机制打开(coords, gamma=0.25),A 半 58.90,与父节点差 −0.2,在噪声(约 2 分)以内;如实报告:本方法族在代理上没有超过父节点 mix。-- 未验证:final 括号(E8.25+E8.75,31 个共有类型、对齐更准、t=0.5)上机制可能没那么有害——代理的结论不一定迁移,但本节点无 final 视图可测。-- 表达插值(mode="both"/"expr")已验证为负收益,不建议后续再试。-- 生物学知识来源:无外部数据/prior 使用;只用了「同一细胞类型在两阶段间连续变形」这一通用发育连续性假设。+| (25, 10, 1.0) seed0(提交) | **62.37** | 64.6 | 66.8 | 53.5 | 64.6 |+| (25, 10) seed1 / seed2 | 61.66 / 62.31 | | | | |+| (15, 10) | 62.25 | 64.8 | 66.6 | 53.5 | 64.1 |+| (8, 15) | 62.06 | 64.1 | 66.5 | 53.8 | 63.9 |+| (10, 10) | 62.00 | 64.3 | 66.9 | 52.9 | 63.9 |+| (15, 20) | 61.90 | 63.4 | 66.5 | 54.4 | 63.3 |+| (25, 7) | 61.72 | 65.0 | 66.8 | 49.8 | 65.3 |+| (15, 5) | 61.80 | 64.1 | 66.2 | 52.3 | 64.7 |+| (25, 15) | 61.41 | 64.1 | 66.2 | 52.1 | 63.4 |+| (25, 10, frac 0.85) | 62.05 | 64.4 | 66.4 | 54.0 | 63.4 |+| (25, 10, frac 0.7) | 60.87 | 63.9 | 65.8 | 52.6 | 61.2 |+| (15, 5) / (10, 5) / (40, 10) | 61.80 / 62.00* / 62.37* | | | | |++*k=40 与 k=25 输出逐位相同(生长顺序对 k 饱和);(10,5) 未单独查分,(10,10)=62.00。++- 趋势:n_clusters=10 附近是峰(5 与 15 都更低);k≥25 饱和;部分随机回填(cluster_frac<1)单调稀释 local_spatial 收益,不取。+- 提交配置 3 种子 A 半均值 62.11,比对照 59.07 高 +3.0,超出 T2 约 1 分噪声。++## 未验证 / 风险++- 只在代理括号(E8.25_late↔E9.5,5 个共有类型)上验证;final 括号(E8.25+E8.75,31 个共有类型、t=0.5)上细胞更多、类型更细,簇大小=配额/10 的绝对值会变,但机制不依赖跨阶段对齐(只用单侧坐标建图),预期方向一致。+- occupancy_dice 略降(簇间留空洞);若 final 上 shape_scale 受损明显,可增大 n_clusters 或用 frac 0.85 折中(代理上 frac 0.85 = 62.05,shape 54.0 最高)。+- 知识来源:无外部生物知识;仅用视图内坐标与标签。diff --git a/solution/cluster_sample.py b/solution/cluster_sample.pynew file mode 100644index 0000000..270d1de--- /dev/null+++ b/solution/cluster_sample.py@@ -0,0 +1,120 @@+"""Spatially contiguous cluster sampling (T2HI-04).++Same per-type quotas as ``stratified_choice`` (deterministic allocation), but+within each type the cells are drawn as greedy k-NN region-growing clusters+instead of i.i.d. random picks, so the output inherits the local neighbourhood+structure of the source stage.+"""++from __future__ import annotations++import heapq++import numpy as np+from scipy.spatial import cKDTree+++def type_alloc(labels: np.ndarray, n: int):+ """Per-type quotas, identical arithmetic to transport.stratified_choice."""+ labels = np.asarray(labels).astype(str)+ n = int(min(max(n, 1), len(labels)))+ if n >= len(labels):+ return None, None, None+ types, counts = np.unique(labels, return_counts=True)+ raw = counts / counts.sum() * n+ alloc = np.floor(raw).astype(int)+ rem = int(n - alloc.sum())+ order = np.argsort(-(raw - alloc))+ for i in range(rem):+ alloc[order[i % len(order)]] += 1+ alloc = np.minimum(alloc, counts)+ deficit = int(n - alloc.sum())+ if deficit > 0:+ spare = counts - alloc+ for i in np.argsort(-spare):+ take = int(min(deficit, spare[i]))+ alloc[i] += take+ deficit -= take+ if deficit == 0:+ break+ return types, counts, alloc+++def grow_clusters(coords: np.ndarray, q: int, rng: np.random.Generator,+ k: int = 15, n_clusters: int = 5) -> np.ndarray:+ """Select q of the m points as ~n_clusters contiguous region-growing clusters."""+ m = int(coords.shape[0])+ if q >= m:+ return np.arange(m)+ coords = np.asarray(coords, dtype=np.float64)+ tree = cKDTree(coords)+ kk = int(min(max(k, 1), m - 1))+ nbrs = tree.query(coords, k=kk + 1)[1][:, 1:]+ cluster_target = max(1, int(np.ceil(q / max(1, int(n_clusters)))))+ selected = np.zeros(m, dtype=bool)+ order: list[int] = []+ n_sel = 0+ while n_sel < q:+ avail = np.flatnonzero(~selected)+ seed = int(avail[int(rng.integers(len(avail)))])+ heap = [(0.0, seed)]+ size = 0+ while heap and size < cluster_target and n_sel < q:+ d, c = heapq.heappop(heap)+ if selected[c]:+ continue+ selected[c] = True+ order.append(c)+ size += 1+ n_sel += 1+ pc = coords[c]+ for nb in nbrs[c]:+ nb = int(nb)+ if not selected[nb]:+ diff = pc - coords[nb]+ heapq.heappush(heap, (float(np.dot(diff, diff)), nb))+ return np.asarray(order[:q], dtype=int)+++def cluster_take(labels: np.ndarray, coords: np.ndarray, n: int,+ rng: np.random.Generator, k: int = 15, n_clusters: int = 5,+ cluster_frac: float = 1.0) -> np.ndarray:+ labels = np.asarray(labels).astype(str)+ n = int(min(max(n, 1), len(labels)))+ if n >= len(labels):+ return np.arange(len(labels))+ types, _counts, alloc = type_alloc(labels, n)+ picks = []+ for lab, kk in zip(types, alloc):+ if kk <= 0:+ continue+ idx = np.flatnonzero(labels == lab)+ q_cl = int(np.clip(round(int(kk) * float(cluster_frac)), 0, int(kk)))+ sel = grow_clusters(coords[idx], q_cl, rng, k=k, n_clusters=n_clusters)+ chosen = idx[sel]+ if chosen.size < int(kk):+ rest = np.setdiff1d(idx, chosen, assume_unique=False)+ extra = rng.choice(rest, int(kk) - chosen.size, replace=False)+ chosen = np.concatenate([chosen, extra])+ picks.append(chosen)+ if not picks:+ return np.arange(n)+ out = np.concatenate(picks)+ if out.size > n:+ out = out[:n]+ return out+++def cluster_mix_indices(labels_a, labels_b, coords_a, coords_b, t: float, n: int,+ rng: np.random.Generator, k: int = 15, n_clusters: int = 5,+ cluster_frac: float = 1.0):+ """Drop-in replacement for sample.mix_indices using cluster sampling."""+ n_b = int(round(float(t) * n))+ n_b = int(np.clip(n_b, 0, n))+ n_a = n - n_b+ if n_a == 0:+ return np.array([], dtype=int), cluster_take(labels_b, coords_b, n_b, rng, k, n_clusters, cluster_frac)+ if n_b == 0:+ return cluster_take(labels_a, coords_a, n_a, rng, k, n_clusters, cluster_frac), np.array([], dtype=int)+ return (cluster_take(labels_a, coords_a, n_a, rng, k, n_clusters, cluster_frac),+ cluster_take(labels_b, coords_b, n_b, rng, k, n_clusters, cluster_frac))diff --git a/solution/ot_interp.py b/solution/ot_interp.pydeleted file mode 100644index 5b2f083..0000000--- a/solution/ot_interp.py+++ /dev/null@@ -1,144 +0,0 @@-"""Per-type entropy-regularised OT (Sinkhorn) barycentric interpolation.--For every cell type present in both bracketing stages, cells of the two sides-are matched by an entropy-regularised transport plan computed on squared-Euclidean expression cost. Each output cell keeps its own side's value with-weight (1-t) / t and moves the rest of the way toward the barycentric-projection of its coupling row, in expression and in coordinates. Types that-exist on only one side (or with too few cells) pass through unchanged, exactly-as in the parent ``mix`` method.-"""--from __future__ import annotations--import numpy as np-from scipy.spatial.distance import cdist--EPS_FRAC = 0.05-ITERS = 60-CAP = 2500-MIN_PER_TYPE = 5---def _sinkhorn_weights(C: np.ndarray, eps: float, iters: int) -> np.ndarray:- """Row-normalised coupling rows W (n_a x n_b), W @ 1 = 1, uniform marginals."""- na, nb = C.shape- K = np.exp(-C / eps)- K += 1e-300- u = np.full(na, 1.0 / na)- v = np.full(nb, 1.0 / nb)- for _ in range(iters):- u = 1.0 / (na * (K @ v) + 1e-300)- v = 1.0 / (nb * (K.T @ u) + 1e-300)- P = u[:, None] * K * v[None, :]- rs = P.sum(axis=1, keepdims=True)- rs[rs <= 0] = 1.0- return P / rs---def _bary_for_targets(X_side: np.ndarray, sub_idx: np.ndarray, W: np.ndarray,- targets_local: np.ndarray) -> np.ndarray:- """Coupling rows (as weight matrices) for target cells of this type.-- ``targets_local`` are positions within the full per-type side matrix. Cells- in the OT subset use their own row; the rest borrow the row of their- nearest subset cell in expression space.- """- pos = -np.ones(X_side.shape[0], dtype=int)- pos[sub_idx] = np.arange(sub_idx.size)- rows = np.empty((targets_local.size, W.shape[1]))- need = pos[targets_local] < 0- have = ~need- if have.any():- rows[have] = W[pos[targets_local[have]]]- if need.any():- d = cdist(X_side[targets_local[need]], X_side[sub_idx], metric="sqeuclidean")- rows[need] = W[d.argmin(axis=1)]- return rows---def ot_interpolate(xa_all: np.ndarray, xb_all: np.ndarray,- ca: np.ndarray, cb: np.ndarray,- labels_a: np.ndarray, labels_b: np.ndarray,- ia: np.ndarray, ib: np.ndarray, t: float,- rng: np.random.Generator,- eps_frac: float = EPS_FRAC, iters: int = ITERS,- cap: int = CAP, min_per_type: int = MIN_PER_TYPE,- mode: str = "both", gamma: float = 1.0):- """Interpolate sampled cells (ia from side a, ib from side b).-- ``xa_all`` / ``xb_all`` are dense expression matrices (cells x genes),- ``ca`` / ``cb`` the aligned, RMS-scaled coordinates of both full stages.- Returns (expr, coords, evidence dict).- """- expr_out = np.empty((ia.size + ib.size, xa_all.shape[1]), dtype=np.float64)- coords_out = np.empty((ia.size + ib.size, 3), dtype=np.float64)- expr_out[:ia.size] = xa_all[ia]- expr_out[ia.size:] = xb_all[ib]- coords_out[:ia.size] = ca[ia]- coords_out[ia.size:] = cb[ib]-- lab_a = np.asarray(labels_a).astype(str)- lab_b = np.asarray(labels_b).astype(str)- shared = sorted(set(lab_a[ia]) & set(lab_b[ib]))- n_moved = 0- disp_dev = []- for typ in shared:- rows_a = np.flatnonzero(lab_a == typ)- rows_b = np.flatnonzero(lab_b == typ)- if rows_a.size < min_per_type or rows_b.size < min_per_type:- continue- sub_a = rows_a if rows_a.size <= cap else np.sort(rng.choice(rows_a, cap, replace=False))- sub_b = rows_b if rows_b.size <= cap else np.sort(rng.choice(rows_b, cap, replace=False))- Xa = xa_all[sub_a]- Xb = xb_all[sub_b]- C = cdist(Xa, Xb, metric="sqeuclidean")- med = float(np.median(C))- eps = max(eps_frac * med, 1e-12)- W = _sinkhorn_weights(C, eps, iters)-- # positions of sampled side-a cells of this type inside ia- sel_a = np.flatnonzero(lab_a[ia] == typ)- sel_b = np.flatnonzero(lab_b[ib] == typ)- # local indices within the per-type full row arrays- loc_a = np.searchsorted(rows_a, ia[sel_a])- loc_b = np.searchsorted(rows_b, ib[sel_b])- suba_loc = np.searchsorted(rows_a, sub_a)- subb_loc = np.searchsorted(rows_b, sub_b)-- Wa = _bary_for_targets(xa_all[rows_a], suba_loc, W, loc_a)- Wb = _bary_for_targets(xb_all[rows_b], subb_loc, W.T.copy(), loc_b)-- bary_x_a = Wa @ xb_all[sub_b]- bary_p_a = Wa @ cb[sub_b]- bary_x_b = Wb @ xa_all[sub_a]- bary_p_b = Wb @ ca[sub_a]-- te = gamma * t if mode in ("both", "expr") else 0.0- tc = gamma * t if mode in ("both", "coords") else 0.0- if sel_a.size:- expr_out[sel_a] = (1.0 - te) * xa_all[ia[sel_a]] + te * bary_x_a- coords_out[sel_a] = (1.0 - tc) * ca[ia[sel_a]] + tc * bary_p_a- d = coords_out[sel_a] - ca[ia[sel_a]]- dm = d.mean(axis=0)- rms = float(np.sqrt((d * d).sum(axis=1).mean())) + 1e-9- disp_dev.append(float(np.mean(np.sqrt(((d - dm) ** 2).sum(axis=1)) > 0.1 * rms)))- n_moved += sel_a.size- if sel_b.size:- k = sel_b + ia.size- ue = gamma * (1.0 - t) if mode in ("both", "expr") else 0.0- uc = gamma * (1.0 - t) if mode in ("both", "coords") else 0.0- expr_out[k] = ue * bary_x_b + (1.0 - ue) * xb_all[ib[sel_b]]- coords_out[k] = uc * bary_p_b + (1.0 - uc) * cb[ib[sel_b]]- d = coords_out[k] - cb[ib[sel_b]]- dm = d.mean(axis=0)- rms = float(np.sqrt((d * d).sum(axis=1).mean())) + 1e-9- disp_dev.append(float(np.mean(np.sqrt(((d - dm) ** 2).sum(axis=1)) > 0.1 * rms)))- n_moved += sel_b.size-- ev = {- "n_shared_types_moved": len(disp_dev),- "n_cells_moved": int(n_moved),- "frac_disp_dev_from_typemean": float(np.mean(disp_dev)) if disp_dev else 0.0,- }- return expr_out, coords_out, evdiff --git a/solution/run.py b/solution/run.pyindex 5b07816..dcdeaf7 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,5 +1,5 @@ #!/usr/bin/env python3-"""Per-type Sinkhorn OT barycentric interpolation (T2 interpolation).+"""Mix interpolation with spatially contiguous cluster sampling (T2HI-04). Brackets the target with the nearest inputs before and after it, puts both clouds in one frame (procrustes, z kept as slice axis), rescales both to the@@ -7,17 +7,14 @@ log-linear RMS exp(log r_a + scale_damp*t*(log r_b - log r_a)), and draws cells stratified by type: round(t*n) from the later stage, the rest from the earlier one (same composition as the parent ``mix`` method). -Mechanism (OT_ENABLED=True): for every cell type present on both sides, an-entropy-regularised transport plan (Sinkhorn, eps = 0.05*median squared-Euclidean expression cost, 60 iterations, uniform marginals, per-side cap-2500 cells with nearest-expression borrowing for the rest) matches cells-between the stages. Each sampled cell then moves to-(1-t)*own + t*barycentric projection of its coupling row, in expression AND-in coordinates, so output cells are genuine intermediates instead of a binary-mixture of endpoint cells. Types on one side only pass through unchanged.--Control (OT_ENABLED=False): identical code path to parent node 2 (``mix``),-bit-identical output.+Mechanism (CLUSTER_SAMPLING=True, solution/cluster_sample.py): within each+type the quota of cells is drawn as ~n_clusters spatially contiguous clusters+(greedy k-NN region growing on the aligned coordinates) instead of i.i.d.+random picks, so the output preserves the local neighbourhood structure of+the source stages. Expression and coordinates still come from real cells.++Control (CLUSTER_SAMPLING=False): identical code path to parent node 2+(``mix``), bit-identical output. """ from __future__ import annotations@@ -30,15 +27,15 @@ import numpy as np from src.task2_spatial.frame import log_interp, rms_radius, scale_to_rms, align_pair from src.task2_spatial.methods import _jitter, _limits, interpolate-from src.task2_spatial.sample import mix_indices, take+from src.task2_spatial.sample import take from src.task2_spatial.transport import as_dense from src.task2_spatial.view_io import board_params, interp_bracket, load_manifest, panel_genes, read_stage, write_t2 -from ot_interp import ot_interpolate+from cluster_sample import cluster_mix_indices, grow_clusters -OT_ENABLED = True+CLUSTER_SAMPLING = True PARAMS = {"align": "procrustes", "scale_damp": 1.0}-OT_PARAMS = {"eps_frac": 0.05, "iters": 60, "cap": 2500, "min_per_type": 5, "mode": "coords", "gamma": 0.25}+CLUSTER_PARAMS = {"k": 25, "n_clusters": 10} def main() -> None:@@ -61,9 +58,9 @@ def main() -> None: stage_b = read_stage(args.data, b, genes) params = board_params(manifest, "mix", PARAMS, args.seed) - if not OT_ENABLED:+ if not CLUSTER_SAMPLING: expr, coords, info = interpolate(stage_a, stage_b, t, params)- ev = {"ot": False}+ ev = {"cluster": False} else: t = float(t) rng = np.random.default_rng(int(params["seed"]))@@ -75,17 +72,21 @@ def main() -> None: ca = scale_to_rms(aligned_a, target_rms) cb = scale_to_rms(aligned_b, target_rms) n = _limits(params, stage_a.n, stage_b.n, t, "interp")- ia, ib = mix_indices(stage_a.labels, stage_b.labels, t, n, rng)- xa_all = np.asarray(stage_a.X.todense(), dtype=np.float64)- xb_all = np.asarray(stage_b.X.todense(), dtype=np.float64)- expr, coords, ev = ot_interpolate(- xa_all, xb_all, ca, cb, stage_a.labels, stage_b.labels, ia, ib, t, rng, **OT_PARAMS)- expr = np.clip(expr, 0.0, None).astype(np.float32)- coords = _jitter(coords, rng)+ ia, ib = cluster_mix_indices(+ stage_a.labels, stage_b.labels, ca, cb, t, n, rng, **CLUSTER_PARAMS)+ x_parts, p_parts = [], []+ if ia.size:+ x_parts.append(as_dense(stage_a.X, ia))+ p_parts.append(np.asarray(ca[ia], dtype=np.float64))+ if ib.size:+ x_parts.append(as_dense(stage_b.X, ib))+ p_parts.append(np.asarray(cb[ib], dtype=np.float64))+ expr = np.clip(np.vstack(x_parts), 0.0, None).astype(np.float32)+ coords = _jitter(np.vstack(p_parts), rng) coords = scale_to_rms(coords, target_rms).astype(np.float32) info.update(t=t, n=int(expr.shape[0]), rms_a=rms_a, rms_b=rms_b, out_rms=rms_radius(coords), n_from_a=int(ia.size), n_from_b=int(ib.size))- ev["ot"] = True+ ev = {"cluster": True, **{f"cl_{kk}": vv for kk, vv in CLUSTER_PARAMS.items()}} keep = {k: info.get(k) for k in ("t", "n", "rms_a", "rms_b", "out_rms", "n_shared_types", "z_dot", "z_flipped", "align")} print(json.dumps({"bracket": [a["stage"], b["stage"]], **keep, **ev}, default=float), file=sys.stderr)
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
用到的知识库条目
| 编号 | 标题 | 出处 |
|---|---|---|
| k001 | Waddington-OT: unbalanced entropic OT between snapshot time points | 10.1016/j.cell.2019.01.006 |
| k003 | Fused Gromov-Wasserstein mapping for spatial snapshots | 10.1038/s41586-024-08453-2 |
| k007 | Interval staging and held-out-window filtering of external data | notes/official/来件/virtualembryo.ai/rules.md |
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
| 改了什么 | 删除父节点无效的 OT 位移,回退到父 mix 路径,并新增 cluster_sample.py:每类型配额不变,但把类型内 i.i.d. 随机抽样换成 k-NN 贪心区域生长的空间连续簇抽样(k=25, n_clusters=10),表达与坐标仍取自真实细胞。 |
|---|---|
| 各组分数的变化 | cell_state:噪声内(+0.88,67.58 vs 66.70) expression_change:噪声内(+1.00,64.90 vs 63.90;T2 噪声约 1 分,处于边界) local_spatial:显著变好(+11.35,65.20 vs 53.85,远超噪声) shape_scale:噪声内略降(-0.84,52.78 vs 53.62;与 METHOD.md 记录的 occupancy_dice 因簇间空洞略降一致) |
| family_id | T2HI-04 |
| 假设是否成立 | 是 |
| 经验 |
|
| mechanism_active | 是 |
| 下一步建议 |
|
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。
分析员
| 角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。 | 分析员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数?大模型一共回复了几轮(每轮可以调用多个工具)。 | 2 |
| 工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。 | 共 1 次:write 1 |
| 用时 | 不到 1 分 |
| token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。 | 输入 11,378 · 输出 1,042 · 思考 173 |
| 任务(第一行) | 分析节点 n7 的打分结果,写 ANALYSIS.json |
| 最后的回答(摘录) | ANALYSIS.json written |
| 原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。 | /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/analyst.jsonl 5 KB /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/analyst.stderr |
工程师
| 角色 | 工程师 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 25 |
| 工具调用 | 共 27 次:bash 21、write 3、read 1、todowrite 1、edit 1 |
| 用时 | 11 分 |
| token 数 | 输入 54,931 · 输出 10,255 · 思考 8,033 |
| 任务(第一行) | 按计划修改程序:空间连续簇抽样替代随机抽样(T2HI-04 保邻域) |
| 最后的回答(摘录) | 完成。最终状态:`solution/` 包含 `run.py`(CLUSTER_SAMPLING=True, k=25, n_clusters=10)、`cluster_sample.py`、`METHOD.md`、`EXECUTION.json`(纯 CPU)。
- **做了什么**:按 PLAN(T2HI-04)删除父节点无效的 OT 位移,把 mix 路径中每类型内的 i.i.d. 随机抽样换成 k-NN 贪心区域生长的空间连续簇抽样(组成配额、表达、坐标系全部不变);对照开关 `CLUSTER_SAMPLING=False` 逐位回退父 mix 路径。
- **查分**(A 半,proxy,8/20 次):对照 59.07 → 提交配置 **62.37**(local_spatial 53.5→64.6,shape 53.1→53.5,expr/state 微升);3 种子均值 62.11,seed0 重跑逐位确定,vec-check 通过。机制证据:输出平均 5-NN 距离 21.1→14.2(−33%)。
- **下一步**:n_clusters=10 是代理峰,final(31 类型、t=0.5)上可扫 n_clusters∈{8,12} 与 cluster_frac=0.85(shape 最高的折中);也可把簇抽样与节点 4 的 matched-align(shape 55.1)组合。 |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/engineer.jsonl 131 KB /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/engineer.stderr |
调研员
| 角色 | 调研员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 1 |
| 工具调用 | 共 0 次:— |
| 用时 | 1 分 |
| token 数 | 输入 9,478 · 输出 900 · 思考 3,762 |
| 任务(第一行) | 为节点 n7(改进,父节点 n5)写改进计划 PLAN.json |
| 最后的回答(摘录) | ```json
{
"name": "空间连续簇抽样替代随机抽样(T2HI-04 保邻域)",
"motivation": "父节点 5 的 OT 位移已证实单调有害(γ→0 收敛回父分),四组分全在噪声内。最弱组 local_spatial 53.85、shape_scale 53.62。当前 mix 的分层抽样对每个类型做独立随机抽取,打散了源阶段的空间邻域结构,导致输出中同类型细胞散布、局部密度不一致,是 local_spatial 偏低的结构原因。节点 4 已证明对齐改进可提升 shape_scale(+1.3),但抽样策略尚未被改进。",
"approach": "步骤:1) 设 OT_ENABLED=False,回退到已验证的 mix 路径(bit-identical)。2) 新增 solution/cluster_sample.py:对每个类型,用 scipy.spatial.cKDTree 在已对齐坐标上建 k-NN(k=15);用贪心扩展法生成空间连续簇——随机选种子细胞,迭代加入当前簇边界的最近未选邻居,直到簇大小达到该类型需抽样数的 1/5(即约 5 个簇覆盖配额);从这些簇中取全部细胞填满配额(若簇总大小超过配额则截断最远边界细胞)。3) 替换 run.py 中 mix_indices 的调用为新的 cluster_mix_indices,保持每类型抽样总数不变(组成不变),仅改变选哪些细胞。4) 单输入阶段(b is None)走原有单阶段回退,不受影响。5) 参数:k=15(搜索 10-25),簇目标大小=配额/5(搜索 /3 到 /8)。先用 proxy 查 1 次(A 半),若 local_spatial ≥55 且 cell_state 不降,再查第 2 次确认。预计实现+跑分 <20 min。",
"expected_groups": ["local_spatial", "shape_scale"],
"risks": "1) 若评分器对局部邻域的度量不敏感于抽样方式(例如只看全局密度),改善可能为零——Engineer 第一次查分即可发现(local_spatial 无变化)。2) 簇太小(/8)退化为随机抽样,簇太大(/3)丧失多样性——用参数搜索覆盖。3) 代理括号仅 5 个共有类型、跨度大,空间结构本身可能不可靠;但本机制不依赖跨阶段对齐质量(只用单侧坐标建图),风险低于 OT。4) 耗时增加(KDTree 建树),但细胞数 <5000 时 <1s,在时限内。",
"family_id": "T2HI-04",
"mechanism": "用贪心 k-NN 扩展生成空间连续簇,以簇为单位抽样替代独立随机抽样,使输出的局部邻域结构继承自源阶段真实空间排列。",
"vs_constant_shift": "不移动任何细胞坐标;改变的是被选中细胞的集合——从随机散点变为空间连续片段。每型内不同簇位于不同空间位置,非统一位移。",
"mechanism_evidence": "1) 输出中同类型细胞的平均 k-NN 距离应显著低于随机抽样版本(预期降 20-40%);2) 输出中每个细胞的 5-NN 中同类型占比应高于随机版本;3) 四组分对比:local_spatial 和 shape_scale 提升而 cell_state / expression_change 不变(因为组成和表达不变,只改了选哪些细胞)。",
"mechanism_off_control": "CLUSTER_SAMPLING=False 时走原始 mix_indices 随机分层抽样,输出与父节点 mix 路径 bit-identical(已有 OT_ENABLED=False 验证路径可复用)。预期差别:关闭后 local_spatial 回到 54.0 左右,开启后若机制有效应 ≥55。",
"sources": []
}
``` |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/researcher.jsonl 4 KB /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/researcher.stderr |
审查员
| 角色 | 审查员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 7 |
| 工具调用 | 共 12 次:read 8、bash 3、write 1 |
| 用时 | 1 分 |
| token 数 | 输入 19,648 · 输出 1,664 · 思考 3,555 |
| 任务(第一行) | 审查节点 n7 的程序是否越界读取、写死目标、钻评分器漏洞 |
| 最后的回答(摘录) | REVIEW.json written |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/reviewer.jsonl 85 KB /home/spark-longxinyang/vec/runs/formal/20261003-043412-search-t2-heart-interp-g24-D-s1/nodes/7/reviewer.stderr |