总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s1
节点 n12
VAE潜空间1-NN耦合→基因空间k=5潜距加权位移(去掉CFM与神经解码器,λ=0.35、gate=0.05、×T外推、VAE固定60 epochs);单输入退路=组成重加权(src reweight心/边权重, Neural Tube剔除)。
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261003-093415-search-t1-r2-D-s1 |
|---|---|
| 父节点 | n8 |
| 子节点 | — |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 改进 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 50.26(+1.5) · X3 45.73(-1.0) · proxy10 59.32(+6.6) · 3 次复测均分 50.65 |
| 审查 | 未审查 |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 35 分 |
| 程序版本 | c35d950067d954f8db8f088e25b7791d9668491b (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git c35d950067:solution/METHOD.md
VAE潜空间1-NN耦合→基因空间k=5潜距加权位移(去掉CFM与神经解码器,λ=0.35、gate=0.05、×T外推、VAE固定60 epochs);单输入退路=组成重加权(src reweight心/边权重, Neural Tube剔除)。
方法族与机制(family: generative_latent,按 PLAN)
- 保留父节点 VAE encoder(512-512,d=20,HVG=2500,stage one-hot 协变量,KL 1e-3,CPU,torch 单线程;训练改固定 epoch 数(墙钟预算会随机器负载变 epoch 数→不确定,已修复并双跑逐位验证一致))。
- 保留潜空间 1-NN 耦合;删除 CFM 场与神经解码器(父节点已证明单转移下场贡献≈0、解码器砸 variogram)。
- 新解码:对末阶段每个输出细胞,在耦合源潜位置中找 k=5 近邻,δ = Σw_j(x_tgt_j−x_src_j)/Σw_j(w=exp(−d²/2σ²),σ=耦合中位距),|δ|>0.05 门控、clip ±2,输出 = x + λ·T·δ,T=(t_target−t_last)/(t_last−t_prev)(仅用时间差,视图无关)。HVG 受两阶段实测基因交集限制(covered mask)。
- 单输入退路(proxy10):src.task1_temporal.reweight 的 type_weights(心脏类×1.6、边缘类×0.25、剔 Neural Tube)+ largest_remainder 分配,从真实细胞重采样;不发明新逻辑。
机制对照(--off / VEC_OFF=1,与实际提交版本一致)
X3(E8.75+E9.0→E9.5,seed 0,A半):
- off(=精确复制末输入):47.62(de_score −0.2429,variogram 0.001619)
- on(λ0.35, gate0.05, k5, T=2→λ_eff 0.70):45.51(提交版 60ep:45.31)(de_score −0.2000,variogram 0.002577,mmd_u 0.0353)
- 机制确实在跑(23% HVG 条目被改,δ 范数 mean 15.3 / sd 2.35,同型细胞位移不同,非每型常数),但在 X3 上净伤害(−2.1,主要来自 variogram 0.00162→0.00258)。λ0.5/gate0.01 更差(44.51);type-level pseudobulk shift(β≤1)最差(44.36,variogram 0.0059)。如实报告:单转移外部数据上任何表达位移目前都跑不赢复制。
查分记录(A半,seed 0)
| 配置 | proxy10 | X3 |
|---|---|---|
| 父节点 8 | 52.75 | 46.75 |
| 组成重加权(maturity=none,提交版) | 59.04(de_dir 0.316/16.85, mmd 0.0339/17.73, vario 0.0011/10.75, de_score 0.134/13.72) | — |
| 组成+lib-size 成熟度选择 top50% | 45.89(有害,弃用) | — |
| latent 位移 λ0.5 g0.01 | — | 44.50 |
| latent 位移 λ0.35 g0.05(71ep,旧墙钟版) | — | 45.51 |
| latent 位移 λ0.35 g0.05,60ep(提交版,确定) | — | 45.31 |
| latent off(复制) | — | 47.62 |
| type shift β1 | — | 44.36 |
预计节点分 ≈ (59.04 + 2×45.31)/3 ≈ 49.8(父 48.75)。proxy10 收益 +6.3 来自组成重加权(复现 node 3 方向,未达其 63.66:缺其成熟度选择实现,lib-size 代理反而有害)。
已验证 / 未验证
- 已验证:两视图跑通 + vec-check 通过;X3 双跑逐位一致(torch.set_num_threads(1) 修复多线程非确定性);off/on 对照同版本代码;默认参数即已查分配置(LAM=0.35、GATE=0.05、K=5、VAE_EPOCHS=60、MATURITY=none);X3 双跑逐位一致,proxy 路径与已查分预测逐位一致。
- 未验证:λ/k 更细扫描;proxy2 视图(本节点不评分,代码路径同两阶段位移);node 3 的真实成熟度选择逻辑(无源码,未能移植)。
- 局限:X3 上机制净负但按 PLAN 提交保持打开;若终选护栏复跑,X3 位移有约 −2 分风险,proxy10 组成收益 +6.3 更大。
知识来源
- 心脏/边缘类权重与 Neural Tube 剔除:直接 import
src.task1_temporal.reweight(harness 提供的 run2 winner 模块,基于已发表阶段 E8.5/E9.5 的解剖学常识,非保留阶段测量)。 - VAE/耦合结构沿用父节点(scVI、CFGen 文献,仅方法)。未使用保留阶段/基因型任何测量信息;未读 external/;未使用 uns.celltype_palette;程序只依赖视图数据与时间差,视图无关。
调研员的计划
| 名称 | VAE latent-guided gene-space coupling displacement, no neural decoder |
|---|---|
| 动机 | Parent node 8 lost -9.46 covariation (X3 variogram 0.001377→0.002022, score -2.20) because dec(z_pred)−dec(z) introduces dense reconstruction noise that destroys gene co-variation. The CFM field itself contributed nothing: off/on de_score identical (-0.2571), X3 45.37 vs 45.40 within noise. proxy10 uniform-copy fallback scored 52.75 vs node 3's 63.66 (-10.9). The VAE latent space IS useful for finding cross-stage correspondences (coupling displacement median 11.52, sd 5.21 shows state-dependent variation), but the decoder and field are the broken components. |
| 做法 | 1. Keep VAE encoder (512-512, d=20, HVG=2500, stage covariate, KL 1e-3, same as parent). Train ≤60s. 2. Keep 1-NN coupling in latent space (same as parent): for each cell in later stage, find nearest in earlier stage → pairs (z_src_i, z_tgt_i, x_src_i, x_tgt_i). 3. REMOVE CFM field entirely (proven zero contribution with single transition). 4. NEW decoding (replaces neural decoder): for each cell c in the last input stage, find its k=5 nearest coupling-pair sources in latent space (by ||z_c - z_src_j||). Compute gene-space displacement δ_c = Σ_j w_j·(x_tgt_j - x_src_j) / Σ_j w_j, where w_j = exp(-||z_c - z_src_j||²/2σ²), σ = median coupling distance. Apply only elements where |δ_c| > 0.01 (sparsity gate), clip to [-2, 2]. Output = x_c + λ·δ_c. λ init 0.5, search {0.3, 0.5, 0.8, 1.0} on X3. 5. proxy10 fallback (single input): implement composition_trend logic from node 3/7 — type-level frequency reweighting by observed trend + within-type maturity selection (reuse node 7's λ=0.15 expression program). This is critical to recover the -10.9 proxy10 deficit. 6. Extrapolation for ≥2 inputs: scale displacement by T = (t_target - t_last)/(t_last - t_prev), same as parent. 7. Run X3 first … |
| 风险 | 1. k=5 latent neighbors may still average out meaningful within-type variation → check displacement sd within types; if < 0.5× between-type sd, reduce k to 3 or 1. 2. VAE latent space may not separate states well enough for meaningful coupling → check coupling displacement distribution; if median < 2, the latent is degenerate. 3. proxy10 composition_trend implementation must match node 3/7 logic exactly or it won't recover the 10.9 pts → Engineer should port node 7's code directly. 4. Gene-space displacement from real pairs is bounded by observed changes; if true target requires extrapolation beyond observed range, this will underperform → λ>1 search partially addresses this. |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:父节点版本 adf9221b97。改动的文件:solution/METHOD.md +33 −14、solution/run.py +248 −156
diff --git a/solution/METHOD.md b/solution/METHOD.mdindex ab98dbf..8bcd609 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,22 +1,41 @@-高斯VAE潜空间+跨阶段1-NN耦合训练CFM位移场,从末阶段潜位置外推T=(Δt目标/Δt训练)步,残差=dec(z_pred)−dec(z)加回输入(软零门、±2截断、范数截断1.5×耦合中位数),λ=0.5阻尼;单输入阶段回退为均匀复制。+VAE潜空间1-NN耦合→基因空间k=5潜距加权位移(去掉CFM与神经解码器,λ=0.35、gate=0.05、×T外推、VAE固定60 epochs);单输入退路=组成重加权(src reweight心/边权重, Neural Tube剔除)。 -## 方法族与机制(family: generative_latent)+## 方法族与机制(family: generative_latent,按 PLAN) -- VAE:encoder/decoder 512-512,latent d=20,stage 作 one-hot 批次协变量,KL 权重 1e-3,高斯损失(数据为 log1p(CP10k),无 counts,故非 scVI)。HVG=2500(scanpy flavor='seurat';scikit-misc 不可用,禁 seurat_v3),非 HVG 基因直接复制输入。-- 耦合:对后一输入阶段每个细胞在前一阶段潜空间取 1-NN,得 (z_src, z_tgt) 对;CFM(MLP 128×2,t 拼接输入)回归速度场 v(z_t,t)≈z_tgt−z_src。-- 外推:只用时间差 T=(t_target−t_last)/(t_last−t_prev),t 夹在 [0,1],Euler 积分,位移×λ=0.5 并截断到 1.5×耦合位移中位数。-- 机制生效证据(X3,E8.75→E9.0,2174 对):耦合位移中位数 11.52,sd 5.21(型内离散度显著非零,非每型常数差);λ=1 时预测位移均值 17.1(被截断);实际改变:约 11% 的 HVG 条目非零残差,每细胞平均 |dx|≈0.11(log 空间)。位移依赖潜位置(同型细胞位移不同)。-- 对照(--off ≡ λ=0):跳过场,z_pred=z。第一版残差用 dec(z_pred)−x(off=重建),X3 off=45.37、on(λ1)=45.40、de_score 两者完全相同(−0.2571)→ 场对基因排序贡献≈0,而解码器重建本身跌破 copy_last(47.92),主要砸 covariation(41.5) 与 de。据此改为残差 = dec(z_pred)−dec(z)(off 时输出=精确复制),把解码偏差从输出中消掉,只保留场在基因空间的像。+- 保留父节点 VAE encoder(512-512,d=20,HVG=2500,stage one-hot 协变量,KL 1e-3,CPU,torch 单线程;训练改固定 epoch 数(墙钟预算会随机器负载变 epoch 数→不确定,已修复并双跑逐位验证一致))。+- 保留潜空间 1-NN 耦合;**删除 CFM 场与神经解码器**(父节点已证明单转移下场贡献≈0、解码器砸 variogram)。+- 新解码:对末阶段每个输出细胞,在耦合源潜位置中找 k=5 近邻,δ = Σw_j(x_tgt_j−x_src_j)/Σw_j(w=exp(−d²/2σ²),σ=耦合中位距),|δ|>0.05 门控、clip ±2,输出 = x + λ·T·δ,T=(t_target−t_last)/(t_last−t_prev)(仅用时间差,视图无关)。HVG 受两阶段实测基因交集限制(covered mask)。+- 单输入退路(proxy10):src.task1_temporal.reweight 的 type_weights(心脏类×1.6、边缘类×0.25、剔 Neural Tube)+ largest_remainder 分配,从真实细胞重采样;不发明新逻辑。++## 机制对照(--off / VEC_OFF=1,与实际提交版本一致)++X3(E8.75+E9.0→E9.5,seed 0,A半):+- off(=精确复制末输入):**47.62**(de_score −0.2429,variogram 0.001619)+- on(λ0.35, gate0.05, k5, T=2→λ_eff 0.70):45.51(提交版 60ep:45.31)(de_score −0.2000,variogram 0.002577,mmd_u 0.0353)+- 机制确实在跑(23% HVG 条目被改,δ 范数 mean 15.3 / sd 2.35,同型细胞位移不同,非每型常数),但在 X3 上**净伤害**(−2.1,主要来自 variogram 0.00162→0.00258)。λ0.5/gate0.01 更差(44.51);type-level pseudobulk shift(β≤1)最差(44.36,variogram 0.0059)。如实报告:单转移外部数据上任何表达位移目前都跑不赢复制。++## 查分记录(A半,seed 0)++| 配置 | proxy10 | X3 |+|---|---|---|+| 父节点 8 | 52.75 | 46.75 |+| 组成重加权(maturity=none,提交版) | **59.04**(de_dir 0.316/16.85, mmd 0.0339/17.73, vario 0.0011/10.75, de_score 0.134/13.72) | — |+| 组成+lib-size 成熟度选择 top50% | 45.89(有害,弃用) | — |+| latent 位移 λ0.5 g0.01 | — | 44.50 |+| latent 位移 λ0.35 g0.05(71ep,旧墙钟版) | — | 45.51 |+| latent 位移 λ0.35 g0.05,60ep(**提交版,确定**) | — | **45.31** |+| latent off(复制) | — | 47.62 |+| type shift β1 | — | 44.36 |++预计节点分 ≈ (59.04 + 2×45.31)/3 ≈ 49.8(父 48.75)。proxy10 收益 +6.3 来自组成重加权(复现 node 3 方向,未达其 63.66:缺其成熟度选择实现,lib-size 代理反而有害)。 ## 已验证 / 未验证 -- 已验证:X3 跑通、vec-check 通过(652 细胞);proxy(单输入)走复制回退,vec-check 通过(5118 细胞);seed 确定(default_rng/torch.manual_seed,无全局随机依赖)。第一版残差的 off/on X3 查分各 1 次。-- 时间耗尽前来不及对第二版残差做完整 λ 扫描;λ 默认取保守 0.5(外推方向押错时四组齐降的风险,见 PLAN risks #3)。-- 未验证:第二版残差在 X3 的最终分数是否超过 copy_last;proxy10 分数(=均匀复制水平,节点 5 显示此回退不吃亏)。-- 已知局限:只有一次转移,场弱辨识;文献先验(k030)警告潜动力学向均值坍缩,故强阻尼+截断。+- 已验证:两视图跑通 + vec-check 通过;X3 双跑逐位一致(torch.set_num_threads(1) 修复多线程非确定性);off/on 对照同版本代码;默认参数即已查分配置(LAM=0.35、GATE=0.05、K=5、VAE_EPOCHS=60、MATURITY=none);X3 双跑逐位一致,proxy 路径与已查分预测逐位一致。+- 未验证:λ/k 更细扫描;proxy2 视图(本节点不评分,代码路径同两阶段位移);node 3 的真实成熟度选择逻辑(无源码,未能移植)。+- 局限:X3 上机制净负但按 PLAN 提交保持打开;若终选护栏复跑,X3 位移有约 −2 分风险,proxy10 组成收益 +6.3 更大。 ## 知识来源 -- scVI(10.1038/s41592-018-0229-2)、scvi-tools(10.1038/s41587-021-01206-w):VAE 结构与 stage 协变量用法(仅方法,无预训练权重)。-- CFGen(arXiv:2407.11734):潜空间条件流匹配生成细胞;本实现为自写的等价 CFM 损失(t~U[0,1],线性插值,回归耦合差)。-- 未使用任何保留阶段/保留基因型的测量数据;未读 external/(proxy 为空,X3 的 external 与输入重复,未使用);未使用 uns.celltype_palette。程序不读视图路径/名称、不读绝对时间(只用时间差),视图无关。+- 心脏/边缘类权重与 Neural Tube 剔除:直接 import `src.task1_temporal.reweight`(harness 提供的 run2 winner 模块,基于已发表阶段 E8.5/E9.5 的解剖学常识,非保留阶段测量)。+- VAE/耦合结构沿用父节点(scVI、CFGen 文献,仅方法)。未使用保留阶段/基因型任何测量信息;未读 external/;未使用 uns.celltype_palette;程序只依赖视图数据与时间差,视图无关。diff --git a/solution/run.py b/solution/run.pyindex 79e0971..92586b4 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,7 +1,14 @@-"""VAE latent space + conditional flow matching displacement field + sparse residual decode.+"""VAE latent-guided gene-space coupling displacement (no CFM, no neural decoder). -View-agnostic: depends only on view data (expression, stage time differences) and --seed.-Single-input fallback: uniform copy of the last input stage (no field can be trained).+Two+ input stages: train a Gaussian VAE (HVG, stage covariate), couple cells of+the later stage to their 1-NN in the earlier stage's latent space, and displace+each output cell in GENE space by a latent-distance-weighted average of the+coupled pairs' real expression differences (k nearest coupling sources).+Single input stage: composition reweighting (heart-focused dissection prior+from src.task1_temporal.reweight) + optional within-type maturity selection.++View-agnostic: depends only on view data (expression, labels, time differences)+and --seed. """ import argparse import json@@ -23,13 +30,16 @@ def read_view(data_dir): return man, genes, inputs -def load_stage(data_dir, entry):- import anndata as ad- a = ad.read_h5ad(os.path.join(data_dir, entry["path"]))- X = a.X- if not isinstance(X, np.ndarray):- X = X.toarray()- return np.asarray(X, dtype=np.float32)+def covered_mask(data_dir, genes, entries):+ """Genes measured in ALL of the given stage entries."""+ gidx = {g: i for i, g in enumerate(genes)}+ mask = np.ones(len(genes), dtype=bool)+ for e in entries:+ gf = os.path.join(data_dir, e["genes_file"])+ measured = set(open(gf).read().split())+ m = np.fromiter((g in measured for g in genes), dtype=bool, count=len(genes))+ mask &= m+ return mask def pick_hvg(Xs, seed, n_top=2500):@@ -75,7 +85,7 @@ def make_vae(n_in, n_stage, d_lat, seed): return Enc(), Dec() -def train_vae(Xh, stage_id, n_stage, d_lat, seed, t_budget):+def train_vae(Xh, stage_id, n_stage, d_lat, seed, n_epochs): import torch torch.manual_seed(seed) enc, dec = make_vae(Xh.shape[1], n_stage, d_lat, seed)@@ -86,9 +96,9 @@ def train_vae(Xh, stage_id, n_stage, d_lat, seed, t_budget): rng = np.random.default_rng(seed + 1) n = Xt.shape[0] bs = 256- t0 = time.time() ep = 0- while ep < 100 and time.time() - t0 < t_budget:+ tot = float("nan")+ while ep < n_epochs: perm = rng.permutation(n) tot = 0.0 for i in range(0, n, bs):@@ -97,74 +107,226 @@ def train_vae(Xh, stage_id, n_stage, d_lat, seed, t_budget): mu, lv = enc(xb, sb) z = mu + torch.randn_like(mu) * (0.5 * lv).exp() rec = dec(z, sb)- loss = ((rec - xb) ** 2).sum(1).mean() + 1e-3 * (-0.5 * (1 + lv - mu.pow(2) - lv.exp()).sum(1).mean())+ loss = ((rec - xb) ** 2).sum(1).mean() + 1e-3 * (-0.5 * (1 + lv - mu.pow(2) - lv.exp()).sum(1)).mean() opt.zero_grad() loss.backward() opt.step() tot += float(loss) * len(b) ep += 1- log(f"VAE: {ep} epochs, loss={tot / n:.4f}, {time.time() - t0:.1f}s")+ log(f"VAE: {ep} epochs, loss={tot / n:.4f}") return enc, dec -def couple(zs, zt, seed):+def displacement_path(args, man, genes, inputs, data_dir, rng):+ """Two+ inputs: latent-guided gene-space coupling displacement."""+ import anndata as ad+ import scipy.sparse as sp+ import torch from sklearn.neighbors import NearestNeighbors- nn = NearestNeighbors(n_neighbors=1).fit(zs)- _, idx = nn.kneighbors(zt)- return zs[idx[:, 0]], zt + t_last = inputs[-1]["time"]+ t_target = man["target"]["time"]+ dt_train = inputs[-1]["time"] - inputs[-2]["time"]+ dt_pred = t_target - t_last+ T = float(dt_pred / dt_train) if dt_train > 0 else 1.0 -def train_cfm(z0, z1, d_lat, seed, t_budget):- import torch- import torch.nn as nn- torch.manual_seed(seed)- net = nn.Sequential(nn.Linear(d_lat + 1, 128), nn.ReLU(),- nn.Linear(128, 128), nn.ReLU(),- nn.Linear(128, d_lat))- opt = torch.optim.Adam(net.parameters(), lr=1e-3)- Z0 = torch.from_numpy(z0)- Z1 = torch.from_numpy(z1)- U = Z1 - Z0- rng = np.random.default_rng(seed + 2)- n = Z0.shape[0]- bs = 256- t0 = time.time()- ep = 0- while ep < 300 and time.time() - t0 < t_budget:- perm = rng.permutation(n)- for i in range(0, n, bs):- b = torch.from_numpy(perm[i:i + bs])- tt = torch.rand(len(b), 1)- zt = (1 - tt) * Z0[b] + tt * Z1[b]- v = net(torch.cat([zt, tt], 1))- loss = ((v - U[b]) ** 2).sum(1).mean()- opt.zero_grad()- loss.backward()- opt.step()- ep += 1- log(f"CFM: {ep} epochs, loss={float(loss):.4f}, {time.time() - t0:.1f}s")- return net+ n_genes = len(genes)+ min_cells = int(man.get("min_cells", 1))+ max_cells = int(man.get("max_cells", 10 ** 9)) + if args.method == "shift":+ # type-level pseudobulk displacement extrapolation (node-7 style)+ a_prev = ad.read_h5ad(os.path.join(data_dir, inputs[-2]["path"]))+ a_last = ad.read_h5ad(os.path.join(data_dir, inputs[-1]["path"]))+ Xp = a_prev.X.tocsr() if sp.issparse(a_prev.X) else sp.csr_matrix(a_prev.X)+ Xl = a_last.X.tocsr() if sp.issparse(a_last.X) else sp.csr_matrix(a_last.X)+ lp = a_prev.obs["celltype"].astype(str).to_numpy()+ ll = a_last.obs["celltype"].astype(str).to_numpy()+ n_last = Xl.shape[0]+ n_out = int(min(max(n_last, min_cells), max_cells))+ if n_last >= n_out:+ rows = np.sort(rng.choice(n_last, n_out, replace=False))+ else:+ rows = np.sort(rng.choice(n_last, n_out, replace=True))+ scale = 1.0 if args.off else min(T, args.beta_cap)+ deltas = {}+ for t in np.unique(ll):+ if t not in set(lp.tolist()):+ continue+ m_last = np.asarray(Xl[ll == t].mean(axis=0), dtype=np.float32).ravel()+ m_prev = np.asarray(Xp[lp == t].mean(axis=0), dtype=np.float32).ravel()+ d = m_last - m_prev+ deltas[str(t)] = np.where(np.abs(d) > args.gate, np.clip(d, -2.0, 2.0), 0.0).astype(np.float32)+ Xout = np.asarray(Xl[rows].toarray(), dtype=np.float32)+ tl = ll[rows]+ changed = 0+ for t, d in deltas.items():+ m = tl == t+ if not m.any():+ continue+ if not args.off:+ Xout[m] = np.maximum(Xout[m] + scale * d, 0.0)+ changed += int(m.sum())+ log(f"type-shift: T={T:.2f}, scale={scale:.2f}, {len(deltas)} types with delta, {changed}/{n_out} cells shifted")+ a = ad.AnnData(X=sp.csr_matrix(Xout))+ a.var_names = genes+ return a, n_out++ cov = covered_mask(data_dir, genes, inputs[-2:])+ log(f"covered genes in both stages: {int(cov.sum())}")++ As = [ad.read_h5ad(os.path.join(data_dir, e["path"])) for e in inputs[-2:]]+ Xsp = [a.X.tocsr() if sp.issparse(a.X) else sp.csr_matrix(a.X) for a in As]+ nA, nB = Xsp[0].shape[0], Xsp[1].shape[0]++ caps = 4000+ sub_idx = []+ sub_dense = []+ for i, M in enumerate(Xsp):+ if M.shape[0] > caps:+ idx = np.sort(rng.choice(M.shape[0], caps, replace=False))+ else:+ idx = np.arange(M.shape[0])+ sub_idx.append(idx)+ sub_dense.append(np.asarray(M[idx].toarray(), dtype=np.float32))++ cov_idx = np.flatnonzero(cov)+ hvg_local = pick_hvg([d[:, cov_idx] for d in sub_dense], args.seed, n_top=2500)+ hvg = cov_idx[hvg_local]+ log(f"HVG: {len(hvg)} genes")++ Xh = np.vstack([d[:, hvg] for d in sub_dense]).astype(np.float32)+ stage_id = np.concatenate([np.full(d.shape[0], i, dtype=np.int64) for i, d in enumerate(sub_dense)])+ d_lat = 20+ enc, _ = train_vae(Xh, stage_id, 2, d_lat, args.seed, n_epochs=args.vae_epochs) -def integrate(net, z, T, lam, max_norm, n_steps=None):- import torch- zc = z.clone()- if n_steps is None:- n_steps = int(max(4, np.ceil(T * 8)))- h = T / n_steps- disp = torch.zeros_like(zc) with torch.no_grad():- for k in range(n_steps):- tcur = min(k * h, 1.0)- tt = torch.full((zc.shape[0], 1), tcur)- v = net(torch.cat([zc, tt], 1))- step = v * h- disp = disp + step- zc = zc + step- d = lam * disp- nrm = d.norm(dim=1, keepdim=True)- scale = torch.clamp(max_norm / (nrm + 1e-8), max=1.0)- return z + d * scale+ s0 = torch.zeros(sub_dense[0].shape[0], 2); s0[:, 0] = 1.0+ s1 = torch.zeros(sub_dense[1].shape[0], 2); s1[:, 1] = 1.0+ zs = enc(torch.from_numpy(sub_dense[0][:, hvg]), s0)[0].numpy()+ zt = enc(torch.from_numpy(sub_dense[1][:, hvg]), s1)[0].numpy()++ # 1-NN coupling: each later-stage cell -> nearest earlier-stage latent+ nn = NearestNeighbors(n_neighbors=1).fit(zs)+ _, src = nn.kneighbors(zt)+ src = src[:, 0]+ z_src = zs[src]+ D = sub_dense[1][:, hvg] - sub_dense[0][src][:, hvg] # gene-space displacements+ cd = np.linalg.norm(zt - z_src, axis=1)+ med = float(np.median(cd))+ sd = float(np.std(cd))+ log(f"coupling: n={len(src)}, median disp={med:.3f}, sd={sd:.3f}")++ # output cells sampled from the last input stage+ n_out = int(min(max(nB, min_cells), max_cells))+ if nB >= n_out:+ rows = np.sort(rng.choice(nB, n_out, replace=False))+ else:+ rows = np.sort(rng.choice(nB, n_out, replace=True))+ Xout_sp = Xsp[1][rows]+ Xout_h = np.asarray(Xout_sp[:, hvg].toarray(), dtype=np.float32)++ if args.off:+ delta = np.zeros_like(Xout_h)+ lam_eff = 0.0+ else:+ with torch.no_grad():+ s_one = torch.zeros(n_out, 2); s_one[:, 1] = 1.0+ z_out = enc(torch.from_numpy(Xout_h), s_one)[0].numpy()+ k = int(args.k)+ knn = NearestNeighbors(n_neighbors=k).fit(z_src)+ dist, jdx = knn.kneighbors(z_out)+ sigma = max(med, 1e-6)+ w = np.exp(-(dist ** 2) / (2.0 * sigma ** 2))+ w = w / w.sum(axis=1, keepdims=True)+ delta = np.einsum("ij,ijg->ig", w, D[jdx])+ lam_eff = args.lam * T+ gate = np.abs(delta) > args.gate+ delta = np.where(gate, np.clip(delta, -2.0, 2.0), 0.0).astype(np.float32)+ Xout_h_new = np.maximum(Xout_h + lam_eff * delta, 0.0).astype(np.float32)++ nz = float(np.mean(delta != 0)) if not args.off else 0.0+ log(f"delta: lam_eff={lam_eff:.3f}, frac nonzero={nz:.4f}, mean|d|={np.abs(delta).mean():.4f}")+ # within-cell dispersion check (mechanism evidence)+ if not args.off and n_out > 10:+ norms = np.linalg.norm(delta, axis=1)+ log(f"delta norms: mean={norms.mean():.3f}, sd={norms.std():.3f}")++ Xout = np.asarray(Xout_sp.toarray(), dtype=np.float32)+ Xout[:, hvg] = Xout_h_new+ a = ad.AnnData(X=sp.csr_matrix(Xout))+ a.var_names = genes+ return a, n_out+++def composition_path(args, man, genes, inputs, data_dir, rng):+ """Single input: composition reweighting + within-type maturity selection."""+ import anndata as ad+ import scipy.sparse as sp+ from src.task1_temporal.reweight import (DROP_TYPES, largest_remainder,+ type_weights)++ min_cells = int(man.get("min_cells", 1))+ max_cells = int(man.get("max_cells", 10 ** 9))+ a_last = ad.read_h5ad(os.path.join(data_dir, inputs[-1]["path"]))+ X = a_last.X if sp.issparse(a_last.X) else sp.csr_matrix(a_last.X)+ X = X.tocsr().astype(np.float32)+ labels = a_last.obs["celltype"].astype(str).to_numpy()+ n_last = X.shape[0]+ n_out = int(min(max(n_last, min_cells), max_cells))++ if args.off:+ if n_last >= n_out:+ rows = np.sort(rng.choice(n_last, n_out, replace=False))+ else:+ rows = np.sort(rng.choice(n_last, n_out, replace=True))+ out = X[rows]+ a = ad.AnnData(X=out)+ a.var_names = genes+ return a, n_out++ # per-cell maturity score (higher = more mature); modes: none | lib | prolif+ mode = args.maturity+ score = None+ if mode == "lib":+ score = np.asarray(X.sum(axis=1)).ravel()+ elif mode == "prolif":+ prolif = ["Mki67", "Top2a", "Ccnb1", "Cdk1", "Ccna2", "Birc5", "Ccnb2", "Aurkb"]+ vnames = list(map(str, a_last.var_names))+ vidx = {g: i for i, g in enumerate(vnames)}+ cols = [vidx[g] for g in prolif if g in vidx]+ if cols:+ score = -np.asarray(X[:, cols].mean(axis=1)).ravel()+ else:+ log("prolif genes not found; falling back to uniform maturity")+ if score is not None:+ # global rank transform so scores are comparable across types+ order = np.argsort(np.argsort(score))+ score = order / max(len(score) - 1, 1)++ types = [str(t) for t in np.unique(labels) if str(t) not in DROP_TYPES]+ counts = np.array([(labels == t).sum() for t in types], dtype=np.float64)+ alloc = largest_remainder(counts * type_weights(types, 1.6, 0.25), n_out)+ blocks = []+ for t, n in zip(types, alloc):+ n = int(n)+ if n <= 0:+ continue+ pool = np.flatnonzero(labels == t)+ if score is not None and args.mat_frac < 1.0 and len(pool) > 20:+ s_t = score[pool]+ keep = int(max(20, np.ceil(len(pool) * args.mat_frac)))+ if keep < len(pool):+ top = np.argpartition(-s_t, keep - 1)[:keep]+ pool = np.sort(pool[top])+ choice = rng.choice(pool, size=n, replace=pool.size < n)+ blocks.append(X[choice])+ out = sp.vstack(blocks, format="csr").astype(np.float32)+ out.eliminate_zeros()+ log(f"composition: {len(types)} types, n_out={out.shape[0]}, maturity={mode}({args.mat_frac})")+ a = ad.AnnData(X=out)+ a.var_names = genes+ return a, out.shape[0] def main():@@ -172,105 +334,35 @@ def main(): ap.add_argument("--data", required=True) ap.add_argument("--out", required=True) ap.add_argument("--seed", type=int, required=True)- ap.add_argument("--lam", type=float, default=float(os.environ.get("LAM", "0.5")))+ ap.add_argument("--lam", type=float, default=float(os.environ.get("LAM", "0.35")))+ ap.add_argument("--k", type=int, default=int(os.environ.get("KNN_K", "5")))+ ap.add_argument("--vae-epochs", type=int, default=int(os.environ.get("VAE_EPOCHS", "60")))+ ap.add_argument("--maturity", default=os.environ.get("MATURITY", "none"))+ ap.add_argument("--mat-frac", type=float, default=float(os.environ.get("MAT_FRAC", "1.0")))+ ap.add_argument("--gate", type=float, default=float(os.environ.get("GATE", "0.05")))+ ap.add_argument("--method", default=os.environ.get("METHOD", "latent"))+ ap.add_argument("--beta-cap", type=float, default=float(os.environ.get("BETA_CAP", "1.0"))) ap.add_argument("--off", action="store_true", default=os.environ.get("VEC_OFF", "0") == "1") args = ap.parse_args()- seed = args.seed- np.random.seed(seed)- import torch- torch.manual_seed(seed) man, genes, inputs = read_view(args.data)- n_genes = len(genes)- min_cells = int(man.get("min_cells", 1))- max_cells = int(man.get("max_cells", 10 ** 9))- t_last = inputs[-1]["time"]- t_target = man["target"]["time"]-- Xs = [load_stage(args.data, e) for e in inputs]- X_last = Xs[-1]- n_last = X_last.shape[0]-- rng = np.random.default_rng(seed)- n_out = int(min(max(n_last, min_cells), max_cells))- if n_last >= n_out:- rows = np.sort(rng.choice(n_last, n_out, replace=False))- else:- rows = np.sort(rng.choice(n_last, n_out, replace=True))- Xout = X_last[rows].copy()-- two_stage = len(inputs) >= 2- if two_stage and not args.off:- # time ratios from differences only (view-independent)- dt_train = inputs[-1]["time"] - inputs[-2]["time"]- dt_pred = t_target - t_last- T = float(dt_pred / dt_train) if dt_train > 0 else 1.0- else:- T = 1.0-- if two_stage:- d_lat = 20- # subsample for training- caps = 4000- sub = []- for X in Xs[-2:]:- if X.shape[0] > caps:- idx = rng.choice(X.shape[0], caps, replace=False)- sub.append(X[idx])- else:- sub.append(X)- hvg = pick_hvg(sub, seed, n_top=2500)- log(f"HVG: {len(hvg)} genes")- Xh = np.vstack([s[:, hvg] for s in sub]).astype(np.float32)- stage_id = np.concatenate([np.full(s.shape[0], i, dtype=np.int64) for i, s in enumerate(sub)])- enc, dec = train_vae(Xh, stage_id, 2, d_lat, seed, t_budget=300)+ rng = np.random.default_rng(args.seed)+ import torch+ torch.set_num_threads(1)+ torch.manual_seed(args.seed) - with torch.no_grad():- zs = enc(torch.from_numpy(sub[0][:, hvg]),- torch.zeros(sub[0].shape[0], 2).index_fill_(1, torch.tensor([0]), 1.0))[0].numpy()- zt = enc(torch.from_numpy(sub[1][:, hvg]),- torch.zeros(sub[1].shape[0], 2).index_fill_(1, torch.tensor([1]), 1.0))[0].numpy()- z0, z1 = couple(zs, zt, seed)- med = float(np.median(np.linalg.norm(z1 - z0, axis=1)))- within = np.std(np.linalg.norm(z1 - z0, axis=1))- log(f"coupling: median disp={med:.3f}, sd={within:.3f}, n={len(z0)}")- cfm = train_cfm(z0, z1, d_lat, seed, t_budget=150)-- # encode output rows at last stage, integrate, decode residual- Xr = Xout[:, hvg].astype(np.float32)- s_one = torch.zeros(n_out, 2)- s_one[:, 1] = 1.0- with torch.no_grad():- z = enc(torch.from_numpy(Xr), s_one)[0]- if args.off:- z_pred = z- else:- z_pred = integrate(cfm, z, T, args.lam, 1.5 * med)- x_hat = dec(z_pred, s_one).numpy()- recon0 = dec(z, s_one).numpy()-- resid = np.clip(x_hat - recon0, -2.0, 2.0).astype(np.float32)- # soft zero gate: genes at 0 in input only move if decoder is confident- expo = np.clip(-(x_hat - 0.05) / 0.05, -60, 60)- gate = (1.0 / (1.0 + np.exp(expo))).astype(np.float32)- zero_mask = Xr == 0- resid[zero_mask] *= gate[zero_mask]- Xout[:, hvg] = np.maximum(Xr + resid, 0.0)- moved = np.abs(resid).sum(1)- log(f"residual: mean |dx| per cell={np.abs(resid).mean():.4f}, frac nonzero changed={np.mean(resid != 0):.4f}")- log(f"disp norms: mean={float((z_pred - z).norm(dim=1).mean()):.3f}")+ t0 = time.time()+ if len(inputs) >= 2:+ log(f"two-stage path: {len(inputs)} inputs")+ a, n_out = displacement_path(args, man, genes, inputs, args.data, rng) else:- log("single input stage: uniform copy fallback")+ log("single-input path: composition reweighting")+ a, n_out = composition_path(args, man, genes, inputs, args.data, rng) - import anndata as ad- import scipy.sparse as sp- obs = {}- a = ad.AnnData(X=sp.csr_matrix(Xout.astype(np.float32)))- a.var_names = genes if man.get("needs_coords"): a.obsm["spatial_3D"] = np.zeros((n_out, 3), dtype=np.float32) a.write_h5ad(args.out)- log(f"wrote {args.out}: {a.shape}")+ log(f"wrote {args.out}: {a.shape}, total {time.time() - t0:.1f}s") if __name__ == "__main__":
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
用到的知识库条目
| 编号 | 标题 | 出处 |
|---|---|---|
| k041 | Within-stage pseudotime and graph toolkit offline: scanpy DPT/PAGA/Leiden, Palantir, CellRank 2 | 10.1186/s13059-019-1663-x (PAGA); 10.1038/s41587-019-0068-4 (Palantir); 10.1038/s41592-024-02303-9 (CellRank 2) |
| k031 | Offline OT toolkit in the sandbox: moscot TemporalProblem, wot OTModel, POT, geomloss | 10.1038/s41586-024-08453-2 (moscot); 10.1016/j.cell.2019.01.006 (Waddington-OT) |
| k038 | RNA velocity family (scVelo, dynamo, CellRank velocity kernel): not applicable to T1 files; substitutes | 10.1038/s41587-020-0591-3 (scVelo); 10.1016/j.cell.2021.12.045 (dynamo); 10.1038/s41592-024-02303-9 (CellRank 2) |
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
| 改了什么 | 删除 CFM 场与神经解码器;X3(两输入)路径改为 VAE 潜空间 1-NN 耦合 + k=5 潜距加权的真实表达差位移(λ=0.35、gate=0.05、×T 外推、VAE 固定 60 epochs);proxy10(单输入)路径由均匀复制改为直接调用 src.task1_temporal.reweight 的组成重加权(心脏类×1.6、边缘×0.25、剔 Neural Tube、largest_remainder 重采样,maturity=none)。耗时 466→120s,内存 4.41→3.10GB。 |
|---|---|
| 各组分数的变化 | cell_state:变好 +2.00:proxy10 mmd_u 0.0405→0.0338(得分 +2.31);X3 −0.25 在噪声内 covariation:变坏 −2.75:X3 variogram 0.002022→0.002551(得分 −1.23),位移仍破坏共变结构;proxy10 variogram 0.001233→0.001098(+0.80)小幅变好。总榜分 +1.51(proxy10 +6.57 远超噪声,X3 −1.02 在噪声内),总分变化本身在 T1 约 2 分噪声边缘 de_recovery:变好 +2.07:proxy10 de_score 0.016→0.148(得分 +1.20),来自组成重加权;X3 +0.18 在噪声内 direction:变好 +3.77:proxy10 de_direction 0.1785→0.3165(得分 +2.27);X3 +0.28 在噪声内 |
| family_id | other |
| 假设是否成立 | 否 |
| 经验 |
|
| mechanism_active | 是 |
| 下一步建议 |
|
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。
分析员
| 角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。 | 分析员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数?大模型一共回复了几轮(每轮可以调用多个工具)。 | 6 |
| 工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。 | 共 6 次:bash 4、read 1、write 1 |
| 用时 | 1 分 |
| token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。 | 输入 22,043 · 输出 1,547 · 思考 3,559 |
| 任务(第一行) | 分析节点 n12 的打分结果,写 ANALYSIS.json |
| 最后的回答(摘录) | ANALYSIS.json written |
| 原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。 | /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/12/analyst.jsonl 21 KB /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/12/analyst.stderr |
工程师
| 角色 | 工程师 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次(失败 1 次) |
| 轮数 | 30 |
| 工具调用 | 共 46 次:bash 39、edit 4、write 2、read 1 |
| 用时 | 30 分 |
| token 数 | 输入 74,005 · 输出 13,487 · 思考 14,662 |
| 任务(第一行) | 按计划修改程序:VAE latent-guided gene-space coupling displacement, no neural decoder |
| 最后的回答(摘录) | Found it: the VAE loop is wall-clock budgeted (`time.time()-t0 < t_budget`) — epoch count varies with machine load. Fixing to fixed epochs. |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/12/engineer.jsonl 182 KB /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/12/engineer.stderr |
调研员
| 角色 | 调研员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 1 |
| 工具调用 | 共 0 次:— |
| 用时 | 1 分 |
| token 数 | 输入 18,756 · 输出 1,441 · 思考 2,859 |
| 任务(第一行) | 为节点 n12(改进,父节点 n8)写改进计划 PLAN.json |
| 最后的回答(摘录) | 3. proxy10 composition_trend implementation must match node 3/7 logic exactly or it won't recover the 10.9 pts → Engineer should port node 7's code directly. 4. Gene-space displacement from real pairs is bounded by observed changes; if true target requires extrapolation beyond observed range, this will underperform → λ>1 search partially addresses this.", "family_id": "generative_latent", "mechanism": "VAE latent space identifies cross-stage cell correspondences; gene-space displacement is computed as a latent-distance-weighted average of real expression differences from coupled pairs, preserving covariation structure that a neural decoder destroys.", "vs_constant_shift": "Displacement depends on each cell's position in latent space: cells of the same type get different displacements based on which coupling pairs are their latent neighbors. Evidence from parent: coupling displacement sd=5.21 vs median=11.52 shows substantial within-type variation. A per-type constant shift would give identical displacement to all cells of a type.", "mechanism_evidence": "1. Within-type displacement dispersion: compute std of ||δ_c|| within each cell type; must be >20% of between-type mean difference. 2. Fraction of genes actually changed per cell (expect 5-15% non-zero residuals, similar to parent's 11%). 3. Compare X3 variogram score to parent (0.002022) and node 3 (0.001377): must be ≤0.0015 to confirm decoder removal fixed covariation. 4. de_score and de_direction on X3 must improve over parent's -0.1948 and -0.004 respectively.", "mechanism_off_control": "λ=0 (via --off flag or VEC_OFF=1): skip displacement entirely, output = exact copy of last input stage. Expected difference: with λ=0.5, X3 de_score should differ from -0.1948 (parent's value) and variogram should improve; if λ=0 and λ=0.5 give identical scores, the mechanism is not running. Also run with k=1 vs k=5 to confirm latent-weighted averaging (not just nearest-neighbor copy) matters.", "sources": [] } ``` |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/12/researcher.jsonl 6 KB /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/12/researcher.stderr |