总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s1
节点 n15
METHOD
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261003-093415-search-t1-r2-D-s1 |
|---|---|
| 父节点 | (种子,没有父节点) |
| 子节点 | — |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 草稿 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 49.55 · X3 47.92 · proxy10 52.81 |
| 审查 | 未审查 |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 17 分 |
| 程序版本 | 5a1b518006ad98c8c1c2a77ea26292469eafda71 (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git 5a1b518006:solution/METHOD.md
METHOD
lineage_branching:Isl1+→Nkx2-5+ 承诺边上,逐细胞向其 kNN 中已承诺邻居的均值做稀疏表达位移(幅度=λ·readiness),双输入时先按观测时间差校验方向,不支持则跳过。
方法族与机制
PLAN family_id = lineage_branching(最小版本)。机制(与 PLAN 一致,逐细胞、非常数位移):
- 取最后一个输入阶段全部细胞;承诺分数 = mean(Nkx2-5, Tnnt2, Myh6, Actc1, Tnnt1, Tbx5) − mean(Isl1, Mef2c)(log 空间)。
- 在心脏/中胚层系标签细胞(关键词匹配 CM/SHF/Pericardium/Endocardium/JCF/Mesoderm/Blood/Endothelium/NCC/Unknown 等,对 X3 的 Unknown 也生效)中取承诺分数前 20% 且承诺标记有表达者为 committed;不足 30 个 → 不修改(support check)。
- HVG 2500 → TruncatedSVD 25 维 → kNN k=20;readiness = 邻居中 committed 的比例;readiness > 0.15 且自身非 committed 者为 branching。
- 每个 branching 细胞的 target = 它自己的 committed 邻居的表达均值(逐细胞不同);位移 step = λ·readiness·(target−x),只作用于 mask = (target>0.5) ∧ (delta>0.25) 的基因(保持稀疏结构,避免把零整体抬高),clip ≥0;λ=0.5(环境变量
LAMBDA可覆盖)。非 branching 细胞原样保留。 - 双输入方向校验(X3 / final 生效):按共同标签加权计算前一阶段→末阶段的伪批量差 d,与平均局部位移方向 mean_shift 求余弦;cos<0.1 → 跳过位移(输出=输入复制),0.1≤cos<0.3 → λ 减半(比 PLAN 的"减半"多加了更硬的丢弃门,依据是未加门时 X3 分数下降,见下)。
- 输出细胞数 = min(末阶段细胞数, max_cells),
default_rng(seed)无放回均匀抽样。
关键参数
k=20,HVG=2500,PCA=25,committed_frac=0.20,readiness_thr=0.15(PLAN 风险节建议试 0.15;0.25 时 proxy 只有 461 个 branching,0.15 → 582),λ=0.5(试 λ=1.0 得 53.14 vs 53.10,噪声内,保留 0.5),mask: target>0.5, delta>0.25。
off 对照(mechanism_off_control)
同一程序 --off(或 LAMBDA=0)→ λ=0,输出=输入的均匀子抽样(copy_last 等价)。提交版保持打开(λ=0.5,无 off 开关生效)。
机制生效证据(proxy E8.5→E9.5,seed 0)
- committed 2265 / pool 11323;branching(被修改)582(3.5%)。
- readiness 分布 q10–q90 = 0,0,0,0,0.9(重尾:只有心脏区细胞有 committed 邻居)。
- 逐细胞位移范数 mean=4.48, q10=2.36, q50=4.17, q90=7.09(异质,非单一常数)。
- 平均位移最大的基因:Myh6, Kcnh7, Ankrd1, Myl3, Slc2a1, Acta1 —— 心脏成熟程序方向,符合 Isl1+→Nkx2-5+ 边的生物学。
- X3(E8.75+E9.0 心脏):branching 候选 407(18.7%),但方向校验 cos=−0.074 < 0.1 → 按规则跳过位移,on 输出与 off 逐字节相同(门在起作用,不是机制没跑)。
查分结果(A 半,seed 0)
| 视图 | on | off(对照) | 分项(on) de_score/de_dir/mmd/vario |
|---|---|---|---|
| proxy10 | 53.10 | 52.97 | 12.90/14.69/15.55/9.96 |
| X3 | 47.62(=off,门跳过) | 47.62 | 10.77/12.28/14.85/9.72 |
早期无门版本在 X3 上:v1(dense mask target>0.05)45.69,variogram 掉到 8.02;v2(sparse mask)46.31,variogram 8.74 —— 说明"向邻居均值位移"在 cos≈0(方向未验证)时损害共变结构,因此加了丢弃门。proxy 上 on−off = +0.13,在 ±2 噪声内:机制在 proxy 上方向正确(top 基因是心脏成熟基因)但幅度上未证明有效,如实报告。
验证过 / 没验证
- 验证过:proxy 与 X3 视图跑通、vec-check ok、seed 0 确定性(两次运行逐字节一致)、off 对照、X3 门触发后 on==off。
- 没验证:proxy2 视图(本节点无该视图目录,但代码只依赖 manifest schema,双输入路径已在 X3 走过);伪装视图重跑(代码不读绝对时间、路径、board/mode 字段);final(E8.5+E9.5,cos 校验可能通过并实际位移)。
- λ、readiness 阈值只在 proxy10 上粗试,未做网格。
知识来源
- 标记基因角色(Nkx2-5/Tnnt2/Myh6/Actc1/Tnnt1/Tbx5 = 心脏承诺/成熟;Isl1/Mef2c = SHF 祖细胞):通用谱系知识与 PLAN motivation 引用的 k063(Jia 2018, de Soysa 2019, Ameen 2022,全部 ≤E9.5 已发表阶段)。
- 未使用任何保留阶段(E10.5/E12.5、禁窗 9.5<E≤13.5)或保留基因型的测量信息;未读外部数据文件(proxy 视图 external 为空,X3 的 external 未使用);未读 prior/ 文件;未读
uns.celltype_palette(视图中已剥掉)。
调研员的计划
| 名称 | lineage_branching: local-neighborhood commitment shift along Isl1→Nkx2-5 edge |
|---|---|
| 动机 | Node 6 (graph_fate draft, 47.66) failed because successor sampling produced unmodified cells; node 3/7/13 show that per-cell expression modification along a biological direction is needed for DE and mmd_u gains. k063 (Jia 2018, de Soysa 2019, Ameen 2022, all ≤E9.5) documents the Isl1+→Nkx2-5+ commitment edge in SHF, active during E7.5–E9.5. The weakest groups in most nodes are de_recovery (~50) and covariation (~48-54); a local, per-cell shift that preserves single-cell residuals should help mmd_u without destroying variogram. |
| 做法 | Steps: (1) Load input stage(s), select HVGs (top 2500 by variance). (2) Build kNN graph (k=20, scanpy sc.pp.neighbors on PCA 25-dim). (3) Define 'committed' cells: score = mean expression of commitment markers (Nkx2-5, Tnnt2, Myh6, Actc1) minus mean of progenitor markers (Isl1, Mef2c at low level); cells in top 20% of this score among mesoderm/cardiac-typed cells are 'committed'. (4) Support check: if fewer than 30 committed cells exist, edge unsupported → output = input copy (identical to off control). (5) For each non-committed cell, compute readiness = fraction of its k=20 neighbors that are committed. (6) Cells with readiness > 0.25 are 'branching'; for each branching cell, identify its committed neighbors (typically 3-8 cells), compute target = mean expression of those neighbors. (7) Shift: x_new = x_old + λ·readiness·(target − x_old), applied only to genes where target−x_old ≠ 0 (sparse shift); clip to ≥0 in log space. λ initial = 0.5, search [0.3, 0.7]. (8) Non-branching cells are unchanged. Output same cell count as input. Single-input (proxy10, X3): use within-stage mechanism as above. Two-input (final E8.5+E9.5): additionally validate direction by checking that the E8.5→… |
| 风险 | 1) Committed-cell identification fails if marker genes are lowly expressed or cell type labels don't include cardiac types (especially on proxy10 whole-embryo where cardiac cells may be <5%); Engineer should log the number of committed cells detected and branching cells count. 2) λ too high collapses diversity → variogram worsens; check variogram score after first run. 3) The edge may not be supported in X3 if cell types are named differently; the support check handles this but means mechanism doesn't run. 4) Readiness threshold 0.25 may be too strict/loose; Engineer should report the distribution of readiness scores and try 0.15/0.35 if too few/many cells branch. Early detection: after first vec-score, if branching cell count is 0 or <5% of input, the mechanism isn't running. |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:这个提交的上一版(种子程序:相对空仓库)。改动的文件:solution/EXECUTION.json +1 −0、solution/METHOD.md +50 −0、solution/run.py +250 −0
diff --git a/solution/EXECUTION.json b/solution/EXECUTION.jsonnew file mode 100644index 0000000..9d5125c--- /dev/null+++ b/solution/EXECUTION.json@@ -0,0 +1 @@+{"gpu": false}diff --git a/solution/METHOD.md b/solution/METHOD.mdnew file mode 100644index 0000000..89beaee--- /dev/null+++ b/solution/METHOD.md@@ -0,0 +1,50 @@+# METHOD++lineage_branching:Isl1+→Nkx2-5+ 承诺边上,逐细胞向其 kNN 中已承诺邻居的均值做稀疏表达位移(幅度=λ·readiness),双输入时先按观测时间差校验方向,不支持则跳过。++## 方法族与机制++PLAN family_id = `lineage_branching`(最小版本)。机制(与 PLAN 一致,逐细胞、非常数位移):++1. 取最后一个输入阶段全部细胞;承诺分数 = mean(Nkx2-5, Tnnt2, Myh6, Actc1, Tnnt1, Tbx5) − mean(Isl1, Mef2c)(log 空间)。+2. 在心脏/中胚层系标签细胞(关键词匹配 CM/SHF/Pericardium/Endocardium/JCF/Mesoderm/Blood/Endothelium/NCC/Unknown 等,对 X3 的 Unknown 也生效)中取承诺分数前 20% 且承诺标记有表达者为 committed;不足 30 个 → 不修改(support check)。+3. HVG 2500 → TruncatedSVD 25 维 → kNN k=20;readiness = 邻居中 committed 的比例;readiness > 0.15 且自身非 committed 者为 branching。+4. 每个 branching 细胞的 target = 它自己的 committed 邻居的表达均值(逐细胞不同);位移 step = λ·readiness·(target−x),只作用于 mask = (target>0.5) ∧ (delta>0.25) 的基因(保持稀疏结构,避免把零整体抬高),clip ≥0;λ=0.5(环境变量 `LAMBDA` 可覆盖)。非 branching 细胞原样保留。+5. 双输入方向校验(X3 / final 生效):按共同标签加权计算前一阶段→末阶段的伪批量差 d,与平均局部位移方向 mean_shift 求余弦;cos<0.1 → 跳过位移(输出=输入复制),0.1≤cos<0.3 → λ 减半(比 PLAN 的"减半"多加了更硬的丢弃门,依据是未加门时 X3 分数下降,见下)。+6. 输出细胞数 = min(末阶段细胞数, max_cells),`default_rng(seed)` 无放回均匀抽样。++## 关键参数++k=20,HVG=2500,PCA=25,committed_frac=0.20,readiness_thr=0.15(PLAN 风险节建议试 0.15;0.25 时 proxy 只有 461 个 branching,0.15 → 582),λ=0.5(试 λ=1.0 得 53.14 vs 53.10,噪声内,保留 0.5),mask: target>0.5, delta>0.25。++## off 对照(mechanism_off_control)++同一程序 `--off`(或 `LAMBDA=0`)→ λ=0,输出=输入的均匀子抽样(copy_last 等价)。提交版保持打开(λ=0.5,无 off 开关生效)。++## 机制生效证据(proxy E8.5→E9.5,seed 0)++- committed 2265 / pool 11323;branching(被修改)582(3.5%)。+- readiness 分布 q10–q90 = 0,0,0,0,0.9(重尾:只有心脏区细胞有 committed 邻居)。+- 逐细胞位移范数 mean=4.48, q10=2.36, q50=4.17, q90=7.09(异质,非单一常数)。+- 平均位移最大的基因:Myh6, Kcnh7, Ankrd1, Myl3, Slc2a1, Acta1 —— 心脏成熟程序方向,符合 Isl1+→Nkx2-5+ 边的生物学。+- X3(E8.75+E9.0 心脏):branching 候选 407(18.7%),但方向校验 cos=−0.074 < 0.1 → 按规则跳过位移,on 输出与 off 逐字节相同(门在起作用,不是机制没跑)。++## 查分结果(A 半,seed 0)++| 视图 | on | off(对照) | 分项(on) de_score/de_dir/mmd/vario |+|---|---|---|---|+| proxy10 | 53.10 | 52.97 | 12.90/14.69/15.55/9.96 |+| X3 | 47.62(=off,门跳过) | 47.62 | 10.77/12.28/14.85/9.72 |++早期无门版本在 X3 上:v1(dense mask target>0.05)45.69,variogram 掉到 8.02;v2(sparse mask)46.31,variogram 8.74 —— 说明"向邻居均值位移"在 cos≈0(方向未验证)时损害共变结构,因此加了丢弃门。proxy 上 on−off = +0.13,在 ±2 噪声内:**机制在 proxy 上方向正确(top 基因是心脏成熟基因)但幅度上未证明有效**,如实报告。++## 验证过 / 没验证++- 验证过:proxy 与 X3 视图跑通、vec-check ok、seed 0 确定性(两次运行逐字节一致)、off 对照、X3 门触发后 on==off。+- 没验证:proxy2 视图(本节点无该视图目录,但代码只依赖 manifest schema,双输入路径已在 X3 走过);伪装视图重跑(代码不读绝对时间、路径、board/mode 字段);final(E8.5+E9.5,cos 校验可能通过并实际位移)。+- λ、readiness 阈值只在 proxy10 上粗试,未做网格。++## 知识来源++- 标记基因角色(Nkx2-5/Tnnt2/Myh6/Actc1/Tnnt1/Tbx5 = 心脏承诺/成熟;Isl1/Mef2c = SHF 祖细胞):通用谱系知识与 PLAN motivation 引用的 k063(Jia 2018, de Soysa 2019, Ameen 2022,全部 ≤E9.5 已发表阶段)。+- 未使用任何保留阶段(E10.5/E12.5、禁窗 9.5<E≤13.5)或保留基因型的测量信息;未读外部数据文件(proxy 视图 external 为空,X3 的 external 未使用);未读 prior/ 文件;未读 `uns.celltype_palette`(视图中已剥掉)。diff --git a/solution/run.py b/solution/run.pynew file mode 100644index 0000000..5912529--- /dev/null+++ b/solution/run.py@@ -0,0 +1,250 @@+"""lineage_branching: local-neighborhood commitment shift along Isl1+ -> Nkx2-5+ edge.++For each cell of the latest input stage, compute a commitment score+(mean of cardiac commitment markers - mean of progenitor markers). Cells in the+top fraction of this score among cardiac/mesodermal-lineage cells are+"committed". Each non-committed cell gets readiness = fraction of its k nearest+neighbors that are committed; cells above a readiness threshold are "branching"+and are shifted toward the mean expression of their own committed neighbors,+magnitude scaled by readiness and lambda. Non-branching cells are unchanged.++Off control: LAMBDA env var or --off sets lambda=0 (pure copy of subsampled+input). Mechanism uses only within-stage data + generic marker knowledge+(cardiac commitment program: Nkx2-5/Tnnt2/Myh6/Actc1 up, Isl1/Mef2c progenitor),+no reserved-stage measurements.+"""+import argparse+import json+import os+import sys++import numpy as np+import scipy.sparse as sp+++# Generic cardiac-lineage knowledge (marker roles from public developmental+# biology literature on SHF commitment, active E7.5-E9.5; not stage-specific+# measurements):+COMMITMENT_MARKERS = ["Nkx2-5", "Tnnt2", "Myh6", "Actc1", "Tnnt1", "Tbx5"]+PROGENITOR_MARKERS = ["Isl1", "Mef2c"]+# Keywords for cardiac / lateral-plate-mesoderm lineage labels (official vocab).+CARDIAC_KEYWORDS = ["CM", "SHF", "Pericardium", "Endocardium", "JCF",+ "Mesoderm", "Blood", "Endothelium", "NCC", "Unknown",+ "Heart", "Epicard", "Proepicardium"]++K_NEIGHBORS = 20+N_HVG = 2500+N_PCS = 25+COMMITTED_FRAC = 0.20+READINESS_THRESHOLD = 0.15+MIN_COMMITTED = 30+TARGET_HI = 0.5+DELTA_MIN = 0.25+DEFAULT_LAMBDA = 0.5+++def log(msg):+ print(f"[lineage_branching] {msg}", file=sys.stderr, flush=True)+++def gene_idx(genes, names):+ pos = {g: i for i, g in enumerate(genes)}+ return [pos[n] for n in names if n in pos]+++def is_cardiac_label(labels):+ mask = np.zeros(len(labels), dtype=bool)+ for i, lab in enumerate(labels):+ s = str(lab)+ for kw in CARDIAC_KEYWORDS:+ if kw in s:+ mask[i] = True+ break+ return mask+++def marker_score(X, genes, up_names, dn_names):+ """mean(log expr of up markers) - mean(log expr of dn markers), per cell."""+ up = gene_idx(genes, up_names)+ dn = gene_idx(genes, dn_names)+ if not up:+ return None+ Xu = X[:, up].toarray() if sp.issparse(X) else np.asarray(X)[:, up]+ s = Xu.mean(axis=1)+ if dn:+ Xd = X[:, dn].toarray() if sp.issparse(X) else np.asarray(X)[:, dn]+ s = s - Xd.mean(axis=1)+ return s.astype(np.float64)+++def main():+ ap = argparse.ArgumentParser()+ ap.add_argument("--data", required=True)+ ap.add_argument("--out", required=True)+ ap.add_argument("--seed", type=int, required=True)+ ap.add_argument("--off", action="store_true", help="off-control: lambda=0")+ args = ap.parse_args()++ from src.task1_temporal import view_io++ view = args.data+ manifest = view_io.load_manifest(view)+ genes = view_io.panel_genes(view, manifest)+ rng = np.random.default_rng(args.seed)++ lam = float(os.environ.get("LAMBDA", DEFAULT_LAMBDA))+ if args.off:+ lam = 0.0+ log(f"lambda={lam}")++ inputs = view_io.inputs_by_time(manifest)+ last = inputs[-1]+ adata = view_io.read_stage(view, last, genes)+ X = adata.X.tocsr().astype(np.float32)+ n_cells, n_genes = X.shape+ labels = view_io.labels_of(adata)+ log(f"last stage {last.get('stage')}: {n_cells} cells x {n_genes} genes")++ # ---- mechanism: local commitment shift ----+ score = marker_score(X, genes, COMMITMENT_MARKERS, PROGENITOR_MARKERS)+ n_committed = 0+ n_branching = 0+ Xout = X+ if score is not None and lam > 0:+ card = is_cardiac_label(labels)+ pool = np.flatnonzero(card)+ if len(pool) >= MIN_COMMITTED:+ thr = np.quantile(score[pool], 1.0 - COMMITTED_FRAC)+ committed = np.zeros(n_cells, dtype=bool)+ sel = pool[score[pool] >= thr]+ # require some commitment-marker signal to avoid noise-only picks+ up_idx = gene_idx(genes, COMMITMENT_MARKERS)+ if up_idx:+ m = np.asarray(X[:, up_idx].sum(axis=1)).ravel() > 0+ sel = sel[m[sel]]+ committed[sel] = True+ n_committed = int(committed.sum())+ log(f"committed cells: {n_committed} / pool {len(pool)}")++ if n_committed >= MIN_COMMITTED:+ # HVG + PCA + kNN+ Xd = X+ mu = np.asarray(Xd.mean(axis=0)).ravel()+ sq = np.asarray(Xd.multiply(Xd).mean(axis=0)).ravel()+ var = np.maximum(sq - mu * mu, 0)+ hvg = np.argsort(-var)[:N_HVG]+ hvg.sort()+ from sklearn.decomposition import TruncatedSVD+ from sklearn.neighbors import NearestNeighbors+ Z = Xd[:, hvg].toarray().astype(np.float64)+ Z -= Z.mean(axis=0)+ svd = TruncatedSVD(n_components=min(N_PCS, Z.shape[1] - 1, Z.shape[0] - 1),+ random_state=args.seed)+ P = svd.fit_transform(Z)+ nn = NearestNeighbors(n_neighbors=min(K_NEIGHBORS + 1, n_cells))+ nn.fit(P)+ dist, idx = nn.kneighbors(P)+ idx = idx[:, 1:] # drop self+ k = idx.shape[1]++ committed_nbrs = committed[idx]+ readiness = committed_nbrs.mean(axis=1)+ branching = (readiness > READINESS_THRESHOLD) & (~committed)+ n_branching = int(branching.sum())+ log(f"branching cells: {n_branching} ({100.0*n_branching/n_cells:.1f}%)")+ log(f"readiness quantiles: {np.quantile(readiness,[0.1,0.25,0.5,0.75,0.9]).round(3)}")++ if n_branching > 0:+ # per-cell target = mean of its committed neighbors (chunked)+ rows = np.flatnonzero(branching)+ CHUNK = 256+ tgts = np.zeros((len(rows), n_genes), dtype=np.float32)+ for s0 in range(0, len(rows), CHUNK):+ rr = rows[s0:s0 + CHUNK]+ nbr = idx[rr]+ cb = committed_nbrs[rr]+ for j in range(len(rr)):+ cn = nbr[j][cb[j]]+ if len(cn) == 0:+ continue+ tgts[s0 + j] = np.asarray(X[cn].mean(axis=0)).ravel()+ blk = X[rows].toarray().astype(np.float64)+ tgt = tgts.astype(np.float64)+ delta = tgt - blk+ mask = (tgt > TARGET_HI) & (delta > DELTA_MIN)+ masked_delta = delta * mask++ # mean local shift direction (lam-invariant; cosine is scale-free)+ mean_shift = (readiness[rows][:, None] * masked_delta).mean(axis=0)++ # ---- two-input direction validation (before applying) ----+ lam_eff = lam+ if len(inputs) >= 2:+ prev = view_io.read_stage(view, inputs[-2], genes)+ lab0 = view_io.labels_of(prev)+ common = sorted(set(lab0.tolist()) & set(labels.tolist()))+ X0 = prev.X.tocsr().astype(np.float32)+ d = np.zeros(n_genes, dtype=np.float64)+ wsum = 0.0+ for c in common:+ m0 = np.flatnonzero(lab0 == c)+ m1 = np.flatnonzero(labels == c)+ if len(m0) < 5 or len(m1) < 5:+ continue+ pb0 = np.asarray(X0[m0].mean(axis=0)).ravel()+ pb1 = np.asarray(X[m1].mean(axis=0)).ravel()+ w = min(len(m0), len(m1))+ d += w * (pb1 - pb0)+ wsum += w+ if wsum > 0 and np.linalg.norm(d) > 0 and np.linalg.norm(mean_shift) > 0:+ d /= wsum+ cos = float(np.dot(d, mean_shift) /+ (np.linalg.norm(d) * np.linalg.norm(mean_shift)))+ log(f"two-input direction cosine={cos:.3f}")+ if cos < 0.1:+ log("cosine < 0.1: local edge direction not supported by observed temporal delta -> skip shift")+ lam_eff = 0.0+ elif cos < 0.3:+ log("0.1 <= cosine < 0.3: halving lambda")+ lam_eff = lam / 2.0++ if lam_eff > 0:+ step = lam_eff * readiness[rows][:, None] * masked_delta+ new = np.maximum(blk + step, 0.0)+ sn = np.linalg.norm(step, axis=1)+ log(f"shift norms: mean={sn.mean():.3f} q10={np.quantile(sn,0.1):.3f} "+ f"q50={np.quantile(sn,0.5):.3f} q90={np.quantile(sn,0.9):.3f}")+ if np.linalg.norm(mean_shift) > 0:+ log(f"genes moved (>0.01 mean): {int((mean_shift>0.01).sum())}, "+ f"top: {[genes[i] for i in np.argsort(-mean_shift)[:8]]}")+ Xl = Xout.tolil()+ for j, i in enumerate(rows):+ nz = np.flatnonzero(new[j] > 0)+ Xl.rows[i] = nz.tolist()+ Xl.data[i] = new[j][nz].astype(np.float32).tolist()+ Xout = Xl.tocsr()+ n_modified = len(rows)+ else:+ n_modified = 0+ log(f"modified cells: {n_modified}")+ else:+ log("mechanism not applied (no markers or lambda=0)")++ # ---- subsample to allowed range ----+ n_out = view_io.target_n_cells(manifest, Xout.shape[0])+ if n_out < Xout.shape[0]:+ take = rng.choice(Xout.shape[0], size=n_out, replace=False)+ take.sort()+ Xout = Xout[take]+ log(f"output cells: {Xout.shape[0]}")++ if sp.issparse(Xout):+ Xout = Xout.tocsr()+ else:+ Xout = sp.csr_matrix(Xout)+ view_io.write_prediction(Xout, genes, args.out, seed=args.seed)+ log(f"wrote {args.out}")+++if __name__ == "__main__":+ main()
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
用到的知识库条目
| 编号 | 标题 | 出处 |
|---|---|---|
| k041 | Within-stage pseudotime and graph toolkit offline: scanpy DPT/PAGA/Leiden, Palantir, CellRank 2 | 10.1186/s13059-019-1663-x (PAGA); 10.1038/s41587-019-0068-4 (Palantir); 10.1038/s41592-024-02303-9 (CellRank 2) |
| k063 | Cardiac progenitor transitions: Isl1+ multipotent attractor state vs Nkx2-5-driven cardiomyocyte commitment (E7.5-E9.5) | 10.1038/s41467-018-07307-6 (Jia 2018); 10.1038/s41586-019-1414-x (de Soysa 2019); 10.1016/j.cell.2022.11.028 (Ameen 2022) |
| k018 | Damped per-type shift: shrinkage alpha on the observed delta | notes/plan/cards/T1.md |
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
| 改了什么 | 全新 draft(相对 best_seed node 3 重写):实现 lineage_branching——心脏/中胚层标签细胞中按承诺分数(Nkx2-5/Tnnt2/Myh6/Actc1/Tnnt1/Tbx5 减 Isl1/Mef2c)取前 20% 为 committed,kNN(k=20, HVG2500+SVD25) 内 readiness>0.15 的非 committed 细胞向其自身 committed 邻居均值做稀疏位移(λ=0.5·readiness,mask target>0.5∧delta>0.25);双输入时按前后阶段伪批量差做方向校验,cos<0.1 丢弃位移、<0.3 减半。 |
|---|---|
| 各组分数的变化 | cell_state:变坏(proxy10):mmd_u 0.0298→0.0403,得分 -3.97;X3 mmd_u 0.0356→0.0332,得分 +0.73,噪声内。 covariation:变坏:proxy10 variogram 0.001028→0.001228,得分 -1.24(skill 0.560→0.497,跌破地板);X3 0.001377→0.001627,得分 -0.98。 de_recovery:变坏:proxy10 de_score 原始值 0.268→0.013(几乎落到地板),得分 -2.51;X3 -0.19→-0.19,得分 -0.09,噪声内。 direction:变坏:proxy10 de_direction 0.360→0.179,得分 -3.13;X3 0.026→-0.021,得分 -0.46,噪声内。 |
| family_id | lineage_branching |
| 假设是否成立 | 否 |
| 经验 |
|
| mechanism_active | 是 |
| 下一步建议 |
|
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。
分析员
| 角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。 | 分析员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数?大模型一共回复了几轮(每轮可以调用多个工具)。 | 4 |
| 工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。 | 共 4 次:read 2、bash 1、write 1 |
| 用时 | 1 分 |
| token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。 | 输入 18,223 · 输出 1,493 · 思考 1,487 |
| 任务(第一行) | 分析节点 n15 的打分结果,写 ANALYSIS.json |
| 最后的回答(摘录) | ANALYSIS.json written |
| 原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。 | /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/15/analyst.jsonl 18 KB /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/15/analyst.stderr |
工程师
| 角色 | 工程师 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 26 |
| 工具调用 | 共 35 次:bash 30、write 3、read 2 |
| 用时 | 13 分 |
| token 数 | 输入 55,573 · 输出 11,408 · 思考 9,300 |
| 任务(第一行) | 按计划修改程序:lineage_branching: local-neighborhood commitment shift along Isl1→Nkx2-5 edge |
| 最后的回答(摘录) | 完成。draft 节点 15:实现 PLAN 的 `lineage_branching` 最小版——Isl1+→Nkx2-5+ 承诺边上逐细胞向 committed kNN 邻居均值做稀疏位移(λ=0.5·readiness,mask target>0.5∧delta>0.25),双输入时按观测时间差做方向校验(cos<0.1 丢弃、<0.3 减半),off 对照用 `--off`/`LAMBDA=0`。 查分(seed 0,A 半,用 9/20 次):proxy10 on=53.10 vs off=52.97(+0.13,噪声内;机制生效:582 细胞被改,top 位移基因 Myh6/Myl3/Ankrd1);X3 方向校验 cos=−0.07 触发丢弃门 → on=off=47.62(无门版本 X3 掉到 45.7–46.3,variogram 受损,门是必要的)。两视图 vec-check 通过、确定性验证通过。 下一步最值得试:(1) 位移幅度/覆盖太小——proxy 只有 3.5% 细胞被改,可对 committed 邻居均值做"保留残差"式插值并扩大 branching 比例;(2) X3 上门被关说明局部边方向与整体时间差不一致,值得改用型内(per-label)方向校验而不是全局加权差;(3) 与节点 7/13 的组成趋势重加权组合(本 draft 未做组成变化)。 |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/15/engineer.jsonl 132 KB /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/15/engineer.stderr |
调研员
| 角色 | 调研员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 1 |
| 工具调用 | 共 0 次:— |
| 用时 | 2 分 |
| token 数 | 输入 13,875 · 输出 1,433 · 思考 4,452 |
| 任务(第一行) | 为节点 n15(草稿,父节点 nNone)写改进计划 PLAN.json,思路 lineage_branching |
| 最后的回答(摘录) | iogram score after first run. 3) The edge may not be supported in X3 if cell types are named differently; the support check handles this but means mechanism doesn't run. 4) Readiness threshold 0.25 may be too strict/loose; Engineer should report the distribution of readiness scores and try 0.15/0.35 if too few/many cells branch. Early detection: after first vec-score, if branching cell count is 0 or <5% of input, the mechanism isn't running.", "family_id": "lineage_branching", "mechanism": "Per-cell branching along a reviewed lineage edge (Isl1+ progenitor → Nkx2-5+ committed): each cell's shift direction is toward its own committed kNN neighbors, magnitude scales with local readiness score, preserving single-cell residuals for non-branching cells.", "vs_constant_shift": "Direction and magnitude vary per cell: each cell moves toward its own local committed neighbors (not a global or per-type mean delta); magnitude is scaled by the cell-specific readiness score (fraction of committed neighbors); only cells with local evidence are modified. A constant per-type shift would apply the same vector to all cells of a type regardless of local context.", "mechanism_evidence": "Engineer should report: (1) number of committed cells detected and branching cells modified (if 0, mechanism not running); (2) distribution of readiness scores; (3) per-cell shift norm distribution (should be heterogeneous, not a single value); (4) cosine similarity between individual cell shifts and the type-level mean delta (should be <0.95 if truly per-cell varying); (5) four-group score breakdown comparing on vs off.", "mechanism_off_control": "Same program with λ=0 (or equivalently, skip the shift step): all cells output unchanged from input. Expected difference: off-control produces copy_last-equivalent output; on-run modifies a subset of cells. If on and off outputs are identical (branching count = 0), mechanism is not running and must be debugged before scoring.", "sources": [] } ``` |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/15/researcher.jsonl 6 KB /home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s1/nodes/15/researcher.stderr |