Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s0

节点 n6

composition_trend — type-level composition reweighting + within-type maturity selection

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-093415-search-t1-r2-D-s0
父节点n3
子节点n9
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 54.01(+0.3) · X3 48.92(+0.2) · proxy10 64.19(+0.5) · 3 次复测均分 53.91
审查通过 1 越界读取:未发现问题。run.py 只通过 src.task1_temporal.view_io 的 load_manifest/panel_genes/read_stage/write_prediction 访问 --data 视图内数据(run.py:27-34, 216-219),无绝对路径、'..'、外部目录或网络访问;METHOD.md 声明未用 external 数据集和 uns.celltype_palette。; 2 硬编码目标统计量:未发现问题。代码中的常量仅为机制超参数(A_TP=-0.55、A_CM=1.2、A_FATE=0.6、K_STRAT=5、MIN_TYPE…
用时?从运行开始到结束(或到现在)的挂钟时间。37 分
程序版本911056ceda4a30975ee1cf6ab5c62135c08b2c02 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 911056ceda:solution/METHOD.md

composition_trend — type-level composition reweighting + within-type maturity selection

Provenance. Derived from agent-produced node 33 of run 20261002-034201-search-t1-abc-r1-A-era (programs.git refs/nodes/33, commit 10400ee; the program behind the official T1:val 50.1, submissions.tsv 2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed, for the seed contract only (2026-10-03):

  • the external-maturity axis A_EXT (read manifest source == "external"; a view-identity read; it never fired on final / proxy / patched X3);
  • the multi-stage pooling POOL, the composition-trajectory axis A_TR, the kNN expression smoothing A_S, the apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode;
  • the inputs_by_time(include_external=False) base selection: the base is simply the latest input stage by time;
  • the VEC_* environment overrides (constants are fixed in the code).

On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission record for the digest check. The mechanism-split study (notes/reports/dev/2026-10-02_mech_split_50.md, candidate r1A:CUS = r1A:CU) is the analysis of exactly this procedure.

Naming. "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain (+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.

Contract: python run.py --data <view> --out <pred.h5ad> --seed <int>; CPU only (EXECUTION.json {"gpu": false}).

Method

Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own celltype labels (whatever vocabulary the view provides; Unknown is an ordinary type).

  1. Cell scores (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle genes), metabolic maturity z_met = OXPHOS (13 genes) − glycolysis (10 genes).
  2. Within-type commitment z_fate: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean removed, z-scored again (only the order inside a type matters).
  3. Weights. Type layer w_t = exp(−0.55 · z(type-mean z_prolif)) (types with < 5 cells: 1). Cell layer w_i = exp(1.2 · z_met_i + 0.6 · z_fate_i). w = clip(w_t · w_i, 1e-6, 1e6).
  4. Sampling. n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5 equal-frequency z_met strata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling without replacement inside a stratum. Expression values are copied unchanged.

Mechanism-off controls (for analysis; not code switches): A_TP = 0 leaves the input composition (mech_split "U"); A_CM = A_FATE = 0 with the same per-type quotas is uniform within type (mech_split "C").

Data / knowledge used

Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets (pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external dataset, no uns.celltype_palette, no frozen probe, no pre-trained weights.

Hyper-parameters

NameValueWhere it came from
A_TP−0.55node 33's default (tuned by the agent on the old proxy / proxy2)
A_CM1.2node 33's default (same)
A_FATE0.6node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6)
K_STRAT, MIN_TYPE_CELLS5, 5node 33's defaults
HVG / PCA / kNN / diffmap2000 / 30 / 30 / 15node 33's defaults

All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input rulers the mechanism is near neutral (admission record).

Known failure modes

  • The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).
  • Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was −2.4 points (MMD −1.8) on top of the composition (mech_split §3).
  • Real cells only: no new expression states, no new cell types.

调研员的计划

名称composition_trend + size-shrunk type weights + direction-consistency guard
动机Node 3 X3 de_score raw −0.1818 (skill 0.445, below floor 0.5), X3 mmd_u skill 0.472 (below floor). X3 weighted ×2, so these below-floor metrics dominate the node score (48.71 vs proxy10 63.66). The fixed A_TP=−0.55 type reweighting shifts composition in a direction anti-correlated with X3's true DE pattern. Node 4 showed expression shifts along pseudotime hurt X3 de_direction (−1.6 pts). Node 5 OT displacement also failed. The structural gap is: no mechanism to detect and shrink harmful composition shifts.
做法Modify node 3's run.py with two additive mechanisms, keeping all existing code as the base:

1. Size-based empirical Bayes shrinkage of type weights. Replace z_prolif_type with z_shrunk = (n_t/(n_t+k))·z_prolif_type, k=30 (search k∈{10,30,60}). Types with few cells get z_shrunk≈0 → w_t≈1 (no reweighting). This is a continuous generalization of the existing MIN_TYPE_CELLS=5 hard cutoff. Implementation: one line change in the type-weight computation.

2. Direction-consistency guard. After computing the reweighted sample indices but before writing output:
a. Compute per-gene within-stage trend: for each gene g, r_g = Spearman(expression_g, z_fate) across all input cells (using the already-computed z_fate). This gives an 'expected DE direction' vector e∈R^G.
b. Compute the pseudobulk shift dp = mean(reweighted cells) − mean(all input cells).
c. Compute cosine c = cos(dp, e). If c < 0, the reweighting pushes the pseudobulk against the within-stage trend.
d. Apply a global damping factor: if c < 0, multiply all type-layer weights w_t by exp(c·γ) (γ=2.0, search γ∈{1,2,4}), which shrinks them toward 1. If c ≥ 0, no damping.
This is computed once per view, adds <2 s …
风险1. If X3 heart types have coherent proliferation-maturity signals but in the correct direction already, shrinkage reduces a beneficial reweighting and proxy10 drops. Engineer detects by checking proxy10 mmd_u skill after first X3 run. 2. The direction guard's gene-level trend (Spearman with z_fate) may itself be noisy for low-expression genes; mitigate by using only genes with detection rate >10%. 3. If the cosine c is near 0 on all views, the guard never fires and the node equals the parent (zero progress). Engineer checks the printed cosine value. 4. Runtime: Spearman over ~20k genes × ~17k cells adds ~5 s; acceptable within 30 min.

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 6bb7e1fdf1。改动的文件:solution/run.py +83 −5

diff --git a/solution/run.py b/solution/run.pyindex b7ef52b..2d139e2 100644--- a/solution/run.py+++ b/solution/run.py@@ -20,6 +20,7 @@ Deterministic for a given --seed. CPU only. from __future__ import annotations  import argparse+import os  import numpy as np @@ -32,12 +33,17 @@ from src.task1_temporal.view_io import (     write_prediction, ) -A_TP = -0.55      # type layer: lower mean proliferation -> larger share-A_CM = 1.2        # cell layer: metabolic maturity (OXPHOS - glycolysis)-A_FATE = 0.6      # cell layer: within-type diffusion pseudotime (more committed cells)+A_TP = float(os.environ.get("VEC_A_TP", "-0.55"))   # type layer: lower mean proliferation -> larger share+A_CM = float(os.environ.get("VEC_A_CM", "1.2"))     # cell layer: metabolic maturity (OXPHOS - glycolysis)+A_FATE = float(os.environ.get("VEC_A_FATE", "0.6")) # cell layer: within-type diffusion pseudotime K_STRAT = 5       # z_met strata per type MIN_TYPE_CELLS = 5 +# --- node 6 mechanisms (defaults ON; VEC_K_SHRINK=0 VEC_GAMMA=0 recovers node 3 exactly) ---+K_SHRINK = float(os.environ.get("VEC_K_SHRINK", "30"))   # empirical-Bayes shrink of type-mean proliferation+GAMMA = float(os.environ.get("VEC_GAMMA", "2.0"))        # direction-consistency guard strength (0 = off)+DET_THRESH = 0.10                                        # guard uses genes detected in >10% of cells+ CYCLE = [     "Mki67", "Top2a", "Pcna", "Ccna2", "Ccnb1", "Ccnb2", "Ccnd1", "Ccneg", "Ccne1",     "Cdk1", "Cdk2", "Cdk4", "Cdk6", "Mcm2", "Mcm3", "Mcm4", "Mcm5", "Mcm6", "Mcm7",@@ -154,6 +160,52 @@ def stratified_sample(w, z_strat, inv, n_types, n_out, K, rng):     return np.sort(np.concatenate(out)) if out else np.empty(0, dtype=np.int64)  +def gene_trend_spearman(X, z_fate, det_thresh, chunk=2000):+    """Per-gene Spearman correlation of expression with z_fate, only for genes detected+    in > det_thresh of cells (0 elsewhere). Ordinal ranks; chunked dense columns."""+    from scipy.stats import rankdata++    n, g = X.shape+    det = np.asarray((X != 0).sum(axis=0), dtype=np.float64).ravel() / max(n, 1)+    cols = np.where(det > det_thresh)[0]+    e = np.zeros(g, dtype=np.float64)+    if len(cols) == 0:+        return e+    rz = rankdata(z_fate).astype(np.float64)+    rz -= rz.mean()+    rzn = float(np.linalg.norm(rz))+    Xcsc = X.tocsc()+    ar = np.arange(n)+    for s in range(0, len(cols), chunk):+        cc = cols[s:s + chunk]+        D = np.asarray(Xcsc[:, cc].todense(), dtype=np.float32)+        order = np.argsort(D, axis=0, kind="stable")+        R = np.empty(D.shape, dtype=np.int32)+        np.put_along_axis(R, order, np.broadcast_to(ar[:, None], D.shape).astype(np.int32), axis=0)+        R = R.astype(np.float64)+        R -= R.mean(axis=0)+        num = R.T @ rz+        den = np.linalg.norm(R, axis=0) * rzn+        e[cc] = np.where(den > 1e-12, num / np.maximum(den, 1e-12), 0.0)+    return e+++def direction_cosine(X, idx, e):+    """Cosine between the pseudobulk shift dp = mean(sampled) - mean(all cells) and the+    within-stage maturity trend e, restricted to genes where e is defined."""+    mask = e != 0.0+    if not mask.any():+        return 0.0+    pb_all = np.asarray(X.mean(axis=0), dtype=np.float64).ravel()+    pb_sel = np.asarray(X[idx].mean(axis=0), dtype=np.float64).ravel()+    dp = (pb_sel - pb_all)[mask]+    ee = e[mask]+    na, ne = np.linalg.norm(dp), np.linalg.norm(ee)+    if na < 1e-12 or ne < 1e-12:+        return 0.0+    return float(dp @ ee / (na * ne))++ def main() -> None:     ap = argparse.ArgumentParser()     ap.add_argument("--data", required=True)@@ -181,13 +233,39 @@ def main() -> None:     raw = raw - tmean(raw)[inv]          # only the within-type order matters     z_fate = zscore(raw) -    w_type = np.exp(A_TP * z_prolif_t)+    # mechanism 1: empirical-Bayes size shrinkage of the type-level proliferation z-score+    shrink = counts.astype(np.float64) / (counts.astype(np.float64) + K_SHRINK)+    z_tp = z_prolif_t * shrink++    w_type = np.exp(A_TP * z_tp)     w_type[counts < MIN_TYPE_CELLS] = 1.0-    w = np.clip(w_type[inv] * np.exp(A_CM * z_met + A_FATE * z_fate), 1e-6, 1e6)+    w_cell = np.exp(A_CM * z_met + A_FATE * z_fate)+    w = np.clip(w_type[inv] * w_cell, 1e-6, 1e6)      n_out = target_n_cells(manifest, adata.n_obs)     rng = np.random.default_rng(args.seed)     idx = stratified_sample(w, z_met, inv, len(uniq), n_out, K_STRAT, rng)++    # mechanism 2: direction-consistency guard on the type layer+    if GAMMA > 0.0:+        e = gene_trend_spearman(X, z_fate, DET_THRESH)+        c = direction_cosine(X, idx, e)+        print(f"[guard] cosine(dp, trend) = {c:.4f}, genes_used = {int((e != 0).sum())}", flush=True)+        if c < 0.0:+            damp = float(np.exp(c * GAMMA))     # in (0, 1): shrinks w_type toward 1+            w_type = np.exp(A_TP * z_tp * damp)+            w_type[counts < MIN_TYPE_CELLS] = 1.0+            w = np.clip(w_type[inv] * w_cell, 1e-6, 1e6)+            idx = stratified_sample(w, z_met, inv, len(uniq), n_out, K_STRAT, rng)+            print(f"[guard] fired: damping = {damp:.4f}", flush=True)++    print(f"[shrink] k = {K_SHRINK}; type weights (raw -> shrunk):", flush=True)+    for t in range(len(uniq)):+        wr = float(np.exp(A_TP * z_prolif_t[t]))+        ws = float(w_type[t])+        if abs(wr - ws) > 1e-6:+            print(f"  {uniq[t]}: n={int(counts[t])} w_raw={wr:.3f} w_final={ws:.3f}", flush=True)+     write_prediction(X[idx], genes, args.out, seed=args.seed)  

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k020Correcting sampling-scope (dissection) bias in compositionnotes/guides/modeling_and_evaluation_guide.html
k036Composition forecasting and mixture models for population predictionnotes/competition/05_lineage_graph.md; notes/competition/09_t1_census_lineage.md; 10.1038/s41586-024-08453-2 (growth rates)
k015Composition x conditional-expression decomposition p(x|t) = sum_z p(z|t) p(x|z,t)notes/handover/03_当前方案与Agent系统设计.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么在 node 3 的组成重加权基础上加两个机制:类型权重的经验贝叶斯尺寸收缩(k=30,z_tp = z_prolif·n/(n+k))和方向一致性守卫(若 cos(dp, 基因-z_fate Spearman 趋势) < 0 则以 exp(c·γ) 阻尼类型层,γ=2);三个幅度系数 A_TP/A_CM/A_FATE 改为可用环境变量覆盖。实际提交打分的是默认配置(A_FATE=0.6)。
各组分数的变化cell_state:噪声内偏正:+0.63(proxy10 mmd_u 0.0298→0.0291,+0.35 分;X3 mmd_u 0.0356→0.0352,+0.11 分)
covariation:噪声内:+0.19(proxy10 variogram 0.001028→0.000998,+0.21 分;X3 −0.05 分)
de_recovery:噪声内偏正:+0.48(X3 de_score 原始值 −0.1818→−0.1558,skill 0.445→0.452,+0.18 分;proxy10 de_score 持平 15.11)
direction:噪声内:−0.13(proxy10 de_direction 0.360→0.358,−0.03 分;X3 0.0264→0.0231,−0.03 分)
family_idcomposition_program
假设是否成立unclear
经验
  1. 对类型层组成重加权做数据自适应收缩/守卫时:k=30 收缩 + γ=2 守卫使 X3 de_score 原始值从 −0.1818 升到 −0.1558、四组全部小幅同向或持平,但总榜分只 +0.32,远小于 T1 约 2 分噪声——机制方向正确但效应量太小,不足以成为独立增益点。
  2. Engineer 最后一条消息声称 A_FATE=0 使 X3 ≈52.2(+3.5),但变化量表显示实际提交的是默认 A_FATE=0.6 配置(X3 48.92,+0.21);调参过程中的最好配置没有成为提交物,声称与打分结果冲突时以变化量表为准。
  3. PLAN 要求的机制关闭对照(k=0, γ=0 应与 node 3 字节一致)和守卫余弦值的证据(每视图 c、阻尼是否触发)都没有留下,无法把 +0.32 归因到收缩还是守卫;加了打印就要把打印值写进交付说明。
mechanism_activeunclear
下一步建议
  1. 针对 direction/de_recovery(尤其 X3):正式提交并打分 VEC_A_FATE=0 的完整配置,验证 Engineer 声称的 X3 +3.5;若成立再扫 A_FATE∈{0, 0.3} × A_CM∈{0.8, 1.2},每配置两把尺子都打分。
  2. 针对 cell_state:在 A_FATE=0 的基础上保留 k=30 收缩,单独跑 VEC_GAMMA=0 对照并记录每视图打印的余弦 c,确认守卫是否真的触发、触发时 X3 与 proxy10 的 c 符号是否如 PLAN 预期相反。
  3. 任何后续提交前先跑机制关闭对照(VEC_K_SHRINK=0 VEC_GAMMA=0 VEC_A_FATE=0.6)确认与父节点输出一致,并把提交配置的环境变量写进 METHOD/最终消息,避免再次提交非最优配置。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。4
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 3 次:bash 2、write 1
用时不到 1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 11,528 · 输出 1,126 · 思考 1,298
任务(第一行)分析节点 n6 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
ANALYSIS.json written
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/analyst.jsonl 9 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/analyst.stderr

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次(失败 1 次)
轮数27
工具调用共 31 次:bash 22、edit 6、read 2、todowrite 1
用时30 分
token 数输入 59,115 · 输出 6,419 · 思考 18,249
任务(第一行)按计划修改程序:composition_trend + size-shrunk type weights + direction-consistency guard
最后的回答(摘录)
A_FATE=0 gives X3 ≈ 52.2 (+3.5 over parent). Verifying on proxy10 now while preparing METHOD.md:
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/engineer.jsonl 205 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/engineer.stderr

调研员

角色调研员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数1
工具调用共 0 次:—
用时3 分
token 数输入 15,707 · 输出 1,530 · 思考 8,105
任务(第一行)为节点 n6(改进,父节点 n3)写改进计划 PLAN.json
最后的回答(摘录)
 is near 0 on all views, the guard never fires and the node equals the parent (zero progress). Engineer checks the printed cosine value. 4. Runtime: Spearman over ~20k genes × ~17k cells adds ~5 s; acceptable within 30 min.",
  "family_id": "composition_program",
  "mechanism": "Data-adaptive shrinkage of type-level composition reweighting: (a) empirical Bayes shrinkage of type-mean proliferation toward zero based on type size, preventing small/noisy types from driving composition shifts; (b) a direction-consistency guard that dampens the composition shift when the resulting pseudobulk change is anti-correlated with the within-stage gene-level maturity trend.",
  "vs_constant_shift": "Constant displacement shifts every cell's expression by the same vector per type, changing expression values. This mechanism changes which cells are sampled (composition), never modifies expression. The shrinkage and guard are data-driven (depend on type sizes and gene-level correlations in the input), not fixed offsets. The direction guard is a feedback mechanism absent from any constant-shift or fixed-reweighting scheme.",
  "mechanism_evidence": "Engineer should report: (1) the cosine c value per view (X3 vs proxy10) — if c<0 on X3 but c>0 on proxy10, the guard is selectively protecting X3; (2) the effective type weights before and after shrinkage — show which types had weights changed; (3) per-metric breakdown: X3 de_score skill should rise from 0.445, X3 mmd_u skill from 0.472; proxy10 metrics should stay within noise; (4) the number of cells whose sampling probability changed by >20%.",
  "mechanism_off_control": "Set k=0 (no size shrinkage, recover raw type-mean z_prolif) and γ=0 (direction guard disabled, damping factor always 1). This recovers node 3's exact weights and sampling. The output should be byte-identical to node 3. If it is not, there is a code bug. If it is, any score difference in the mechanism-on run is attributable to the shrinkage+guard.",
  "sources": []
}
```
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/researcher.jsonl 7 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/researcher.stderr

审查员

角色审查员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数6
工具调用共 10 次:read 8、bash 1、write 1
用时1 分
token 数输入 20,823 · 输出 1,562 · 思考 2,326
任务(第一行)审查节点 n6 的程序是否越界读取、写死目标、钻评分器漏洞
最后的回答(摘录)
REVIEW.json written
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/reviewer.jsonl 97 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/6/reviewer.stderr