Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

← 返回总览

单 Agent 运行

20261003-155754-t3-gata4-komix-b

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-155754-t3-gata4-komix-b
方式?自动搜索:ERA 式搜索树,多个节点不断改进程序;单 Agent:一个 Agent 会话从头做到尾。单 Agent
框架 / 模型opencode / alibaba-token-plan-cn/qwen3.8-max
题目?比赛的哪道题、哪个阶段,例如 T1:val 是第 1 题的验证阶段。T3:gata4
状态已结束
后台服务?在 Spark 上以 systemd 用户服务运行的程序,例如每次运行和这个看板本身。vec-t3-gata4-komix-b-20261003-155754
开始 / 结束10-03 15:57 / 10-03 16:59
预算4 小时,最多 400 轮
启动时代码有未提交改动?启动时代码仓库有未提交的修改,这次运行不能完全按 git 版本复现。否
最终文件T3__gata4.h5ad

最终选择说明 SELECTION.md

SELECTION — run 20261003-155754-t3-gata4-komix-b

Board in this run's harness header: T3:gata4 only.

T3:gata4 — chosen submission

final/T3__gata4.h5ad

python scripts/t3_komix.py \
  --wt data/raw/official/t3/WT_E8.75.h5ad \
  --ko data/raw/official/t3/E9.5_mab21l2_ko.h5ad \
  --out final/T3__gata4.h5ad \
  --beta 0.40 --k 0.55 --seed 0 --max-cells 7449 --zero-gene Gata4
python scripts/verify_submission.py --input final/T3__gata4.h5ad --board T3:gata4   # passed

Method: T3 method card v2 main line komix — real cells of the observed training knockout (Mab21l2 KO E9.5, official training data) mixed into the stage-matched WT carrier (WT E8.75), each cell keeping its own expression and its own spatial_3D; no averaging, no shift transfer, no per-gene editing, no densification. Plus the card's optional --zero-gene component for the board's knocked-out gene.

outputvalue
cells7 449 (= manifest max_cells, ≥ min_cells 1 000)
WT E8.75 cells / Mab21l2 KO E9.5 cells4 097 / 3 352
ko_fraction (uns["t3_komix"])0.4500
nonzero fraction vs WT input0.940 (hard check 0.8–1.25 → ok)
genes500, panel order, obsm["spatial_3D"] present, no NaN
paramsβ=0.40, k=0.55, seed=0, zero_gene=Gata4, ko_exclude_half=none

Disclosure required by the method card: the output contains real cells of the training knockout sample (Mab21l2 KO E9.5, official training data); they are 45.0 % of the submitted cells (uns["t3_komix"]: n_ko 3 352 / n 7 449, ko_fraction 0.4500), and their coordinates come from the same file.

Why this variant
  1. The card's main line at this run's start point. The harness gave variant b the start β=0.40, k=0.55; the a/b/c variants differ only there, and the comparison across runs is by official quota, so the start point is kept unchanged (see 3: the knobs are second-order).
  2. --zero-gene Gata4 is on. It is the only component whose sign is verified by allowed training data: in the official Mab21l2 KO file the targeted gene's transcript is essentially gone (pseudobulk 0.0081 and 0.19 % of cells nonzero, versus 0.976 / 18.7 % in WT E9.5), so a knockout of Gata4 should make Gata4 the board's largest down-response. Setting it to 0 in every output cell only turns nonzeros into zeros (never densifies; density ratio 0.940). In the hypothesis scoring below it is worth +0.2 to +0.6 response-group points and is never negative, under every hypothesis about the unseen truth that I could build from allowed files.
  3. β and k are second-order. dp = (1−k)·d_mix + k·ε with std(ε) ≈ 0.023 versus std(d_mix) ≈ 0.427, and the two rank-based response metrics are invariant to the scale of dp: over β∈{0.40,0.55,0.80,1.00} × k∈{0.45,0.55,0.65} the response groups move by ≤1.5 points (all of it through severity_slope) and the 20-point cell-state group by ≤1 point (analysis/t3_grid.py, analysis/t3_cellstate.py, analysis/LOG.md). No local ruler score was used to pick them (method card v2: lgo is documented as rank-inverted against the official board; lgo875's target is komix's own carrier difference, so scoring komix with it is circular — record only, see below).
Honest statement of the residual risk

d_mix = pb(KO E9.5) − pb(WT E8.75) = d_dev + d_ko is dominated by the E8.75→E9.5 progression (||d_dev|| = 9.84 vs ||d_ko|| = 5.64; corr(d_mix,d_dev) = +0.83, corr(d_mix,d_ko) = +0.28, and the two components themselves correlate −0.30). If the Gata4 truth were a pure developmental arrest (dt ≈ −d_dev), komix would land ~7 points below the floor on the response groups; if it is shaped like the observed knockout (dt ≈ d_ko), ~14 points above it (analysis/LOG.md §2). This cannot be settled locally.

Two things decide the bet in favour of keeping the komix line:

  • k is not a hedge. Because de_score and de_direction are rank-based and std(ε) is 18× smaller than std(d_mix), the direction of dp — and therefore the whole downside under H_delay — is essentially identical for every k ≤ 0.75 (32.8 / 32.8 / 32.9 response points at k = 0.45 / 0.55 / 0.65). Raising k does not buy protection; it only shrinks severity_slope towards its floor and, in the limit k = 1, becomes a WT copy, whose dp is pure sampling / batch noise (the case the card flags as "照抄 WT 不是 50 分").
  • The curators' official-score calibration behind card v2 ranks programs by exactly this carrier direction. scripts/ruler_calibration_t2t3.py (in this workspace) documents three official T3:gata4 points for pre-komix programs — 45.9 / 53.4 / 41.0 — with the most conservative one at 41.0, and states that the lgo875 ruler ("rewards moving from WT E8.75 towards the knockout sample, penalises developmental delay") reproduced their ordering and sign, while lgo was inverted. That is the evidence card v2 rests on; it is three points, it is not independent of the scoring service's view, and I could not reproduce the calibration here (it re-runs programs from other runs' workspaces), so it is cited as corroboration, not proof.
Local ruler (record only, not a selection criterion, not an estimate of the official score)

Pipeline check (task book §5.1): scripts/score_proxy.py --board T3:gata4 --baselines reproduces the local ruler's floor / ceiling — wt_identity 50.00, split_half 100.00 — and the submitted file is reproducible from the command above (re-ran it and compared X, spatial_3D and the labels: identical).

final/proxy/T3__gata4.h5ad is the identical method and parameters, generated with --ko-exclude-half B so that the cells the local ruler uses as truth (the second split_half(seed=0) half of the Mab21l2 KO file) are not in the file — without that flag the recording would be circular. Scored with scripts/score_proxy.py --board T3:gata4 (T3:gata4/proxy_mab21l2, reference t2_heart/E9.5, truth = Mab21l2 KO half B, 25 147 cells):

proxy_score T3:gata4 69.71

Groups: response_genes 62.0, direction 69.2, magnitude 96.1, cell_state 48.8 (modeling/experiments/logs/proxy/T3:gata4__komix_b040_k055_zgata4_exclB.json; ledger rows in notes/changelogs/experiments.tsv).

This is a local proxy ruler on a different gene and a different stage, and it is partly circular even with half B excluded: the knockout cells in the prediction come from half A of the very sample whose half B is the ruler's truth (same embryo, same slide, same batch), which is why magnitude reads 96.1. The method card also documents this ruler family (lgo/lgo10) as rank-inverted against the official board. Per the card it is recorded here for the harness's re-scoring check only; the submission above was chosen on mechanism, not on this number.

External data disclosure (外部数据披露)

无 / none. Only the official training files under data/raw/official/t3/ (WT_E8.75.h5ad, WT_E9.5.h5ad, E9.5_mab21l2_ko.h5ad) and data/reference/panels/ were read. No network access, no external dataset, no pretrained model. No held-out genotype (Gata4 or β-catenin knockout, any allele, any phenotypically similar perturbation) and no forbidden-window data was read or used.

Files

pathwhat
final/T3__gata4.h5adsubmission (β=0.40, k=0.55, seed 0, --zero-gene Gata4)
final/proxy/T3__gata4.h5adsame method on the local-ruler input (--ko-exclude-half B)
analysis/LOG.mdmeasurements behind every statement above
analysis/t3_decompose.py, analysis/t3_cellstate.py, analysis/t3_grid.py, analysis/t3_grid.jsonanalysis scripts / results
out/komix_b040_k055_s0.h5ad, out/komix_b040_k055_s0_zgata4.h5adgenerated candidates
notes/changelogs/experiments.tsvone row per experiment

出处?这次运行用的代码、配置、数据和模型的指纹,靠它们可以原样找回并复现。

代码版本?运行锁定时代码仓库的 git 提交号;带 +dirty 表示当时有未提交的改动。bc1f100a4d3d9e4d4e1ee7eb6f5c2826c75710db
版本标签?可提交运行在锁定的提交上打的 git 标签 run/<运行编号>,以后能原样找回代码。—
配置文件agent/configs/experiments/g24_t2_t3/t3_komix_b.yaml
配置指纹 a92a81fb9e6e499b
锁定指纹?启动时把配置、代码、提示词等全部锁定后算出的总指纹。3065ff3af1a426fe 锁定格式 v4
运行专用代码副本?启动时为这次运行单独检出的一份只读代码;运行全程只读它,不受主副本更新影响。没有(这次运行早于运行专用代码副本功能,读的是启动时的代码副本)
大模型?每个角色用的大模型;括号里是接口实际报告的模型名。Agent:alibaba-token-plan-cn/qwen3.8-max
运行目录/home/spark-longxinyang/vec/runs/formal/20261003-155754-t3-gata4-komix-b 主机 spark-ad3f
提交账本?notes/changelogs/submissions.tsv:每次上传官网的记录和返回的官网分。这次运行没有提交过官网
迭代报告?运行结束后自动生成的中文复盘报告(notes/reports/runs)。还没有迭代报告