单 Agent 运行
20261002-215432-t3-gata4-g24q
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261002-215432-t3-gata4-g24q |
|---|---|
| 方式?自动搜索:ERA 式搜索树,多个节点不断改进程序;单 Agent:一个 Agent 会话从头做到尾。 | 单 Agent |
| 框架 / 模型 | opencode / alibaba-token-plan-cn/qwen3.8-max |
| 题目?比赛的哪道题、哪个阶段,例如 T1:val 是第 1 题的验证阶段。 | T3:gata4 |
| 状态 | 已结束 |
| 后台服务?在 Spark 上以 systemd 用户服务运行的程序,例如每次运行和这个看板本身。 | vec-t3-gata4-g24q-20261002-215432 |
| 开始 / 结束 | 10-02 21:54 / 10-02 22:36 |
| 预算 | 1 小时 30 分,最多 150 轮 |
| 启动时代码有未提交改动?启动时代码仓库有未提交的修改,这次运行不能完全按 git 版本复现。 | 否 |
| 最终文件 | T3__gata4.h5ad |
最终选择说明 SELECTION.md
SELECTION — run 20261002-215432-t3-gata4-g24q
Board in scope: T3:gata4 (predict the E8.75 Gata4 knockout from WT E8.75 / WT E9.5 / E9.5 Mab21l2 KO).
Selected submission
final/T3__gata4.h5ad — 7449 cells x 500 genes, obsm["spatial_3D"] = the WT coordinates of the
kept cells, no cell-type column submitted. Format check: scripts/verify_submission.py → SUCCESS.
Method: prior_ko = self-zero + curated regulatory prior + lineage-program term.
Code modeling/src/task3_perturbation/prior_ko.py, driver scripts/t3_prior.py,
parameters modeling/src/task3_perturbation/gata4_prior_params.json.
python scripts/t3_prior.py --gene Gata4 \
--wt data/raw/official/t3/WT_E8.75.h5ad \
--out final/T3__gata4.h5ad \
--params modeling/src/task3_perturbation/gata4_prior_params.jsonThree channels, in order:
- Self-zero (
self_mix=1). Every cell'sGata4is set to 0. This is the only channel the method card could validate locally (Mab21l2 detection 0.19% in KO vs 18.7% in WT), and it is what carries the submission above the wt_identity floor. - Curated prior (
prior_scale=5.5, 110 genes). A hand-written table of expected log-fold-change direction/weight for Gata4 loss, applied multiplicatively in linear (CP10k) space, so zeros stay zeros and the sparsity/covariance structure of the wild type is preserved:- weight −1.0 — sarcomere / contractile / excitation–contraction / natriuretic program (Ttn, Myl7, Myl4, Myh7, Lmod1, Nebl, Csrp3, Tcap, Ankrd1, Rbm24, Fbxo32, Fbxl22, Mybphl, Ablim3, Sorbs2, Mtus2, Thsd4, Cpne5, Sln, Pvalb, Ckmt2, Pln, Casq1, Ryr3, Scn5a, Kcnj3, Cacna1d, Cacnb2, Cacna2d2, Hcn4, Gja5, Gja4, Trpc3/6, Nppa, Nppb, Timp4).
- weight −0.5 — the cardiac transcription-factor / signalling network Gata4 sits in (Tbx5, Mef2c, Hand2, Hand1, Isl1, Bmp10, Hey2, Nfatc1, Gata5, Tbx2, Tbx3, Prdm6, Bmp2, Bmp4, Wnt2, Nrg1).
- weight −0.3 — Gata4-dependent non-myocardial compartments: proepicardium/septum transversum (Msln, Wt1, Tbx18, Upk3b, Upk1b, Postn), endocardium (Cdh5, Pecam1), foregut/gut endoderm (Afp, Apoa1, Apob, Apoc2, Fgb, Serpina1b, Aldob, Ahsg, Itih5, Colec10, Hhex, Foxa2, Muc16).
- weight +0.15 — mesenchyme / non-cardiac lineages that relatively expand when the myocardial program fails, plus paralogue compensation (Col1a1/2, Col3a1, Col5a1, Vcan, Fmod, Cthrc1, Tgfb2, Pdgfrb, Twist1, Msx2, Sox10, Ngfr, Tbx1, Six1, Sox2, Ptn, Crabp1, Sfrp1, Foxc1, Gata6).
- Lineage-program term (
generic_down=0.45,generic_up=0.12,generic_ref=2.5). Purely data-driven and gene-agnostic: a gene is pulled down in proportion to how strongly it is enriched in the compartment that expresses the knocked gene, and a gene is nudged up in proportion to how strongly it is enriched in the complementary compartment. Enrichment must also be large relative to the gene's own level (generic_min_rel=0.5) so broadly expressed genes (H19, Sptbn1, Ezr) are not swept in. This adds graded coverage of cardiac genes outside the hand-written list (e.g. Thsd4, Arhgap31, Agtr2, Gyg, Sfrp5, Ccdc80) which is what givesde_score's ranked down-set its tail.
All three channels are gated per cell type by mean(Gata4)/1.0, clipped to [0,1] and zeroed
below 0.30 (gate_ref=1.0, gate_min=0.30). Cell-type means rather than per-cell values because
single-cell dropout would switch the effect off in most of a type's genuinely positive cells. The
gate is ≈1.0 for IFT-CM / V-CM / Endo / JCF / Gut-Endoderm / Unknown, ≈0.6–0.7 for Peri / pPHM /
ExEM / Allantois, and ≈0 for Neural Tube / PAM / CSE / NCC, so the perturbation stays inside the
Gata4 domain.
Cell cap: WT E8.75 has 24826 cells > max_cells=7449, so the output is a cell-type stratified
without-replacement subsample to the cap (method-card safety rule; never random, never below the
cap). Coordinates follow the kept cells. Cell-type labels are used only inside the model (for the
gate, the lineage-program term and the stratification); both submitted files ship an empty
obs — no cell-type column — as the task book requires, and verify_submission.py passes on
both.
Resulting WT→predicted pseudobulk delta: ‖dp‖ = 2.52, 24 genes past the |Δ| ≥ 0.25 DE threshold, largest down = Gata4 −0.92, Myl7 −0.78, Ttn −0.70, Thsd4 −0.66, Lmod1 −0.59, Myh7 −0.57, Myl4 −0.56, Hcn4/Csrp3 −0.55; largest up = Col5a1 +0.18, Twist1 +0.09, Pdgfrb +0.07. For reference, the Mab21l2 KO's own delta has ‖dt‖ = 5.64 with rms 0.56 over its 90 DE genes, so the submitted delta is the right order of magnitude for a real perturbation rather than a rounding artefact.
Proxy evaluation
proxy_score T3:gata4 53.94final/proxy/T3__gata4.h5ad = the same code path on the proxy inputs
(--gene Mab21l2 --wt data/raw/official/t3/WT_E9.5.h5ad, generic_down=generic_up=0,
curated prior empty, cap 7449). It reproduces the method card's naive_zero_n7449 = 53.94
(floor wt_identity = 50.00, full-cell naive_zero = 55.74).
Why the proxy file is not the full Gata4 configuration. The prior for a gene is knowledge about that gene. Mab21l2 has no curated entry and no lineage program to gate on, so for Mab21l2 the method degenerates exactly to self-zero + cap — which is what the proxy file contains. Turning the curated Gata4 table on for Mab21l2 would be a different, deliberately mismatched experiment, not "the same method".
Why the proxy cannot measure the part of this submission that matters. The proxy target is the E9.5 Mab21l2 KO, whose WT→KO delta is dominated by single-embryo batch and by composition: the cardiac genes move up in it (Myl7 +1.23, Meg3 +1.37, Thsd4 +1.02, Homer2 +1.00, Ahnak +0.94, Myl4 +0.93, Ankrd1 +0.88, Csrp3 +0.87), i.e. the KO slide simply contains a larger cardiomyocyte fraction than the WT slide. Any prediction that models Gata4 loss as "the cardiac program goes down" is anti-correlated with that delta by construction and would score below the floor on the proxy while being the biologically correct prediction for Gata4. The method card already records the same conclusion from the other direction (co-expression propagation 52.75 → 48.5; composition 48.16). The proxy is therefore used here only for what it can test — format, the self-zero channel, the stratified cap, and the scoring pipeline — which it passes at 53.94.
That batch observation also killed one candidate channel: a deplete term that drops cells of
strongly-affected types to model a hypoplastic heart (it raises ‖dp‖ from 1.5 to 3.5 and is the
cheapest way to buy severity_slope). The only local KO example shows the composition difference
between a KO slide and a WT slide going the other way, so the bet was dropped and deplete=0.0
in the submission.
Why this over the method card's self_zero baseline
The board's four ranked groups (de_score 30%, de_direction 25%, severity_slope 25%, cell_state 20%)
are all computed on the WT→mutant delta. A one-gene delta (self_zero) leaves 499 genes tied at
zero, so de_score and de_direction sit at chance and severity_slope lands at log β ≈ −3.4
(the prediction is ~30× too weak: β = dt_self·dp_self / ‖dt‖² with a single contributing gene).
Spreading a correctly-signed delta over ~110 Gata4-program genes at realistic magnitude is the only
route from "slightly above the floor" to the 64–78 band the top-10 reached, and the metric design
makes it low-variance: severity_slope skill only needs corr(dp, dt) > 0.1 and β within ~[0.3, 3.4]
to score 85+, and de_score is rank-based so the exact scale does not matter. Downside is bounded:
if the prior is wrong the three response groups fall back to roughly their floor values (~45/45/50)
while self-zero is retained, i.e. ≈ −7 from the baseline rather than a disqualification-grade
failure. Upside if it is right is ≈ +15 to +20.
Risks, stated plainly: (i) the curated table is knowledge, not measurement, and cannot be validated locally; (ii) the magnitude is calibrated against the Mab21l2 delta norm, which is itself batch-inflated, so β may land at 0.3–0.6 rather than 1; (iii) the +0.15 up-weights are the weakest part of the prior (a hypoplastic heart's up-genes may be a pure composition effect, which this submission deliberately does not model); (iv) single embryo, single slide.
External data disclosure
None. No network was used and no external file, dataset or pretrained model was read. The only
inputs are data/raw/official/t3/{WT_E8.75,WT_E9.5,E9.5_mab21l2_ko}.h5ad and
data/reference/panels/T3__gata4.genes.txt. The Gata4 prior table is the model's own
domain knowledge of published Gata4 biology (Kuo et al. 1997; Molkentin et al. 1997; Watt et al.
2004 and the cardiac gene-regulatory-network literature) — no Gata4 mutant, β-catenin mutant or
phenotype-similar perturbation measurement was used, and nothing from the reserved E9.5→E13.5
window was used. data/raw/official/t3/E9.5_mab21l2_ko.h5ad was read only to characterise the
proxy target's delta (it is released training data for this board).
出处?这次运行用的代码、配置、数据和模型的指纹,靠它们可以原样找回并复现。
| 代码版本?运行锁定时代码仓库的 git 提交号;带 +dirty 表示当时有未提交的改动。 | db589274e0080508aa963cf0e77e8f05dd2147ef |
|---|---|
| 版本标签?可提交运行在锁定的提交上打的 git 标签 run/<运行编号>,以后能原样找回代码。 | — |
| 配置文件 | agent/configs/experiments/g24_quick/t3_spark.yaml 配置指纹 c499cedd733f97b4 |
| 锁定指纹?启动时把配置、代码、提示词等全部锁定后算出的总指纹。 | 09692a17fb561597 锁定格式 v4 |
| 运行专用代码副本?启动时为这次运行单独检出的一份只读代码;运行全程只读它,不受主副本更新影响。 | 没有(这次运行早于运行专用代码副本功能,读的是启动时的代码副本) |
| 大模型?每个角色用的大模型;括号里是接口实际报告的模型名。 | Agent:alibaba-token-plan-cn/qwen3.8-max |
| 运行目录 | /home/spark-longxinyang/vec/runs/formal/20261002-215432-t3-gata4-g24q 主机 spark-ad3f |
| 提交账本?notes/changelogs/submissions.tsv:每次上传官网的记录和返回的官网分。 | 第 23 行 · 2026-10-02 · T3:gata4 · 预测文件指纹 a60e3820e6150626 · 官网分 45.9 |
| 迭代报告?运行结束后自动生成的中文复盘报告(notes/reports/runs)。 | 还没有迭代报告 |