Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s1

节点 n3

composition_trend — type-level composition reweighting + within-type maturity selection

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-093415-search-t1-r2-D-s1
父节点(种子,没有父节点)
子节点n4、n7
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。种子
状态已打分
分数搜索目标分 53.69 · X3 48.71 · proxy10 63.66 · 3 次复测均分 53.91
审查通过 seed (human-written, locked)
用时?从运行开始到结束(或到现在)的挂钟时间。2 分
程序版本ba77be1df19d877439f08d0b9fd9cac9c2f80b75 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git ba77be1df1:solution/METHOD.md

composition_trend — type-level composition reweighting + within-type maturity selection

Provenance. Derived from agent-produced node 33 of run 20261002-034201-search-t1-abc-r1-A-era (programs.git refs/nodes/33, commit 10400ee; the program behind the official T1:val 50.1, submissions.tsv 2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed, for the seed contract only (2026-10-03):

  • the external-maturity axis A_EXT (read manifest source == "external"; a view-identity read; it never fired on final / proxy / patched X3);
  • the multi-stage pooling POOL, the composition-trajectory axis A_TR, the kNN expression smoothing A_S, the apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode;
  • the inputs_by_time(include_external=False) base selection: the base is simply the latest input stage by time;
  • the VEC_* environment overrides (constants are fixed in the code).

On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission record for the digest check. The mechanism-split study (notes/reports/dev/2026-10-02_mech_split_50.md, candidate r1A:CUS = r1A:CU) is the analysis of exactly this procedure.

Naming. "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain (+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.

Contract: python run.py --data <view> --out <pred.h5ad> --seed <int>; CPU only (EXECUTION.json {"gpu": false}).

Method

Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own celltype labels (whatever vocabulary the view provides; Unknown is an ordinary type).

  1. Cell scores (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle genes), metabolic maturity z_met = OXPHOS (13 genes) − glycolysis (10 genes).
  2. Within-type commitment z_fate: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean removed, z-scored again (only the order inside a type matters).
  3. Weights. Type layer w_t = exp(−0.55 · z(type-mean z_prolif)) (types with < 5 cells: 1). Cell layer w_i = exp(1.2 · z_met_i + 0.6 · z_fate_i). w = clip(w_t · w_i, 1e-6, 1e6).
  4. Sampling. n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5 equal-frequency z_met strata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling without replacement inside a stratum. Expression values are copied unchanged.

Mechanism-off controls (for analysis; not code switches): A_TP = 0 leaves the input composition (mech_split "U"); A_CM = A_FATE = 0 with the same per-type quotas is uniform within type (mech_split "C").

Data / knowledge used

Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets (pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external dataset, no uns.celltype_palette, no frozen probe, no pre-trained weights.

Hyper-parameters

NameValueWhere it came from
A_TP−0.55node 33's default (tuned by the agent on the old proxy / proxy2)
A_CM1.2node 33's default (same)
A_FATE0.6node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6)
K_STRAT, MIN_TYPE_CELLS5, 5node 33's defaults
HVG / PCA / kNN / diffmap2000 / 30 / 30 / 15node 33's defaults

All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input rulers the mechanism is near neutral (admission record).

Known failure modes

  • The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).
  • Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was −2.4 points (MMD −1.8) on top of the composition (mech_split §3).
  • Real cells only: no new expression states, no new cell types.

调研员的计划

名称seed composition_trend
做法seed composition_trend: r1-A 冠军(run 20261002-034201-search-t1-abc-r1-A-era 节点 33,官网 50.1)的组成部分:按类型平均增殖分数重加权类型配额(增殖低的类型占比升高),型内按代谢成熟度(O

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:这个提交的上一版(种子程序:相对空仓库)。改动的文件:solution/EXECUTION.json +1 −0、solution/METHOD.md +72 −0、solution/README.md +4 −0、solution/run.py +195 −0

diff --git a/solution/EXECUTION.json b/solution/EXECUTION.jsonnew file mode 100644index 0000000..9d5125c--- /dev/null+++ b/solution/EXECUTION.json@@ -0,0 +1 @@+{"gpu": false}diff --git a/solution/METHOD.md b/solution/METHOD.mdnew file mode 100644index 0000000..3bc5139--- /dev/null+++ b/solution/METHOD.md@@ -0,0 +1,72 @@+# composition_trend — type-level composition reweighting + within-type maturity selection++**Provenance.** Derived from agent-produced node 33 of run `20261002-034201-search-t1-abc-r1-A-era`+(programs.git `refs/nodes/33`, commit `10400ee`; the program behind the official T1:val 50.1, submissions.tsv+2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed,+for the seed contract only (2026-10-03):++- the external-maturity axis `A_EXT` (read manifest `source == "external"`; a view-identity read; it never fired on+  final / proxy / patched X3);+- the multi-stage pooling `POOL`, the composition-trajectory axis `A_TR`, the kNN expression smoothing `A_S`, the+  apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode;+- the `inputs_by_time(include_external=False)` base selection: the base is simply the latest input stage by time;+- the `VEC_*` environment overrides (constants are fixed in the code).++On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission+record for the digest check. The mechanism-split study (`notes/reports/dev/2026-10-02_mech_split_50.md`, candidate+`r1A:CUS` = `r1A:CU`) is the analysis of exactly this procedure.++**Naming.** "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose+cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted+pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the+whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain+(+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.++Contract: `python run.py --data <view> --out <pred.h5ad> --seed <int>`; CPU only (`EXECUTION.json {"gpu": false}`).++## Method++Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own+`celltype` labels (whatever vocabulary the view provides; `Unknown` is an ordinary type).++1. **Cell scores** (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle+   genes), metabolic maturity `z_met` = OXPHOS (13 genes) − glycolysis (10 genes).+2. **Within-type commitment** `z_fate`: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted+   at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean+   removed, z-scored again (only the order inside a type matters).+3. **Weights.** Type layer `w_t = exp(−0.55 · z(type-mean z_prolif))` (types with < 5 cells: 1). Cell layer+   `w_i = exp(1.2 · z_met_i + 0.6 · z_fate_i)`. `w = clip(w_t · w_i, 1e-6, 1e6)`.+4. **Sampling.** n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type+   quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5+   equal-frequency `z_met` strata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling+   without replacement inside a stratum. Expression values are copied unchanged.++Mechanism-off controls (for analysis; not code switches): `A_TP = 0` leaves the input composition (mech_split "U");+`A_CM = A_FATE = 0` with the same per-type quotas is uniform within type (mech_split "C").++## Data / knowledge used++Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets+(pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external+dataset, no `uns.celltype_palette`, no frozen probe, no pre-trained weights.++## Hyper-parameters++| Name | Value | Where it came from |+|---|---|---|+| `A_TP` | −0.55 | node 33's default (tuned by the agent on the old proxy / proxy2) |+| `A_CM` | 1.2 | node 33's default (same) |+| `A_FATE` | 0.6 | node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6) |+| `K_STRAT`, `MIN_TYPE_CELLS` | 5, 5 | node 33's defaults |+| HVG / PCA / kNN / diffmap | 2000 / 30 / 30 / 15 | node 33's defaults |++All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input+rulers the mechanism is near neutral (admission record).++## Known failure modes++- The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next+  step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).+- Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was+  −2.4 points (MMD −1.8) on top of the composition (mech_split §3).+- Real cells only: no new expression states, no new cell types.diff --git a/solution/README.md b/solution/README.mdnew file mode 100644index 0000000..7d4b398--- /dev/null+++ b/solution/README.md@@ -0,0 +1,4 @@+# composition_trend++r1-A 冠军(run 20261002-034201-search-t1-abc-r1-A-era 节点 33,官网 50.1)的组成部分:按类型平均增殖分数重加权类型配额(增殖低的类型占比升高),型内按代谢成熟度(OXPHOS − 糖酵解)分 5 层、再按扩散伪时间偏向更“承诺”的细胞,加权无放回抽取最新输入阶段的真实细胞。表达不改;去掉了 A_EXT(读 manifest `source` 的视图身份分支)和所有环境变量开关。+纯 CPU,final 约 75 s、峰值内存约 7.3 GB(2 线程)。准入记录:`agent/seeds/T1__val/ADMISSION_2026-10-03_composition_trend.md`。diff --git a/solution/run.py b/solution/run.pynew file mode 100644index 0000000..b7ef52b--- /dev/null+++ b/solution/run.py@@ -0,0 +1,195 @@+#!/usr/bin/env python3+"""composition_trend: type-level composition reweighting + within-type maturity selection of real cells (seed).++Derived from agent-produced node 33 of run 20261002-034201-search-t1-abc-r1-A-era (commit 10400ee, official T1:val+50.1); METHOD.md has the provenance and what was removed. Expression is never changed: the output is a weighted,+stratified sample of the latest input stage's cells.++  type layer   w_t = exp(A_TP * z(mean z_prolif over the type's cells))        A_TP = -0.55+  cell layer   w_i = exp(A_CM * z_met_i + A_FATE * z_fate_i)                   A_CM = 1.2, A_FATE = 0.6+               z_met = z(mean OXPHOS - mean glycolysis), z_fate = within-type centred diffusion pseudotime+  sampling     per-type quota by weight sum (largest remainder, capped by type size, overflow re-apportioned);+               inside a type K = 5 equal-frequency z_met strata, stratum quota by weight sum, Efraimidis-Spirakis+               weighted sampling without replacement inside a stratum.+  n            the latest stage's cell count clipped to [min_cells, max_cells] (as copy_last).++Same code path on every view: the latest input stage by time; the program reads no manifest identity field+(mode / source / board / dataset), no file or directory name, no absolute stage time, no external dataset.+Deterministic for a given --seed. CPU only.+"""+from __future__ import annotations++import argparse++import numpy as np++from src.task1_temporal.view_io import (+    labels_of,+    load_manifest,+    panel_genes,+    read_stage,+    target_n_cells,+    write_prediction,+)++A_TP = -0.55      # type layer: lower mean proliferation -> larger share+A_CM = 1.2        # cell layer: metabolic maturity (OXPHOS - glycolysis)+A_FATE = 0.6      # cell layer: within-type diffusion pseudotime (more committed cells)+K_STRAT = 5       # z_met strata per type+MIN_TYPE_CELLS = 5++CYCLE = [+    "Mki67", "Top2a", "Pcna", "Ccna2", "Ccnb1", "Ccnb2", "Ccnd1", "Ccneg", "Ccne1",+    "Cdk1", "Cdk2", "Cdk4", "Cdk6", "Mcm2", "Mcm3", "Mcm4", "Mcm5", "Mcm6", "Mcm7",+    "Orc1", "Cdc6", "Cdt1", "Rrm1", "Rrm2", "Tyms", "Dtl", "Cenpf", "Cenpe",+    "Birc5", "Aurkb", "Plk1", "Kif20a", "Kif11", "Nusap1",+]+OXPHOS = [+    "Ndufa4", "Ndufb8", "Uqcrb", "Uqcrc1", "Cox5a", "Cox6b1", "Cox7a2", "Atp5a1",+    "Atp5b", "Atp5f1b", "Atp5pb", "Sdha", "Sdhb",+]+GLYC = ["Slc2a1", "Slc2a3", "Hk1", "Hk2", "Pfkp", "Pgk1", "Pgam1", "Eno1", "Ldha", "Pkma"]+++def group_score(X, genes, names):+    index = {g: i for i, g in enumerate(genes)}+    cols = [index[g] for g in names if g in index]+    if not cols:+        return np.zeros(X.shape[0], dtype=np.float64)+    return np.asarray(X[:, cols].mean(axis=1), dtype=np.float64).ravel()+++def zscore(x):+    m, s = x.mean(), x.std()+    if not np.isfinite(s) or s < 1e-9:+        return np.zeros_like(x)+    return np.clip((x - m) / s, -3.0, 3.0)+++def fate_pseudotime(adata, z_prolif, z_met, seed):+    """Diffusion pseudotime rooted at the most progenitor-like cell (argmax 2 z_prolif - z_met, first index wins)."""+    import anndata as ad+    import scanpy as sc++    tmp = ad.AnnData(X=np.asarray(adata.X.todense(), dtype=np.float32), obs=adata.obs.copy())+    sc.pp.highly_variable_genes(tmp, n_top_genes=2000)+    tmp = tmp[:, tmp.var["highly_variable"]].copy()+    sc.pp.pca(tmp, n_comps=30, random_state=seed)+    sc.pp.neighbors(tmp, n_neighbors=30, random_state=seed)+    root = int(np.argmax(2.0 * z_prolif - z_met))+    tmp.uns["iroot"] = root+    sc.tl.diffmap(tmp, n_comps=15, random_state=seed)+    tmp.uns["iroot"] = root+    sc.tl.dpt(tmp, n_branchings=0)+    return np.asarray(tmp.obs["dpt_pseudotime"], dtype=np.float64)+++def apportion(wsum, caps, n):+    """Largest-remainder apportionment of n proportional to wsum, capped by caps, overflow redistributed."""+    alloc = np.zeros(len(wsum), dtype=np.int64)+    rem = int(n)+    for _ in range(len(wsum) + 5):+        if rem <= 0:+            break+        room = caps - alloc+        act = room > 0+        if not act.any():+            break+        tot = float(wsum[act].sum())+        target = np.zeros(len(wsum))+        if tot > 0:+            target[act] = rem * wsum[act] / tot+        else:+            target[act] = rem / float(act.sum())+        add = np.minimum(np.floor(target).astype(np.int64), room)+        if add.sum() == 0:+            cand = np.where(act)[0]+            cand = cand[np.argsort(-target[cand], kind="stable")][:rem]+            add = np.zeros(len(wsum), dtype=np.int64)+            add[cand] = 1+            add = np.minimum(add, room)+        alloc += add+        rem -= int(add.sum())+    return alloc+++def es_sample(rows, w, q, rng):+    """Efraimidis-Spirakis weighted sample without replacement."""+    if q <= 0 or len(rows) == 0:+        return np.empty(0, dtype=np.int64)+    if q >= len(rows):+        return rows+    u = rng.random(len(rows))+    keys = np.log(np.maximum(u, 1e-300)) / np.maximum(w[rows], 1e-12)+    return rows[np.argpartition(-keys, q - 1)[:q]]+++def stratified_sample(w, z_strat, inv, n_types, n_out, K, rng):+    total = len(w)+    if n_out >= total:+        return np.arange(total)+    counts = np.bincount(inv, minlength=n_types)+    n_t = apportion(np.bincount(inv, weights=w, minlength=n_types), counts, n_out)+    out = []+    for t in range(n_types):+        nt = int(n_t[t])+        if nt <= 0:+            continue+        rows = np.where(inv == t)[0]+        if nt >= len(rows) or K < 2 or len(rows) < 2 * K:+            out.append(es_sample(rows, w, nt, rng))+            continue+        srows = rows[np.argsort(z_strat[rows], kind="stable")]+        base, extra = divmod(len(srows), K)+        pos, bounds = 0, []+        for k in range(K):+            sz = base + (1 if k < extra else 0)+            bounds.append(srows[pos:pos + sz])+            pos += sz+        sizes = np.array([len(b) for b in bounds], dtype=np.int64)+        q = apportion(np.array([w[b].sum() for b in bounds]), sizes, nt)+        for k in range(K):+            if q[k] > 0:+                out.append(es_sample(bounds[k], w, int(q[k]), rng))+    return np.sort(np.concatenate(out)) if out else np.empty(0, dtype=np.int64)+++def main() -> None:+    ap = argparse.ArgumentParser()+    ap.add_argument("--data", required=True)+    ap.add_argument("--out", required=True)+    ap.add_argument("--seed", type=int, default=0)+    args = ap.parse_args()++    manifest = load_manifest(args.data)+    genes = panel_genes(args.data, manifest)+    last = sorted(manifest["inputs"], key=lambda e: float(e["time"]))[-1]   # every input stage treated alike+    adata = read_stage(args.data, last, genes, missing="fill")+    X = adata.X+    labels = labels_of(adata) if "celltype" in adata.obs.columns else np.full(adata.n_obs, "all")++    z_prolif = zscore(group_score(X, genes, CYCLE))+    z_met = zscore(group_score(X, genes, OXPHOS) - group_score(X, genes, GLYC))+    uniq, inv = np.unique(labels, return_inverse=True)+    counts = np.bincount(inv, minlength=len(uniq))++    def tmean(x):+        return np.bincount(inv, weights=x, minlength=len(uniq)) / np.maximum(counts, 1)++    z_prolif_t = zscore(tmean(z_prolif))+    raw = zscore(fate_pseudotime(adata, z_prolif, z_met, args.seed))+    raw = raw - tmean(raw)[inv]          # only the within-type order matters+    z_fate = zscore(raw)++    w_type = np.exp(A_TP * z_prolif_t)+    w_type[counts < MIN_TYPE_CELLS] = 1.0+    w = np.clip(w_type[inv] * np.exp(A_CM * z_met + A_FATE * z_fate), 1e-6, 1e6)++    n_out = target_n_cells(manifest, adata.n_obs)+    rng = np.random.default_rng(args.seed)+    idx = stratified_sample(w, z_met, inv, len(uniq), n_out, K_STRAT, rng)+    write_prediction(X[idx], genes, args.out, seed=args.seed)+++if __name__ == "__main__":+    main()

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

没有记录调研来源。

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

没有分析结果(ANALYSIS.json)。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。里列出的文件看。

这个节点没有大模型对话记录。