总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s1
节点 n3
composition_trend — type-level composition reweighting + within-type maturity selection
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261003-093415-search-t1-r2-D-s1 |
|---|---|
| 父节点 | (种子,没有父节点) |
| 子节点 | n4、n7 |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 种子 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 53.69 · X3 48.71 · proxy10 63.66 · 3 次复测均分 53.91 |
| 审查 | 通过 seed (human-written, locked) |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 2 分 |
| 程序版本 | ba77be1df19d877439f08d0b9fd9cac9c2f80b75 (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git ba77be1df1:solution/METHOD.md
composition_trend — type-level composition reweighting + within-type maturity selection
Provenance. Derived from agent-produced node 33 of run 20261002-034201-search-t1-abc-r1-A-era
(programs.git refs/nodes/33, commit 10400ee; the program behind the official T1:val 50.1, submissions.tsv
2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed,
for the seed contract only (2026-10-03):
- the external-maturity axis
A_EXT(read manifestsource == "external"; a view-identity read; it never fired on final / proxy / patched X3); - the multi-stage pooling
POOL, the composition-trajectory axisA_TR, the kNN expression smoothingA_S, the apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode; - the
inputs_by_time(include_external=False)base selection: the base is simply the latest input stage by time; - the
VEC_*environment overrides (constants are fixed in the code).
On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission
record for the digest check. The mechanism-split study (notes/reports/dev/2026-10-02_mech_split_50.md, candidate
r1A:CUS = r1A:CU) is the analysis of exactly this procedure.
Naming. "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain (+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.
Contract: python run.py --data <view> --out <pred.h5ad> --seed <int>; CPU only (EXECUTION.json {"gpu": false}).
Method
Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own
celltype labels (whatever vocabulary the view provides; Unknown is an ordinary type).
- Cell scores (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle genes), metabolic maturity
z_met= OXPHOS (13 genes) − glycolysis (10 genes). - Within-type commitment
z_fate: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean removed, z-scored again (only the order inside a type matters). - Weights. Type layer
w_t = exp(−0.55 · z(type-mean z_prolif))(types with < 5 cells: 1). Cell layerw_i = exp(1.2 · z_met_i + 0.6 · z_fate_i).w = clip(w_t · w_i, 1e-6, 1e6). - Sampling. n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5 equal-frequency
z_metstrata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling without replacement inside a stratum. Expression values are copied unchanged.
Mechanism-off controls (for analysis; not code switches): A_TP = 0 leaves the input composition (mech_split "U");
A_CM = A_FATE = 0 with the same per-type quotas is uniform within type (mech_split "C").
Data / knowledge used
Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets
(pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external
dataset, no uns.celltype_palette, no frozen probe, no pre-trained weights.
Hyper-parameters
| Name | Value | Where it came from |
|---|---|---|
A_TP | −0.55 | node 33's default (tuned by the agent on the old proxy / proxy2) |
A_CM | 1.2 | node 33's default (same) |
A_FATE | 0.6 | node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6) |
K_STRAT, MIN_TYPE_CELLS | 5, 5 | node 33's defaults |
| HVG / PCA / kNN / diffmap | 2000 / 30 / 30 / 15 | node 33's defaults |
All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input rulers the mechanism is near neutral (admission record).
Known failure modes
- The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).
- Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was −2.4 points (MMD −1.8) on top of the composition (mech_split §3).
- Real cells only: no new expression states, no new cell types.
调研员的计划
| 名称 | seed composition_trend |
|---|---|
| 做法 | seed composition_trend: r1-A 冠军(run 20261002-034201-search-t1-abc-r1-A-era 节点 33,官网 50.1)的组成部分:按类型平均增殖分数重加权类型配额(增殖低的类型占比升高),型内按代谢成熟度(O |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:这个提交的上一版(种子程序:相对空仓库)。改动的文件:solution/EXECUTION.json +1 −0、solution/METHOD.md +72 −0、solution/README.md +4 −0、solution/run.py +195 −0
diff --git a/solution/EXECUTION.json b/solution/EXECUTION.jsonnew file mode 100644index 0000000..9d5125c--- /dev/null+++ b/solution/EXECUTION.json@@ -0,0 +1 @@+{"gpu": false}diff --git a/solution/METHOD.md b/solution/METHOD.mdnew file mode 100644index 0000000..3bc5139--- /dev/null+++ b/solution/METHOD.md@@ -0,0 +1,72 @@+# composition_trend — type-level composition reweighting + within-type maturity selection++**Provenance.** Derived from agent-produced node 33 of run `20261002-034201-search-t1-abc-r1-A-era`+(programs.git `refs/nodes/33`, commit `10400ee`; the program behind the official T1:val 50.1, submissions.tsv+2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed,+for the seed contract only (2026-10-03):++- the external-maturity axis `A_EXT` (read manifest `source == "external"`; a view-identity read; it never fired on+ final / proxy / patched X3);+- the multi-stage pooling `POOL`, the composition-trajectory axis `A_TR`, the kNN expression smoothing `A_S`, the+ apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode;+- the `inputs_by_time(include_external=False)` base selection: the base is simply the latest input stage by time;+- the `VEC_*` environment overrides (constants are fixed in the code).++On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission+record for the digest check. The mechanism-split study (`notes/reports/dev/2026-10-02_mech_split_50.md`, candidate+`r1A:CUS` = `r1A:CU`) is the analysis of exactly this procedure.++**Naming.** "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose+cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted+pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the+whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain+(+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.++Contract: `python run.py --data <view> --out <pred.h5ad> --seed <int>`; CPU only (`EXECUTION.json {"gpu": false}`).++## Method++Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own+`celltype` labels (whatever vocabulary the view provides; `Unknown` is an ordinary type).++1. **Cell scores** (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle+ genes), metabolic maturity `z_met` = OXPHOS (13 genes) − glycolysis (10 genes).+2. **Within-type commitment** `z_fate`: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted+ at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean+ removed, z-scored again (only the order inside a type matters).+3. **Weights.** Type layer `w_t = exp(−0.55 · z(type-mean z_prolif))` (types with < 5 cells: 1). Cell layer+ `w_i = exp(1.2 · z_met_i + 0.6 · z_fate_i)`. `w = clip(w_t · w_i, 1e-6, 1e6)`.+4. **Sampling.** n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type+ quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5+ equal-frequency `z_met` strata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling+ without replacement inside a stratum. Expression values are copied unchanged.++Mechanism-off controls (for analysis; not code switches): `A_TP = 0` leaves the input composition (mech_split "U");+`A_CM = A_FATE = 0` with the same per-type quotas is uniform within type (mech_split "C").++## Data / knowledge used++Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets+(pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external+dataset, no `uns.celltype_palette`, no frozen probe, no pre-trained weights.++## Hyper-parameters++| Name | Value | Where it came from |+|---|---|---|+| `A_TP` | −0.55 | node 33's default (tuned by the agent on the old proxy / proxy2) |+| `A_CM` | 1.2 | node 33's default (same) |+| `A_FATE` | 0.6 | node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6) |+| `K_STRAT`, `MIN_TYPE_CELLS` | 5, 5 | node 33's defaults |+| HVG / PCA / kNN / diffmap | 2000 / 30 / 30 / 15 | node 33's defaults |++All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input+rulers the mechanism is near neutral (admission record).++## Known failure modes++- The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next+ step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).+- Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was+ −2.4 points (MMD −1.8) on top of the composition (mech_split §3).+- Real cells only: no new expression states, no new cell types.diff --git a/solution/README.md b/solution/README.mdnew file mode 100644index 0000000..7d4b398--- /dev/null+++ b/solution/README.md@@ -0,0 +1,4 @@+# composition_trend++r1-A 冠军(run 20261002-034201-search-t1-abc-r1-A-era 节点 33,官网 50.1)的组成部分:按类型平均增殖分数重加权类型配额(增殖低的类型占比升高),型内按代谢成熟度(OXPHOS − 糖酵解)分 5 层、再按扩散伪时间偏向更“承诺”的细胞,加权无放回抽取最新输入阶段的真实细胞。表达不改;去掉了 A_EXT(读 manifest `source` 的视图身份分支)和所有环境变量开关。+纯 CPU,final 约 75 s、峰值内存约 7.3 GB(2 线程)。准入记录:`agent/seeds/T1__val/ADMISSION_2026-10-03_composition_trend.md`。diff --git a/solution/run.py b/solution/run.pynew file mode 100644index 0000000..b7ef52b--- /dev/null+++ b/solution/run.py@@ -0,0 +1,195 @@+#!/usr/bin/env python3+"""composition_trend: type-level composition reweighting + within-type maturity selection of real cells (seed).++Derived from agent-produced node 33 of run 20261002-034201-search-t1-abc-r1-A-era (commit 10400ee, official T1:val+50.1); METHOD.md has the provenance and what was removed. Expression is never changed: the output is a weighted,+stratified sample of the latest input stage's cells.++ type layer w_t = exp(A_TP * z(mean z_prolif over the type's cells)) A_TP = -0.55+ cell layer w_i = exp(A_CM * z_met_i + A_FATE * z_fate_i) A_CM = 1.2, A_FATE = 0.6+ z_met = z(mean OXPHOS - mean glycolysis), z_fate = within-type centred diffusion pseudotime+ sampling per-type quota by weight sum (largest remainder, capped by type size, overflow re-apportioned);+ inside a type K = 5 equal-frequency z_met strata, stratum quota by weight sum, Efraimidis-Spirakis+ weighted sampling without replacement inside a stratum.+ n the latest stage's cell count clipped to [min_cells, max_cells] (as copy_last).++Same code path on every view: the latest input stage by time; the program reads no manifest identity field+(mode / source / board / dataset), no file or directory name, no absolute stage time, no external dataset.+Deterministic for a given --seed. CPU only.+"""+from __future__ import annotations++import argparse++import numpy as np++from src.task1_temporal.view_io import (+ labels_of,+ load_manifest,+ panel_genes,+ read_stage,+ target_n_cells,+ write_prediction,+)++A_TP = -0.55 # type layer: lower mean proliferation -> larger share+A_CM = 1.2 # cell layer: metabolic maturity (OXPHOS - glycolysis)+A_FATE = 0.6 # cell layer: within-type diffusion pseudotime (more committed cells)+K_STRAT = 5 # z_met strata per type+MIN_TYPE_CELLS = 5++CYCLE = [+ "Mki67", "Top2a", "Pcna", "Ccna2", "Ccnb1", "Ccnb2", "Ccnd1", "Ccneg", "Ccne1",+ "Cdk1", "Cdk2", "Cdk4", "Cdk6", "Mcm2", "Mcm3", "Mcm4", "Mcm5", "Mcm6", "Mcm7",+ "Orc1", "Cdc6", "Cdt1", "Rrm1", "Rrm2", "Tyms", "Dtl", "Cenpf", "Cenpe",+ "Birc5", "Aurkb", "Plk1", "Kif20a", "Kif11", "Nusap1",+]+OXPHOS = [+ "Ndufa4", "Ndufb8", "Uqcrb", "Uqcrc1", "Cox5a", "Cox6b1", "Cox7a2", "Atp5a1",+ "Atp5b", "Atp5f1b", "Atp5pb", "Sdha", "Sdhb",+]+GLYC = ["Slc2a1", "Slc2a3", "Hk1", "Hk2", "Pfkp", "Pgk1", "Pgam1", "Eno1", "Ldha", "Pkma"]+++def group_score(X, genes, names):+ index = {g: i for i, g in enumerate(genes)}+ cols = [index[g] for g in names if g in index]+ if not cols:+ return np.zeros(X.shape[0], dtype=np.float64)+ return np.asarray(X[:, cols].mean(axis=1), dtype=np.float64).ravel()+++def zscore(x):+ m, s = x.mean(), x.std()+ if not np.isfinite(s) or s < 1e-9:+ return np.zeros_like(x)+ return np.clip((x - m) / s, -3.0, 3.0)+++def fate_pseudotime(adata, z_prolif, z_met, seed):+ """Diffusion pseudotime rooted at the most progenitor-like cell (argmax 2 z_prolif - z_met, first index wins)."""+ import anndata as ad+ import scanpy as sc++ tmp = ad.AnnData(X=np.asarray(adata.X.todense(), dtype=np.float32), obs=adata.obs.copy())+ sc.pp.highly_variable_genes(tmp, n_top_genes=2000)+ tmp = tmp[:, tmp.var["highly_variable"]].copy()+ sc.pp.pca(tmp, n_comps=30, random_state=seed)+ sc.pp.neighbors(tmp, n_neighbors=30, random_state=seed)+ root = int(np.argmax(2.0 * z_prolif - z_met))+ tmp.uns["iroot"] = root+ sc.tl.diffmap(tmp, n_comps=15, random_state=seed)+ tmp.uns["iroot"] = root+ sc.tl.dpt(tmp, n_branchings=0)+ return np.asarray(tmp.obs["dpt_pseudotime"], dtype=np.float64)+++def apportion(wsum, caps, n):+ """Largest-remainder apportionment of n proportional to wsum, capped by caps, overflow redistributed."""+ alloc = np.zeros(len(wsum), dtype=np.int64)+ rem = int(n)+ for _ in range(len(wsum) + 5):+ if rem <= 0:+ break+ room = caps - alloc+ act = room > 0+ if not act.any():+ break+ tot = float(wsum[act].sum())+ target = np.zeros(len(wsum))+ if tot > 0:+ target[act] = rem * wsum[act] / tot+ else:+ target[act] = rem / float(act.sum())+ add = np.minimum(np.floor(target).astype(np.int64), room)+ if add.sum() == 0:+ cand = np.where(act)[0]+ cand = cand[np.argsort(-target[cand], kind="stable")][:rem]+ add = np.zeros(len(wsum), dtype=np.int64)+ add[cand] = 1+ add = np.minimum(add, room)+ alloc += add+ rem -= int(add.sum())+ return alloc+++def es_sample(rows, w, q, rng):+ """Efraimidis-Spirakis weighted sample without replacement."""+ if q <= 0 or len(rows) == 0:+ return np.empty(0, dtype=np.int64)+ if q >= len(rows):+ return rows+ u = rng.random(len(rows))+ keys = np.log(np.maximum(u, 1e-300)) / np.maximum(w[rows], 1e-12)+ return rows[np.argpartition(-keys, q - 1)[:q]]+++def stratified_sample(w, z_strat, inv, n_types, n_out, K, rng):+ total = len(w)+ if n_out >= total:+ return np.arange(total)+ counts = np.bincount(inv, minlength=n_types)+ n_t = apportion(np.bincount(inv, weights=w, minlength=n_types), counts, n_out)+ out = []+ for t in range(n_types):+ nt = int(n_t[t])+ if nt <= 0:+ continue+ rows = np.where(inv == t)[0]+ if nt >= len(rows) or K < 2 or len(rows) < 2 * K:+ out.append(es_sample(rows, w, nt, rng))+ continue+ srows = rows[np.argsort(z_strat[rows], kind="stable")]+ base, extra = divmod(len(srows), K)+ pos, bounds = 0, []+ for k in range(K):+ sz = base + (1 if k < extra else 0)+ bounds.append(srows[pos:pos + sz])+ pos += sz+ sizes = np.array([len(b) for b in bounds], dtype=np.int64)+ q = apportion(np.array([w[b].sum() for b in bounds]), sizes, nt)+ for k in range(K):+ if q[k] > 0:+ out.append(es_sample(bounds[k], w, int(q[k]), rng))+ return np.sort(np.concatenate(out)) if out else np.empty(0, dtype=np.int64)+++def main() -> None:+ ap = argparse.ArgumentParser()+ ap.add_argument("--data", required=True)+ ap.add_argument("--out", required=True)+ ap.add_argument("--seed", type=int, default=0)+ args = ap.parse_args()++ manifest = load_manifest(args.data)+ genes = panel_genes(args.data, manifest)+ last = sorted(manifest["inputs"], key=lambda e: float(e["time"]))[-1] # every input stage treated alike+ adata = read_stage(args.data, last, genes, missing="fill")+ X = adata.X+ labels = labels_of(adata) if "celltype" in adata.obs.columns else np.full(adata.n_obs, "all")++ z_prolif = zscore(group_score(X, genes, CYCLE))+ z_met = zscore(group_score(X, genes, OXPHOS) - group_score(X, genes, GLYC))+ uniq, inv = np.unique(labels, return_inverse=True)+ counts = np.bincount(inv, minlength=len(uniq))++ def tmean(x):+ return np.bincount(inv, weights=x, minlength=len(uniq)) / np.maximum(counts, 1)++ z_prolif_t = zscore(tmean(z_prolif))+ raw = zscore(fate_pseudotime(adata, z_prolif, z_met, args.seed))+ raw = raw - tmean(raw)[inv] # only the within-type order matters+ z_fate = zscore(raw)++ w_type = np.exp(A_TP * z_prolif_t)+ w_type[counts < MIN_TYPE_CELLS] = 1.0+ w = np.clip(w_type[inv] * np.exp(A_CM * z_met + A_FATE * z_fate), 1e-6, 1e6)++ n_out = target_n_cells(manifest, adata.n_obs)+ rng = np.random.default_rng(args.seed)+ idx = stratified_sample(w, z_met, inv, len(uniq), n_out, K_STRAT, rng)+ write_prediction(X[idx], genes, args.out, seed=args.seed)+++if __name__ == "__main__":+ main()
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
没有记录调研来源。
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
没有分析结果(ANALYSIS.json)。
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。里列出的文件看。
这个节点没有大模型对话记录。