Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-153851-search-t1-D-smoke2h

节点 n4

composition_program:组成层(收缩信号 + TV 预算)与型内成熟选择分层,wsum 耦合配额

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-153851-search-t1-D-smoke2h
父节点n3
子节点—
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 53.71(+0.0) · X3 48.70(-0.0) · proxy10 63.74(+0.1) · 3 次复测均分 53.69
审查不通过 1 未发现问题:run.py 只通过 view_io 的 load_manifest、panel_genes、read_stage 读取 manifest["inputs"] 里的条目(inputs[-1],两输入时另读 inputs[-2])。没有绝对路径、'..'、data/raw、downloads 或评分器路径,没有读取 X3 视图的 external 条目,也没有联网。CP_OVERRIDE 只是本地配置用的环境变量,不读数据。; 2 未发现目标统计量硬编码:类型份额、计数、δ_t、r_t 和 TV 预算都由输入阶段的 celltype 计数现场算出。代码里没有按类型名写死的比例表…
用时?从运行开始到结束(或到现在)的挂钟时间。1 小时 1 分
程序版本10bd2c912063e3183f34ba9cc89190c4eb232ad3 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 10bd2c9120:solution/METHOD.md

composition_program:组成层(收缩信号 + TV 预算)与型内成熟选择分层,wsum 耦合配额

组成层 π'∝π·exp(β·s):单输入 s=−z(收缩增殖均值),两输入 s=r·(δ_logshare×区间比+0.3·增殖先验);型内 A_CM=1.2、A_FATE=0.6;配额按型权重和耦合分配。只重采样真实细胞,不改表达。

方法(与父节点 3 的差异)

  • 组成层重构(PLAN 步骤 1):
    • 单输入(proxy10):s_t = −zscore(类型均值 z_prolif × r_t),r_t = n_t/(n_t+κ),κ=50(类型均值向全体均值收缩,小类型不再被大幅推动,替代父节点的 MIN_TYPE_CELLS=5 硬截断)。β=0.55,TV(π',π) ≤ τ=0.25(超出时 β 按 ×0.85 迭代缩小;proxy10 实测 β_used=0.55,TV=0.222)。
    • 两输入(X3 / final / proxy2;父节点完全不用第一个输入):δ_t = Laplace 平滑 log 份额差(E_prev→E_last),乘区间比 clip((t_target−t_last)/(t_last−t_prev), 0.5, 2)(只用相对时间差,视图无关);只用于两阶段都存在的类型,缺席类型 δ=0(范围丢失不外推到零);s_t = r_t·(δ_t + λ·(−z_prolif_t)),λ=0.3。预算 eff_τ = min(τ, 0.5·TV_obs)(X3 实测 TV_obs=0.319,β_used=0.287,TV(out,in)=0.145)。
  • 型内层:保留父节点轴(z_met = OXPHOS−糖酵解,z_fate = 型内居中 dpt),A_CM=1.2、A_FATE=0.6、K=5 分层、Efraimidis-Spirakis。配额按型权重和(wsum)耦合分配(组成 × 成熟度),与父节点相同。
  • 程序只按输入阶段数分支;不读任何视图身份字段、绝对时间、文件名。seed 确定(X3 上 seed 0 复跑逐字节一致,seed 1 输出不同)。CP_OVERRIDE 环境变量仅本地筛选用(harness 环境不会设置,默认值即提交配置)。

被证否 / 被验证的部件(A 半查分,18 次,见下表)

被证否(都在 X3 上明显劣化,>噪声):

  1. PLAN 的宽度保护(输出 z_met std ≥0.85×输入,否则减半 A_CM 重抽):X3 46.72 vs 无保护 48.65;proxy10 上 +0.9(64.18 vs 63.32,噪声内)。净效应为负(X3 权重 2),弃用。
  2. PLAN 的 A_CM=0.6:X3 mmd_u 12.2–12.9 分 vs A_CM=1.2 的 14.3;v1 全配置 X3 44.16。保持 1.2。
  3. 解耦的 shares 配额(配额严格按 π',型内权重只选细胞):X3 44.55 vs wsum 耦合 48.88。耦合本身是父节点在 X3 上的关键部件。
  4. 激进趋势外推(区间比 3、预算放宽到 1.0·TV_obs,TV(out)=0.30):X3 45.36 vs 保守 48.65。
  5. τ=0.15 收紧组成(proxy10 β 被压到 0.338):61.27 vs τ=0.25 的 64.18/63.32。

被验证:

  • 两输入趋势层与父节点增殖规则在 X3 上打平且 DE 项略好(趋势 48.65:de_score 11.42 / de_dir 12.59;父复制 48.88:11.14 / 12.61;差异在 ±2 噪声内),而趋势层是 final 视图(E8.5+E9.5→E10.5)唯一有数据方向依据的组成信号。
  • κ=50 收缩在 X3 上代价 0.3 分(48.65 vs κ→0 的 48.96,噪声内),保留(PLAN 机制,压小类型噪声)。

查分记录(board / points:de_score, de_direction, mmd_u, variogram)

#配置X3proxy10
1v1 = PLAN 初值(trend, κ50, τ.15, shares, A_CM0.6, guard.85)44.16:10.69, 11.80, 12.18, 9.5159.23:13.28, 16.88, 18.79, 10.29
2v1 --ablate composition45.30:10.77, 11.94, 12.94, 9.65—
3v1 --ablate within45.36:10.86, 12.30, 12.91, 9.29—
4prolif 组成 + within off(shares)44.36:10.60, 12.19, 12.26, 9.31—
5prolif 组成 + A_CM1.2(shares)44.55:10.52, 11.71, 12.23, 10.10—
6父节点原输出(A 半校准)48.48:11.14, 12.61, 14.09, 10.64—(官方 B 半 63.66)
7父复制(prolif, κ→0, 无τ, wsum, A_CM1.2, 无guard)48.88:11.14, 12.65, 14.42, 10.6863.20:14.95, 17.65, 19.42, 11.18
8trend κ→0 dtf2 wsum A_CM1.248.96:11.52, 12.73, 14.26, 10.45—
9最终(trend κ50 τ.25 dtf自动 wsum A_CM1.2 无guard)48.65:11.42, 12.59, 14.34, 10.3063.32:14.95, 17.67, 19.46, 11.25
10最终 + guard(被证否 1)46.72:11.23, 12.34, 13.51, 9.6464.18:14.92, 17.69, 20.13, 11.44
11最终但 τ=0.15(被证否 5)—61.27:14.92, 17.15, 18.37, 10.82
12激进外推(被证否 4)45.36:10.95, 12.12, 12.28, 10.01—
13最终 --ablate composition48.83:11.73, 12.68, 14.54, 9.89—
14最终 --ablate within45.54:10.95, 12.28, 13.05, 9.26—

估计节点分(A 半,seed 0):(63.32 + 2×48.65)/3 ≈ 53.5,与父节点 53.69(B 半)在噪声内打平。替代尺子上没有净得分收益;本节点的交付是结构性修复:两层分离、可单独关闭、带收缩与预算,且两输入趋势分支首次在 X3 上验证为不劣于父节点的固定规则——final 视图有两个官方输入,父节点在那里只能退回单输入规则,本程序能用上 E8.5→E9.5 的观测趋势。这一收益无法在 proxy10(单输入)上体现,X3 是唯一能验证它的尺子。

机制生效证据(--ablate 对照,X3,与主配置同 seed 同流程)

  • 组成层(最终 vs --ablate composition,输出内容不同,β_used 0.287→0):48.65 vs 48.83。在 X3 上组成层近似中性(−0.18,远小于 ±2 噪声);分项上 de_score −0.31、de_direction −0.09、mmd_u −0.19、variogram +0.41。在 proxy10(单输入分支)上组成层的量级证据来自父/复制系:copy_last 52.75(B 半,节点 1)→ 含组成层 63.2–64.2,主要落在 mmd_u 与 de_direction。
  • 型内层(最终 vs --ablate within):48.65 vs 45.54,+3.1,主要在 mmd_u(14.34 vs 13.05)与 de_score(11.42 vs 10.95)——型内成熟选择在 X3 上是有效杠杆。
  • --ablate all ≈ 均匀分层子采样(≈copy_last 子样本),未单独查分(额度用尽);参考节点 1 copy_last X3 47.92(B 半)。
  • 每型份额(X3,δ 驱动):IFT-CM 0.225→0.325(s=+1.77,r=0.91)、AVC-CM 0.059→0.103(s=+2.45,r=0.72)、aSHF 0.015→0.016(s=+0.74,r=0.40)、SV-CM 0.081→0.049(s=−1.27)、Unknown 0.201→0.155(s=−0.41);微小类型(BEC n=3、NCC-derived n=1、ST n=1)r≤0.06,s≈0,几乎不动——收缩系数与份额变化幅度正相关,符合设计。实际 TV(out,in)=0.145 < 预算 0.16(预算生效:β 0.55→0.287)。proxy10(增殖规则):AVC-CM 0.036→0.102、IFT-CM 0.057→0.115、OFT/RV-CM 0.070→0.124 增,pSHF 0.131→0.075、Neural Tube 0.046→0.022、Paraxial Mesoderm 0.056→0.027 减;TV=0.222 ≤ τ=0.25。被改动最多的 5 型(proxy10,细胞数):pSHF 2192→~383、AVC-CM 598→~523(+)、IFT-CM 949→~587(+)、Neural Tube 773→~112、OFT/RV-CM 1176→~635(+)。

知识来源

  • 增殖(细胞周期基因程序)、代谢成熟(OXPHOS/糖酵解)、扩散伪时间:通用细胞状态注释,继承自父节点(种子 composition_trend);不涉及任何保留阶段的测量值。
  • 趋势层 δ_t、收缩 r_t、TV 预算:全部由视图输入现场计算,无任何外部/文献数值、无硬编码类型名单或比例。未使用发育事件表行(组成方向不取自先验事件,只取自观测趋势 + 增殖弱先验)。
  • 合规:只用 view 内数据;未读 E9.5+(T1 禁窗)任何测量;未读 uns.celltype_palette。

未验证

  • 两输入趋势分支在 final(E8.5+E9.5→E10.5)与 proxy2 上的表现(无尺子可查);参数取保守值(λ=0.3,eff_τ ≤ 0.5·TV_obs,区间比 clip ≤2)。
  • seed 1/2 的 X3 复跑稳定性(额度用尽);seed 0 复现性已本地验证。
  • 运行时约 93 s(proxy10)/ 11 s(X3),峰值内存远低于 28 GB 上限;GPU 未使用(EXECUTION.json: gpu=false)。

调研员的计划

名称composition_program:组成层改为收缩 logit 趋势,型内成熟选择分层消融并限强度
动机父节点 3(composition_trend)的榜分 53.69,几乎全部收益来自 proxy10(63.66,四项都高于地板,例如 mmd_u skill 0.651、de_direction 0.713)。X3 的权重是 2,但 X3 只有 48.71,只比 copy_last(节点 1:47.92)高 0.8,还低于 OT 种子(节点 2:50.65)。X3 上有两项低于地板:de_score −0.18(skill 0.445,11.11 / 12.5 分)和 mmd_u(skill 0.472,14.17 / 15 分)。按评分简报 §3.5,组成押错方向时四项一起下降。父节点的结构问题有三点:(1) 组成方向由一条单快照固定规则决定(增殖低的类型份额变大,强度 A_TP = −0.55 固定),没有任何向“不变”的收缩,也没有不确定性处理;(2) 小类型的类型均值噪声大,却按同样强度推动份额;(3) 型内成熟选择 A_CM = 1.2 很强,会收窄型内状态宽度,可能正是 X3 上 mmd_u 低于地板的原因。另外,父节点在有两个输入阶段时(final:E8.5 + E9.5)完全不用第一阶段,而组成趋势是方向库指出的主要杠杆。
做法在父节点代码上改成 composition_program 结构。组成层与型内层分开,各自带收缩、可单独关闭。
0) 诊断(先做,1–2 次查分):2×2 析因,组成层开/关 × 型内层开/关,在 X3 与 proxy10 上各查一次 score_parts,确认 X3 的损失来自哪一层。只看分项,不看总分。
1) 组成层:目标份额 π'_t ∝ π_t · exp(β · s_t),s_t 是经验贝叶斯收缩后的类型信号。
· 单输入阶段(proxy10,以及 X3 若只有一个输入;退路):s_t = −z(类型均值 z_prolif)。类型均值先按 n_t / (n_t + κ) 向全体均值收缩,κ 初值 50,范围 {20, 50, 150}。这样小类型不再被大幅推动。
· 两输入阶段(final;proxy2 这类视图若出现):s_t = 收缩后的观测 logit 份额变化 δ_t = logit π_t(last) − logit π_t(prev),同样按计数可靠性收缩(Dirichlet 伪计数 κ)。只用两阶段都有的类型;按 k020,前期有、后期消失的类型视为取材范围丢失,δ_t 置 0,不外推到零。增殖规则只作弱先验:s_t = δ_t + λ · (−z_prolif),λ 初值 0.3。程序只按“输入阶段数”分支(数据属性),不读任何视图身份字段。
· 总变化预算:限制 TV(π', π) ≤ τ,超出时把 β 等比缩小。τ 初值 0.15,范围 {0.08, 0.15, 0.25};两输入时 τ 取观测 TV(π_last, π_prev) × 0.5 与上述值中的较小者。β 初值 0.55(与父节点同量级),范围 {0.3, 0.55, 0.8}。
2) 型内层:A_CM 从 1.2 降到 0.6(范围 {0, 0.3, 0.6, 1.2}),A_FATE 保持 0.6,保留 5 层 z_met 分层。另加一条宽度保护:每型输出的 z_met 标准差不得低于输入的 0.85 倍,否则把该型的 A_CM 减半重抽。表达值仍原样拷贝真实细胞。
3) 筛选流程:先在 3,000 细胞子集上跑通,再全量。每个配置在 X3 和 proxy10 上各查 1 次(共约 12–14 次)。最好的 2 个配置用 seed 1、2 复查 X3(X3 权重 2,噪声约 2 分),剩余查分留给消融对照。选择标准:X3 不低于父节点,且 X3 的 de_score、mmd_u 至少回到地板;proxy10 下降不超过 2 分。不要只为 proxy10 调参。
风险(1) 单输入时组成方向仍只来自增殖规则,若这条规则在 X3 上方向本身就错,收缩只能把损失减到接近 copy_last,拿不到收益。早期判断:β = 0 与 β = 0.55 在 X3 上比较 score_parts,若 de_direction 随 β 单调下降,说明方向错,应把组成层收缩到很小,把收益寄托在型内宽度修复上。(2) 降低 A_CM 可能使 proxy10 掉 2–4 分;用析因结果权衡(节点分数中 X3 占 2/3)。(3) 两输入分支在 X3 和 proxy10 上都验证不了(禁做清单:不能只靠单输入尺子选两输入程序),必须在 METHOD 里写明它未经尺子验证,参数取保守值(τ 小、λ 小)。(4) 伪装视图重跑时,平移阶段时间不得改变输出:只用阶段的相对顺序,不用绝对时间。

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 7bef8fd09a。改动的文件:solution/METHOD.md +64 −72、solution/README.md +3 −3、solution/run.py +203 −70

diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 3bc5139..b1be7e7 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,72 +1,64 @@-# composition_trend — type-level composition reweighting + within-type maturity selection--**Provenance.** Derived from agent-produced node 33 of run `20261002-034201-search-t1-abc-r1-A-era`-(programs.git `refs/nodes/33`, commit `10400ee`; the program behind the official T1:val 50.1, submissions.tsv-2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed,-for the seed contract only (2026-10-03):--- the external-maturity axis `A_EXT` (read manifest `source == "external"`; a view-identity read; it never fired on-  final / proxy / patched X3);-- the multi-stage pooling `POOL`, the composition-trajectory axis `A_TR`, the kNN expression smoothing `A_S`, the-  apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode;-- the `inputs_by_time(include_external=False)` base selection: the base is simply the latest input stage by time;-- the `VEC_*` environment overrides (constants are fixed in the code).--On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission-record for the digest check. The mechanism-split study (`notes/reports/dev/2026-10-02_mech_split_50.md`, candidate-`r1A:CUS` = `r1A:CU`) is the analysis of exactly this procedure.--**Naming.** "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose-cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted-pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the-whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain-(+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.--Contract: `python run.py --data <view> --out <pred.h5ad> --seed <int>`; CPU only (`EXECUTION.json {"gpu": false}`).--## Method--Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own-`celltype` labels (whatever vocabulary the view provides; `Unknown` is an ordinary type).--1. **Cell scores** (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle-   genes), metabolic maturity `z_met` = OXPHOS (13 genes) − glycolysis (10 genes).-2. **Within-type commitment** `z_fate`: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted-   at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean-   removed, z-scored again (only the order inside a type matters).-3. **Weights.** Type layer `w_t = exp(−0.55 · z(type-mean z_prolif))` (types with < 5 cells: 1). Cell layer-   `w_i = exp(1.2 · z_met_i + 0.6 · z_fate_i)`. `w = clip(w_t · w_i, 1e-6, 1e6)`.-4. **Sampling.** n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type-   quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5-   equal-frequency `z_met` strata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling-   without replacement inside a stratum. Expression values are copied unchanged.--Mechanism-off controls (for analysis; not code switches): `A_TP = 0` leaves the input composition (mech_split "U");-`A_CM = A_FATE = 0` with the same per-type quotas is uniform within type (mech_split "C").--## Data / knowledge used--Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets-(pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external-dataset, no `uns.celltype_palette`, no frozen probe, no pre-trained weights.--## Hyper-parameters--| Name | Value | Where it came from |-|---|---|---|-| `A_TP` | −0.55 | node 33's default (tuned by the agent on the old proxy / proxy2) |-| `A_CM` | 1.2 | node 33's default (same) |-| `A_FATE` | 0.6 | node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6) |-| `K_STRAT`, `MIN_TYPE_CELLS` | 5, 5 | node 33's defaults |-| HVG / PCA / kNN / diffmap | 2000 / 30 / 30 / 15 | node 33's defaults |--All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input-rulers the mechanism is near neutral (admission record).--## Known failure modes--- The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next-  step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).-- Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was-  −2.4 points (MMD −1.8) on top of the composition (mech_split §3).-- Real cells only: no new expression states, no new cell types.+# composition_program:组成层(收缩信号 + TV 预算)与型内成熟选择分层,wsum 耦合配额++组成层 π'∝π·exp(β·s):单输入 s=−z(收缩增殖均值),两输入 s=r·(δ_logshare×区间比+0.3·增殖先验);型内 A_CM=1.2、A_FATE=0.6;配额按型权重和耦合分配。只重采样真实细胞,不改表达。++## 方法(与父节点 3 的差异)++- **组成层重构**(PLAN 步骤 1):+  - 单输入(proxy10):s_t = −zscore(类型均值 z_prolif × r_t),r_t = n_t/(n_t+κ),κ=50(类型均值向全体均值收缩,小类型不再被大幅推动,替代父节点的 MIN_TYPE_CELLS=5 硬截断)。β=0.55,TV(π',π) ≤ τ=0.25(超出时 β 按 ×0.85 迭代缩小;proxy10 实测 β_used=0.55,TV=0.222)。+  - 两输入(X3 / final / proxy2;父节点完全不用第一个输入):δ_t = Laplace 平滑 log 份额差(E_prev→E_last),乘区间比 clip((t_target−t_last)/(t_last−t_prev), 0.5, 2)(只用相对时间差,视图无关);只用于两阶段都存在的类型,缺席类型 δ=0(范围丢失不外推到零);s_t = r_t·(δ_t + λ·(−z_prolif_t)),λ=0.3。预算 eff_τ = min(τ, 0.5·TV_obs)(X3 实测 TV_obs=0.319,β_used=0.287,TV(out,in)=0.145)。+- **型内层**:保留父节点轴(z_met = OXPHOS−糖酵解,z_fate = 型内居中 dpt),A_CM=1.2、A_FATE=0.6、K=5 分层、Efraimidis-Spirakis。**配额按型权重和(wsum)耦合分配**(组成 × 成熟度),与父节点相同。+- 程序只按输入阶段数分支;不读任何视图身份字段、绝对时间、文件名。seed 确定(X3 上 seed 0 复跑逐字节一致,seed 1 输出不同)。`CP_OVERRIDE` 环境变量仅本地筛选用(harness 环境不会设置,默认值即提交配置)。++## 被证否 / 被验证的部件(A 半查分,18 次,见下表)++被证否(都在 X3 上明显劣化,>噪声):+1. **PLAN 的宽度保护**(输出 z_met std ≥0.85×输入,否则减半 A_CM 重抽):X3 46.72 vs 无保护 48.65;proxy10 上 +0.9(64.18 vs 63.32,噪声内)。净效应为负(X3 权重 2),弃用。+2. **PLAN 的 A_CM=0.6**:X3 mmd_u 12.2–12.9 分 vs A_CM=1.2 的 14.3;v1 全配置 X3 44.16。保持 1.2。+3. **解耦的 shares 配额**(配额严格按 π',型内权重只选细胞):X3 44.55 vs wsum 耦合 48.88。耦合本身是父节点在 X3 上的关键部件。+4. **激进趋势外推**(区间比 3、预算放宽到 1.0·TV_obs,TV(out)=0.30):X3 45.36 vs 保守 48.65。+5. **τ=0.15 收紧组成**(proxy10 β 被压到 0.338):61.27 vs τ=0.25 的 64.18/63.32。++被验证:+- 两输入趋势层与父节点增殖规则在 X3 上打平且 DE 项略好(趋势 48.65:de_score 11.42 / de_dir 12.59;父复制 48.88:11.14 / 12.61;差异在 ±2 噪声内),而趋势层是 final 视图(E8.5+E9.5→E10.5)唯一有数据方向依据的组成信号。+- κ=50 收缩在 X3 上代价 0.3 分(48.65 vs κ→0 的 48.96,噪声内),保留(PLAN 机制,压小类型噪声)。++## 查分记录(board / points:de_score, de_direction, mmd_u, variogram)++| # | 配置 | X3 | proxy10 |+|---|---|---|---|+| 1 | v1 = PLAN 初值(trend, κ50, τ.15, shares, A_CM0.6, guard.85) | 44.16:10.69, 11.80, 12.18, 9.51 | 59.23:13.28, 16.88, 18.79, 10.29 |+| 2 | v1 --ablate composition | 45.30:10.77, 11.94, 12.94, 9.65 | — |+| 3 | v1 --ablate within | 45.36:10.86, 12.30, 12.91, 9.29 | — |+| 4 | prolif 组成 + within off(shares) | 44.36:10.60, 12.19, 12.26, 9.31 | — |+| 5 | prolif 组成 + A_CM1.2(shares) | 44.55:10.52, 11.71, 12.23, 10.10 | — |+| 6 | 父节点原输出(A 半校准) | 48.48:11.14, 12.61, 14.09, 10.64 | —(官方 B 半 63.66) |+| 7 | 父复制(prolif, κ→0, 无τ, wsum, A_CM1.2, 无guard) | 48.88:11.14, 12.65, 14.42, 10.68 | 63.20:14.95, 17.65, 19.42, 11.18 |+| 8 | trend κ→0 dtf2 wsum A_CM1.2 | 48.96:11.52, 12.73, 14.26, 10.45 | — |+| 9 | **最终**(trend κ50 τ.25 dtf自动 wsum A_CM1.2 无guard) | **48.65**:11.42, 12.59, 14.34, 10.30 | **63.32**:14.95, 17.67, 19.46, 11.25 |+| 10 | 最终 + guard(被证否 1) | 46.72:11.23, 12.34, 13.51, 9.64 | 64.18:14.92, 17.69, 20.13, 11.44 |+| 11 | 最终但 τ=0.15(被证否 5) | — | 61.27:14.92, 17.15, 18.37, 10.82 |+| 12 | 激进外推(被证否 4) | 45.36:10.95, 12.12, 12.28, 10.01 | — |+| 13 | 最终 --ablate composition | 48.83:11.73, 12.68, 14.54, 9.89 | — |+| 14 | 最终 --ablate within | 45.54:10.95, 12.28, 13.05, 9.26 | — |++估计节点分(A 半,seed 0):(63.32 + 2×48.65)/3 ≈ **53.5**,与父节点 53.69(B 半)在噪声内打平。替代尺子上没有净得分收益;本节点的交付是结构性修复:两层分离、可单独关闭、带收缩与预算,且**两输入趋势分支首次在 X3 上验证为不劣于父节点的固定规则**——final 视图有两个官方输入,父节点在那里只能退回单输入规则,本程序能用上 E8.5→E9.5 的观测趋势。这一收益无法在 proxy10(单输入)上体现,X3 是唯一能验证它的尺子。++## 机制生效证据(--ablate 对照,X3,与主配置同 seed 同流程)++- 组成层(最终 vs --ablate composition,输出内容不同,β_used 0.287→0):48.65 vs 48.83。在 X3 上组成层近似中性(−0.18,远小于 ±2 噪声);分项上 de_score −0.31、de_direction −0.09、mmd_u −0.19、variogram +0.41。在 proxy10(单输入分支)上组成层的量级证据来自父/复制系:copy_last 52.75(B 半,节点 1)→ 含组成层 63.2–64.2,主要落在 mmd_u 与 de_direction。+- 型内层(最终 vs --ablate within):48.65 vs 45.54,+3.1,主要在 mmd_u(14.34 vs 13.05)与 de_score(11.42 vs 10.95)——型内成熟选择在 X3 上是有效杠杆。+- --ablate all ≈ 均匀分层子采样(≈copy_last 子样本),未单独查分(额度用尽);参考节点 1 copy_last X3 47.92(B 半)。+- 每型份额(X3,δ 驱动):IFT-CM 0.225→0.325(s=+1.77,r=0.91)、AVC-CM 0.059→0.103(s=+2.45,r=0.72)、aSHF 0.015→0.016(s=+0.74,r=0.40)、SV-CM 0.081→0.049(s=−1.27)、Unknown 0.201→0.155(s=−0.41);微小类型(BEC n=3、NCC-derived n=1、ST n=1)r≤0.06,s≈0,几乎不动——收缩系数与份额变化幅度正相关,符合设计。实际 TV(out,in)=0.145 < 预算 0.16(预算生效:β 0.55→0.287)。proxy10(增殖规则):AVC-CM 0.036→0.102、IFT-CM 0.057→0.115、OFT/RV-CM 0.070→0.124 增,pSHF 0.131→0.075、Neural Tube 0.046→0.022、Paraxial Mesoderm 0.056→0.027 减;TV=0.222 ≤ τ=0.25。被改动最多的 5 型(proxy10,细胞数):pSHF 2192→~383、AVC-CM 598→~523(+)、IFT-CM 949→~587(+)、Neural Tube 773→~112、OFT/RV-CM 1176→~635(+)。++## 知识来源++- 增殖(细胞周期基因程序)、代谢成熟(OXPHOS/糖酵解)、扩散伪时间:通用细胞状态注释,继承自父节点(种子 composition_trend);不涉及任何保留阶段的测量值。+- 趋势层 δ_t、收缩 r_t、TV 预算:全部由视图输入现场计算,无任何外部/文献数值、无硬编码类型名单或比例。未使用发育事件表行(组成方向不取自先验事件,只取自观测趋势 + 增殖弱先验)。+- 合规:只用 view 内数据;未读 E9.5+(T1 禁窗)任何测量;未读 uns.celltype_palette。++## 未验证++- 两输入趋势分支在 final(E8.5+E9.5→E10.5)与 proxy2 上的表现(无尺子可查);参数取保守值(λ=0.3,eff_τ ≤ 0.5·TV_obs,区间比 clip ≤2)。+- seed 1/2 的 X3 复跑稳定性(额度用尽);seed 0 复现性已本地验证。+- 运行时约 93 s(proxy10)/ 11 s(X3),峰值内存远低于 28 GB 上限;GPU 未使用(EXECUTION.json: gpu=false)。diff --git a/solution/README.md b/solution/README.mdindex 7d4b398..e862d3f 100644--- a/solution/README.md+++ b/solution/README.md@@ -1,4 +1,4 @@-# composition_trend+# composition_program -r1-A 冠军(run 20261002-034201-search-t1-abc-r1-A-era 节点 33,官网 50.1)的组成部分:按类型平均增殖分数重加权类型配额(增殖低的类型占比升高),型内按代谢成熟度(OXPHOS − 糖酵解)分 5 层、再按扩散伪时间偏向更“承诺”的细胞,加权无放回抽取最新输入阶段的真实细胞。表达不改;去掉了 A_EXT(读 manifest `source` 的视图身份分支)和所有环境变量开关。-纯 CPU,final 约 75 s、峰值内存约 7.3 GB(2 线程)。准入记录:`agent/seeds/T1__val/ADMISSION_2026-10-03_composition_trend.md`。+在父节点 composition_trend(r1-A 冠军,官网 50.1)基础上重构为两层可分别消融的结构:组成层 π'∝π·exp(β·s)(单输入:收缩增殖规则;两输入:观测 log 份额趋势 × 区间比 + 弱增殖先验,TV 预算约束),型内层保留代谢成熟度 + 扩散伪时间加权(A_CM=1.2, A_FATE=0.6),配额按型权重和耦合分配。只重采样最新输入阶段的真实细胞,表达不改。+`--ablate composition|within|all` 分别关闭组成层 / 型内层 / 两层。纯 CPU;proxy10 约 93 s,X3 约 11 s。方法与查分证据见 METHOD.md。`CP_OVERRIDE` 环境变量仅本地筛选用,harness 环境不设置,默认值即提交配置。diff --git a/solution/run.py b/solution/run.pyindex dbddc68..dc304e9 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,23 +1,33 @@ #!/usr/bin/env python3-"""composition_trend: type-level composition reweighting + within-type maturity selection of real cells (seed).--Derived from agent-produced node 33 of run 20261002-034201-search-t1-abc-r1-A-era (commit 10400ee, official T1:val-50.1); METHOD.md has the provenance and what was removed. Expression is never changed: the output is a weighted,-stratified sample of the latest input stage's cells.--  type layer   w_t = exp(A_TP * z(mean z_prolif over the type's cells))        A_TP = -0.55-  cell layer   w_i = exp(A_CM * z_met_i + A_FATE * z_fate_i)                   A_CM = 1.2, A_FATE = 0.6-               z_met = z(mean OXPHOS - mean glycolysis), z_fate = within-type centred diffusion pseudotime-  sampling     per-type quota by weight sum (largest remainder, capped by type size, overflow re-apportioned);-               inside a type K = 5 equal-frequency z_met strata, stratum quota by weight sum, Efraimidis-Spirakis-               weighted sampling without replacement inside a stratum.-  n            the latest stage's cell count clipped to [min_cells, max_cells] (as copy_last).----ablate (G39.7): type_layer | cell_layer (maturity, fate) | anything else = both layers off (uniform weights).--Same code path on every view: the latest input stage by time; the program reads no manifest identity field-(mode / source / board / dataset), no file or directory name, no absolute stage time, no external dataset.-Deterministic for a given --seed. CPU only.+"""composition_program: shrunk composition layer + within-type maturity selection of real cells.++Improvement over the composition_trend seed (parent node 3). Two independent layers, each shrunk towards+"no change" and separately ablatable:++  composition layer   target shares  pi'_t  prop  pi_t * exp(beta * s_t),  beta throttled so that the+                      budget  eff_tau = tau  (one input)  or  min(tau, 0.5 * observed TV)  (two inputs)  holds+      one input stage : s_t = -z( type-mean z_prolif shrunk by r_t = n_t/(n_t+kappa) )   (kappa = 50)+      two+ input stages: s_t = r_t * ( delta_t + lambda * (-z_prolif_t) ),  lambda = 0.3,+                        delta_t = observed log-share trend between the two inputs (Laplace-smoothed counts),+                        scaled by the interval ratio clip((t_target - t_last)/(t_last - t_prev), 0.5, 2)+                        (relative times only); delta_t = 0 for types absent from either stage (no+                        extrapolation to zero; disappeared types are treated as scope losses)+  within-type layer   w_i = exp(A_CM * z_met_i + A_FATE * z_fate_i),  A_CM = 1.2, A_FATE = 0.6+                      z_met = z(OXPHOS - glycolysis), z_fate = within-type centred diffusion pseudotime;+                      K = 5 equal-frequency z_met strata per type, Efraimidis-Spirakis inside a stratum;+                      type quotas apportion by the per-type sum of w (composition x maturity coupling,+                      as in the parent; the decoupled "quota = pi'" variant scored clearly worse on X3)+  n                   the latest stage's cell count clipped to [min_cells, max_cells]++Expression values are never changed: the output is a resample of the latest input stage's cells.++--ablate (G39.7): composition -> beta = 0 (shares = input); within -> A_CM = A_FATE = 0 (uniform within type);+all / anything else -> both off (a stratified uniform subsample of the last input).++Same code path on every view: branches only on the number of input stages; reads no manifest identity field,+no file or directory name, no absolute stage time (only stage order and time differences). The local+screening override CP_OVERRIDE (env var, JSON) is never set by the harness environment; defaults are the+submitted configuration. Deterministic for a given --seed. CPU only. """ from __future__ import annotations @@ -34,11 +44,14 @@ from src.task1_temporal.view_io import (     write_prediction, ) -A_TP = -0.55      # type layer: lower mean proliferation -> larger share-A_CM = 1.2        # cell layer: metabolic maturity (OXPHOS - glycolysis)-A_FATE = 0.6      # cell layer: within-type diffusion pseudotime (more committed cells)+BETA = 0.55       # composition layer strength+KAPPA = 50.0      # reliability shrinkage: r_t = n_t / (n_t + kappa)+TAU = 0.25        # total-variation budget on share change (two inputs: also capped by 0.5 * observed TV)+LAM = 0.3         # weight of the proliferation prior in the two-input signal+A_CM = 1.2        # within-type: metabolic maturity (OXPHOS - glycolysis)+A_FATE = 0.6      # within-type: diffusion pseudotime (more committed cells) K_STRAT = 5       # z_met strata per type-MIN_TYPE_CELLS = 5+WIDTH_FLOOR = 0.0  # width guard disabled (empirically worse on the X3 ruler; see METHOD.md)  CYCLE = [     "Mki67", "Top2a", "Pcna", "Ccna2", "Ccnb1", "Ccnb2", "Ccnd1", "Ccneg", "Ccne1",@@ -126,34 +139,77 @@ def es_sample(rows, w, q, rng):     return rows[np.argpartition(-keys, q - 1)[:q]]  -def stratified_sample(w, z_strat, inv, n_types, n_out, K, rng):-    total = len(w)-    if n_out >= total:-        return np.arange(total)-    counts = np.bincount(inv, minlength=n_types)-    n_t = apportion(np.bincount(inv, weights=w, minlength=n_types), counts, n_out)-    out = []-    for t in range(n_types):-        nt = int(n_t[t])-        if nt <= 0:-            continue-        rows = np.where(inv == t)[0]-        if nt >= len(rows) or K < 2 or len(rows) < 2 * K:-            out.append(es_sample(rows, w, nt, rng))-            continue-        srows = rows[np.argsort(z_strat[rows], kind="stable")]-        base, extra = divmod(len(srows), K)-        pos, bounds = 0, []-        for k in range(K):-            sz = base + (1 if k < extra else 0)-            bounds.append(srows[pos:pos + sz])-            pos += sz-        sizes = np.array([len(b) for b in bounds], dtype=np.int64)-        q = apportion(np.array([w[b].sum() for b in bounds]), sizes, nt)-        for k in range(K):-            if q[k] > 0:-                out.append(es_sample(bounds[k], w, int(q[k]), rng))-    return np.sort(np.concatenate(out)) if out else np.empty(0, dtype=np.int64)+def sample_within_type(rows, w_all, z_strat, nt, K, rng):+    """Stratified weighted sample of nt cells from one type's rows; w_all is indexed by global cell index."""+    if nt <= 0:+        return np.empty(0, dtype=np.int64)+    if nt >= len(rows) or K < 2 or len(rows) < 2 * K:+        return es_sample(rows, w_all, nt, rng)+    srows = rows[np.argsort(z_strat[rows], kind="stable")]+    base, extra = divmod(len(srows), K)+    bounds, pos = [], 0+    for k in range(K):+        sz = base + (1 if k < extra else 0)+        bounds.append(srows[pos:pos + sz])+        pos += sz+    sizes = np.array([len(b) for b in bounds], dtype=np.int64)+    q = apportion(np.array([w_all[b].sum() for b in bounds]), sizes, nt)+    out = [es_sample(bounds[k], w_all, int(q[k]), rng) for k in range(K) if q[k] > 0]+    return np.concatenate(out) if out else np.empty(0, dtype=np.int64)+++def composition_signal(counts_last, labels_last, counts_prev, labels_prev,+                       z_prolif_t, kappa, lam, dtf=1.0):+    """s_t per type of labels_last, plus the observed TV(pi_last, pi_prev) (0 with one input).++    Type-level evidence is shrunk towards "no change" by r_t = n_t / (n_t + kappa)."""+    r = counts_last / (counts_last + kappa)++    if counts_prev is None:+        # single input: shrunk proliferation rule (low proliferation -> growing share)+        return -zscore(z_prolif_t * r), 0.0++    # two+ inputs: observed log-share trend, Laplace-smoothed, only for types present in both stages+    pos_prev = {lab: i for i, lab in enumerate(labels_prev)}+    m = np.array([counts_prev[pos_prev[lab]] if lab in pos_prev else 0 for lab in labels_last], dtype=np.float64)+    N, M = counts_last.sum(), counts_prev.sum()+    K_last, K_prev = len(labels_last), len(labels_prev)+    p_last = (counts_last + 1.0) / (N + K_last)+    p_prev = (m + 1.0) / (M + K_prev)+    both = (counts_last > 0) & (m > 0)+    delta = np.zeros(len(labels_last))+    delta[both] = dtf * (np.log(p_last[both]) - np.log(p_prev[both]))+    # types absent from the last stage are scope losses: they get no quota anyway (counts_last = 0)+    s = r * (delta + lam * (-zscore(z_prolif_t)))+    tv_obs = 0.5 * float(np.abs(p_last - p_prev).sum())+    return s, tv_obs+++def _cfg(default):+    """Local screening override (env var); the harness environment never sets it."""+    import os+    v = os.environ.get("CP_OVERRIDE")+    if not v:+        return default+    import json+    return {**default, **json.loads(v)}+++CFG = _cfg({+    "dt_scale": 0.0,        # 0 = auto: clip((t_target - t_last)/(t_last - t_prev), 0.5, 2); >0 = fixed factor+    "comp_mode": "trend",   # "trend": two-input log-share trend (+ prior); "prolif": proliferation rule always+    "lam": LAM,+    "kappa": KAPPA,+    "beta": BETA,+    "tau": TAU,+    "a_cm": A_CM,+    "a_fate": A_FATE,+    "width_floor": WIDTH_FLOOR,+    "tv_cap_frac": 0.5,     # two-input: eff budget = min(tau, tv_cap_frac * observed TV)+    "alloc": "wsum",        # "wsum": parent-style, quota by per-type weight sum (composition x maturity coupling)+                            # "shares": quota strictly by target shares pi'+    "diag": False,+})   def main() -> None:@@ -161,14 +217,25 @@ def main() -> None:     ap.add_argument("--data", required=True)     ap.add_argument("--out", required=True)     ap.add_argument("--seed", type=int, default=0)-    # G39.7 mechanism-off control: type_layer -> A_TP = 0; cell_layer (or maturity / fate) -> A_CM = A_FATE = 0; any other-    # name -> both layers off (uniform weights: a stratified copy of the latest stage)     ap.add_argument("--ablate", default=None)     args = ap.parse_args() +    beta, kappa, tau, lam = CFG["beta"], CFG["kappa"], CFG["tau"], CFG["lam"]+    a_cm, a_fate = CFG["a_cm"], CFG["a_fate"]+    width_floor = CFG["width_floor"]+    ablate = (args.ablate or "").strip().lower()+    if ablate in ("composition", "comp"):+        beta = 0.0+    elif ablate in ("within", "within_type", "cell", "maturity", "fate"):+        a_cm = a_fate = 0.0+    elif ablate:+        beta = 0.0+        a_cm = a_fate = 0.0+     manifest = load_manifest(args.data)     genes = panel_genes(args.data, manifest)-    last = sorted(manifest["inputs"], key=lambda e: float(e["time"]))[-1]   # every input stage treated alike+    inputs = sorted(manifest["inputs"], key=lambda e: float(e["time"]))+    last = inputs[-1]     adata = read_stage(args.data, last, genes, missing="fill")     X = adata.X     labels = labels_of(adata) if "celltype" in adata.obs.columns else np.full(adata.n_obs, "all")@@ -176,31 +243,97 @@ def main() -> None:     z_prolif = zscore(group_score(X, genes, CYCLE))     z_met = zscore(group_score(X, genes, OXPHOS) - group_score(X, genes, GLYC))     uniq, inv = np.unique(labels, return_inverse=True)-    counts = np.bincount(inv, minlength=len(uniq))+    inv = np.asarray(inv).ravel()+    counts = np.bincount(inv, minlength=len(uniq)).astype(np.float64)      def tmean(x):         return np.bincount(inv, weights=x, minlength=len(uniq)) / np.maximum(counts, 1) -    z_prolif_t = zscore(tmean(z_prolif))+    z_prolif_t = tmean(z_prolif)     raw = zscore(fate_pseudotime(adata, z_prolif, z_met, args.seed))-    raw = raw - tmean(raw)[inv]          # only the within-type order matters+    raw = raw - tmean(raw)[inv]     z_fate = zscore(raw) -    a_tp, a_cm, a_fate = A_TP, A_CM, A_FATE-    if args.ablate:-        if args.ablate in ("type_layer", "type"):-            a_tp = 0.0-        elif args.ablate in ("cell_layer", "cell", "maturity", "fate"):-            a_cm = a_fate = 0.0-        else:-            a_tp = a_cm = a_fate = 0.0-    w_type = np.exp(a_tp * z_prolif_t)-    w_type[counts < MIN_TYPE_CELLS] = 1.0-    w = np.clip(w_type[inv] * np.exp(a_cm * z_met + a_fate * z_fate), 1e-6, 1e6)-+    # ---- composition layer ----+    # trend extrapolation factor: remaining gap relative to the observed input gap (relative times only)+    dt_scale = CFG["dt_scale"]+    if dt_scale <= 0.0:+        dt_scale = 1.0+        if len(inputs) >= 2:+            dt_prev = float(inputs[-1]["time"]) - float(inputs[-2]["time"])+            dt_next = float(manifest["target"]["time"]) - float(inputs[-1]["time"])+            if dt_prev > 1e-9:+                dt_scale = float(np.clip(dt_next / dt_prev, 0.5, 2.0))+    counts_prev, labels_prev = None, None+    comp_mode = CFG["comp_mode"]+    if len(inputs) >= 2 and comp_mode != "prolif":+        prev_entry = inputs[-2]+        prev = read_stage(args.data, prev_entry, genes, missing="fill")+        plabels = labels_of(prev) if "celltype" in prev.obs.columns else np.full(prev.n_obs, "all")+        puniq, pinv = np.unique(plabels, return_inverse=True)+        counts_prev = np.bincount(np.asarray(pinv).ravel(), minlength=len(puniq)).astype(np.float64)+        labels_prev = puniq+        del prev++    s, tv_obs = composition_signal(counts, uniq, counts_prev, labels_prev,+                                   z_prolif_t, kappa, lam, dtf=dt_scale)+    pi = counts / counts.sum()+    if beta <= 0.0:+        pi_target = pi+        b_used = 0.0+    else:+        eff_tau = min(tau, CFG["tv_cap_frac"] * tv_obs) if counts_prev is not None else tau+        b_used = beta+        pi_target = pi+        for _ in range(80):+            q = pi * np.exp(b_used * s)+            pi_target = q / q.sum()+            if 0.5 * float(np.abs(pi_target - pi).sum()) <= eff_tau + 1e-12:+                break+            b_used *= 0.85++    # ---- within-type layer ----     n_out = target_n_cells(manifest, adata.n_obs)+    w_cell = np.clip(np.exp(a_cm * z_met + a_fate * z_fate), 1e-6, 1e6)+    if CFG["alloc"] == "wsum":+        # coupled apportionment: type quota by the sum of its cells' weights (composition x maturity)+        w_tot = (pi_target / np.maximum(pi, 1e-12))[inv] * w_cell+        quota = apportion(np.bincount(inv, weights=w_tot, minlength=len(uniq)),+                          counts.astype(np.int64), n_out)+    else:+        quota = apportion(pi_target * n_out, counts.astype(np.int64), n_out)     rng = np.random.default_rng(args.seed)-    idx = stratified_sample(w, z_met, inv, len(uniq), n_out, K_STRAT, rng)++    if CFG["diag"]:+        import sys+        print(f"[diag] n_inputs={len(inputs)} n_types={len(uniq)} n_out={n_out} "+              f"tv_obs={tv_obs:.4f} beta_used={b_used:.3f} "+              f"TV(out,in)={0.5 * np.abs(pi_target - pi).sum():.4f}", file=sys.stderr)+        order = np.argsort(-counts)+        for t in order:+            print(f"[diag] {uniq[t][:36]:36s} n={int(counts[t]):6d} pi={pi[t]:.4f} "+                  f"pi'={pi_target[t]:.4f} s={s[t]:+.3f} r={counts[t] / (counts[t] + kappa):.2f}", file=sys.stderr)++    z_met_t_std_in = np.array([z_met[inv == t].std() if counts[t] >= 3 else np.inf for t in range(len(uniq))])+    out_idx = []+    for t in range(len(uniq)):+        rows = np.where(inv == t)[0]+        nt = int(quota[t])+        if nt <= 0 or len(rows) == 0:+            continue+        acm = a_cm+        w_all = np.empty(len(z_met))+        for attempt in range(3):+            w_all[:] = 1e-6+            w_all[rows] = np.clip(np.exp(acm * z_met[rows] + a_fate * z_fate[rows]), 1e-6, 1e6)+            sel = sample_within_type(rows, w_all, z_met, nt, K_STRAT, rng)+            if len(sel) < 3 or width_floor <= 0.0 or not np.isfinite(z_met_t_std_in[t]) or acm <= 0.0:+                break+            if z_met[sel].std() >= width_floor * z_met_t_std_in[t]:+                break+            acm *= 0.5+        out_idx.append(sel)+    idx = np.sort(np.concatenate(out_idx)) if out_idx else np.empty(0, dtype=np.int64)     write_prediction(X[idx], genes, args.out, seed=args.seed)  

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k020Correcting sampling-scope (dissection) bias in compositionnotes/guides/modeling_and_evaluation_guide.html
k036Composition forecasting and mixture models for population predictionnotes/competition/05_lineage_graph.md; notes/competition/09_t1_census_lineage.md; 10.1038/s41586-024-08453-2 (growth rates)
k015Composition x conditional-expression decomposition p(x|t) = sum_z p(z|t) p(x|z,t)notes/handover/03_当前方案与Agent系统设计.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么把父节点重构为可分别消融的两层:组成层改为 π'∝π·exp(β·s),加 κ=50 的计数收缩和 TV 预算(τ=0.25);两输入时(X3)组成信号改用观测 log 份额趋势×区间比,再加 0.3×增殖先验,预算取 0.5·TV_obs,实测 β 被压到 0.287。型内层(A_CM=1.2、A_FATE=0.6)和 wsum 耦合配额沿用父节点。PLAN 的三项核心改动都在筛选中被弃用:A_CM 降到 0.6、宽度保护、解耦份额配额。
各组分数的变化cell_state:噪声内:分组 +0.37。X3 mmd_u 0.03556→0.03497,得分 14.17→14.35(+0.17),仍低于地板 15。proxy10 0.02981→0.02984,得分 −0.01。
covariation:噪声内但方向偏坏:分组 −1.24。X3 variogram 0.001377→0.001475,得分 10.67→10.26(−0.41)。proxy10 0.001028→0.001018,得分 +0.07。
de_recovery:噪声内:分组 +0.97,变化全部来自 X3。X3 de_score −0.1818→−0.1299,得分 11.11→11.47(+0.36),仍低于地板 12.5。proxy10 0.2679 不变。
direction:噪声内:分组 −0.35。X3 de_direction 0.0264→0.0121,得分 −0.14,已接近地板 12.5。proxy10 0.3600→0.3609,得分 +0.02。
榜分:噪声内:53.69→53.71(+0.02)。proxy10 63.66→63.74(+0.07),X3 48.71→48.70(−0.01)。
family_idcomposition_program
假设是否成立否
经验
  1. 在两输入的 X3 上,把单快照增殖规则换成观测 log 份额趋势(区间比 clip 0.5–2,预算 0.5·TV_obs):榜分不变(48.71→48.70),de_score +0.36、variogram −0.41,都在噪声内。说明 X3 上组成信号的来源不是主要矛盾。
  2. PLAN 认为 A_CM=1.2 收窄型内宽度、导致 X3 的 mmd_u 低于地板,这一点被否定:Engineer 的 A 半查分中,A_CM=0.6 或加宽度保护后 X3 mmd_u 得分降到 12.2–13.5,而 A_CM=1.2 为 14.3。强成熟选择在 X3 上是正贡献。
  3. 配额严格按目标份额 π' 分配(解耦),在 X3 上比按型权重和分配(wsum 耦合)低约 4 分(44.55 vs 48.88,A 半,未经 harness 复核)。组成×成熟度耦合是父节点在 X3 上的关键部件,不要拆开。
  4. 按 Engineer 的消融,X3 上的有效杠杆是型内层:关掉后降 3.1 分,主要是 mmd_u 14.34→13.05。组成层近似中性(关掉后 +0.18)。
  5. proxy10 单输入分支对 TV 预算敏感:τ 收到 0.15 会把 β 压到 0.338,proxy10 跌到 61.27。τ=0.25 时预算不约束,结果与父节点等价。
  6. harness 的 `--ablate mechanism` 在本代码里落入 else 分支,等于两层全关(≈copy_last)。所以 −4.94 衡量的是两层合计,主要是继承来的型内层,不能证明新组成机制生效。消融开关名应与 harness 约定一致。
  7. 耗时从 117 s 升到 264.6 s,与 Engineer 自报的 proxy10 93 s 不符,以变化量表为准,内存不变(6.86 GB)。
mechanism_activeunclear
下一步建议
  1. 针对 cell_state / de_recovery(X3 的 mmd_u 和 de_score 仍低于地板):固定 A_CM=1.2 和 wsum 耦合,在 X3 上扫 A_FATE ∈ {0.3, 0.6, 1.0} 和 K_STRAT ∈ {3, 5, 8},只看 X3 score_parts,最优配置用 seed 1、2 复查。
  2. 针对机制验证:把 `--ablate mechanism` 映射为只关组成层(β=0),让 harness 对照真正检验收缩趋势组成层,并在 X3 上与主配置配对比较 de_direction 和 mmd_u。
  3. 针对 covariation / direction(X3 variogram −0.41、de_direction −0.14):在 X3 上做 tv_cap_frac ∈ {0.25, 0.5} × λ ∈ {0, 0.3} 的 2×2 小扫描。若趋势分支持续拉低 variogram,就把两输入预算收紧到 0.25·TV_obs。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 vec-relay-claude-cli/claude-opus-5-5
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。1
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 0 次:—
用时不到 1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 2 · 输出 3,481 · 思考 0
任务(第一行)分析节点 n4 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
,X3 48.71→48.70(−0.01)。",
  "de_recovery": "噪声内:分组 +0.97,变化全部来自 X3。X3 de_score −0.1818→−0.1299,得分 11.11→11.47(+0.36),仍低于地板 12.5。proxy10 0.2679 不变。",
  "direction": "噪声内:分组 −0.35。X3 de_direction 0.0264→0.0121,得分 −0.14,已接近地板 12.5。proxy10 0.3600→0.3609,得分 +0.02。",
  "cell_state": "噪声内:分组 +0.37。X3 mmd_u 0.03556→0.03497,得分 14.17→14.35(+0.17),仍低于地板 15。proxy10 0.02981→0.02984,得分 −0.01。",
  "covariation": "噪声内但方向偏坏:分组 −1.24。X3 variogram 0.001377→0.001475,得分 10.67→10.26(−0.41)。proxy10 0.001028→0.001018,得分 +0.07。"
 },
 "hypothesis_supported": "no",
 "failure_reason": null,
 "lessons": [
  "在两输入的 X3 上,把单快照增殖规则换成观测 log 份额趋势(区间比 clip 0.5–2,预算 0.5·TV_obs):榜分不变(48.71→48.70),de_score +0.36、variogram −0.41,都在噪声内。说明 X3 上组成信号的来源不是主要矛盾。",
  "PLAN 认为 A_CM=1.2 收窄型内宽度、导致 X3 的 mmd_u 低于地板,这一点被否定:Engineer 的 A 半查分中,A_CM=0.6 或加宽度保护后 X3 mmd_u 得分降到 12.2–13.5,而 A_CM=1.2 为 14.3。强成熟选择在 X3 上是正贡献。",
  "配额严格按目标份额 π' 分配(解耦),在 X3 上比按型权重和分配(wsum 耦合)低约 4 分(44.55 vs 48.88,A 半,未经 harness 复核)。组成×成熟度耦合是父节点在 X3 上的关键部件,不要拆开。",
  "按 Engineer 的消融,X3 上的有效杠杆是型内层:关掉后降 3.1 分,主要是 mmd_u 14.34→13.05。组成层近似中性(关掉后 +0.18)。",
  "proxy10 单输入分支对 TV 预算敏感:τ 收到 0.15 会把 β 压到 0.338,proxy10 跌到 61.27。τ=0.25 时预算不约束,结果与父节点等价。",
  "harness 的 `--ablate mechanism` 在本代码里落入 else 分支,等于两层全关(≈copy_last)。所以 −4.94 衡量的是两层合计,主要是继承来的型内层,不能证明新组成机制生效。消融开关名应与 harness 约定一致。",
  "耗时从 117 s 升到 264.6 s,与 Engineer 自报的 proxy10 93 s 不符,以变化量表为准,内存不变(6.86 GB)。"
 ],
 "next_suggestions": [
  "针对 cell_state / de_recovery(X3 的 mmd_u 和 de_score 仍低于地板):固定 A_CM=1.2 和 wsum 耦合,在 X3 上扫 A_FATE ∈ {0.3, 0.6, 1.0} 和 K_STRAT ∈ {3, 5, 8},只看 X3 score_parts,最优配置用 seed 1、2 复查。",
  "针对机制验证:把 `--ablate mechanism` 映射为只关组成层(β=0),让 harness 对照真正检验收缩趋势组成层,并在 X3 上与主配置配对比较 de_direction 和 mmd_u。",
  "针对 covariation / direction(X3 variogram −0.41、de_direction −0.14):在 X3 上做 tv_cap_frac ∈ {0.25, 0.5} × λ ∈ {0, 0.3} 的 2×2 小扫描。若趋势分支持续拉低 variogram,就把两输入预算收紧到 0.25·TV_obs。"
 ],
 "family_id": "composition_program",
 "mechanism_active": "unclear"
}
```

注:本次调用没有文件工具,以上是 ANALYSIS.json 的完整内容,未能写入磁盘。
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/test/20261003-153851-search-t1-D-smoke2h/nodes/4/analyst.jsonl 4 KB

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数60
工具调用共 72 次:bash 47、edit 20、write 3、read 2
用时49 分
token 数输入 131,784 · 输出 26,885 · 思考 29,701
任务(第一行)按计划修改程序:composition_program:组成层改为收缩 logit 趋势,型内成熟选择分层消融并限强度
最后的回答(摘录)
完成。总结:

**做了什么**:把父节点(composition_trend)重构为 PLAN 的 composition_program 两层结构:组成层 π'∝π·exp(β·s),单输入用收缩增殖规则(r=n/(n+κ),κ=50),两输入用观测 log 份额趋势×区间比+弱增殖先验(λ=0.3),受 TV 预算约束(τ=0.25,两输入再受 0.5·TV_obs 限制);型内层保留成熟度+伪时间选择(A_CM=1.2);配额按 wsum 耦合分配。支持 `--ablate composition/within/all`。已验证:seed 确定、伪装视图(时间+1、manifest 键序打乱)输出逐字节一致、两视图 vec-check 通过、proxy10 运行 93s。

**查分结果**(18 次,A 半):最终配置 X3 48.65 / proxy10 63.32,节点估计 ≈53.5,与父节点 53.69 噪声内打平。被证否:PLAN 的宽度保护(X3 −1.9)、A_CM=0.6、解耦 shares 配额(X3 −4.3)、激进趋势外推(X3 −3.3)、τ=0.15。消融:X3 上型内层是主要杠杆(+3.1),组成层近中性但 DE 分项略好于父规则;趋势分支首次在 X3 验证不劣于固定增殖规则,且是 final(两输入)唯一有数据方向依据的组成信号。

**下一步最值得试**:(1) 型内选择轴换成事件表标记程序(心室/心房成熟、EPDC-EMT 轴,行 2/8/9),X3 的 de_score 仍低于地板,是最弱分项;(2) 组成信号加入型内成熟度的类型级聚合(wsum 耦合已证明有效,说明"成熟类型扩张"是真方向,可显式建模);(3) 在 final 型两输入官方数据到位前,用 proxy2 视图检验趋势分支的跨数据集鲁棒性。
原始记录位置/home/spark-longxinyang/vec/runs/test/20261003-153851-search-t1-D-smoke2h/nodes/4/engineer.jsonl 327 KB
/home/spark-longxinyang/vec/runs/test/20261003-153851-search-t1-D-smoke2h/nodes/4/engineer.stderr

调研员

角色调研员 vec-relay-claude-cli/claude-opus-5-5
调用次数1 次
轮数1
工具调用共 0 次:—
用时不到 1 分
token 数输入 2 · 输出 4,347 · 思考 0
任务(第一行)为节点 n4(改进,父节点 n3)写改进计划 PLAN.json
最后的回答(摘录)
小类型不再被大幅推动。\n  · 两输入阶段(final;proxy2 这类视图若出现):s_t = 收缩后的观测 logit 份额变化 δ_t = logit π_t(last) − logit π_t(prev),同样按计数可靠性收缩(Dirichlet 伪计数 κ)。只用两阶段都有的类型;按 k020,前期有、后期消失的类型视为取材范围丢失,δ_t 置 0,不外推到零。增殖规则只作弱先验:s_t = δ_t + λ · (−z_prolif),λ 初值 0.3。程序只按“输入阶段数”分支(数据属性),不读任何视图身份字段。\n  · 总变化预算:限制 TV(π', π) ≤ τ,超出时把 β 等比缩小。τ 初值 0.15,范围 {0.08, 0.15, 0.25};两输入时 τ 取观测 TV(π_last, π_prev) × 0.5 与上述值中的较小者。β 初值 0.55(与父节点同量级),范围 {0.3, 0.55, 0.8}。\n2) 型内层:A_CM 从 1.2 降到 0.6(范围 {0, 0.3, 0.6, 1.2}),A_FATE 保持 0.6,保留 5 层 z_met 分层。另加一条宽度保护:每型输出的 z_met 标准差不得低于输入的 0.85 倍,否则把该型的 A_CM 减半重抽。表达值仍原样拷贝真实细胞。\n3) 筛选流程:先在 3,000 细胞子集上跑通,再全量。每个配置在 X3 和 proxy10 上各查 1 次(共约 12–14 次)。最好的 2 个配置用 seed 1、2 复查 X3(X3 权重 2,噪声约 2 分),剩余查分留给消融对照。选择标准:X3 不低于父节点,且 X3 的 de_score、mmd_u 至少回到地板;proxy10 下降不超过 2 分。不要只为 proxy10 调参。",
 "expected_groups": ["cell_state", "de_recovery", "direction"],
 "risks": "(1) 单输入时组成方向仍只来自增殖规则,若这条规则在 X3 上方向本身就错,收缩只能把损失减到接近 copy_last,拿不到收益。早期判断:β = 0 与 β = 0.55 在 X3 上比较 score_parts,若 de_direction 随 β 单调下降,说明方向错,应把组成层收缩到很小,把收益寄托在型内宽度修复上。(2) 降低 A_CM 可能使 proxy10 掉 2–4 分;用析因结果权衡(节点分数中 X3 占 2/3)。(3) 两输入分支在 X3 和 proxy10 上都验证不了(禁做清单:不能只靠单输入尺子选两输入程序),必须在 METHOD 里写明它未经尺子验证,参数取保守值(τ 小、λ 小)。(4) 伪装视图重跑时,平移阶段时间不得改变输出:只用阶段的相对顺序,不用绝对时间。",
 "family_id": "composition_program",
 "mechanism": "把预测拆成类型份额与型内细胞选择两部分。份额由收缩后的类型信号(两输入时用观测 logit 趋势,单输入时用增殖规则)驱动,并受总变化预算 τ 约束;型内选择限制强度并保护状态宽度。两部分各自向“不变”收缩,可单独关闭。",
 "vs_constant_shift": "不做任何表达位移,只重采样真实细胞。与父节点的全局组成重加权相比:份额变化按类型的证据量(细胞数、两阶段一致性)逐型收缩,有总变化上限,范围丢失的类型不外推到零;两输入时方向来自观测趋势,而不是一条固定规则。",
 "mechanism_evidence": "报告以下内容:每型输入份额、目标份额、输出份额,以及收缩系数 n_t / (n_t + κ) 与份额变化幅度的相关(小类型的变化应更小);实际 TV(π', π) 及预算是否生效;每型输出与输入 z_met 的均值和标准差比;四组分在 X3 和 proxy10 上的分项变化;被改动最多的 5 个类型及其被增减的细胞数。两输入分支用 final 视图做离线自检(不查分):δ_t 的符号与输出份额变化的符号一致。",
 "mechanism_off_control": "--ablate composition:β = 0,份额等于输入,只剩型内层。预期 X3 和 proxy10 的 de_direction 与 mmd_u 发生变化;若输出与开启时完全相同,说明组成层没有运行。--ablate within:A_CM = A_FATE = 0,型内均匀抽样,预期主要改变 mmd_u 与 variogram。--ablate all:等价于 copy_last 的子样本。三个对照各在 X3 上查一次,与主配置配对比较。"}
```
原始记录位置/home/spark-longxinyang/vec/runs/test/20261003-153851-search-t1-D-smoke2h/nodes/4/researcher.jsonl 6 KB

审查员

角色审查员 vec-relay-claude-cli/claude-opus-5-5
调用次数1 次
轮数1
工具调用共 0 次:—
用时不到 1 分
token 数输入 2 · 输出 2,321 · 思考 0
任务(第一行)审查节点 n4 的程序是否越界读取、写死目标、钻评分器漏洞
最后的回答(摘录)
{"verdict": "fail", "reasons": ["1 未发现问题:run.py 只通过 view_io 的 load_manifest、panel_genes、read_stage 读取 manifest[\"inputs\"] 里的条目(inputs[-1],两输入时另读 inputs[-2])。没有绝对路径、'..'、data/raw、downloads 或评分器路径,没有读取 X3 视图的 external 条目,也没有联网。CP_OVERRIDE 只是本地配置用的环境变量,不读数据。", "2 未发现目标统计量硬编码:类型份额、计数、δ_t、r_t 和 TV 预算都由输入阶段的 celltype 计数现场算出。代码里没有按类型名写死的比例表,也没有细胞数或表达量常量。run.py 顶部的 CYCLE、OXPHOS、GLYC 基因列表属于细胞状态标记程序,按第 5 条处理。", "3 未发现钻评分器漏洞:输出是按型重采样的真实细胞,表达值不变。末尾的 np.sort(idx) 只保留输入的原始行序,没有为按行号抽样的指标专门排列细胞,也没有缩放或平均细胞。", "4 未发现问题:rng = np.random.default_rng(args.seed);scanpy 的 pca、neighbors、diffmap 都传入 random_state=seed;root 用 argmax 确定性选取;apportion 和分层排序都用 stable 排序。CP_OVERRIDE 在 harness 环境中不会设置,不影响同 seed 复现。", "5 fail:run.py 顶部写死了三组细胞状态标记程序(CYCLE 细胞周期 34 个基因、OXPHOS 13 个、GLYC 10 个),用于计算 z_prolif 和 z_met,后者驱动组成层信号与型内权重。这类引用必须注明来源,但 METHOD.md 的“知识来源”一节只写了“通用细胞状态注释,继承自父节点(种子 composition_trend)”,run.py 注释里也没有任何知识条目编号或 DOI,属于缺来源。另外,未发现读取保留阶段、禁窗或保留基因型实测数据的代码。", "6 未发现问题:程序只按 len(inputs) 分支;inputs 按 time 排序后用相对时间差(dt_next/dt_prev,并做 clip)。没有写死阶段名、文件名或输入数,基因缺失用 missing=\"fill\" 补齐,可以在 final 的两输入视图上运行。", "7 未发现问题:PLAN.json 中没有 [compliance: removed] 标记,也没有 _compliance 字段。METHOD.md 和代码中没有出现保留阶段的测量值、比例或细胞数。"]}
原始记录位置/home/spark-longxinyang/vec/runs/test/20261003-153851-search-t1-D-smoke2h/nodes/4/reviewer.jsonl 3 KB