Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s0

节点 n21 终选程序?按该运行锁定的规则最终选出的程序;可能是候选节点,也可能由护栏回退到基线。在终选来历上

composition_trend + α=3 重锚定 + 置信度加权稀疏型内位移:按 (型,基因) 两阶段 Welch t 连续缩放 conf=|t|/(|t|+3),每型只位移 |t| 前 200 基因(β=3;仅双输入视图激活;X3 A 半 54.88→55.67,proxy10 与父逐字节一致)

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-093415-search-t1-r2-D-s0
父节点n14
子节点—
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 56.83(+0.4) · X3 55.84(+0.7) · proxy10 58.80(+0.0) · 3 次复测均分 56.77
审查通过 1 越界读取:未发现问题——所有数据读取均通过 view_io 的 load_manifest/read_stage/panel_genes(run.py:26-33, 577-580, 629-643),无绝对路径、无 '..'、无 prior/external/评分器路径访问,无网络代码。; 2 硬编码目标统计量:未发现问题——常量仅为规则权重(A_TP/A_CM/A_FATE/α/β/T0/TMIN/TOPG,run.py:35-80),CYCLE/OXPHOS/GLYC(run.py:82-92)是通用通路基因集而非目标阶段测量;所有细胞比例、均值、t 统计量均由 read_stag…
用时?从运行开始到结束(或到现在)的挂钟时间。27 分
程序版本a6e52bb75fdce9ef1697504980bbcaacb8b30ba7 (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git a6e52bb75f:solution/METHOD.md

composition_trend + α=3 重锚定 + 置信度加权稀疏型内位移:按 (型,基因) 两阶段 Welch t 连续缩放 conf=|t|/(|t|+3),每型只位移 |t| 前 200 基因(β=3;仅双输入视图激活;X3 A 半 54.88→55.67,proxy10 与父逐字节一致)

Parent (node 14, 56.39): composition_trend weighted sampling + per-type re-anchoring α=3 + uniform per-type two-stage pseudobulk displacement (β=1, median gate). PLAN asked to replace the uniform displacement kernel with a confidence-weighted one: scale each (type,gene) displacement by its two-stage t-statistic and truncate to the strongest-evidence genes per type.

Method (family: composition_program; mechanism = confidence-weighted sparse typed displacement)

Active only when the view has ≥2 input stages (uses the last two by time; no absolute time, no view identity read). Per cell type t and gene g:

  1. dmu = pb_t(later) − pb_t(earlier) (log space), EB-shrunk by within-type detection count min(n1,n2)/(min(n1,n2)+K), K=30; genes with <20 detections in either stage of the type fall back to the global shrunk delta (same fallback as parent).
  2. Welch t: t = dmu / sqrt(var1/n1 + var2/n2) from the type's cells in the two stages (variance via E[x²]−E[x]², sparse-safe); fallback genes take the global t.
  3. Confidence conf = |t|/(|t|+T0), T0=3. Gate: |t| ≥ T_MIN=2. Per-type truncation: keep only the TOPG=200 strongest-|t| genes.
  4. Displacement D = clip(β·conf·dmu_shrunk, ±MAX_SHIFT=1.0), β=3, added at existing nonzeros only, clipped ≥0 (sparsity and covariation preserved), after the α=3 recentering.

Ships: MODE=tconf, β=3, T0=3, T_MIN=2, TOPG=200. Env knobs: VEC_SHIFT_{BETA,K,ZMIN,MAX,MODE, T0,TMIN,TOPG}. Single-input views (proxy10): block never runs → output byte-identical to parent. Deterministic (no RNG in the block); CPU only (EXECUTION.json gpu:false).

Screen (vec-score A-half, X3, seed 0; parent type β=1 = 54.88, de −0.043, dir 0.075, mmd 0.0259)

variantX3 boardde_scorede_dirmmd_uvariogram
b1 tmin2 g20055.110.0140.08350.026620.001208
b2 tmin2 g055.23−0.0140.08110.025680.001213
b2 tmin2 g200 (=tmin3)55.660.0570.08220.026090.001214
b2 tmin3 g055.570.0430.08120.025960.001217
b3 tmin2 g200 (ships)55.670.0430.08260.025720.001221
b4 tmin2 g055.710.0140.07520.024780.001236
b4 tmin2 g200 (=tmin3)55.470.0140.08120.025520.001229
b2 tmin2 g10055.460.0430.08410.026340.001212

Seed robustness (same-seed paired, X3 A-half): seeds 0/1/2 → tconf 55.67/54.87/55.26 vs parent 54.88/54.68/55.29 (Δ +0.79/+0.19/−0.03, mean +0.32). de_direction improves in 3/3 seeds (0.0826/0.0738/0.0738 vs 0.0748/0.0683/0.0652); mmd_u improves in 3/3; de_score improves at seed 0 only (0.043 vs −0.043; seeds 1/2 both −0.014 vs −0.029/0.0). Honest read: total gain is inside the ~2-pt noise; the consistent parts are de_direction and mmd_u, not de_score. β=3+TOPG=200 chosen over β=4 g0 (higher seed-0 board) because g0's gain is mmd-only with de_direction falling to parent level and variogram degrading, and it lacks the PLAN truncation mechanism. PLAN's literal acceptance thresholds (de_dir≥0.095, mmd≤0.0252) match B-half values and are unreachable on A-half; applied relatively (≥ parent on dir and mmd) instead.

Mechanism evidence

  • Gated pairs: 20,816 (type,gene) pass |t|≥2 before truncation (per-type min/med/max 1422/1755/2326 — far above PLAN's ≥30 floor, no T_MIN relaxation needed); TOPG=200 truncates to 2,000 pairs, touching 36,808 nonzero entries (all changed), mean |shift| 0.28 (β=3).
  • Spearman(|D|, |t|) over displaced pairs = 0.78 (β=3): displacement magnitude scales with statistical confidence, unlike the parent's uniform kernel (by construction 0 correlation).
  • Off control VEC_SHIFT_BETA=0: shift block skipped entirely (no [gene_shift] line), output = recenter-only path = node-10-equivalent. Uniform-mode control (T0=0⇒conf≡1, TMIN=0, TOPG=0, β=3): touches 1,187,380 entries vs shipped 36,808 (32×), confirming the shipped sparsity comes from the confidence gate + truncation, and the shipped-vs-uniform difference is exactly the new kernel.
  • Parent-path regression: VEC_SHIFT_MODE=type β=1 seed 0 reproduces parent A-half exactly (54.877, de −0.0429, dir 0.0748, mmd 0.02589, var 0.001206 vs parent's reported 54.88/−0.043/0.075/0.0259/0.001206) → all deltas above come from the new kernel, no collateral code change.
  • proxy10 (single input): ran seed 0, n_input_stages=1, block inactive, vec-check ok; identical code path to parent ⇒ proxy10 stays 58.80 (not re-scored, query budget).

Verified / not verified

  • Verified: determinism (seed 0 twice byte-identical); vec-check ok on proxy and X3; off controls above; 3-seed paired A-half comparison on X3.
  • Not verified: B-half scores (all screening is A-half); proxy10 score not re-queried (argued byte-identical); no biological prior used (sources: none — the kernel is purely statistical); gain may not replicate given it is within noise on total board.

调研员的计划

名称置信度加权稀疏型内位移:按 |t| 缩放并截尾,替代均匀 β 位移
动机父节点 14 最弱组 de_recovery 50.14(X3 de_score 原始值 -0.026,仍在地板下),direction 也仅 58.17。父节点筛选显示均匀 β 位移呈平顶(β=1/2/3 同为 54.88)且 β=5 起全面崩坏(53.71),说明中位数门控放进来的噪声基因与有把握基因获得同样幅度的位移,放大 β 时噪声被同步放大——这是可解释的结构问题。数据驱动 Δμ(节点 14)与 CollecTRI 网络方向(节点 17,56.12<56.39)均已试且未抬起 de 两项,本次不再换方向来源,而是改位移核本身:按 (型,基因) 的两阶段 t 统计量连续缩放位移并只保留每型证据最强的基因。
做法在父节点 SHIFT 块内(仅 ≥2 输入阶段激活,proxy10 单输入不激活、保持逐字节一致):1) 对每个 (型,基因) 用该型两阶段细胞算伪批量差 dmu 与合并标准误 se,t=dmu/se;2) 置信因子 conf=|t|/(|t|+T0),T0 初值 3(搜索 1–10);3) 位移 = clip(β·conf·dmu_shrunk, ±MAX_SHIFT),仅作用于 |t|≥T_MIN(初值 2.0,搜索 2/3)的基因;可选每型只取 |t| 前 TOPG 个基因(初值 200,0=关闭),使被位移的基因成为预测中的主导变化基因,直接对应 de_score 看排序的机制;4) 保留原有非零位置施加、clip≥0、EB 收缩 K=30 作二级收缩。筛选:X3 A 半 seed 0,小网格 β∈{2,4}×T_MIN∈{2,3}×TOPG∈{0,200},约 6–8 次查询;入选标准不是总分而是 de_direction≥0.095 且 mmd_u≤0.0252(不许掉 cell_state 收益)。最佳配置用 seed 1、2 复跑确认(差距 <2 分视为噪声)。总查询 ≤12,余量留给 off control 与完全一致验证。
风险1) X3 型内细胞少(~65/型),t 统计量本身嘈杂,|t| 排序可能仍是噪声——Engineer 应在筛选前先打印过门控的 (型,基因) 数,若每型 <30 个基因则放宽 T_MIN 到 1.5;2) 稀疏化过头退化成无操作(输出≈父),用 off control 对比位移条目数即可发现;3) 收益可能仍在 ~2 分噪声内,必须多子复确认,禁止单次 +0.5 就出厂;4) 不改善 proxy10(单输入),节点分提升只能来自 X3(权重 2/3),预期上限低。

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 ba44b165dc。改动的文件:solution/METHOD.md +63 −66、solution/run.py +120 −3

diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 2a3e363..9c88fd9 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,78 +1,75 @@-# composition_trend + α=3 重锚定 + 两阶段逐基因方向位移(按细胞型 EB 收缩,β=1;仅双输入视图激活,X3 A 半 54.51→54.88,proxy10 与父逐字节一致)+# composition_trend + α=3 重锚定 + 置信度加权稀疏型内位移:按 (型,基因) 两阶段 Welch t 连续缩放 conf=|t|/(|t|+3),每型只位移 |t| 前 200 基因(β=3;仅双输入视图激活;X3 A 半 54.88→55.67,proxy10 与父逐字节一致) -Parent (node 10, 55.86): composition_trend weighted sampling + per-type re-anchoring α=3.0.-PLAN asked for a two-stage gene-level EB-shrunk direction shift to lift X3 de_recovery-(de_score −0.052, below floor). Implemented as specified, plus a per-type variant found by screening.+Parent (node 14, 56.39): composition_trend weighted sampling + per-type re-anchoring α=3 + uniform+per-type two-stage pseudobulk displacement (β=1, median gate). PLAN asked to replace the uniform+displacement kernel with a confidence-weighted one: scale each (type,gene) displacement by its+two-stage t-statistic and truncate to the strongest-evidence genes per type. -## Method (family: composition_program; mechanism = two-stage pseudobulk delta direction shift)+## Method (family: composition_program; mechanism = confidence-weighted sparse typed displacement) -Active only when the view has ≥2 input stages (sorted by time; uses the last two):+Active only when the view has ≥2 input stages (uses the last two by time; no absolute time, no+view identity read). Per cell type t and gene g: -1. Δμ_g = pb(later) − pb(earlier), log space. Two modes:-   - `global`: one Δμ per gene; EB shrinkage w_g = n_g/(n_g+K), K=30, n_g = detected-cell count-     over both stages;-   - `type` (ships): per-cell-type Δμ_t using the type's cells in each stage, shrunk by-     min(n1_g,n2_g)/(min+K); genes with <20 detections in either stage of that type fall back to-     the global shrunk Δμ.-2. Sparse gate: keep |Δμ_s| > ZMIN·median(|Δμ_s|), ZMIN=1.5.-3. Apply to each sampled cell after α=3 recentering: x[i,g] += clip(β·Δμ_s, ±MAX_SHIFT), MAX_SHIFT=1.0,-   only at existing nonzeros, clipped ≥0 (sparsity and covariation preserved). β=1.0 ships.-4. Single-input views (proxy10): block never runs → output byte-identical to parent node 10-   (verified via h5 X/data comparison). No RNG in the block; deterministic.+1. `dmu = pb_t(later) − pb_t(earlier)` (log space), EB-shrunk by within-type detection count+   `min(n1,n2)/(min(n1,n2)+K)`, K=30; genes with <20 detections in either stage of the type fall+   back to the global shrunk delta (same fallback as parent).+2. Welch t: `t = dmu / sqrt(var1/n1 + var2/n2)` from the type's cells in the two stages+   (variance via E[x²]−E[x]², sparse-safe); fallback genes take the global t.+3. Confidence `conf = |t|/(|t|+T0)`, T0=3. Gate: `|t| ≥ T_MIN=2`. Per-type truncation: keep only+   the TOPG=200 strongest-|t| genes.+4. Displacement `D = clip(β·conf·dmu_shrunk, ±MAX_SHIFT=1.0)`, β=3, added at existing nonzeros+   only, clipped ≥0 (sparsity and covariation preserved), after the α=3 recentering. -Env knobs: VEC_SHIFT_BETA / VEC_SHIFT_K / VEC_SHIFT_ZMIN / VEC_SHIFT_MAX / VEC_SHIFT_MODE.+Ships: MODE=tconf, β=3, T0=3, T_MIN=2, TOPG=200. Env knobs: VEC_SHIFT_{BETA,K,ZMIN,MAX,MODE,+T0,TMIN,TOPG}. Single-input views (proxy10): block never runs → output byte-identical to parent.+Deterministic (no RNG in the block); CPU only (EXECUTION.json gpu:false). -## Screen (vec-score A-half, seed 0; X3 parent = 54.51)+## Screen (vec-score A-half, X3, seed 0; parent type β=1 = 54.88, de −0.043, dir 0.075, mmd 0.0259)  | variant | X3 board | de_score | de_dir | mmd_u | variogram | |---|---|---|---|---|---|-| global β=0.5 | 54.27 | −0.014 | 0.076 | 0.0278 | 0.001215 |-| global β=5 | 47.46 | −0.229 | −0.008 | 0.0360 | 0.001601 |-| global β=20 | 38.53 | −0.157 | −0.061 | 0.0581 | 0.003925 |-| type β=5 | 53.71 | −0.057 | 0.022 | 0.0250 | 0.001412 |-| type β=15 | 48.35 | −0.129 | −0.028 | 0.0296 | 0.002165 |-| **type β=1 (ships)** | **54.88** | −0.043 | 0.075 | 0.0259 | 0.001206 |-| type β=2 | 54.88 | −0.057 | 0.057 | 0.0250 | 0.001232 |-| type β=3 | 54.88 | −0.029 | 0.043 | 0.0246 | 0.001279 |--- **Global mode is actively harmful when amplified** (PLAN risk #3 realized): the whole-embryo-  E8.75→E9.0 Qiu delta, scaled up, moves X3 in the wrong direction on all four metrics. Ships OFF.-- **Typed mode gives a small, flat-topped gain** (β=1/2/3 all 54.88, +0.37 over parent, within the-  ~2-pt noise): the gain is mmd_u (0.0272→0.0259, +0.6 pts), de_score/de_direction ~unchanged at-  β=1, variogram unchanged. **The PLAN's target group (de_recovery) was NOT lifted** — X3 de_score-  stays below floor. The honest read: per-type deltas mostly reposition cell states slightly toward-  the later stage (helps MMD), not a true DE direction recovery.-- β=1 chosen over 2/3 (identical board) because it preserves de_direction best (0.075 vs 0.043).--## Mechanism evidence (off control & diagnostics, X3 seed 0)--- Off control VEC_SHIFT_BETA=0: X3 output byte-identical to parent run (verified).-- Active (typed β=1): 186,400 (type,gene) pairs pass the gate (median gate over |Δμ_s|);-  1.18M nonzero entries touched, all changed; mean |shift| ≈ 0.041 log units; shifts are-  type-specific (not a constant vector): per-type Δμ differ by construction; >0 within-type-  dispersion since gate+fallback differ per type. Cells changed: all 652 sampled X3 cells.-- Proxy10 default run byte-identical to parent (1 input → block inactive), so proxy10 score is-  unchanged by construction (no query spent re-verifying).+| b1 tmin2 g200 | 55.11 | 0.014 | 0.0835 | 0.02662 | 0.001208 |+| b2 tmin2 g0 | 55.23 | −0.014 | 0.0811 | 0.02568 | 0.001213 |+| b2 tmin2 g200 (=tmin3) | 55.66 | 0.057 | 0.0822 | 0.02609 | 0.001214 |+| b2 tmin3 g0 | 55.57 | 0.043 | 0.0812 | 0.02596 | 0.001217 |+| **b3 tmin2 g200 (ships)** | **55.67** | **0.043** | **0.0826** | **0.02572** | 0.001221 |+| b4 tmin2 g0 | 55.71 | 0.014 | 0.0752 | 0.02478 | 0.001236 |+| b4 tmin2 g200 (=tmin3) | 55.47 | 0.014 | 0.0812 | 0.02552 | 0.001229 |+| b2 tmin2 g100 | 55.46 | 0.043 | 0.0841 | 0.02634 | 0.001212 |++Seed robustness (same-seed paired, X3 A-half): seeds 0/1/2 → tconf 55.67/54.87/55.26 vs parent+54.88/54.68/55.29 (Δ +0.79/+0.19/−0.03, mean +0.32). de_direction improves in 3/3 seeds+(0.0826/0.0738/0.0738 vs 0.0748/0.0683/0.0652); mmd_u improves in 3/3; de_score improves at+seed 0 only (0.043 vs −0.043; seeds 1/2 both −0.014 vs −0.029/0.0). **Honest read: total gain is+inside the ~2-pt noise; the consistent parts are de_direction and mmd_u, not de_score.**+β=3+TOPG=200 chosen over β=4 g0 (higher seed-0 board) because g0's gain is mmd-only with+de_direction falling to parent level and variogram degrading, and it lacks the PLAN truncation+mechanism. PLAN's literal acceptance thresholds (de_dir≥0.095, mmd≤0.0252) match B-half values+and are unreachable on A-half; applied relatively (≥ parent on dir and mmd) instead.++## Mechanism evidence++- Gated pairs: 20,816 (type,gene) pass |t|≥2 before truncation (per-type min/med/max+  1422/1755/2326 — far above PLAN's ≥30 floor, no T_MIN relaxation needed); TOPG=200 truncates+  to 2,000 pairs, touching 36,808 nonzero entries (all changed), mean |shift| 0.28 (β=3).+- Spearman(|D|, |t|) over displaced pairs = 0.78 (β=3): displacement magnitude scales with+  statistical confidence, unlike the parent's uniform kernel (by construction 0 correlation).+- Off control VEC_SHIFT_BETA=0: shift block skipped entirely (no [gene_shift] line), output =+  recenter-only path = node-10-equivalent. Uniform-mode control (T0=0⇒conf≡1, TMIN=0, TOPG=0,+  β=3): touches 1,187,380 entries vs shipped 36,808 (32×), confirming the shipped sparsity comes+  from the confidence gate + truncation, and the shipped-vs-uniform difference is exactly the+  new kernel.+- Parent-path regression: VEC_SHIFT_MODE=type β=1 seed 0 reproduces parent A-half exactly+  (54.877, de −0.0429, dir 0.0748, mmd 0.02589, var 0.001206 vs parent's reported+  54.88/−0.043/0.075/0.0259/0.001206) → all deltas above come from the new kernel, no+  collateral code change.+- proxy10 (single input): ran seed 0, n_input_stages=1, block inactive, vec-check ok; identical+  code path to parent ⇒ proxy10 stays 58.80 (not re-scored, query budget).  ## Verified / not verified -- Verified: vec-check ok on proxy and X3; off-control byte-identity; proxy byte-identity;-  default == VEC_SHIFT_BETA=1 VEC_SHIFT_MODE=type run; determinism (no RNG in block).-- Not verified: B-half (A-half gain +0.37 is within noise — may not replicate); seeds ≠0; K/ZMIN-  grids (only β scanned); behaviour on the final view (E8.5+E9.5 official, larger types) — the-  typed delta there is official-data-based, expected better-powered than X3's 2.2k-cell Qiu stages.-- Queries used: 9 of 20.--## Knowledge used--None beyond parent (textbook cell-cycle/OXPHOS/glycolysis sets). No external files, no held-out-stage info, no forbidden-window measurements; the two-stage delta is computed live from the view's-own inputs. View-independent: uses only input order by time and time differences implicitly (none-hard-coded), no manifest identity fields, no absolute times, no file names.--## Next direction--Typed deltas help mmd_u only; de_recovery on X3 remains sub-floor for every variant tried. A-direction signal from prior/ regulatory knowledge (CollecTRI TF targets, Reactome cardiac-pathways) constrained to genes with two-stage support is the untested lever; alternatively accept-de_recovery as floor-bound on X3 and push mmd_u (composition) further.+- Verified: determinism (seed 0 twice byte-identical); vec-check ok on proxy and X3; off+  controls above; 3-seed paired A-half comparison on X3.+- Not verified: B-half scores (all screening is A-half); proxy10 score not re-queried (argued+  byte-identical); no biological prior used (sources: none — the kernel is purely statistical);+  gain may not replicate given it is within noise on total board.diff --git a/solution/run.py b/solution/run.pyindex bcfe293..61a4064 100644--- a/solution/run.py+++ b/solution/run.py@@ -62,7 +62,7 @@ RECENTER_K = 0.0       # per-type shrinkage of alpha: a_t = alpha * n_sel(t)/(n_                        # alpha, proxy types large (~284) — shrinkage would pull X3 back toward                        # alpha=2 while barely helping proxy. Kept as an off control. -SHIFT_BETA = 1.0       # per-type two-stage direction shift (only when >=2 input stages; ships MODE=type).+SHIFT_BETA = 3.0       # confidence-weighted per-type two-stage shift (>=2 input stages; ships MODE=tconf).                        # Per gene: dmu_g = pb(last) - pb(prev); EB shrinkage w_g = n_g/(n_g+K) with                        # n_g = detected-cell count over both stages; sparse gate |dmu_s| > ZMIN*median;                        # x[i,g] += clip(BETA*dmu_s_g, +-MAX_SHIFT) at nonzero entries, clipped >=0.@@ -75,6 +75,10 @@ SHIFT_K = 30.0         # EB shrinkage constant (detection-count based) SHIFT_ZMIN = 1.5       # sparse gate: multiple of median |dmu_shrunk| SHIFT_MAX = 1.0        # per-gene shift cap (log space) +SHIFT_T0 = 3.0         # tconf mode: conf = |t|/(|t|+T0); displacement = clip(beta*conf*dmu_s, +-MAX)+SHIFT_TMIN = 2.0       # tconf gate: only (type,gene) pairs with |t| >= TMIN are displaced+SHIFT_TOPG = 200       # per-type truncation: keep only the TOPG strongest-|t| genes (0 = off)+ CYCLE = [     "Mki67", "Top2a", "Pcna", "Ccna2", "Ccnb1", "Ccnb2", "Ccnd1", "Ccneg", "Ccne1",     "Cdk1", "Cdk2", "Cdk4", "Cdk6", "Mcm2", "Mcm3", "Mcm4", "Mcm5", "Mcm6", "Mcm7",@@ -405,6 +409,109 @@ def gene_delta_by_type(view, entries, genes, labels, inv, k_shrink, min_det=20):     return out  +def gene_delta_t_typed(view, entries, genes, labels, inv, k_shrink, min_det=20):+    """Per (type,gene) two-stage delta with Welch t-statistic.++    Returns (dmus_t, t_t): EB-shrunk per-type delta (same fallback-to-global rule as+    gene_delta_by_type) and t = dmu / sqrt(var1/n1 + var2/n2) computed from the type's cells in+    the last two input stages. Low-detection genes (<min_det in either stage of the type) fall+    back to the global delta/t."""+    a1 = read_stage(view, entries[-2], genes, missing="fill")+    a2 = read_stage(view, entries[-1], genes, missing="fill")+    lab1 = labels_of(a1) if "celltype" in a1.obs.columns else np.full(a1.n_obs, "all")+    lab2 = labels_of(a2) if "celltype" in a2.obs.columns else np.full(a2.n_obs, "all")+    X1, X2 = a1.X.tocsr(), a2.X.tocsr()++    def stats(Xr):+        n = max(Xr.shape[0], 1)+        m = np.asarray(Xr.mean(axis=0), dtype=np.float64).ravel()+        sq = np.asarray(Xr.multiply(Xr).mean(axis=0), dtype=np.float64).ravel()+        v = np.maximum(sq - m * m, 0.0)+        c = np.asarray(Xr.getnnz(axis=0)).ravel()+        return m, v, c, n++    gm1, gv1, gc1, gn1 = stats(X1)+    gm2, gv2, gc2, gn2 = stats(X2)+    gdmu = gm2 - gm1+    gse = np.sqrt(gv1 / gn1 + gv2 / gn2)+    with np.errstate(divide="ignore", invalid="ignore"):+        gt = np.where(gse > 1e-12, gdmu / gse, 0.0)+    gnn = gc1 + gc2+    gdmus = gdmu * (gnn / (gnn + float(k_shrink)))++    uniq = np.unique(labels)+    out_d = np.tile(gdmus, (len(uniq), 1))+    out_t = np.tile(gt, (len(uniq), 1))+    pos = {t: i for i, t in enumerate(uniq)}+    for t in np.unique(np.concatenate([lab1, lab2])):+        if t not in pos:+            continue+        r1 = np.where(lab1 == t)[0]+        r2 = np.where(lab2 == t)[0]+        if len(r1) == 0 or len(r2) == 0:+            continue+        m1, v1, c1, n1 = stats(X1[r1])+        m2, v2, c2, n2 = stats(X2[r2])+        d = m2 - m1+        se = np.sqrt(v1 / n1 + v2 / n2)+        with np.errstate(divide="ignore", invalid="ignore"):+            tt = np.where(se > 1e-12, d / se, 0.0)+        cmin = np.minimum(c1, c2)+        ds = d * (cmin / (cmin + float(k_shrink)))+        ok = (c1 >= min_det) & (c2 >= min_det)+        out_d[pos[t]] = np.where(ok, ds, gdmus)+        out_t[pos[t]] = np.where(ok, tt, gt)+    return out_d, out_t+++def apply_gene_shift_tconf(Xsel, dmus_t, tstat_t, inv_out, beta, t0, tmin, topg, max_shift):+    """Confidence-weighted sparse typed displacement.++    conf = |t|/(|t|+T0); D = clip(beta*conf*dmu_s, +-max_shift) on (type,gene) pairs with+    |t| >= tmin, optionally truncated to the topg strongest-|t| genes per type. Applied at+    existing nonzeros only, clipped >=0 (sparsity/covariation preserved). beta=0 -> untouched."""+    from scipy import sparse as sp++    at = np.abs(tstat_t)+    conf = at / (at + t0) if t0 > 0 else np.ones_like(at)+    keep = at >= tmin+    D = np.zeros(dmus_t.shape, dtype=np.float64)+    if topg and topg > 0:+        for i in range(D.shape[0]):+            ki = np.where(keep[i])[0]+            if ki.size == 0:+                continue+            if topg < ki.size:+                ki = ki[np.argsort(-at[i, ki], kind="stable")[:topg]]+            D[i, ki] = np.clip(beta * conf[i, ki] * dmus_t[i, ki], -max_shift, max_shift)+    else:+        D = np.where(keep, np.clip(beta * conf * dmus_t, -max_shift, max_shift), 0.0)+    D = D.astype(np.float32)+    kept_per_type = [int(k.sum()) for k in keep]+    Xc = sp.csr_matrix(Xsel).copy()+    row_of = np.repeat(np.asarray(inv_out), np.diff(Xc.indptr))+    dvec = D[row_of, Xc.indices]+    m = dvec != 0+    before = Xc.data[m].copy()+    Xc.data[m] = np.maximum(Xc.data[m] + dvec[m], 0.0)+    changed = Xc.data[m] != before+    Xc.eliminate_zeros()+    nz = D[keep] if keep.any() else np.array([0.0])+    diag = {+        "pairs_kept_total": int(keep.sum()),+        "kept_per_type_min_med_max": [min(kept_per_type), int(np.median(kept_per_type)), max(kept_per_type)],+        "entries_touched": int(m.sum()),+        "entries_changed": int(changed.sum()),+        "mean_abs_shift": float(np.abs(dvec[m]).mean()) if m.any() else 0.0,+        "mean_abs_D_kept": float(np.abs(nz).mean()),+        "max_abs_t_kept": float(at[keep].max()) if keep.any() else 0.0,+    }+    if keep.any() and np.std(D[keep]) > 0 and np.std(at[keep]) > 0:+        from scipy.stats import spearmanr+        diag["spearman_absD_vs_abst"] = float(spearmanr(np.abs(D[keep]), at[keep]).statistic)+    return Xc, diag++ def apply_gene_shift_typed(Xsel, dmus_t, inv_out, beta, zmin, max_shift):     from scipy import sparse as sp @@ -514,10 +621,20 @@ def main() -> None:     Xsel, diag = recenter(X, inv, Xsel, inv[idx], alpha, k_adapt=k_adapt)     n_inputs = len(manifest["inputs"])     print(f"[inputs] n_input_stages={n_inputs} shift_beta={shift_beta}", file=sys.stderr)-    shift_mode = os.environ.get("VEC_SHIFT_MODE", "type")   # global | type (screening knob)+    shift_mode = os.environ.get("VEC_SHIFT_MODE", "tconf")   # global | type | tconf (screening knob)+    shift_t0 = float(os.environ.get("VEC_SHIFT_T0", SHIFT_T0))+    shift_tmin = float(os.environ.get("VEC_SHIFT_TMIN", SHIFT_TMIN))+    shift_topg = int(float(os.environ.get("VEC_SHIFT_TOPG", SHIFT_TOPG)))     if shift_beta != 0.0 and n_inputs >= 2:         entries = sorted(manifest["inputs"], key=lambda e: float(e["time"]))-        if shift_mode == "type":+        if shift_mode == "tconf":+            dmus_t, tstat_t = gene_delta_t_typed(args.data, entries, genes, labels, inv, shift_k)+            Xsel, sdiag = apply_gene_shift_tconf(+                Xsel, dmus_t, tstat_t, inv[idx], shift_beta, shift_t0, shift_tmin, shift_topg, shift_max+            )+            sdiag["mode"] = "tconf"+            sdiag["params"] = {"beta": shift_beta, "t0": shift_t0, "tmin": shift_tmin, "topg": shift_topg}+        elif shift_mode == "type":             dmus_t = gene_delta_by_type(args.data, entries, genes, labels, inv, shift_k)             Xsel, sdiag = apply_gene_shift_typed(Xsel, dmus_t, inv[idx], shift_beta, shift_zmin, shift_max)             sdiag["mode"] = "type"

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k041Within-stage pseudotime and graph toolkit offline: scanpy DPT/PAGA/Leiden, Palantir, CellRank 210.1186/s13059-019-1663-x (PAGA); 10.1038/s41587-019-0068-4 (Palantir); 10.1038/s41592-024-02303-9 (CellRank 2)
k031Offline OT toolkit in the sandbox: moscot TemporalProblem, wot OTModel, POT, geomloss10.1038/s41586-024-08453-2 (moscot); 10.1016/j.cell.2019.01.006 (Waddington-OT)
k018Damped per-type shift: shrinkage alpha on the observed deltanotes/plan/cards/T1.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么把父节点(14)的均匀型内位移核(β=1,中位数门控)换成置信度加权稀疏核:每 (型,基因) 算两阶段 Welch t,conf=|t|/(|t|+3),仅位移 |t|≥2 且每型 |t| 前 200 的基因,D=clip(3·conf·dmu_shrunk, ±1),保留 EB 收缩 K=30;仅 ≥2 输入阶段视图激活,proxy10 代码路径不变。
各组分数的变化cell_state:噪声内:X3 mmd_u 0.02514→0.02498,得分 +0.07,组 +0.17(3/3 种子一致小幅改善但幅度极小)。
covariation:噪声内略降:X3 variogram 0.001211→0.001225,得分 -0.07,组 -0.23;稀疏化(触 36.8K 条目 vs 父 1.18M)基本保住了共变结构。
de_recovery:变好但幅度小:X3 de_score 原始值 -0.026→0.039(由地板下转正到地板上),得分 +0.56,组 skill +1.51;proxy10 完全不变。Engineer 自己的 3 种子数据显示 de_score 仅 seed 0 改善,跨种子不稳。
direction:噪声内偏好:X3 de_direction 0.0916→0.099,得分 +0.08,组 +0.22;Engineer 报告 de_direction 3/3 种子一致改善(0.0826/0.0738/0.0738 vs 0.0748/0.0683/0.0652),是本改动最一致的正向项。
family_idcomposition_program
假设是否成立unclear
经验
  1. 在 X3 型内两阶段位移上,用 Welch t 连续缩放 + 每型截尾(top-200)替代均匀 β 位移,使 de_score 原始值首次由地板下转正(-0.026→0.039),且位移稀疏 32 倍(36.8K vs 1.18M 条目)后 variogram 几乎不受损——统计门控比放大幅度更有效。
  2. X3 单种子 de_score 变化不可信:出厂配置 seed 0 榜分 55.67 vs 父 54.88,但 seeds 1/2 配对差仅 +0.19/-0.03,de_score 只在 seed 0 转正;任何以 de_score 为主的收益必须 ≥3 种子配对确认。
  3. 总榜分 +0.44(56.39→56.83)在 ~2 分噪声内;proxy10 因单输入完全不动,本方法族的收益上限被锁在 X3 的 2/3 权重里,单项零点几分的改善不会推动榜分。
  4. 对照设计有效可复用:β=0 off control 跳过整块、均匀模式 control(T0=0/TMIN=0/TOPG=0)触条目数差 32 倍、type β=1 路径精确回归复现父 A 半 54.877,三者共同把差异归因到新核而非附带代码改动。
mechanism_active是
下一步建议
  1. 针对 X3 de_recovery/direction:把 conf 核与 CollecTRI 等先验方向做一致性过滤(只在网络支持方向上位移),并用 3 种子配对验证 de_score 是否跨种子稳定为正——这是让本改动脱离噪声的唯一途径。
  2. 针对 X3 cell_state:β=4 无截尾(g0)在筛选中给出全场最好 mmd_u 0.02478,可做 de 分支(tconf top-200)与 mmd 分支(g0)的组合或按型加权,但需先确认两者位移向量不互相抵消。
  3. 在投入更多位移核变体前,先用一次 B 半或多种子查询验证 tconf 相对父的差异是否可复现(当前 A 半 3 种子均值仅 +0.32);若不可复现,应转向组成/mmd_u 杠杆而非继续调 DE 位移参数。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。2
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 1 次:write 1
用时不到 1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 15,098 · 输出 1,190 · 思考 952
任务(第一行)分析节点 n21 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
ANALYSIS.json written
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/analyst.jsonl 6 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/analyst.stderr

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数33
工具调用共 38 次:bash 26、edit 8、read 3、write 1
用时19 分
token 数输入 68,586 · 输出 11,620 · 思考 13,210
任务(第一行)按计划修改程序:置信度加权稀疏型内位移:按 |t| 缩放并截尾,替代均匀 β 位移
最后的回答(摘录)
proxy2 harness module unavailable here; the 2-input path is already proven on X3 (same code path). Final state is committed in `solution/`.

**总结**

- **做了什么**:按 PLAN 把父节点(14)的均匀型内位移核换成置信度加权稀疏核:每 (型,基因) 用两阶段 Welch t 统计量算 conf=|t|/(|t|+3),仅位移 |t|≥2 且每型 |t| 前 200 的基因,D=clip(3·conf·dmu_shrunk, ±1),EB 收缩 K=30 保留;仅双输入视图激活,proxy10 与父逐字节一致(type 路径回归验证精确复现父 A 半 54.877)。
- **查分**(17/20 次):X3 A 半 seed 0 出厂配置 55.67(父 54.88),de_score −0.043→+0.043;seeds 1/2 配对 +0.19/−0.03(3 种子均值 +0.32,在 ~2 分噪声内);de_direction 与 mmd_u 在 3/3 种子一致改善。off control:β=0 块整体跳过;均匀模式触 1.19M 条目 vs 出厂 36.8K(32×),Spearman(|D|,|t|)=0.78 证明位移幅度随置信度缩放。vec-check 两视图通过,确定性验证通过。
- **下一步最值得试**:出厂增益在噪声内且 de_score 仅 seed 0 转正——B 半可能不认。若继续,方向应是 (1) 把 conf 核与 prior/ CollecTRI 方向做一致性过滤(只在网络支持的方向上位移,可能让 de_score 跨种子稳定为正);(2) 或承认 X3 型内信号已到上限,把预算转向 mmd_u(β=4 无截尾 g0 的 mmd 0.02478 是全场最好,可作为 cell_state 专用分支与 de 分支加权组合)。
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/engineer.jsonl 181 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/engineer.stderr

调研员

角色调研员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数1
工具调用共 0 次:—
用时6 分
token 数输入 21,419 · 输出 6,818 · 思考 8,192
任务(第一行)为节点 n21(改进,父节点 n14)写改进计划 PLAN.json
最后的回答(摘录)
模式开关(VEC_SHIFT_CONF=off → conf≡1、t 门控关闭、均匀)应复现父节点。
5. proxy10:不激活 → 完全一致,无需查询。

报告:位移基因数、位移大小的分布、位移与 |t| 的相关性、四个指标各自的变化。

风险:小类型中嘈杂的 t;稀疏化过度 → 无操作(检查通过门控的基因数,应 ≥ 每类型 50 个基因……);收益在噪声内;确认需要种子重复。

写。
</think>

```json
{"name": "置信度加权稀疏型内位移:按 |t| 缩放并截尾,替代均匀 β 位移",
 "motivation": "父节点 14 最弱组 de_recovery 50.14(X3 de_score 原始值 -0.026,仍在地板下),direction 也仅 58.17。父节点筛选显示均匀 β 位移呈平顶(β=1/2/3 同为 54.88)且 β=5 起全面崩坏(53.71),说明中位数门控放进来的噪声基因与有把握基因获得同样幅度的位移,放大 β 时噪声被同步放大——这是可解释的结构问题。数据驱动 Δμ(节点 14)与 CollecTRI 网络方向(节点 17,56.12<56.39)均已试且未抬起 de 两项,本次不再换方向来源,而是改位移核本身:按 (型,基因) 的两阶段 t 统计量连续缩放位移并只保留每型证据最强的基因。",
 "approach": "在父节点 SHIFT 块内(仅 ≥2 输入阶段激活,proxy10 单输入不激活、保持逐字节一致):1) 对每个 (型,基因) 用该型两阶段细胞算伪批量差 dmu 与合并标准误 se,t=dmu/se;2) 置信因子 conf=|t|/(|t|+T0),T0 初值 3(搜索 1–10);3) 位移 = clip(β·conf·dmu_shrunk, ±MAX_SHIFT),仅作用于 |t|≥T_MIN(初值 2.0,搜索 2/3)的基因;可选每型只取 |t| 前 TOPG 个基因(初值 200,0=关闭),使被位移的基因成为预测中的主导变化基因,直接对应 de_score 看排序的机制;4) 保留原有非零位置施加、clip≥0、EB 收缩 K=30 作二级收缩。筛选:X3 A 半 seed 0,小网格 β∈{2,4}×T_MIN∈{2,3}×TOPG∈{0,200},约 6–8 次查询;入选标准不是总分而是 de_direction≥0.095 且 mmd_u≤0.0252(不许掉 cell_state 收益)。最佳配置用 seed 1、2 复跑确认(差距 <2 分视为噪声)。总查询 ≤12,余量留给 off control 与完全一致验证。",
 "expected_groups": ["de_recovery", "direction"],
 "risks": "1) X3 型内细胞少(~65/型),t 统计量本身嘈杂,|t| 排序可能仍是噪声——Engineer 应在筛选前先打印过门控的 (型,基因) 数,若每型 <30 个基因则放宽 T_MIN 到 1.5;2) 稀疏化过头退化成无操作(输出≈父),用 off control 对比位移条目数即可发现;3) 收益可能仍在 ~2 分噪声内,必须多子复确认,禁止单次 +0.5 就出厂;4) 不改善 proxy10(单输入),节点分提升只能来自 X3(权重 2/3),预期上限低。",
 "family_id": "composition_program",
 "mechanism": "型内程序变化部分的位移核从'均匀 β × 硬中位数门控'改为'按 (型,基因) 两阶段 t 统计量连续缩放 + 每型截尾保留最强证据基因',位移幅度与统计置信度成正比。",
 "vs_constant_shift": "每型不是统一常向量:同一型内不同基因的位移随各自 |t| 连续变化(0 到 β·dmu),且每型只作用证据最强的基因子集;组成部分(加权抽样、重锚定)完全不动,位移可整体关闭。",
 "mechanism_evidence": "Engineer 报告:过门控 (型,基因) 对数与每型分布、位移幅度与 |t| 的秩相关(应显著为正)、实际被改的非零条目数、四分项各自变化(重点 X3 de_score/de_direction 与 mmd_u);β=4+conf 与父 β=1 的位移向量差异应集中在高 |t| 基因。",
 "mechanism_off_control": "同一程序 VEC_SHIFT_BETA=0 再跑一次,预期与父节点输出逐字节一致;另设 conf≡1 且门控关闭的'均匀模式'复现父节点位移,确认新核相对父的差异全部来自置信度加权。",
 "sources": []}
```
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/researcher.jsonl 26 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/researcher.stderr

审查员

角色审查员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数8
工具调用共 10 次:bash 8、read 1、write 1
用时2 分
token 数输入 32,356 · 输出 1,146 · 思考 2,698
任务(第一行)审查节点 n21 的程序是否越界读取、写死目标、钻评分器漏洞
最后的回答(摘录)
REVIEW.json written
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/reviewer.jsonl 129 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/21/reviewer.stderr