Virtual Embryo Challenge更新于 10-03 18:47(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-093415-search-t1-r2-D-s0

节点 n4 在终选来历上

composition_trend + 型内拟时序表达平移(已实现并筛查;λ>0 均降 X3 加权分,默认 λ=0 出厂)

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-093415-search-t1-r2-D-s0
父节点n3
子节点n8
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。改进
状态已打分
分数搜索目标分 53.69(+0.0) · X3 48.71(+0.0) · proxy10 63.66(+0.0)
审查未审查
用时?从运行开始到结束(或到现在)的挂钟时间。30 分
程序版本26100abfc56a8be2376009196bd7131419e2917e (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 26100abfc5:solution/METHOD.md

composition_trend + 型内拟时序表达平移(已实现并筛查;λ>0 均降 X3 加权分,默认 λ=0 出厂)

Parent (node 3, composition_trend): type-level composition reweighting + within-type maturity selection of real cells; expression never modified. This node adds, per PLAN (family composition_program), a gene-specific, type-specific expression shift along within-type pseudotime trends with global shrinkage — implemented, verified working, and screened on both rulers. The mechanism failed the PLAN's own success criterion and ships disabled (LAM=0 default, VEC_LAM env re-enables); with λ=0 the program is byte-identical to the parent on both views.

Method (delta vs parent)

Cell selection, weights, sampling and n are unchanged. After selecting the output cells, when λ≠0:

  1. Within-type slopes. For each cell type with ≥10 cells in the latest input stage, OLS slope β[t,g] of gene g against the parent's z_fate (within-type centred diffusion pseudotime). Genes kept only if detected in ≥2% of the type's cells and standardized slope |r| = |β·σ_g/σ_z| ≥ 0.1; others β=0. Types with <10 cells fall back to the global all-cell slope (same filters).
  2. Shift. X[i,g] += λ·Δ·β[t,g] where X[i,g]>0 (ACT_TOPQ=0), clipped at ≥0. Δ=0.3. Zero-activation of top-5% positive-slope genes was implemented (ACT_TOPQ) and tested.
  3. Variance guard. Genes whose output std collapses below 0.5× unshifted std are reverted.
  4. Single-stage. Slopes use only the latest input stage; identical code path on single- and multi-input views. No absolute times, no view identity, no file names read.

Mechanism-off control and evidence the mechanism runs

  • VEC_LAM=0 output is byte-identical to parent node 3 on both views (verified: X3 and proxy matrix diffs = 0 nonzero).
  • λ=0.15/ACT_TOPQ=0.05: 0.82M (X3) / 9.1M (proxy) matrix entries changed; ~0.5% of genes move pseudobulk by >0.01 (below PLAN's hoped 20% — the detection+|r| filters are restrictive).
  • λ=0.3/ACT_TOPQ=0: 0.46M / 6.8M entries changed; pseudobulk cosine with parent 0.99998.
  • DE metrics are scale-invariant, so λ only matters through the mix of shift vs composition in dp; de_score was identical at λ=0.3 and λ=1.0 (−0.2143), i.e. the shift dominates the dp ranking from λ≈0.3 up and moves it the wrong way on X3.

Screen results (vec-score A-half, seed 0; parent: X3 48.71 / proxy10 63.66 → node 53.69)

configX3 pts (de_score / de_dir / mmd / vario)proxy10 ptsnode score (X3×2+proxy)/3
λ=0 (= parent)48.71 (−0.182 / 0.026 / .472 / .533 skill)63.6653.69
λ=0.15, ACT_TOPQ=0.05(de_dir skill .505, mmd .468, vario .478)~63.6< parent
λ=0.3, ACT_TOPQ=048.20 (−0.214 / 0.011 / .470 / .527)63.9553.45
λ=1.0, ACT_TOPQ=047.93 (−0.214 / 0.008 / .470 / .516)—< 53.45

Findings: (a) zero-activation (ACT_TOPQ>0) clearly hurts variogram → shipped 0; (b) on X3 (weight ×2) the forward pseudotime shift monotonically degrades de_score, de_direction, mmd_u and variogram — the within-type maturity axis re-amplifies exactly the signal composition selection already pushes, and X3 truth does not reward more of it; (c) proxy10 gains +0.3 at λ=0.3 (mmd_u 0.651→0.662), within noise and outweighed 2:1 by X3. PLAN required ≥3-point X3 de_score gain for success; measured direction is negative → mechanism falsified, ships OFF.

Verified / not verified

  • Verified: λ=0 ≡ parent on proxy and X3; final default code runs on both views (proxy 82 s, X3 ~60 s, seed 0) and passes vec-check; determinism (no RNG in the shift path).
  • Not verified: λ<0 (shift toward progenitors — biologically backwards, not screened); Δ independent of λ (only the product enters); seeds >0; the final two-input view (code path identical by construction).
  • Queries used: 5 of 20.

Hyper-parameters

NameValueWhere it came from
LAM (λ)0.0screen above (0.15/0.3/1.0 all net-negative on the weighted score)
DELTA (Δ)0.3PLAN default (only λ·Δ matters)
DET_MIN / r_min / MIN_SLOPE_CELLS0.02 / 0.1 / 10PLAN + risk mitigation 1
ACT_TOPQ0.0PLAN 0.05 tested; degraded X3 variogram skill 0.533→0.478
parent constantsunchangednode 3

Knowledge used: same as parent (textbook cell-cycle / OXPHOS / glycolysis gene sets); the shift uses no external knowledge — slopes are estimated from the view's own data. No held-out stage, no forbidden-window information, no external dataset.

Suggested next direction

The maturity/pseudotime axis is exhausted on X3 (both composition and expression pushes along it degrade de_score below chance). A different signal is needed for X3's de_recovery — e.g. slope estimation against a trajectory axis rooted differently (lineage-branching direction from the prior/TF-network resources), or per-gene shifts constrained by CollecTRI regulator activity rather than pseudotime. Also worth testing: λ<0 was never screened and X3's negative de_score means the parent's dp is anti-correlated with truth — a dampening (not reversal) of the maturity component of dp might lift X3 de_score toward 0, but must be checked against proxy10 where the same component is beneficial.

调研员的计划

名称composition_program: add within-type pseudotime-regression expression shift with shrinkage
动机Node 3 scores 53.69 but de_recovery is 49.78 (below floor). On X3 (weight ×2), de_score = −0.1818 (skill 0.445, below floor) and mmd_u skill 0.472 (below floor). The parent never modifies expression—pseudobulk change comes only from composition reweighting. On X3 (heart-specific, different gene panel and cell types), this composition-only signal mis-ranks genes, producing a negative de_score. Adding a data-driven, gene-specific expression shift within each type should correct the direction signal and lift de_recovery, especially on X3.
做法Keep the parent's composition reweighting and cell selection unchanged. Add a new step after cell selection:

1. Within-type slope estimation. For each cell type with ≥10 selected cells, regress each gene's expression against the already-computed z_fate (within-type centred diffusion pseudotime) using robust rank correlation (Spearman ρ) or simple OLS on the z-scored values. Store slope β_g,type for each gene g. For types with <10 cells or genes with near-zero variance, set β=0.

2. Expression shift. For each selected cell i of type t, shift: X_shifted[i,g] = X[i,g] + λ · β_g,t · Δ, where Δ is a fixed pseudotime advance (start Δ=0.3, range [0.1, 0.5]). λ is the global shrinkage (start λ=0.15, search [0.05, 0.3]). To preserve sparsity: only add the shift where X[i,g]>0 OR where β_g,t is in the top 5% positive slopes (allowing a small number of gene activations); clip shifted values at 0.

3. Variance guard. After shifting, verify per-gene std has not collapsed (std_shifted > 0.5 × std_original); if it has for a gene, revert that gene. This protects variogram.

4. Single-input fallback. The method uses only the latest input stage (pseudotime is computed within it), s…
风险1) Pseudotime slopes are noisy for low-expression genes → noisy shifts could hurt variogram. Mitigation: only shift genes with |Spearman ρ| > 0.1 within type; monitor variogram score. 2) On X3, cell type vocabulary may differ or types may be too small for stable regression → slope estimation fails. Mitigation: fall back to global (all-cell) slope for types with <10 cells. 3) Shift could push mmd_u worse if it moves cells away from true distribution. Mitigation: keep λ small (≤0.3); check mmd_u skill doesn't drop below 0.47. 4) Improvement on X3 de_score may be <2 points (noise level). Engineer should require ≥3-point gain on X3 de_score raw value before declaring success; use 2 seeds to confirm. 5) Runtime: regression on ~20k genes × ~5k cells is fast (<5s), no risk to 30-min budget.

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:父节点版本 6bb7e1fdf1。改动的文件:solution/METHOD.md +80 −67、solution/run.py +114 −1

diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 3bc5139..335ce34 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,72 +1,85 @@-# composition_trend — type-level composition reweighting + within-type maturity selection--**Provenance.** Derived from agent-produced node 33 of run `20261002-034201-search-t1-abc-r1-A-era`-(programs.git `refs/nodes/33`, commit `10400ee`; the program behind the official T1:val 50.1, submissions.tsv-2026-10-02). The mechanism and every constant are the node's own defaults; nothing was re-tuned. What was removed,-for the seed contract only (2026-10-03):--- the external-maturity axis `A_EXT` (read manifest `source == "external"`; a view-identity read; it never fired on-  final / proxy / patched X3);-- the multi-stage pooling `POOL`, the composition-trajectory axis `A_TR`, the kNN expression smoothing `A_S`, the-  apoptosis / cell-level proliferation terms (all 0 in the node's submitted defaults) and the CellRank fate mode;-- the `inputs_by_time(include_external=False)` base selection: the base is simply the latest input stage by time;-- the `VEC_*` environment overrides (constants are fixed in the code).--On every current neutral view (final, proxy, X1, X3–X6) the seed and node 33 select the same cells; see the admission-record for the digest check. The mechanism-split study (`notes/reports/dev/2026-10-02_mech_split_50.md`, candidate-`r1A:CUS` = `r1A:CU`) is the analysis of exactly this procedure.--**Naming.** "trend" is descriptive, not a two-stage model: the type weights come from a single snapshot (types whose-cells are, on average, less proliferative get a larger share). On the official final view this moves the predicted-pseudobulk along the previous step's direction (cosine with E9.5 − E8.5: 0.33 for the composition part, 0.34 for the-whole procedure; mech_split §4), and in the E8.5 → E9.5 rehearsal the composition part carries most of the gain-(+10.0 of +11.4 points; mech_split §2.1). Nothing in the code reads a second stage.--Contract: `python run.py --data <view> --out <pred.h5ad> --seed <int>`; CPU only (`EXECUTION.json {"gpu": false}`).--## Method--Base = the latest input stage of the view (final: E9.5; proxy: E8.5; rulers: their last input), all cells, its own-`celltype` labels (whatever vocabulary the view provides; `Unknown` is an ordinary type).--1. **Cell scores** (generic gene-set means, z-scored over the stage, clipped at ±3): proliferation (34 cell-cycle-   genes), metabolic maturity `z_met` = OXPHOS (13 genes) − glycolysis (10 genes).-2. **Within-type commitment** `z_fate`: HVG 2000 → PCA 30 → kNN 30 → diffusion map 15 → diffusion pseudotime rooted-   at the most progenitor-like cell (argmax 2·z_prolif − z_met, first index on ties); z-scored, the type mean-   removed, z-scored again (only the order inside a type matters).-3. **Weights.** Type layer `w_t = exp(−0.55 · z(type-mean z_prolif))` (types with < 5 cells: 1). Cell layer-   `w_i = exp(1.2 · z_met_i + 0.6 · z_fate_i)`. `w = clip(w_t · w_i, 1e-6, 1e6)`.-4. **Sampling.** n = the stage's cell count clipped to [min_cells, max_cells] (final: 5,118 of 17,057). Per-type-   quota ∝ the type's weight sum (largest remainder, capped by type size, overflow re-apportioned); inside a type 5-   equal-frequency `z_met` strata, stratum quotas ∝ stratum weight sums, Efraimidis–Spirakis weighted sampling-   without replacement inside a stratum. Expression values are copied unchanged.--Mechanism-off controls (for analysis; not code switches): `A_TP = 0` leaves the input composition (mech_split "U");-`A_CM = A_FATE = 0` with the same per-type quotas is uniform within type (mech_split "C").--## Data / knowledge used--Only the latest input stage of the view. Generic knowledge: textbook cell-cycle, OXPHOS and glycolysis gene sets-(pathway membership, not stage measurements). No held-out stage, no information from (E9.5, E13.5], no external-dataset, no `uns.celltype_palette`, no frozen probe, no pre-trained weights.+# composition_trend + 型内拟时序表达平移(已实现并筛查;λ>0 均降 X3 加权分,默认 λ=0 出厂)++Parent (node 3, composition_trend): type-level composition reweighting + within-type maturity+selection of real cells; expression never modified. This node adds, per PLAN+(family `composition_program`), a gene-specific, type-specific expression shift along+within-type pseudotime trends with global shrinkage — implemented, verified working, and+screened on both rulers. **The mechanism failed the PLAN's own success criterion and ships+disabled (LAM=0 default, `VEC_LAM` env re-enables); with λ=0 the program is byte-identical to+the parent on both views.**++## Method (delta vs parent)++Cell selection, weights, sampling and n are unchanged. After selecting the output cells, when λ≠0:++1. **Within-type slopes.** For each cell type with ≥10 cells in the latest input stage, OLS+   slope β[t,g] of gene g against the parent's z_fate (within-type centred diffusion+   pseudotime). Genes kept only if detected in ≥2% of the type's cells and standardized slope+   |r| = |β·σ_g/σ_z| ≥ 0.1; others β=0. Types with <10 cells fall back to the global all-cell+   slope (same filters).+2. **Shift.** X[i,g] += λ·Δ·β[t,g] where X[i,g]>0 (ACT_TOPQ=0), clipped at ≥0. Δ=0.3.+   Zero-activation of top-5% positive-slope genes was implemented (ACT_TOPQ) and tested.+3. **Variance guard.** Genes whose output std collapses below 0.5× unshifted std are reverted.+4. **Single-stage.** Slopes use only the latest input stage; identical code path on single- and+   multi-input views. No absolute times, no view identity, no file names read.++## Mechanism-off control and evidence the mechanism runs++- `VEC_LAM=0` output is **byte-identical to parent node 3** on both views (verified: X3 and+  proxy matrix diffs = 0 nonzero).+- λ=0.15/ACT_TOPQ=0.05: 0.82M (X3) / 9.1M (proxy) matrix entries changed; ~0.5% of genes move+  pseudobulk by >0.01 (below PLAN's hoped 20% — the detection+|r| filters are restrictive).+- λ=0.3/ACT_TOPQ=0: 0.46M / 6.8M entries changed; pseudobulk cosine with parent 0.99998.+- DE metrics are scale-invariant, so λ only matters through the *mix* of shift vs composition+  in dp; de_score was identical at λ=0.3 and λ=1.0 (−0.2143), i.e. the shift dominates the+  dp ranking from λ≈0.3 up and moves it the *wrong way* on X3.++## Screen results (vec-score A-half, seed 0; parent: X3 48.71 / proxy10 63.66 → node 53.69)++| config | X3 pts (de_score / de_dir / mmd / vario) | proxy10 pts | node score (X3×2+proxy)/3 |+|---|---|---|---|+| λ=0 (= parent) | 48.71 (−0.182 / 0.026 / .472 / .533 skill) | 63.66 | 53.69 |+| λ=0.15, ACT_TOPQ=0.05 | (de_dir skill .505, mmd .468, vario **.478**) | ~63.6 | < parent |+| λ=0.3, ACT_TOPQ=0 | 48.20 (−0.214 / 0.011 / .470 / .527) | 63.95 | 53.45 |+| λ=1.0, ACT_TOPQ=0 | 47.93 (−0.214 / 0.008 / .470 / .516) | — | < 53.45 |++Findings: (a) zero-activation (ACT_TOPQ>0) clearly hurts variogram → shipped 0; (b) on X3+(weight ×2) the forward pseudotime shift monotonically degrades de_score, de_direction, mmd_u+and variogram — the within-type maturity axis re-amplifies exactly the signal composition+selection already pushes, and X3 truth does not reward more of it; (c) proxy10 gains +0.3 at+λ=0.3 (mmd_u 0.651→0.662), within noise and outweighed 2:1 by X3. PLAN required ≥3-point X3+de_score gain for success; measured direction is negative → mechanism falsified, ships OFF.++## Verified / not verified++- Verified: λ=0 ≡ parent on proxy and X3; final default code runs on both views (proxy 82 s,+  X3 ~60 s, seed 0) and passes vec-check; determinism (no RNG in the shift path).+- Not verified: λ<0 (shift toward progenitors — biologically backwards, not screened);+  Δ independent of λ (only the product enters); seeds >0; the final two-input view (code path+  identical by construction).+- Queries used: 5 of 20.  ## Hyper-parameters  | Name | Value | Where it came from | |---|---|---|-| `A_TP` | −0.55 | node 33's default (tuned by the agent on the old proxy / proxy2) |-| `A_CM` | 1.2 | node 33's default (same) |-| `A_FATE` | 0.6 | node 33's default (agent scan +0.2 … +1.0 on proxy seed 0, peak at 0.6) |-| `K_STRAT`, `MIN_TYPE_CELLS` | 5, 5 | node 33's defaults |-| HVG / PCA / kNN / diffmap | 2000 / 30 / 30 / 15 | node 33's defaults |--All were chosen on the old single-input proxy (E8.5 → E9.5), which is also where they look best; on the two-input-rulers the mechanism is near neutral (admission record).--## Known failure modes--- The type layer is a single-snapshot heuristic: it bets that low-proliferation types expand next. Where the next-  step's composition change does not go that way, every ranked metric gets worse (mech_split zero-development control).-- Within-type selection pushes the sample towards OXPHOS-high / late-pseudotime cells; on patched X3 this part was-  −2.4 points (MMD −1.8) on top of the composition (mech_split §3).-- Real cells only: no new expression states, no new cell types.+| LAM (λ) | 0.0 | screen above (0.15/0.3/1.0 all net-negative on the weighted score) |+| DELTA (Δ) | 0.3 | PLAN default (only λ·Δ matters) |+| DET_MIN / r_min / MIN_SLOPE_CELLS | 0.02 / 0.1 / 10 | PLAN + risk mitigation 1 |+| ACT_TOPQ | 0.0 | PLAN 0.05 tested; degraded X3 variogram skill 0.533→0.478 |+| parent constants | unchanged | node 3 |++Knowledge used: same as parent (textbook cell-cycle / OXPHOS / glycolysis gene sets); the shift+uses no external knowledge — slopes are estimated from the view's own data. No held-out stage,+no forbidden-window information, no external dataset.++## Suggested next direction++The maturity/pseudotime axis is exhausted on X3 (both composition and expression pushes along+it degrade de_score below chance). A different signal is needed for X3's de_recovery — e.g.+slope estimation against a trajectory axis rooted differently (lineage-branching direction from+the prior/TF-network resources), or per-gene shifts constrained by CollecTRI regulator activity+rather than pseudotime. Also worth testing: λ<0 was never screened and X3's negative de_score+means the parent's dp is anti-correlated with truth — a *dampening* (not reversal) of the+maturity component of dp might lift X3 de_score toward 0, but must be checked against proxy10+where the same component is beneficial.diff --git a/solution/run.py b/solution/run.pyindex b7ef52b..2027299 100644--- a/solution/run.py+++ b/solution/run.py@@ -38,6 +38,15 @@ A_FATE = 0.6      # cell layer: within-type diffusion pseudotime (more committed K_STRAT = 5       # z_met strata per type MIN_TYPE_CELLS = 5 +LAM = 0.0         # global shrinkage of the within-type pseudotime expression shift.+                  # Screened 0.15/0.3/1.0 (ACT_TOPQ 0/0.05, seed 0): every lam>0 lowered the+                  # weighted node score (X3 de_score -0.18 -> -0.21, de_direction 0.026 -> 0.011,+                  # monotone in lam; proxy10 +0.3 at best). Ships OFF; set VEC_LAM to re-enable.+DELTA = 0.3       # pseudotime advance (in z_fate units) for the shift+MIN_SLOPE_CELLS = 10   # types below this fall back to the global (all-cell) slope+DET_MIN = 0.02    # gene must be detected in >=2% of the type's cells to be shifted+ACT_TOPQ = 0.0    # zero-activation OFF: top-5% activation degraded X3 variogram (0.533 -> 0.478)+ CYCLE = [     "Mki67", "Top2a", "Pcna", "Ccna2", "Ccnb1", "Ccnb2", "Ccnd1", "Ccneg", "Ccne1",     "Cdk1", "Cdk2", "Cdk4", "Cdk6", "Mcm2", "Mcm3", "Mcm4", "Mcm5", "Mcm6", "Mcm7",@@ -154,6 +163,101 @@ def stratified_sample(w, z_strat, inv, n_types, n_out, K, rng):     return np.sort(np.concatenate(out)) if out else np.empty(0, dtype=np.int64)  +def type_slopes(X, z_fate, inv, n_types, det_min=DET_MIN, r_min=0.1):+    """OLS slope of each gene vs within-type pseudotime, per cell type.++    beta[t, g] = cov(x_g, z_fate | t) / var(z_fate | t); genes that are too sparse+    (<det_min detected) or whose standardized slope is |r| < r_min get beta = 0.+    Types with < MIN_SLOPE_CELLS cells fall back to the global (all-cell) slope.+    """+    from scipy import sparse as sp++    n_cells, n_genes = X.shape+    Xc = X.tocsr()+    beta = np.zeros((n_types, n_genes), dtype=np.float64)++    def slope_for(rows):+        z = z_fate[rows]+        zc = z - z.mean()+        zz = float(zc @ zc)+        if zz < 1e-9 or len(rows) < 2:+            return None, None+        Xs = Xc[rows]+        b = np.asarray(zc @ Xs).ravel() / zz          # cov/var (zc sums to 0)+        mu = np.asarray(Xs.mean(axis=0)).ravel()+        sq = np.asarray(Xs.multiply(Xs).mean(axis=0)).ravel()+        var_g = np.maximum(sq - mu * mu, 0.0)+        det = np.asarray(Xs.getnnz(axis=0)).ravel() / len(rows)+        return b, (var_g, det)++    g_beta, g_stats = slope_for(np.arange(n_cells))+    if g_beta is None:+        return beta+    g_std_z = float(np.std(z_fate))+    g_r = g_beta * np.sqrt(g_stats[0]) / max(g_std_z, 1e-9)+    g_keep = (g_stats[1] >= det_min) & (np.abs(g_r) >= r_min)+    g_beta = np.where(g_keep, g_beta, 0.0)++    for t in range(n_types):+        rows = np.where(inv == t)[0]+        if len(rows) < MIN_SLOPE_CELLS:+            beta[t] = g_beta+            continue+        b, stats = slope_for(rows)+        if b is None:+            beta[t] = g_beta+            continue+        std_z = float(np.std(z_fate[rows]))+        r = b * np.sqrt(stats[0]) / max(std_z, 1e-9)+        keep = (stats[1] >= det_min) & (np.abs(r) >= r_min)+        beta[t] = np.where(keep, b, 0.0)+    return beta+++def apply_shift(Xsel, inv_out, beta, lam, delta, act_topq=ACT_TOPQ, chunk=1024):+    """X + lam*delta*beta[type] where X>0 or beta is in the top act_topq positive+    slopes of the cell's type; clipped at >=0. Per-gene variance guard: genes whose+    output std collapsed below 0.5x are reverted. lam=0 leaves X untouched."""+    from scipy import sparse as sp++    Xc = sp.csr_matrix(Xsel)+    n_out, n_genes = Xc.shape+    orig_std = np.asarray(Xc.multiply(Xc).mean(axis=0)).ravel() - np.asarray(Xc.mean(axis=0)).ravel() ** 2+    orig_std = np.sqrt(np.maximum(orig_std, 0.0))++    scale = float(lam) * float(delta)+    # per-type activation mask for zero entries (top positive slopes)+    n_act = int(round(act_topq * n_genes))+    act_mask = np.zeros((beta.shape[0], n_genes), dtype=bool)+    if scale != 0.0 and n_act > 0:+        for t in range(beta.shape[0]):+            b = beta[t]+            pos = np.where(b > 0)[0]+            if len(pos) == 0:+                continue+            if len(pos) > n_act:+                pos = pos[np.argsort(-b[pos], kind="stable")[:n_act]]+            act_mask[t, pos] = True++    cols = np.arange(n_genes)+    out = np.empty((n_out, n_genes), dtype=np.float32)+    for s in range(0, n_out, chunk):+        e = min(s + chunk, n_out)+        D = np.asarray(Xc[s:e].todense(), dtype=np.float32)+        types = inv_out[s:e]+        for j, t in enumerate(types):+            b = (scale * beta[t]).astype(np.float32)+            m = (D[j] > 0) | act_mask[t]+            D[j] = np.where(m, np.maximum(D[j] + b, 0.0), D[j])+        out[s:e] = D++    new_std = np.sqrt(np.maximum((out ** 2).mean(axis=0) - out.mean(axis=0) ** 2, 0.0))+    bad = (orig_std > 1e-6) & (new_std < 0.5 * orig_std)+    if bad.any():+        out[:, bad] = np.asarray(Xc[:, bad].todense(), dtype=np.float32)+    return sp.csr_matrix(out)++ def main() -> None:     ap = argparse.ArgumentParser()     ap.add_argument("--data", required=True)@@ -188,7 +292,16 @@ def main() -> None:     n_out = target_n_cells(manifest, adata.n_obs)     rng = np.random.default_rng(args.seed)     idx = stratified_sample(w, z_met, inv, len(uniq), n_out, K_STRAT, rng)-    write_prediction(X[idx], genes, args.out, seed=args.seed)++    import os+    lam = float(os.environ.get("VEC_LAM", LAM))       # mechanism-off control: VEC_LAM=0+    delta = float(os.environ.get("VEC_DELTA", DELTA))+    actq = float(os.environ.get("VEC_ACTQ", ACT_TOPQ))+    Xsel = X[idx]+    if lam != 0.0 and delta != 0.0:+        beta = type_slopes(X, z_fate, inv, len(uniq))+        Xsel = apply_shift(Xsel, inv[idx], beta, lam, delta, act_topq=actq)+    write_prediction(Xsel, genes, args.out, seed=args.seed)   if __name__ == "__main__":

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

用到的知识库条目

编号标题出处
k020Correcting sampling-scope (dissection) bias in compositionnotes/guides/modeling_and_evaluation_guide.html
k036Composition forecasting and mixture models for population predictionnotes/competition/05_lineage_graph.md; notes/competition/09_t1_census_lineage.md; 10.1038/s41586-024-08453-2 (growth rates)
k015Composition x conditional-expression decomposition p(x|t) = sum_z p(z|t) p(x|z,t)notes/handover/03_当前方案与Agent系统设计.md

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

改了什么在父节点(node 3, composition_trend)基础上实现了型内拟时序回归表达平移(OLS斜率+检出率/|r|过滤+方差守卫+零激活开关)并筛查λ∈{0.15,0.3,1.0};所有λ>0净降分,出厂默认λ=0,输出与父节点逐字节相同,榜分53.69不变(耗时81.8s vs 147.5s因缓存/负载差异,非机制)。
各组分数的变化cell_state:不变(53.21):出厂λ=0;筛查中λ=0.3时proxy10 mmd_u skill 0.651→0.662(+0.3,噪声内),X3 mmd_u skill 0.472→0.470(噪声内)
covariation:不变(54.21):出厂λ=0;筛查中零激活ACT_TOPQ=0.05明显伤X3 variogram(skill 0.533→0.478,得分约-1.1),λ=0.3+ACTQ=0时变化在噪声内
de_recovery:不变(49.78):出厂λ=0;筛查中λ=0.3使X3 de_score原始值-0.182→-0.214(skill 0.445不变分以下),proxy10 de_score 0.2679不变
direction:不变(57.77):出厂λ=0;筛查中λ=0.3使X3 de_direction 0.026→0.011(得分-约1.6,超出T1噪声,方向为负)
family_idcomposition_program
假设是否成立否
经验
  1. 在composition_trend已用成熟度轴(z_met+z_fate)做组成/细胞选择的基础上,再沿同一型内拟时序轴做表达平移是重复放大同一信号:X3上de_score/de_direction随λ单调恶化(-0.182→-0.214,0.026→0.011),四项无一改善,proxy10仅+0.3(噪声内)且被X3权重×2压倒。
  2. 把零表达项抬成小正数(零激活top-5%斜率基因)明确伤X3 variogram(skill 0.533→0.478),验证了评分规则简报中'抬零值损害mmd_u和variogram'的警告;平移只在X>0处施加是更安全的选择。
  3. DE两项对dp的幅度缩放不变,λ只通过平移与组成信号在dp排序中的相对权重起作用;λ≥0.3时de_score在λ=0.3和λ=1.0完全相同(-0.2143),说明平移已主导排序,继续加大λ无信息。
  4. 机制证伪后出厂λ=0(与父节点逐字节相同)是正确做法:保住53.69不退步,且METHOD.md留下完整筛查表,避免后续节点重复试同一方向。
  5. 检出率≥2%+|r|≥0.1的过滤使实际被平移的基因只占约0.5%(PLAN期望>20%):稀疏单细胞数据上按基因回归斜率过滤后,能触及的基因远少于预期,机制即使方向对、杠杆也很小。
mechanism_active否
下一步建议
  1. 针对X3(de_recovery):父节点dp与真值反相关(de_score=-0.18低于chance),成熟度轴已耗尽;改用独立信号源做基因方向平移,如用prior/CollecTRI调控网络推断的调节因子活性约束每基因的位移方向,而非型内拟时序斜率。
  2. 针对X3:筛查λ<0(沿成熟度轴回退/衰减dp的成熟度分量,可能把de_score从-0.18拉向0),但必须同时监控proxy10的mmd_u和de_direction(那里成熟度分量是有益的,X3权重×2需净收益>0)。
  3. 针对direction/cell_state:若要动组成,尝试换轨迹根方向(谱系分支方向而非'最祖细胞'启发式argmax 2·z_prolif−z_met)重新估计z_fate后再做型内选择权重,检验X3上组成信号方向是否可翻转。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。

分析员

角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。分析员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数?大模型一共回复了几轮(每轮可以调用多个工具)。2
工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。共 1 次:write 1
用时不到 1 分
token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。输入 13,505 · 输出 1,160 · 思考 177
任务(第一行)分析节点 n4 的打分结果,写 ANALYSIS.json
最后的回答(摘录)
ANALYSIS.json written
原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/4/analyst.jsonl 6 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/4/analyst.stderr

工程师

角色工程师 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数32
工具调用共 37 次:bash 26、edit 7、read 2、write 2
用时26 分
token 数输入 82,733 · 输出 11,643 · 思考 17,839
任务(第一行)按计划修改程序:composition_program: add within-type pseudotime-regression expression shift with shrinkage
最后的回答(摘录)
完成。做了什么、查分结果、下一步:

**做了什么**:按 PLAN 在父节点(node 3, composition_trend)上实现了型内拟时序回归表达平移(OLS 斜率 β×λ·Δ,检出率/标准化斜率过滤,全局斜率回退,方差守卫,零激活开关)。机制对照验证:λ=0 与父节点在两视图逐字节相同;λ=0.15/0.3 确实改变了 0.5M–9M 个矩阵元素。

**查分结果**(5/20 次额度,seed 0):λ=0.3+ACTQ=0 → X3 48.20(父 48.71)、proxy10 63.95(父 63.66),加权节点分 53.45 < 53.69;λ=1.0 更差;零激活(ACT_TOPQ=0.05)明显伤 X3 variogram(skill 0.533→0.478)。X3 上 de_score/de_direction 随 λ 单调恶化——拟时序平移放大的正是组成选择已在推的成熟度轴,X3 真值不奖励更多。**机制按 PLAN 自身标准(X3 de_score ≥+3 分)被证伪,出厂默认 λ=0**(=父节点行为,82s 跑完 proxy,两视图均过 vec-check),METHOD.md 完整记录证据与数据。

**下一步最值得试**:X3 的 de_score 为负(低于 chance)说明父节点 dp 与真值反相关,成熟度轴已耗尽;应换信号源——用 prior/(CollecTRI 调控网络)约束的基因方向平移,或对成熟度分量做衰减(λ<0 方向未筛查,但需同时盯 proxy10,那里该分量是有益的)。
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/4/engineer.jsonl 249 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/4/engineer.stderr

调研员

角色调研员 alibaba-token-plan-cn/qwen3.8-max
调用次数1 次
轮数1
工具调用共 0 次:—
用时1 分
token 数输入 14,054 · 输出 1,568 · 思考 2,228
任务(第一行)为节点 n4(改进,父节点 n3)写改进计划 PLAN.json
最后的回答(摘录)
get.",
  "family_id": "composition_program",
  "mechanism": "Within each cell type, estimate gene-specific expression trends along diffusion pseudotime (already computed by parent). Extrapolate a small step along these trends for selected cells, with global shrinkage λ. This adds a gene-directional signal to the pseudobulk that composition reweighting alone cannot provide.",
  "vs_constant_shift": "A constant per-type shift applies the same vector to all cells of a type regardless of gene identity. Here the shift is gene-specific (each gene has its own slope from pseudotime regression), cell-position-dependent (magnitude can scale with the cell's pseudotime position), and type-specific (slopes are estimated within each type separately). The direction comes from data-driven within-type trajectories, not from a hand-written displacement vector. It also preserves sparsity structure by only activating top-slope genes at zeros.",
  "mechanism_evidence": "Engineer should verify: (1) With λ>0, the per-gene dp (pseudobulk change vs reference) has non-zero correlation with the within-type slopes (confirming the shift is actually changing gene rankings). (2) de_score and de_direction on X3 improve by ≥0.05 raw units. (3) The fraction of genes with |shift| > 0.01 is substantial (>20% of genes), confirming the mechanism touches many genes. (4) Four-group breakdown: de_recovery should rise, mmd_u should not drop >2 points. (5) Compare output pseudobulk cosine with parent pseudobulk: should be <0.99 (expression actually changed, not just composition).",
  "mechanism_off_control": "Set λ=0 in the same program. This zeroes out the expression shift entirely; the output should be identical to parent node 3 (same cells selected, expression unmodified). Engineer runs once with λ=0 and once with λ=0.15, confirms (a) λ=0 score matches parent ≈53.69, (b) λ=0.15 changes de_score/de_direction. If outputs are identical for λ=0 and λ=0.15, the mechanism is not running.",
  "sources": []
}
```
原始记录位置/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/4/researcher.jsonl 7 KB
/home/spark-longxinyang/vec/runs/formal/20261003-093415-search-t1-r2-D-s0/nodes/4/researcher.stderr