总览 · ← 返回运行 20261002-202907-search-t1-scr-A
节点 n5 在终选来历上
在父节点分型 PCA 均值外推上叠加"形状演化":按细胞型估计各 PC 两阶段方差比,收缩后乘时间比、只允许扩张(r≥1),逐细胞缩放其在 PC 空间偏离型心的残差,使型内分布展宽、共变结构随时间演化。
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261002-202907-search-t1-scr-A |
|---|---|
| 父节点 | n4 |
| 子节点 | n7、n9 |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 改进 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 58.56(+2.4) · X3 58.56(+2.4) · 3 次复测均分 59.11 |
| 审查 | 通过 1 越界读取:未发现问题——run.py 只通过 src.task1_temporal.view_io 的 load_manifest/read_stage/panel_genes/covered_mask 读取 --data 视图内文件(run.py:31-41,77-100),无绝对路径、..、/mnt、/home、data/raw、downloads、打分器或 src/common/evaluation 引用,无联网。; 2 硬编码目标统计量:未发现问题——常量仅为超参数(α=1.8、τ=0.3、β=3、clip 等,run.py:65-98);细胞类型全部从 labels_of(pr… |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 22 分 |
| 程序版本 | a9e37a676c193f8943b3e2fde7bc2cd9e88a4a3d (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git a9e37a676c:solution/METHOD.md
在父节点分型 PCA 均值外推上叠加"形状演化":按细胞型估计各 PC 两阶段方差比,收缩后乘时间比、只允许扩张(r≥1),逐细胞缩放其在 PC 空间偏离型心的残差,使型内分布展宽、共变结构随时间演化。
方法
父节点(node 4)流程不变:两输入阶段共同覆盖基因取前 2500 HVG,合并中心化后 svds(v0=ones,确定性)取 25 PC;位移幅度 s = α·(t_target−t_last)/(t_last−t_prev),α=1.8,clip [0,4](X3 上 dt_in=0.25、dt_out=0.5 → s=3.6);对两阶段都可配对(≥5 细胞)的同名类型在型内求 PC 均值差,λ_k=var_k/(var_k+0.3) 收缩后逆变换回基因空间,只加到该型细胞的已有非零元素上,clip≥0;单输入视图精确退化为 copy_last。
新增形状演化(family: lowrank_shape,默认开,LOWRANK_SHAPE_BETA=3.0):
- 对每个配对类型(形状部分要求两阶段各 ≥10 细胞,
LOWRANK_SHAPE_MINCELLS=10),算型内每 PC 方差 var_prev_k、var_last_k,比值 r_k=var_last_k/var_prev_k。 - 向 1 收缩:r_s = 1+(r−1)·var_last/(var_last+τ_shape),τ_shape=0.5(
LOWRANK_SHAPE_TAU)。 - 按同一时间比外推:r_t = clip(1+s·(r_s−1), 1.0, 4.0)。下界 1.0 = 只允许扩张(
LOWRANK_SHAPE_RLO):收缩方向被冻结(数据里型内方差随时间总体增大,且收缩外推在 s=3.6 下把大量 PC 推成负值被 clip,实测拖累 cell_state)。 - 逐细胞:delta_pc = β·(sqrt(r_t)−1)·(cell_pc − 型心_pc),β=3.0,作用于全部 25 PC(
LOWRANK_SHAPE_NPCS=25),逆变换到基因空间,与均值位移相加后只加到非零元素上,clip≥0,eliminate_zeros。 - 不满足配对/细胞数条件的类型:只做均值位移或完全不动(
LOWRANK_FALLBACK=none,与父一致)。
β 扫描(X3 A 半,seed0):0→54.90(=父)、1→55.88、1.5→56.39、2→56.90、3→57.33、4→57.33、5→56.92、8→53.65(崩塌)。β∈[3,4] 平台,取 β=3(covariation 更高、离崩塌区更远)。
机制生效证据(对照 SHAPE_BETA=0)
- 对照:
LOWRANK_SHAPE_BETA=0输出与父节点逐元素 maxdiff=0.0(完全复现父节点,其 X3 A 半 seed0=54.90);LOWRANK_TAU=1e6时 λ≈0 → copy_last(父节点已验证,等价 47.92 基线)。 - 改变了哪些细胞:5 个类型被形状化(AVC-CM、IFT-CM、OFT/RV-CM、SV-CM、Unknown);Endocardium、V-CM 等不可配对类型位移差为 0(未动)。
- 非常数位移:形状开启后同型细胞间 |位移|(全基因求和)的 std = 63–115(常数位移应为 0)——同型不同 PC 位置的细胞得到不同位移。
- 型内方差变化:预测的型内平均基因方差 var_on/var_off = 1.03–1.13(5 个被形状化类型全部 >5% 变化的有 3 个:AVC-CM +12.6%、Unknown +5.9%、IFT-CM +5.6%)。
- 方差比非平凡:各型 mean|r_k−1| = 0.13–1.77,sqrt(r_t) ∈ [1, 2.0](上界 clip 生效),机制不在退化区(|r−1|≫0.05 放弃阈值)。
- 四组分变化(β=0 → β=3,seed0 A 半):cell_state 66.1→75.5(+9.4),de_recovery 52.0→52.5(+0.5),direction 50.2→50.1(−0.1),covariation 47.7→45.2(−2.5)。总分 54.90→57.33(+2.4)。mmd_u 0.0218→0.0176:目标阶段的分布确实比最后输入阶段更展宽,扩张型形状演化直接改善分布匹配。
X3 查分记录(A 半,seed0,除注明外)
| 配置 | 分 | cell_state | covar | de_rec | dir |
|---|---|---|---|---|---|
| 父 α=1.8(β=0 复现) | 54.90 | 66.1 | 47.7 | 52.0 | 50.2 |
| β=1 τ_s=0.5 mc5 npc10 clip[0.25,4] | 55.02 | 66.79 | 47.22 | 51.96 | 50.22 |
| mc50 τ_s=2 clip[0.5,2] | 55.00 | 66.71 | 47.21 | 51.96 | 50.20 |
| mc50 τ_s=0.5 clip[0.5,2] | 54.99 | 66.71 | 47.19 | 51.96 | 50.20 |
| 只扩张 RLO=1 mc50 npc10 | 55.78 | 69.50 | 47.02 | 51.96 | 50.16 |
| 只扩张 mc10 npc10 | 55.85 | 69.74 | 46.97 | 51.96 | 50.16 |
| 只扩张 mc50 npc25 | 55.88 | 69.90 | 46.92 | 51.96 | 50.15 |
| 只扩张 mc10 npc25 | 55.94 | 70.16 | 46.86 | 51.96 | 50.14 |
| β=1.5 | 56.39 | 71.96 | 46.43 | 51.96 | 50.10 |
| β=1.5 α=2.2 | 56.49 | 72.52 | 45.99 | 51.96 | 50.16 |
| β=2 | 56.90 | 73.47 | 46.03 | 52.47 | 50.12 |
| β=2 α=2.2 | 56.92 | 73.83 | 45.57 | 52.47 | 50.15 |
| β=3(提交) | 57.33 / 57.10(seed1) | 75.52 / 73.76 | 45.21 / 45.15 | 52.47 / 51.96 | 50.08 / 51.80 |
| β=4 | 57.33 | 76.09 | 44.45 | 52.47 | 49.98 |
| β=3 τ_s=2 | 57.18 | 74.93 | 45.35 | 52.47 | 50.07 |
| β=5 | 56.92 | 75.28 | 43.72 | 52.47 | 49.90 |
| β=8 | 53.65 | 67.14 | 41.78 | 50.96 | 49.64 |
验证过的
- vec-check ok;seed0 复跑逐元素一致;伪装视图(时间统一 +1 天、manifest 键乱序重排版、换路径)输出与真实视图逐元素 maxdiff=0.0 → 视图无关(s 只用时间差;RLO/clip/β 是常数不是绝对时间)。
- 默认配置输出与查分过的 c11(β=3)逐元素一致。
- 单输入退路(copy_last)逻辑未改动,与父一致。
- 运行 ~10 s、内存 <2 GB(limits: 28 GB / 30 min)。
没验证的 / 风险
- 未在 T1 proxy / proxy2 / final 视图实测(本节点只挂 X3)。final 上 dt_in=1、dt_out=1 → s=1.8,扩张幅度更温和(r_t=1+1.8(r_s−1)),机制同样成立,但幅度是否仍最优未测。
- β=3 是 A 半调出;A 半增益 +2.4(seed0)/ +1.7(seed1)略超 T1 噪声 2 分,B 半可能缩水。β∈[3,4] 平坦、[2,5] 都 >56.9,对 β 误差不敏感。
- covariation 随 β 单调下降(47.7→45.2),若 B 半 covariation 权重更高或地板更低,净收益会小于 A 半;但 4 组加权下 β=3 仍是最优平台。
- 方差比估计在小类型上噪声大(AVC-CM n_prev=18),已用 mc=10 + τ_shape 收缩 + 只扩张 + clip≤4 四重防护;mc10 与 mc50 分差仅 0.06。
知识来源
未使用任何保留阶段/基因型的测量信息;未读禁窗数据;未用 uns.celltype_palette。只用 view 内两个输入阶段的表达、标签、时间差。"细胞状态多样性随发育时间增加"是通用发育生物学常识(谱系 progressively 分化),不针对任何禁窗阶段;其余为统计常识(方差比收缩、PCA),无外部数据注入。
调研员的计划
| 名称 | Per-type variance-ratio shape evolution on top of mean shift |
|---|---|
| 动机 | Node 4 covariation 47.97 is the weakest group and flat vs parent (-0.47, noise). direction 50.37 only +1.17 (noise). ANALYSIS attributes covariation stagnation to additive shift + clip≥0 destroying gene-gene covariance, and notes 'low-rank shrinkage contribution unproven'—the actual lowrank_shape mechanism (covariance/shape evolution) was never implemented. The parent only shifts per-type means; evolving per-type variance structure in PC space directly targets covariation and adds per-cell heterogeneity for direction. |
| 做法 | Keep the parent's per-type mean-shift pipeline (PCA 25-dim, 2500 HVG, α=1.8, τ=0.3, sparsity-preserving). Add a shape-evolution step per matched cell type (≥5 cells both stages): 1. Compute per-type, per-PC variance in each stage: var_prev_k, var_last_k. 2. Variance ratio r_k = var_last_k / var_prev_k, shrunk toward 1: r_k_shrunk = 1 + (r_k - 1) * var_last_k / (var_last_k + τ_shape), τ_shape=0.5 (env LOWRANK_SHAPE_TAU, search 0.1–2.0). This prevents noisy variance estimates from dominating. 3. Extrapolate: r_target_k = 1 + s * (r_k_shrunk - 1), clip to [0.25, 4.0] (env LOWRANK_SHAPE_CLIP). 4. Per-cell affine in PC space: new_pc_k = (mean_last_k + s·δ_mean_k) + sqrt(r_target_k) · (cell_pc_k - mean_last_k). The first term is the existing mean shift; the second scales each cell's residual from the type centroid by the extrapolated variance change. 5. Inverse-transform to gene space. For each cell, compute delta_gene = new_reconstruction - old_reconstruction (HVG subspace only). Add delta_gene only to already-nonzero entries (preserve sparsity support), clip ≥ 0, eliminate_zeros. 6. Types without a match or <5 cells: no shift (LOWRANK_FALLBACK=none, same as parent). Single-input view… |
| 风险 | 1) Variance ratios noisy for small types (<20 cells): mitigated by τ_shape shrinkage toward 1 and minimum 5-cell threshold; Engineer should check if types with <15 cells have extreme r_k and consider raising threshold to 10. 2) Scaling up residuals may push cells away from expected positions, hurting cell_state: monitor cell_state in first vec-score run; if it drops >3 from 67.95, reduce SHAPE_CLIP to 2.0 or restrict to top 5 PCs. 3) If variance is roughly constant between stages (r_k≈1), the shape component is a no-op and adds nothing: check mean |r_k - 1| across types; if <0.05, the mechanism is not operating and the approach should be abandoned early (after 2 queries). 4) clip≥0 after inverse transform may still truncate: compare with multiplicative application x*exp(δ/x) on a subset; if covariation differs by >1 point, switch to multiplicative. 5) Engineer should verify SHAPE_BETA=0 reproduces parent score within noise before proceeding. |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:父节点版本 52b391699b。改动的文件:solution/METHOD.md +46 −31、solution/README.md +2 −2、solution/run.py +82 −23
diff --git a/solution/METHOD.md b/solution/METHOD.mdindex 9a1de51..6923e6c 100644--- a/solution/METHOD.md+++ b/solution/METHOD.md@@ -1,48 +1,63 @@-按细胞型在 PCA(25维,2500HVG) 空间估计两输入阶段均值差,方差收缩后乘时间比逐细胞加到最后阶段非零表达上(clip≥0);单输入退路 copy_last。+在父节点分型 PCA 均值外推上叠加"形状演化":按细胞型估计各 PC 两阶段方差比,收缩后乘时间比、只允许扩张(r≥1),逐细胞缩放其在 PC 空间偏离型心的残差,使型内分布展宽、共变结构随时间演化。 ## 方法 -- 读最后两个输入阶段(`read_stage` 默认 fill),取两阶段共同覆盖基因,按合并方差选前 2500 HVG。-- 合并矩阵中心化后 `scipy.sparse.linalg.svds`(固定 v0=ones,确定性)取前 25 个 PC,λ_k = var_k/(var_k+τ),τ=0.3(环境变量 `LOWRANK_TAU`)。-- 位移幅度 s = α·(t_target − t_last)/(t_last − t_prev),α=1.8,clip 到 [0,4]。只用时间差,不用绝对时间;X3 上 dt_in=0.25、dt_out=0.5 → s=3.6。-- **按细胞型位移**(默认开,`LOWRANK_PER_TYPE=1`):对两个阶段都有 ≥5 个细胞的同名类型,分别在该类型内部计算 PC 空间均值差,逆变换回基因空间,只加到该类型的细胞上;`Unknown` / 无法配对的类型不位移(`LOWRANK_FALLBACK=none`,实测优于用全局差兜底)。这一步是关键:全局均值差被细胞组成变化污染,直接加会把 de_direction 打成负分。-- 位移只加到已有非零元素上(保持稀疏支撑不变),clip≥0 后 eliminate_zeros。早期版本把 HVG 列写满成稠密块,covariation 从 48 掉到 27,已弃用。-- 单输入阶段视图:精确输出 copy_last(与父节点一致)。+父节点(node 4)流程不变:两输入阶段共同覆盖基因取前 2500 HVG,合并中心化后 `svds`(v0=ones,确定性)取 25 PC;位移幅度 s = α·(t_target−t_last)/(t_last−t_prev),α=1.8,clip [0,4](X3 上 dt_in=0.25、dt_out=0.5 → s=3.6);对两阶段都可配对(≥5 细胞)的同名类型在型内求 PC 均值差,λ_k=var_k/(var_k+0.3) 收缩后逆变换回基因空间,只加到该型细胞的已有非零元素上,clip≥0;单输入视图精确退化为 copy_last。 -## 机制关闭对照(mechanism_off_control)+新增形状演化(family: lowrank_shape,默认开,`LOWRANK_SHAPE_BETA=3.0`):+1. 对每个配对类型(形状部分要求两阶段各 ≥10 细胞,`LOWRANK_SHAPE_MINCELLS=10`),算型内每 PC 方差 var_prev_k、var_last_k,比值 r_k=var_last_k/var_prev_k。+2. 向 1 收缩:r_s = 1+(r−1)·var_last/(var_last+τ_shape),τ_shape=0.5(`LOWRANK_SHAPE_TAU`)。+3. 按同一时间比外推:r_t = clip(1+s·(r_s−1), 1.0, 4.0)。**下界 1.0 = 只允许扩张**(`LOWRANK_SHAPE_RLO`):收缩方向被冻结(数据里型内方差随时间总体增大,且收缩外推在 s=3.6 下把大量 PC 推成负值被 clip,实测拖累 cell_state)。+4. 逐细胞:delta_pc = β·(sqrt(r_t)−1)·(cell_pc − 型心_pc),β=3.0,作用于全部 25 PC(`LOWRANK_SHAPE_NPCS=25`),逆变换到基因空间,与均值位移相加后只加到非零元素上,clip≥0,eliminate_zeros。+5. 不满足配对/细胞数条件的类型:只做均值位移或完全不动(`LOWRANK_FALLBACK=none`,与父一致)。 -`LOWRANK_TAU=1e6` → 所有 λ_k≈0 → 位移≈0。本地对比关闭版与父节点 copy_last 预测:最大绝对差 3e-4(float32 舍入),等价于 copy_last(父节点 X3 = 47.92)。机制开启后 X3 seed0 = 54.90,机制确实生效:2331 个 HVG |位移|>0.01,shift 最大 1.1(log 空间),cell_state 49.7→66.1、de_recovery 44.1→52.0。+β 扫描(X3 A 半,seed0):0→54.90(=父)、1→55.88、1.5→56.39、2→56.90、3→**57.33**、4→57.33、5→56.92、8→53.65(崩塌)。β∈[3,4] 平台,取 β=3(covariation 更高、离崩塌区更远)。 -## X3 查分记录(A 半)+## 机制生效证据(对照 SHAPE_BETA=0)++- **对照**:`LOWRANK_SHAPE_BETA=0` 输出与父节点逐元素 maxdiff=0.0(完全复现父节点,其 X3 A 半 seed0=54.90);`LOWRANK_TAU=1e6` 时 λ≈0 → copy_last(父节点已验证,等价 47.92 基线)。+- **改变了哪些细胞**:5 个类型被形状化(AVC-CM、IFT-CM、OFT/RV-CM、SV-CM、Unknown);Endocardium、V-CM 等不可配对类型位移差为 0(未动)。+- **非常数位移**:形状开启后同型细胞间 |位移|(全基因求和)的 std = 63–115(常数位移应为 0)——同型不同 PC 位置的细胞得到不同位移。+- **型内方差变化**:预测的型内平均基因方差 var_on/var_off = 1.03–1.13(5 个被形状化类型全部 >5% 变化的有 3 个:AVC-CM +12.6%、Unknown +5.9%、IFT-CM +5.6%)。+- **方差比非平凡**:各型 mean|r_k−1| = 0.13–1.77,sqrt(r_t) ∈ [1, 2.0](上界 clip 生效),机制不在退化区(|r−1|≫0.05 放弃阈值)。+- **四组分变化**(β=0 → β=3,seed0 A 半):cell_state 66.1→75.5(+9.4),de_recovery 52.0→52.5(+0.5),direction 50.2→50.1(−0.1),covariation 47.7→45.2(−2.5)。总分 54.90→57.33(+2.4)。mmd_u 0.0218→0.0176:目标阶段的分布确实比最后输入阶段更展宽,扩张型形状演化直接改善分布匹配。++## X3 查分记录(A 半,seed0,除注明外) | 配置 | 分 | cell_state | covar | de_rec | dir | |---|---|---|---|---|---|-| 父 copy_last | 47.92 | 49.68 | 48.44 | 44.09 | 49.20 |-| 全局位移 τ=1(稠密块) | 41.62 | 42.95 | 27.65 | 43.80 | 49.02 |-| 全局位移 τ=1(保稀疏) | 46.83 | 47.04 | 47.66 | 43.80 | 48.94 |-| 分型位移均值合并 τ=1 | 47.37 | 44.33 | 47.72 | 48.62 | 49.50 |-| 分型逐细胞 fb=global τ=1 | 52.33 | 58.79 | 48.25 | 50.48 | 49.71 |-| 分型逐细胞 fb=none τ=1 | 52.82 | 59.67 | 48.52 | 50.96 | 49.88 |-| τ=0.3 / 0.1 | 52.92 / 52.95 | | | | |-| τ=3 / 10 | 52.31 / 51.64 | | | | |-| α=1.4 / 0.7 | 54.24 / 51.69 | | | | |-| **α=1.8 τ=0.3(提交)** | **54.90**(seed0)/ 55.44(seed1) | 66.1 | 47.7 | 52.0 | 50.2 |-| α=2.2 τ=0.3 | 54.94 | 66.8 | 47.4 | 51.5 | 50.3 |--α 曲线在 1.8–2.2 变平(s 达到 cap=4 附近),covariation 随 α 缓慢下降(48.5→47.4),α=1.8 取平衡点。+| 父 α=1.8(β=0 复现) | 54.90 | 66.1 | 47.7 | 52.0 | 50.2 |+| β=1 τ_s=0.5 mc5 npc10 clip[0.25,4] | 55.02 | 66.79 | 47.22 | 51.96 | 50.22 |+| mc50 τ_s=2 clip[0.5,2] | 55.00 | 66.71 | 47.21 | 51.96 | 50.20 |+| mc50 τ_s=0.5 clip[0.5,2] | 54.99 | 66.71 | 47.19 | 51.96 | 50.20 |+| 只扩张 RLO=1 mc50 npc10 | 55.78 | 69.50 | 47.02 | 51.96 | 50.16 |+| 只扩张 mc10 npc10 | 55.85 | 69.74 | 46.97 | 51.96 | 50.16 |+| 只扩张 mc50 npc25 | 55.88 | 69.90 | 46.92 | 51.96 | 50.15 |+| 只扩张 mc10 npc25 | 55.94 | 70.16 | 46.86 | 51.96 | 50.14 |+| β=1.5 | 56.39 | 71.96 | 46.43 | 51.96 | 50.10 |+| β=1.5 α=2.2 | 56.49 | 72.52 | 45.99 | 51.96 | 50.16 |+| β=2 | 56.90 | 73.47 | 46.03 | 52.47 | 50.12 |+| β=2 α=2.2 | 56.92 | 73.83 | 45.57 | 52.47 | 50.15 |+| **β=3(提交)** | **57.33** / 57.10(seed1) | 75.52 / 73.76 | 45.21 / 45.15 | 52.47 / 51.96 | 50.08 / 51.80 |+| β=4 | 57.33 | 76.09 | 44.45 | 52.47 | 49.98 |+| β=3 τ_s=2 | 57.18 | 74.93 | 45.35 | 52.47 | 50.07 |+| β=5 | 56.92 | 75.28 | 43.72 | 52.47 | 49.90 |+| β=8 | 53.65 | 67.14 | 41.78 | 50.96 | 49.64 | ## 验证过的 -- vec-check ok;seed 0/1 各自复跑逐元素一致(确定性);伪装视图(时间 +1、manifest 键乱序、换路径)输出与真实视图逐元素相同(max diff 0.0)→ 视图无关。-- 单输入视图(本地构造)退路正常,输出 652×32285。-- 运行 ~5 s、峰值内存 <2 GB,远低于 limits。+- vec-check ok;seed0 复跑逐元素一致;**伪装视图**(时间统一 +1 天、manifest 键乱序重排版、换路径)输出与真实视图逐元素 maxdiff=0.0 → 视图无关(s 只用时间差;RLO/clip/β 是常数不是绝对时间)。+- 默认配置输出与查分过的 c11(β=3)逐元素一致。+- 单输入退路(copy_last)逻辑未改动,与父一致。+- 运行 ~10 s、内存 <2 GB(limits: 28 GB / 30 min)。 -## 没验证的+## 没验证的 / 风险 -- 未在 T1 proxy / proxy2 / final 视图上跑(本节点只挂 X3 视图);final 视图 dt_in=1、dt_out=1 → s=1.8,机制同样成立,但分数未实测。-- α=1.8 是 X3 A 半上调出的,B 半与 final 上幅度可能不同;α 平坦区宽(1.4–2.2 都 >54),风险有限。-- 分型位移依赖两阶段标签可配对;若某视图标签大量为 Unknown,则多数细胞不位移,退化接近 copy_last(安全侧)。+- 未在 T1 proxy / proxy2 / final 视图实测(本节点只挂 X3)。final 上 dt_in=1、dt_out=1 → s=1.8,扩张幅度更温和(r_t=1+1.8(r_s−1)),机制同样成立,但幅度是否仍最优未测。+- β=3 是 A 半调出;A 半增益 +2.4(seed0)/ +1.7(seed1)略超 T1 噪声 2 分,B 半可能缩水。β∈[3,4] 平坦、[2,5] 都 >56.9,对 β 误差不敏感。+- covariation 随 β 单调下降(47.7→45.2),若 B 半 covariation 权重更高或地板更低,净收益会小于 A 半;但 4 组加权下 β=3 仍是最优平台。+- 方差比估计在小类型上噪声大(AVC-CM n_prev=18),已用 mc=10 + τ_shape 收缩 + 只扩张 + clip≤4 四重防护;mc10 与 mc50 分差仅 0.06。 ## 知识来源 -未使用任何保留阶段/基因型的测量信息;未读 E9.5 之后的数据;未用 `uns.celltype_palette`。只用了 view 内输入阶段的表达与标签、时间差,以及通用做法(PCA/收缩估计为统计常识,无生物学先验注入)。+未使用任何保留阶段/基因型的测量信息;未读禁窗数据;未用 `uns.celltype_palette`。只用 view 内两个输入阶段的表达、标签、时间差。"细胞状态多样性随发育时间增加"是通用发育生物学常识(谱系 progressively 分化),不针对任何禁窗阶段;其余为统计常识(方差比收缩、PCA),无外部数据注入。diff --git a/solution/README.md b/solution/README.mdindex d9b8d84..663e2b0 100644--- a/solution/README.md+++ b/solution/README.md@@ -1,4 +1,4 @@-# lowrank_shape: per-celltype PCA low-rank shrunken mean-shift extrapolation+# lowrank_shape: per-type PCA mean-shift + variance-ratio shape evolution See METHOD.md. Entry: `python run.py --data <view> --out <pred.h5ad> --seed <int>` (pure CPU, `EXECUTION.json` gpu=false).-Env knobs (defaults are the submitted config): `LOWRANK_TAU=0.3`, `LOWRANK_ALPHA=1.8`, `LOWRANK_PER_TYPE=1`, `LOWRANK_FALLBACK=none`, `LOWRANK_NHVG=2500`, `LOWRANK_NPC=25`, `LOWRANK_SCAP=4.0`.+Env knobs (defaults are the submitted config): `LOWRANK_TAU=0.3`, `LOWRANK_ALPHA=1.8`, `LOWRANK_PER_TYPE=1`, `LOWRANK_FALLBACK=none`, `LOWRANK_NHVG=2500`, `LOWRANK_NPC=25`, `LOWRANK_SCAP=4.0`, `LOWRANK_SHAPE_BETA=3.0` (0 = parent node 4 exactly), `LOWRANK_SHAPE_TAU=0.5`, `LOWRANK_SHAPE_CLIP=4.0`, `LOWRANK_SHAPE_RLO=1.0` (expansion-only), `LOWRANK_SHAPE_NPCS=25`, `LOWRANK_SHAPE_MINCELLS=10`.diff --git a/solution/run.py b/solution/run.pyindex 28c3a06..295ed69 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,12 +1,23 @@ #!/usr/bin/env python3-"""PCA low-rank shrunken mean-shift extrapolation over copy_last.--Estimate the temporal delta between the last two input stages in a low-rank-PCA space (top HVGs), shrink each PC direction by lambda_k = var_k / (var_k + tau)-so low-variance noise directions are suppressed, scale by the relative time gap-(target - last) / (last - prev), and add the resulting gene-space shift to the-sampled last-stage cells (clipped at >= 0). Single-input views fall back to-copy_last exactly.+"""Per-type PCA mean-shift + variance-ratio shape evolution over copy_last.++Mean part (as parent node 4): estimate the per-cell-type temporal delta between+the last two input stages in a low-rank PCA space (top HVGs), shrink each PC+direction by lambda_k = var_k / (var_k + tau), scale by the relative time gap,+and add the gene-space shift to sampled last-stage cells (nonzeros only, clip+>= 0).++Shape part (family lowrank_shape mechanism): per matched cell type, compute+each PC's within-type variance in both stages, form the ratio+r_k = var_last_k / var_prev_k, shrink it toward 1 with weight+var_last_k / (var_last_k + tau_shape), extrapolate by the same time-ratio+factor s, and rescale each cell's PC-space residual from its type centroid by+sqrt(r_target_k) (top shape_npcs PCs only). This evolves within-type spread /+gene-gene covariation, giving cells of the same type different effective+displacements (unlike a constant shift).++LOWRANK_SHAPE_BETA=0 turns the shape part off and reproduces the parent+exactly. Single-input views fall back to copy_last exactly. """ from __future__ import annotations@@ -56,6 +67,12 @@ def main() -> None: n_pcs = _env_int("LOWRANK_NPC", 25) per_type = _env_int("LOWRANK_PER_TYPE", 1) s_cap = _env_float("LOWRANK_SCAP", 4.0)+ shape_beta = _env_float("LOWRANK_SHAPE_BETA", 3.0)+ shape_tau = _env_float("LOWRANK_SHAPE_TAU", 0.5)+ shape_clip = _env_float("LOWRANK_SHAPE_CLIP", 4.0)+ shape_rlo = _env_float("LOWRANK_SHAPE_RLO", 1.0)+ shape_npcs = _env_int("LOWRANK_SHAPE_NPCS", 25)+ shape_min_cells = _env_int("LOWRANK_SHAPE_MINCELLS", 10) manifest = load_manifest(args.data) genes = panel_genes(args.data, manifest)@@ -108,45 +125,81 @@ def main() -> None: n1 = A_prev.shape[0] Z = Cc @ V # n_tot x k scores- delta_global = Z[n1:].mean(axis=0) - Z[:n1].mean(axis=0)+ Z_prev = Z[:n1]+ Z_last = Z[n1:]+ delta_global = Z_last.mean(axis=0) - Z_prev.mean(axis=0) shift_by_type: dict[str, np.ndarray] | None = None+ shape_scale_by_type: dict[str, np.ndarray] | None = None delta = delta_global+ lab_prev = labels_of(prev)+ lab_last = labels_of(last) if per_type:- lab_prev = labels_of(prev)- lab_last = labels_of(last) shift_by_type = {}+ if shape_beta > 0.0:+ shape_scale_by_type = {}+ debug_rows = [] for t in sorted(set(lab_prev.tolist()) & set(lab_last.tolist())): m1 = lab_prev == t m2 = lab_last == t if m1.sum() < 5 or m2.sum() < 5: continue- dt_pc = Z[n1:][m2].mean(axis=0) - Z[:n1][m1].mean(axis=0)- shift_by_type[t] = ((s * lam * dt_pc) @ V.T).astype(np.float32)+ dt_pc = Z_last[m2].mean(axis=0) - Z_prev[m1].mean(axis=0)+ shift_by_type[t] = (s * lam * dt_pc) @ Vt # hvg space, float64+ if shape_scale_by_type is not None and m1.sum() >= shape_min_cells and m2.sum() >= shape_min_cells:+ vp = Z_prev[m1].var(axis=0)+ vl = Z_last[m2].var(axis=0)+ eps = 1e-8+ r = (vl + eps) / (vp + eps)+ w = vl / (vl + shape_tau)+ r_s = 1.0 + (r - 1.0) * w+ r_t = np.clip(1.0 + s * (r_s - 1.0), shape_rlo, shape_clip)+ scale = np.sqrt(r_t) - 1.0+ if shape_npcs < scale.size:+ scale = scale.copy()+ scale[shape_npcs:] = 0.0+ shape_scale_by_type[t] = shape_beta * scale+ debug_rows.append((t, int(m1.sum()), int(m2.sum()), float(np.mean(np.abs(r - 1.0))),+ float(np.sqrt(r_t).min()), float(np.sqrt(r_t).max()))) fallback = os.environ.get("LOWRANK_FALLBACK", "none")- shift_h = ((s * lam * delta) @ V.T).astype(np.float32) # global shift on hvg+ shift_h = (s * lam * delta) @ Vt # global shift on hvg, float64 cols = hvg Xc = sp.csr_matrix(X_last, dtype=np.float32, copy=True)- lab_last_rows = labels_of(last)[rows]+ lab_last_rows = lab_last[rows] if shift_by_type is not None:- base = shift_h if fallback == "global" else np.zeros_like(shift_h)+ base = shift_h.astype(np.float32) if fallback == "global" else np.zeros(hvg.size, dtype=np.float32) shift_full = np.zeros(Xc.shape[1], dtype=np.float32) shift_full[cols] = base Xc.data += shift_full[Xc.indices]- for t, sh_t in shift_by_type.items():+ for t, sh_t64 in shift_by_type.items(): sel = np.flatnonzero(lab_last_rows == t) if sel.size == 0: continue- sf_t = np.zeros(Xc.shape[1], dtype=np.float32)- sf_t[cols] = sh_t - base # net correction vs base- sub = Xc[sel]- sub.data += sf_t[sub.indices]- Xc[sel] = sub+ sh_t = sh_t64.astype(np.float32)+ if shape_scale_by_type is not None and t in shape_scale_by_type:+ scale_t = shape_scale_by_type[t]+ resid = Z_last[rows[sel]] - Z_last[lab_last == t].mean(axis=0) # n_t x k+ d_gene = (resid * scale_t[None, :]) @ Vt # n_t x hvg+ D_t = (sh_t[None, :] - base[None, :]) + d_gene.astype(np.float32)+ sub = Xc[sel]+ row_ids = np.repeat(np.arange(sub.shape[0]), np.diff(sub.indptr))+ pos = np.searchsorted(cols, sub.indices)+ pos_c = np.clip(pos, 0, cols.size - 1)+ valid = cols[pos_c] == sub.indices+ dvals = D_t[row_ids, pos_c]+ sub.data += np.where(valid, dvals, 0.0).astype(np.float32)+ Xc[sel] = sub+ else:+ sf_t = np.zeros(Xc.shape[1], dtype=np.float32)+ sf_t[cols] = sh_t - base+ sub = Xc[sel]+ sub.data += sf_t[sub.indices]+ Xc[sel] = sub else: shift_full = np.zeros(Xc.shape[1], dtype=np.float32)- shift_full[cols] = shift_h+ shift_full[cols] = shift_h.astype(np.float32) Xc.data += shift_full[Xc.indices] np.clip(Xc.data, 0.0, None, out=Xc.data)@@ -159,6 +212,12 @@ def main() -> None: n_moved = int((np.abs(shift_h) > 0.01).sum()) print(f"s={s:.3f} k={k} genes_moved={n_moved} shift_absmax={np.abs(shift_h).max():.3f} " f"delta_pc_absmax={np.abs(delta).max():.3f} lam[:5]={np.round(lam[:5], 3)}")+ if shape_scale_by_type:+ print(f"shape_beta={shape_beta} tau_shape={shape_tau} npcs={shape_npcs} "+ f"types_shaped={len(shape_scale_by_type)}")+ for t, n1t, n2t, mrl, smin, smax in debug_rows:+ print(f" type={t} n_prev={n1t} n_last={n2t} mean|r-1|={mrl:.3f} "+ f"sqrt(r_t)=[{smin:.3f},{smax:.3f}]") if __name__ == "__main__":
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
用到的知识库条目
| 编号 | 标题 | 出处 |
|---|---|---|
| k041 | Within-stage pseudotime and graph toolkit offline: scanpy DPT/PAGA/Leiden, Palantir, CellRank 2 | 10.1186/s13059-019-1663-x (PAGA); 10.1038/s41587-019-0068-4 (Palantir); 10.1038/s41592-024-02303-9 (CellRank 2) |
| k018 | Damped per-type shift: shrinkage alpha on the observed delta | notes/plan/cards/T1.md |
| k031 | Offline OT toolkit in the sandbox: moscot TemporalProblem, wot OTModel, POT, geomloss | 10.1038/s41586-024-08453-2 (moscot); 10.1016/j.cell.2019.01.006 (Waddington-OT) |
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
| 改了什么 | 在父节点分型 PCA 均值外推(α=1.8, τ=0.3, 2500 HVG, 25 PC)之上叠加 lowrank_shape 形状演化:按配对类型估计各 PC 两阶段方差比 r_k,τ_shape=0.5 收缩、按时间比 s 外推并 clip 到 [1.0, 4.0](只允许扩张,RLO=1),逐细胞将偏离型心的 PC 残差缩放 (sqrt(r_t)-1)·β,β=3.0,作用于全部 25 PC;SHAPE_BETA=0 时逐元素复现父节点(maxdiff=0.0)。耗时 2.7→10.5s,内存不变。 |
|---|---|
| 各组分数的变化 | cell_state:变好,67.95→78.12 (+10.17),远超噪声,是本节点榜分增益的唯一来源 covariation:变坏,47.97→45.62 (-2.35),略超 T1 约 2 分噪声;与 PLAN 预期相反(PLAN 目标是 covariation 改善 ≥2) de_recovery:噪声内,54.37→53.85 (-0.52) direction:噪声内,50.37→50.14 (-0.23) |
| family_id | lowrank_shape |
| 假设是否成立 | 否 |
| 经验 |
|
| mechanism_active | 是 |
| 下一步建议 |
|
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。
分析员
| 角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。 | 分析员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数?大模型一共回复了几轮(每轮可以调用多个工具)。 | 3 |
| 工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。 | 共 2 次:bash 1、write 1 |
| 用时 | 不到 1 分 |
| token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。 | 输入 13,041 · 输出 1,178 · 思考 368 |
| 任务(第一行) | 分析节点 n5 的打分结果,写 ANALYSIS.json |
| 最后的回答(摘录) | ANALYSIS.json written |
| 原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。 | /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/analyst.jsonl 7 KB /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/analyst.stderr |
工程师
| 角色 | 工程师 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 31 |
| 工具调用 | 共 33 次:bash 25、write 3、read 2、edit 2、todowrite 1 |
| 用时 | 18 分 |
| token 数 | 输入 65,655 · 输出 13,727 · 思考 12,916 |
| 任务(第一行) | 按计划修改程序:Per-type variance-ratio shape evolution on top of mean shift |
| 最后的回答(摘录) | 完成。 **做了什么**:在父节点(分型 PCA 均值外推)上实现了 PLAN 的 lowrank_shape 机制——按细胞型估计各 PC 的两阶段方差比,τ_shape=0.5 收缩后乘时间比外推,逐细胞缩放其偏离型心的 PC 残差。关键发现:双向缩放无效(55.0),**只允许扩张**(RLO=1.0)+ 放大系数 β=3 显著改善 cell_state(66.1→75.5,mmd_u 0.0218→0.0176),β∈[3,4] 为平台,β≥8 崩塌。 **查分**(X3 A 半,用 15/20):β=3 seed0 **57.33**、seed1 **57.10**(父 54.90/55.44,+1.8 均值)。对照 SHAPE_BETA=0 与父输出逐元素 diff=0.0。机制证据:同型细胞位移 std 63–115(非常数),5 个被形状化类型型内方差 +3~13%。确定性、伪装视图一致(diff 0.0)、vec-check ok、运行 13 s。 **下一步最值得试**:① covariation 是唯一受损组(47.7→45.2),可试乘性解码(x·exp(δ))或只对扩张 PC 施加、或对 covariation 做各向同性方差校准;② direction 组停滞(~50),可对 Unknown/不可配对类型用 OT 耦合位移;③ β 平台在 final 视图(s=1.8 更温和)上可能位置不同,值得在 proxy/final 类视图验证幅度传递性。 |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/engineer.jsonl 145 KB /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/engineer.stderr |
调研员
| 角色 | 调研员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 1 |
| 工具调用 | 共 0 次:— |
| 用时 | 2 分 |
| token 数 | 输入 13,494 · 输出 1,722 · 思考 4,019 |
| 任务(第一行) | 为节点 n5(改进,父节点 n4)写改进计划 PLAN.json |
| 最后的回答(摘录) | mily_id": "lowrank_shape", "mechanism": "Per-type covariance evolution: estimate how each PC direction's variance changes between the two input stages within each cell type, shrink the variance ratio toward 1, extrapolate by the time-ratio factor s, and rescale each cell's PC-space residual from its type centroid by sqrt(r_target_k). This changes the population shape (spread and covariation) rather than only the mean.", "vs_constant_shift": "Constant per-type shift adds the same vector to every cell of a type, moving the mean but preserving within-type variance and all pairwise covariances unchanged. This method additionally rescales each cell's deviation from the type centroid by the extrapolated variance change per PC direction, so the within-type spread and gene-gene covariation structure evolve. Two cells of the same type at different positions in PC space receive different effective displacements, unlike a constant shift.", "mechanism_evidence": "1) Report per-type mean |r_k - 1| and the distribution of sqrt(r_target_k) values to show variance change is nonzero and nontrivial.\n2) Compare covariation score between SHAPE_BETA=0 (mean shift only, should match parent) and SHAPE_BETA=1 (mean + shape): covariation must improve ≥2 points to claim the shape component works.\n3) Report within-type variance of predicted expression for top 5 types vs last-stage input: shape evolution should change variance by >5% in at least 3 types.\n4) Report per-cell displacement dispersion within each type (std of |delta| across cells of same type): constant shift gives 0, shape evolution gives >0.\n5) Four-group breakdown for each configuration.", "mechanism_off_control": "LOWRANK_SHAPE_BETA=0 → all r_target_k set to 1.0 → sqrt(1)=1 → residuals unchanged → only mean shift remains → output identical to parent node 4 (verify max abs diff < 1e-6). Additionally LOWRANK_TAU=1e6 → all λ_k≈0 → no shift at all → copy_last. Run both controls and report scores.", "sources": [] } ``` |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/researcher.jsonl 7 KB /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/researcher.stderr |
审查员
| 角色 | 审查员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 5 |
| 工具调用 | 共 10 次:read 8、bash 1、write 1 |
| 用时 | 1 分 |
| token 数 | 输入 20,648 · 输出 1,355 · 思考 1,907 |
| 任务(第一行) | 审查节点 n5 的程序是否越界读取、写死目标、钻评分器漏洞 |
| 最后的回答(摘录) | REVIEW.json written |
| 原始记录位置 | /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/reviewer.jsonl 96 KB /home/spark-longxinyang/vec/runs/formal/20261002-202907-search-t1-scr-A/nodes/5/reviewer.stderr |