总览 · ← 返回运行 20261001-024931-search-t1-spark-eng-default
节点 n4
父节点重加权锚(4000 真实细胞)保持不变,新增可按 β 开关的族内单快照成熟方向平移;β=0(默认)逐位复现父节点。
| 运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。 | 20261001-024931-search-t1-spark-eng-default |
|---|---|
| 父节点 | n2 |
| 子节点 | — |
| 操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。 | 改进 |
| 状态 | 已打分 |
| 分数 | 搜索目标分 55.97(+0.0) · proxy 55.97(+0.0) · 3 次复测均分 56.06 |
| 审查 | 通过 1 越界读取:未发现问题——数据访问全部经 view_io(run.py:132-137)且 dpt.py:53 只读视图内 prior/go/gene_sets_bp.gmt(view_manifest 的 prior 清单含 go),无绝对路径/..、无网络下载。; 2 硬编码目标统计量:未发现问题——run.py 仅含族规则权重 heart 1.6/edge 0.25(run.py:127-128,属允许的类型族规则),DROP_TYPES/type_weights 从任务脚手架 src.task1_temporal.reweight 导入(run.py:36-40),deltas 全… |
| 用时?从运行开始到结束(或到现在)的挂钟时间。 | 26 分 |
| 程序版本 | 4da4a5ed26f17fbe79d4bc8dbcb2b1aedeadbcc1 (programs.git) |
方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。
来自 programs.git 4da4a5ed26:solution/METHOD.md
父节点重加权锚(4000 真实细胞)保持不变,新增可按 β 开关的族内单快照成熟方向平移;β=0(默认)逐位复现父节点。
方法
组成(父节点已验证部件,未动)
- 复制
reweight.py的规则:丢Neural Tube,心脏侧类型 ×1.6,Surface Ectoderm / EXEM / Paraxial Mesoderm×0.25,其余 ×1,largest_remainder分配 4000 个细胞。 run.py:resample_with_labels用与heart_reweight完全相同的类型遍历顺序和rng.choice调用序列,因此 β=0 时输出与父节点逐位相同(已用np.array_equal核对:True)。- 抽样保真检查(PLAN 步骤 1):E8.5 每个类型的分配数都小于该类型细胞数(最大 Foregut 518/2405、pSHF 756/2192),
replace=pool.size<n全为 False,重复行 = 0,因此不需要换成加权不放回抽样,省下一次查分。
表达(新增,默认关闭)
solution/dpt.py:对最新输入阶段的每个细胞类型(≥200 个细胞才做)
- 族内 log1p(CP10k) 面板,子采样 ≤3000 细胞;只保留族内平均表达 ≥0.05 的基因;
- 族内中心化并投掉「每细胞 log 总读数」方向(
_detotal,去技术深度轴); randomized_svd(n_components=5, random_state=0)取前 5 个 PC;- 增殖打分 p = 细胞周期基因集在族内的 z 分数均值。基因集来自视图
prior/go/gene_sets_bp.gmt里 GO:0044770 / GO:0044772 / GO:0051301 / GO:0140014(面板命中 470 个基因);prior 缺失或命中 <10 时退回内置通用细胞周期列表,不会崩; - 取 |corr(score_k, p)| 最大的 PC,要求 ≥0.20,否则该族 δ=0;方向指向增殖下降的一侧(时间前进);
- δ = mean(x | 顶三分位) − mean(x | 底三分位)(对数空间);
- 稳定性收缩:族内随机对半划分 4 次,每半重算 (5)(6),逐基因统计两半符号一致次数 → w∈{0,.25,.5,.75,1},δ̃=w·δ,|δ̃| 按族内 99 分位截断;
- 生成:x'=clip(x+β·δ̃,0),可选
--renorm把每细胞重归一到输入总读数中位数(数据是 CP10k,线性总量恒为 1e4,所以 renorm on/off 几乎等价)。
两个及以上输入阶段(final)时,额外算同名族的伪批量差 pseudobulk_delta(同样的表达阈值 + 4 次细胞 bootstrap 符号收缩),合成 δ=β·δ̃_dpt+α·Δ̃_pb,α=0.35 固定,且有符号门:cos(δ_dpt,Δ_pb)<0 的族丢掉 DPT 项只保留 Δ_pb。族归属沿用父节点的精确名集合 + type_weights;该分支在 proxy 上不执行。
验证过什么
- 护栏:
--beta 0输出与父节点逐位相同(已核对),vec-checkok,1.5 s / <1.3 GB。 - 方向合理性(不查分,先看数):8/18 个族 |corr|≥0.20 拿到 δ≠0(AVC-CM .42、Foregut .38、IFT-CM .41、LV-CM .46、OFT/RV-CM .51、RV-CM .41、SV-CM .42、aSHF .21);其余(Blood/EXEM/Endothelium/JCF/NCC/Neural Tube/Paraxial/Pericardium/Surface Ectoderm/pSHF)|corr|<0.20 → δ=0,退回真实细胞。δ 的 top 基因与「增殖下降 + 心肌成熟」一致:IFT-CM 晚期侧升 Nkx2-5、Tnnt2、Tnni1、Myl4、Des、Actn2、Nexn、Crip2;OFT/RV-CM 晚期侧降 Mki67、Top2a、Cenpe、Cks1b、升 Cdkn1c。没有发现技术轴主导的迹象。
- 替代评测(seed 0,1 次/配置):
| 配置 | 榜分 | cell_state | covariation | de_recovery | direction |
|---|---|---|---|---|---|
| 父节点(β=0,复现) | 55.9694 | 56.90 | 54.97 | 53.06 | 58.56 |
| β=0.15 + renorm | 55.9437 | 57.41 | 53.63 | 53.06 | 58.92 |
| β=0.15 − renorm | 55.9335 | 57.43 | 53.57 | 53.06 | 58.90 |
| β=0.30 + renorm | 55.8012 | 57.84 | 52.19 | 53.06 | 58.98 |
趋势单调且清楚:平移让 mmd_u 变小(cell_state +0.5/+0.9)、de_direction 微升(direction +0.36/+0.42),但 variogram 变大(covariation −1.3/−2.8),de_score 一位都没动(0.1111 三次完全相同)。净效果 ≤0,且 β 越大越差。
- 因此默认 β=0,把这套机制作为可开关的代码保留。按 PLAN 步骤 5 的接受判据(4-seed 均值 +1.5 且 ≥3/4 同向),β>0 在 seed 0 上就没有正收益,不值得再花 12 次查分做 4-seed 复核(时间预算也不允许)。
没验证 / 零结果
- 单快照伪时间平移在 T1 proxy 上对 de_recovery 完全无效(de_score 三位小数不变),对 covariation 有害。原因推测:δ 是族内第三分位对比,量级 |δ̃| 中位远小于真实 E8.5→E9.5 的族间 DE,且它把细胞推离真实流形(variogram 变差正是共变结构被破坏的信号)。
- 没验证多 seed(时间不够);没验证 β<0.15;没验证 final 的 Δ_pb 分支(proxy 只有一个输入阶段,该分支不执行,只做了「不崩」级别的代码检查)。
- 没验证只平移 δ≠0 的 8 个族之外的其它组合、也没试「只对增殖最强的族平移」。
下一步最值得试
- covariation 是父节点第二弱的组,而且它对任何表达扰动都很敏感(β=0.15 就掉 1.34)。先做不改变每个细胞流形位置的表达改动:例如按族做「细胞间的重排/混合」(把同族细胞的基因表达按 rank 做轻度混合)而不是加常数向量。
- de_recovery 对常数位移完全不响应,说明它比较的是族间 DE 结构;试试族特异的少量高置信基因(只在 δ 的 top-50 基因、且 w=1 的基因上平移,β 放大到 0.5–1),其余基因不动,避免全基因噪声推高 variogram。
- 组成侧还有空间:父节点的心脏 ×1.6 是在 2 分噪声上选的,但
cell_state56.9 明显受 mmd_u 驱动;用 E8.5 自身的 marker 结构(视图内数据)做一次「组织占比 → 期望细胞数」的重新标定,比手调 1.6/0.25 更可能迁移到 final。
运行
python solution/run.py --data <view> --out pred.h5ad --seed 0 # β=0,= 父节点,1.5 s
python solution/run.py --data <view> --out pred.h5ad --seed 0 --beta 0.15 --verbose # 28 s,打印每族 |corr| 与 top 基因调研员的计划
| 名称 | 重加权锚 + 稳定性收缩的单快照伪时间平移 |
|---|---|
| 动机 | 父节点 2(55.97)四组:cell_state 56.90、covariation 54.97、de_recovery 53.06、direction 58.56,最弱是 de_recovery 与 covariation。全树里节点 1 copy_last 与节点 3 pseudobulk_shift 分数与四组逐位相同(49.77 / 48.07 / 51.54 / 50.00 / 50.16),证明 proxy 只有一个输入阶段、任何两阶段差值位移都退化成复制:至今没有一次真正改变表达的位移被评测过。父节点相对 copy_last 的 +6.20 全部来自组成重加权(cell_state +8.83、direction +8.40、covariation +3.43),表达层一点没动,de_recovery 只涨 +3.06。所以下一个增量必须来自表达层,而且只能用单阶段可算的信号(T1-06 快照内伪时间方向)才能在 proxy 上被真正测到、在 final 上又不退化。covariation 54.97 低于 cell_state 56.90,另一个可疑点是父节点用加权重抽样得到 4000 个细胞:若为有放回抽样,×1.6 的心脏族会产生克隆重复行,直接压低群体多样性与共变结构,这是可零风险修掉的保真问题。 |
| 做法 | 总体:保留父节点全部已验证部件(family map、重加权、N=4000、write_prediction),只加两件事——(A) 抽样保真修复,(B) 每族一个由数据导出、经稳定性收缩的小步长成熟方向平移。β=0 必须逐位复现父节点。 步骤 0 护栏(1 次查分):新 run.py 加开关 --beta(默认 0)、--renorm(默认 on)、--weighted-without-replacement。先以 --beta 0 查分,必须得到 55.97;不符则先修回归,不要往下走。 步骤 1 抽样保真(1 次查分):统计父配置输出中的重复行数。若重复率 > 2%,把抽样换成加权不放回(Efraimidis–Spirakis:key_i = U_i^(1/w_i),取 key 最大的 n 个),族权重与随机种子不变。判据是重复率本身(保真修复),不是单次分数;只要不降分就保留。 步骤 2 单快照成熟方向(主机制,T1-06):对最新输入阶段的每个细胞族(族内 ≥200 个细胞才做,否则该族 δ=0): a. 取 log1p(CP10k) 面板,族内中心化,并投掉 log 总读数方向(去技术轴); b. PCA 取前 5 个主成分(族内子采样 ≤3000 个细胞即可,方向是均值差,子采样不影响结论); c. 增殖打分 p = 细胞周期基因集在族内的 z 分数均值;基因集优先取 data/external/prior 的 Reactome / GO BP 中 cell-cycle 相关集(10–500 基因),prior 缺失时用一份通用细胞周期基因列表兜底,任何情况下不得因缺 prior 而崩; d. 在与 p 的相关绝对值最大的主成分 k* 上定向:要求 |corr(score_k*, p)| ≥ 0.20,否则该族 δ=0(没有可信方向就退回真实细胞);方向取 corr 为负的一侧(时间前进 = 增殖下降); e. δ_c = mean(x | score_k* 顶三分位) − mean(x | 底三分位),在对数空间计算; f. 稳定性收缩:族内随机对半划分 R=4 次,每半各自重算 (d)(e),逐基因统计两半符号一致次数 → w_g ∈ {0,0.25,0.5,0.75,1},δ̃_g = w_g·δ_g;只保留族内平均表达 ≥ 0.05 的基因,|δ̃| 按族内 99 分位截断(防止个别噪声基因主导 de_recovery)。 步骤 3 生成:x'_i = clip(x_i + β·δ̃_{c(i)}, 0),再按 --renorm 把每个细胞重归一到输入阶段的总读数中位数并 log1p。这样每个输出细胞仍锚在一个真实细胞上、族内协方差形状不被破坏、不输出平均细胞、也不做整体缩放(符合 T1-12 原则,且不动坐标尺度… |
| 风险 | 1) 定向错误是主风险:PC 轴可能是技术轴或方向搞反,β>0 会同时拉低 direction 与 de_recovery。Engineer 应最早发现的方式是先不查分,只在 3–5 个大族上打印 |corr(score,p)| 与 δ_c 的 top-20 正/负基因,人工确认是细胞周期下降/分化轴;|corr|<0.20 的族必须 δ=0。2) 增益可能小于 2 分噪声:必须 4 seeds 均值 + ≥3/4 同向才判定有效,禁止凭 seed 0 单次分数选 β(k014:单组 cell_state 变化 <6 分视为噪声)。3) --renorm 与 clip(≥0) 叠加可能把分布压窄、伤 covariation:监控每族方差与输出重复率,若 covariation 掉 >1 就试 --renorm off 再查 1 次。4) prior 基因集在 proxy 视图里可能不存在或覆盖率低(T1 面板 Reactome 仅 27%):必须有通用细胞周期基因兜底路径,且族内基因不足时 δ=0 而不是崩。5) 查分预算超支:严格按步骤 4 的顺序执行,粗筛只用 seed 0。6) 迁移失效:final 的标签名集合与 proxy 不同,若 family map 用精确名匹配,新标签会漏成『其余 ×1』,权重与位移同时失效;提交前用一个含未知标签的小样本做单元检查(只验证归族逻辑,不涉及任何保留阶段信息)。7) 明确不做:不调族权重 1.6/0.25(父节点已验证部件,在 2 分噪声的 proxy 上调它就是过拟合);不加随机低秩噪声(covariation 收益不确定、可能伤分布距离);不引入 OT/VAE/flow(30 分钟内不可靠且 proxy 不可识别)。 8) 合规:只用视图内数据与允许的 external/prior 基因集;不读、不猜、不写任何保留阶段的类型名、比例、标记基因或表达;所有位移量运行时从输入阶段现算,不硬编码阶段统计量;不做坐标旋转/缩放等针对打分器的取巧变换。 |
代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。
对比:父节点版本 2be2358fd9。改动的文件:solution/METHOD.md +56 −0、solution/dpt.py +269 −0、solution/run.py +137 −11
diff --git a/solution/METHOD.md b/solution/METHOD.mdnew file mode 100644index 0000000..77f2eac--- /dev/null+++ b/solution/METHOD.md@@ -0,0 +1,56 @@+父节点重加权锚(4000 真实细胞)保持不变,新增可按 β 开关的族内单快照成熟方向平移;β=0(默认)逐位复现父节点。++# 方法++## 组成(父节点已验证部件,未动)+- 复制 `reweight.py` 的规则:丢 `Neural Tube`,心脏侧类型 ×1.6,`Surface Ectoderm / EXEM / Paraxial Mesoderm` ×0.25,其余 ×1,`largest_remainder` 分配 4000 个细胞。+- `run.py:resample_with_labels` 用与 `heart_reweight` 完全相同的类型遍历顺序和 `rng.choice` 调用序列,因此 β=0 时输出与父节点**逐位相同**(已用 `np.array_equal` 核对:True)。+- 抽样保真检查(PLAN 步骤 1):E8.5 每个类型的分配数都小于该类型细胞数(最大 Foregut 518/2405、pSHF 756/2192),`replace=pool.size<n` 全为 False,**重复行 = 0**,因此不需要换成加权不放回抽样,省下一次查分。++## 表达(新增,默认关闭)+`solution/dpt.py`:对最新输入阶段的每个细胞类型(≥200 个细胞才做)+1. 族内 log1p(CP10k) 面板,子采样 ≤3000 细胞;只保留族内平均表达 ≥0.05 的基因;+2. 族内中心化并投掉「每细胞 log 总读数」方向(`_detotal`,去技术深度轴);+3. `randomized_svd(n_components=5, random_state=0)` 取前 5 个 PC;+4. 增殖打分 p = 细胞周期基因集在族内的 z 分数均值。基因集来自视图 `prior/go/gene_sets_bp.gmt` 里 GO:0044770 / GO:0044772 / GO:0051301 / GO:0140014(面板命中 470 个基因);prior 缺失或命中 <10 时退回内置通用细胞周期列表,不会崩;+5. 取 |corr(score_k, p)| 最大的 PC,要求 ≥0.20,否则该族 δ=0;方向指向增殖下降的一侧(时间前进);+6. δ = mean(x | 顶三分位) − mean(x | 底三分位)(对数空间);+7. 稳定性收缩:族内随机对半划分 4 次,每半重算 (5)(6),逐基因统计两半符号一致次数 → w∈{0,.25,.5,.75,1},δ̃=w·δ,|δ̃| 按族内 99 分位截断;+8. 生成:x'=clip(x+β·δ̃,0),可选 `--renorm` 把每细胞重归一到输入总读数中位数(数据是 CP10k,线性总量恒为 1e4,所以 renorm on/off 几乎等价)。++两个及以上输入阶段(final)时,额外算同名族的伪批量差 `pseudobulk_delta`(同样的表达阈值 + 4 次细胞 bootstrap 符号收缩),合成 δ=β·δ̃_dpt+α·Δ̃_pb,α=0.35 固定,且有符号门:cos(δ_dpt,Δ_pb)<0 的族丢掉 DPT 项只保留 Δ_pb。族归属沿用父节点的精确名集合 + `type_weights`;该分支在 proxy 上不执行。++# 验证过什么++- **护栏**:`--beta 0` 输出与父节点逐位相同(已核对),`vec-check` ok,1.5 s / <1.3 GB。+- **方向合理性**(不查分,先看数):8/18 个族 |corr|≥0.20 拿到 δ≠0(AVC-CM .42、Foregut .38、IFT-CM .41、LV-CM .46、OFT/RV-CM .51、RV-CM .41、SV-CM .42、aSHF .21);其余(Blood/EXEM/Endothelium/JCF/NCC/Neural Tube/Paraxial/Pericardium/Surface Ectoderm/pSHF)|corr|<0.20 → δ=0,退回真实细胞。δ 的 top 基因与「增殖下降 + 心肌成熟」一致:IFT-CM 晚期侧升 Nkx2-5、Tnnt2、Tnni1、Myl4、Des、Actn2、Nexn、Crip2;OFT/RV-CM 晚期侧降 Mki67、Top2a、Cenpe、Cks1b、升 Cdkn1c。没有发现技术轴主导的迹象。+- **替代评测(seed 0,1 次/配置)**:++| 配置 | 榜分 | cell_state | covariation | de_recovery | direction |+|---|---|---|---|---|---|+| 父节点(β=0,复现) | 55.9694 | 56.90 | 54.97 | 53.06 | 58.56 |+| β=0.15 + renorm | 55.9437 | 57.41 | 53.63 | 53.06 | 58.92 |+| β=0.15 − renorm | 55.9335 | 57.43 | 53.57 | 53.06 | 58.90 |+| β=0.30 + renorm | 55.8012 | 57.84 | 52.19 | 53.06 | 58.98 |++ 趋势单调且清楚:平移让 `mmd_u` 变小(cell_state +0.5/+0.9)、`de_direction` 微升(direction +0.36/+0.42),但 `variogram` 变大(covariation −1.3/−2.8),`de_score` 一位都没动(0.1111 三次完全相同)。净效果 ≤0,且 β 越大越差。+- 因此**默认 β=0**,把这套机制作为可开关的代码保留。按 PLAN 步骤 5 的接受判据(4-seed 均值 +1.5 且 ≥3/4 同向),β>0 在 seed 0 上就没有正收益,不值得再花 12 次查分做 4-seed 复核(时间预算也不允许)。++# 没验证 / 零结果++- **单快照伪时间平移在 T1 proxy 上对 de_recovery 完全无效**(de_score 三位小数不变),对 covariation 有害。原因推测:δ 是族内第三分位对比,量级 |δ̃| 中位远小于真实 E8.5→E9.5 的族间 DE,且它把细胞推离真实流形(variogram 变差正是共变结构被破坏的信号)。+- 没验证多 seed(时间不够);没验证 β<0.15;没验证 final 的 Δ_pb 分支(proxy 只有一个输入阶段,该分支不执行,只做了「不崩」级别的代码检查)。+- 没验证只平移 δ≠0 的 8 个族之外的其它组合、也没试「只对增殖最强的族平移」。++# 下一步最值得试++1. covariation 是父节点第二弱的组,而且它对任何表达扰动都很敏感(β=0.15 就掉 1.34)。先做**不改变每个细胞流形位置**的表达改动:例如按族做「细胞间的重排/混合」(把同族细胞的基因表达按 rank 做轻度混合)而不是加常数向量。+2. de_recovery 对常数位移完全不响应,说明它比较的是族间 DE 结构;试试**族特异的少量高置信基因**(只在 δ 的 top-50 基因、且 w=1 的基因上平移,β 放大到 0.5–1),其余基因不动,避免全基因噪声推高 variogram。+3. 组成侧还有空间:父节点的心脏 ×1.6 是在 2 分噪声上选的,但 `cell_state` 56.9 明显受 mmd_u 驱动;用 E8.5 自身的 marker 结构(视图内数据)做一次「组织占比 → 期望细胞数」的重新标定,比手调 1.6/0.25 更可能迁移到 final。++# 运行++```bash+python solution/run.py --data <view> --out pred.h5ad --seed 0 # β=0,= 父节点,1.5 s+python solution/run.py --data <view> --out pred.h5ad --seed 0 --beta 0.15 --verbose # 28 s,打印每族 |corr| 与 top 基因+```diff --git a/solution/dpt.py b/solution/dpt.pynew file mode 100644index 0000000..b5db2bd--- /dev/null+++ b/solution/dpt.py@@ -0,0 +1,269 @@+"""Single-snapshot maturation direction (DPT-like) per cell group.++Everything is computed at run time from the input stage(s) given by the+manifest: no stage names, proportions or measured statistics are baked in.++For one cell group we look for the axis inside the group that best tracks a+decline in proliferation (canonical proxy for embryonic time passing within a+snapshot), take the high-minus-low expression contrast along that axis, and+shrink it gene-wise by a split-half sign-agreement weight so noise genes do+not dominate the shift.+"""++from __future__ import annotations++from pathlib import Path++import numpy as np+from scipy import sparse++GO_CC_TERMS = {+ "GO:0044770", # cell cycle phase transition+ "GO:0044772", # mitotic cell cycle phase transition+ "GO:0051301", # cell division+ "GO:0140014", # mitotic nuclear division+}+GO_CC_NAMES = {+ "cell cycle phase transition",+ "mitotic cell cycle phase transition",+ "cell division",+ "mitotic nuclear division",+}+# Fallback only used when the prior gene sets are unavailable/too sparse.+FALLBACK_CC = [+ "Mki67", "Top2a", "Cdk1", "Ccna2", "Ccnb1", "Ccnb2", "Ccne1", "Ccnd1",+ "Pcna", "Mcm2", "Mcm3", "Mcm4", "Mcm5", "Mcm6", "Mcm7", "Plk1", "Aurka",+ "Aurkb", "Birc5", "Bub1", "Bub1b", "Cdc20", "Cdc6", "Cdt1", "E2f1",+ "Foxm1", "Kif11", "Kif23", "Nusap1", "Rrm2", "Tyms", "Ube2c", "Cenpa",+ "Cenpe", "Cenpf", "Chek1", "Gins1", "Orc1", "Rad51", "Aspm",+]+MIN_EXP = 0.05+MIN_CELLS = 200+MAX_SUB = 3000+N_PC = 5+MIN_CORR = 0.20+N_SPLIT = 4+Q_LO, Q_HI = 1.0 / 3.0, 2.0 / 3.0+++def cell_cycle_genes(view, genes: list[str]) -> np.ndarray:+ """Boolean mask over the panel marking proliferation genes (prior first)."""+ mask = np.zeros(len(genes), dtype=bool)+ idx = {g: i for i, g in enumerate(genes)}+ path = Path(view) / "prior" / "go" / "gene_sets_bp.gmt"+ picked: set[str] = set()+ if path.exists():+ try:+ with path.open("r", encoding="utf-8") as fh:+ for ln in fh:+ parts = ln.rstrip("\n").split("\t")+ if len(parts) < 3:+ continue+ if parts[0].strip() in GO_CC_TERMS or parts[1].strip().lower() in GO_CC_NAMES:+ picked.update(g for g in parts[2:] if g)+ except OSError:+ picked = set()+ if len(picked & set(idx)) < 10:+ picked = set(FALLBACK_CC)+ for g in picked:+ j = idx.get(g)+ if j is not None:+ mask[j] = True+ return mask+++def _detotal(Dc: np.ndarray) -> np.ndarray:+ """Remove the per-cell log-total-reads direction (technical depth axis)."""+ tot = Dc.sum(axis=1, dtype=np.float64)+ t = tot - tot.mean()+ nrm = np.linalg.norm(t)+ if nrm < 1e-8:+ return Dc+ u = (t / nrm).astype(np.float32)+ return Dc - np.outer(u, Dc.T @ u).astype(np.float32)+++def _dense(X, rows: np.ndarray) -> np.ndarray:+ if sparse.issparse(X):+ return np.asarray(X[rows].todense(), dtype=np.float32)+ return np.asarray(X)[rows].astype(np.float32)+++class _Axis:+ """Fitted centring + PC directions of one cell group."""++ __slots__ = ("keep", "mu", "u", "Vt", "D", "cc", "active", "prol", "sc", "k", "corr", "lo", "hi", "late_first")++ def __init__(self, D: np.ndarray, cc_mask: np.ndarray, n_pc: int = N_PC):+ self.D = D+ self.cc = cc_mask+ keep = D.std(axis=0) > 0+ self.keep = keep+ Dk = np.ascontiguousarray(D[:, keep], dtype=np.float32)+ self.mu = Dk.mean(axis=0)+ self.u = None+ Dc = _detotal(Dk - self.mu)+ npc = max(1, min(n_pc, Dc.shape[0] - 1, Dc.shape[1] - 1))+ try:+ from sklearn.utils.extmath import randomized_svd++ _, _, Vt = randomized_svd(Dc, n_components=npc, random_state=0)+ self.Vt = Vt.astype(np.float32)+ except Exception:+ v = np.linalg.eigh((Dc.T @ Dc).astype(np.float64))[1]+ self.Vt = np.ascontiguousarray(v[:, np.argsort(-(Dc.T @ Dc).diagonal())][:, :npc].T.astype(np.float32))+ self.active = np.ones(D.shape[1], dtype=bool)+ self.sc = self.project(None)+ self.prol = self._proliferation(None)+ self.k, self.corr, self.lo, self.hi, self.late_first = self._orient(self.sc, self.prol)++ def project(self, rows):+ D = self.D if rows is None else self.D[rows]+ Dk = np.ascontiguousarray(D[:, self.keep], dtype=np.float32) - self.mu+ return (_detotal(Dk) @ self.Vt.T).astype(np.float64)++ def _proliferation(self, rows: np.ndarray) -> np.ndarray:+ ccm = self.cc+ if ccm.sum() < 5:+ return np.zeros(len(rows), dtype=np.float64)+ Z = (self.D[:, ccm] if rows is None else self.D[np.ix_(rows, ccm)]).astype(np.float64)+ sd = Z.std(axis=0)+ sd[sd < 1e-8] = 1.0+ return ((Z - Z.mean(axis=0)) / sd).mean(axis=1)++ @staticmethod+ def _orient(sc: np.ndarray, prol: np.ndarray):+ if np.allclose(prol, prol[0]):+ return 0, 0.0, -np.inf, np.inf, False+ cs = np.zeros(sc.shape[1])+ for j in range(sc.shape[1]):+ s = sc[:, j]+ if s.std() < 1e-12:+ continue+ cs[j] = np.nan_to_num(np.corrcoef(s, prol)[0, 1])+ k = int(np.argmax(np.abs(cs)))+ s = sc[:, k]+ lo, hi = np.quantile(s, [Q_LO, Q_HI])+ # forward in time = proliferation goes down+ late_first = cs[k] < 0+ return k, float(abs(cs[k])), float(lo), float(hi), bool(late_first)++ def contrast(self, rows: np.ndarray):+ """delta over the *active* columns given by ``self.active``, or None."""+ if self.corr < MIN_CORR:+ return None+ n = self.D.shape[0] if rows is None else rows.size+ if n < 30:+ return None+ sc = self.project(rows)[:, self.k]+ prol = self._proliferation(rows)+ _, c, lo, hi, late_first = self._orient(sc[:, None], prol)+ if c < MIN_CORR:+ return None+ late = sc >= hi if late_first else sc <= lo+ early = sc <= lo if late_first else sc >= hi+ if late.sum() < 10 or early.sum() < 10:+ return None+ Da = self.D[:, self.active] if rows is None else self.D[np.ix_(rows, self.active)]+ return (Da[late].mean(axis=0) - Da[early].mean(axis=0)).astype(np.float64)+++def group_delta(+ X,+ labels: np.ndarray,+ genes: list[str],+ cc_mask: np.ndarray,+ seed: int = 0,+ min_cells: int = MIN_CELLS,+ n_split: int = N_SPLIT,+ verbose: bool = False,+) -> dict[str, np.ndarray]:+ """delta per cell group, on the panel index, shrunk by split-half agreement."""+ out: dict[str, np.ndarray] = {}+ rng = np.random.default_rng(seed + 104729)+ all_rows = np.arange(labels.size)+ for g in [str(x) for x in np.unique(labels)]:+ zero = np.zeros(len(genes), dtype=np.float32)+ rows = all_rows[labels == g]+ if rows.size < min_cells:+ out[g] = zero+ continue+ if rows.size > MAX_SUB:+ rows = np.sort(rng.choice(rows, size=MAX_SUB, replace=False))+ D = _dense(X, rows)+ active = D.mean(axis=0) >= MIN_EXP+ cc_local = cc_mask & active+ try:+ ax = _Axis(D, cc_local)+ except Exception as exc: # pragma: no cover - never fail the run+ if verbose:+ print(f"[dpt] {g}: axis fit failed ({exc})")+ out[g] = zero+ continue+ ax.active = active+ delta = ax.contrast(None)+ if delta is None:+ out[g] = zero+ if verbose:+ print(f"[dpt] {g:24s} n={rows.size:5d} |corr|={ax.corr:.3f} -> delta=0")+ continue+ agree = np.zeros(delta.size, dtype=np.float64)+ m = rows.size+ for _ in range(n_split):+ perm = rng.permutation(m)+ h1, h2 = perm[: m // 2], perm[m // 2:]+ d1 = ax.contrast(h1)+ d2 = ax.contrast(h2)+ if d1 is None or d2 is None:+ continue+ agree += (np.sign(d1) == np.sign(d2)) & (np.abs(d1) > 1e-9) & (np.abs(d2) > 1e-9)+ d = (agree / max(1, n_split)) * delta+ cap = float(np.quantile(np.abs(d), 0.99))+ if cap > 0:+ d = np.clip(d, -cap, cap)+ full = np.zeros(len(genes), dtype=np.float32)+ full[active] = d.astype(np.float32)+ out[g] = full+ if verbose:+ top = np.argsort(-np.abs(full))[:12]+ print(f"[dpt] {g:24s} n={rows.size:5d} |corr|={ax.corr:.3f} late_high={ax.late_first} "+ f"nz={(np.abs(d)>0).sum()} cap={cap:.3f} top=" ++ ",".join(f"{genes[i]}{'+' if full[i]>0 else '-'}" for i in top))+ return out+++def pseudobulk_delta(+ X_prev, lab_prev, X_last, lab_last, genes: list[str], seed: int = 0+) -> dict[str, np.ndarray]:+ """Per-group pseudobulk difference between the two latest input stages."""+ rng = np.random.default_rng(seed + 7919)+ out: dict[str, np.ndarray] = {}+ shared = sorted(set(str(x) for x in np.unique(lab_prev)) & set(str(x) for x in np.unique(lab_last)))+ for g in shared:+ zero = np.zeros(len(genes), dtype=np.float32)+ a = np.flatnonzero(lab_prev == g)+ b = np.flatnonzero(lab_last == g)+ if a.size < 50 or b.size < 50:+ out[g] = zero+ continue+ if a.size > MAX_SUB:+ a = np.sort(rng.choice(a, size=MAX_SUB, replace=False))+ if b.size > MAX_SUB:+ b = np.sort(rng.choice(b, size=MAX_SUB, replace=False))+ Da, Db = _dense(X_prev, a), _dense(X_last, b)+ ma, mb = Da.mean(axis=0).astype(np.float64), Db.mean(axis=0).astype(np.float64)+ active = (ma >= MIN_EXP) | (mb >= MIN_EXP)+ d = np.zeros(len(genes), dtype=np.float64)+ d[active] = (mb - ma)[active]+ boot = np.zeros(len(genes), dtype=np.float64)+ for _ in range(N_SPLIT):+ ia = rng.integers(0, a.size, a.size)+ ib = rng.integers(0, b.size, b.size)+ bb = Db[ib].mean(axis=0).astype(np.float64) - Da[ia].mean(axis=0).astype(np.float64)+ boot += (np.sign(bb) == np.sign(d)) & (np.abs(bb) > 1e-9)+ d = (boot / N_SPLIT) * d+ cap = float(np.quantile(np.abs(d), 0.99))+ if cap > 0:+ d = np.clip(d, -cap, cap)+ out[g] = d.astype(np.float32)+ return outdiff --git a/solution/run.py b/solution/run.pyindex a752b21..b059965 100644--- a/solution/run.py+++ b/solution/run.py@@ -1,18 +1,44 @@ #!/usr/bin/env python3-"""heart_jcf_peri: run2's T1 winner. Reweight the latest input stage by cell type.+"""Reweighted anchors + a stability-shrunk single-snapshot maturation shift. -Neural Tube dropped; heart types (CM, endothelium/endocardium, SHF/PHM, JCF,-pericardium, proepicardium) x1.6; surface ectoderm, EXEM, paraxial mesoderm-x0.25; everything else x1. 4000 real cells, no expression shift. Proxy reads-E8.5, final reads E9.5; the type set already names both stages' labels.+Composition side (verified in the parent, untouched): drop the groups absent+from the later dissection, scale heart-side groups up and the peripheral+edge groups down, then draw ``N_CELLS`` real cells from the latest input+stage without replacement.++Expression side (new): every output cell stays anchored on a real cell, and+its group's maturation contrast is added in log space with a small step+``beta``. The contrast is estimated inside the latest snapshot from the axis+that best tracks declining proliferation, oriented towards "later", and shrunk+gene-wise by split-half sign agreement. With two or more input stages the+group pseudobulk difference between the two latest stages is added as well+(step ``alpha``) behind a sign gate, so nothing degenerates to a plain copy.++``--beta 0`` reproduces the parent prediction bit for bit. """ from __future__ import annotations import argparse+import os+import sys++for _v in ("OMP_NUM_THREADS", "OPENBLAS_NUM_THREADS", "MKL_NUM_THREADS", "NUMEXPR_NUM_THREADS"):+ os.environ.setdefault(_v, "4") # bounded, reproducible BLAS++from pathlib import Path++import numpy as np+from scipy import sparse -from src.task1_temporal.reweight import heart_reweight-from src.task1_temporal.view_io import (+sys.path.insert(0, str(Path(__file__).resolve().parent))++from src.task1_temporal.reweight import ( # noqa: E402+ DROP_TYPES,+ largest_remainder,+ type_weights,+)+from src.task1_temporal.view_io import ( # noqa: E402 inputs_by_time, labels_of, load_manifest,@@ -22,7 +48,71 @@ from src.task1_temporal.view_io import ( write_prediction, ) +import dpt # noqa: E402+ N_CELLS = 4000+ALPHA = 0.35 # conservative: a full-strength constant shift scores below copy_last on T1+++def resample_with_labels(X, labels, n_cells, heart_weight, edge_weight, seed):+ """Same allocation / RNG sequence as ``heart_reweight``, plus output labels."""+ rng = np.random.default_rng(seed)+ types = [str(t) for t in np.unique(labels) if str(t) not in DROP_TYPES]+ counts = np.array([(labels == t).sum() for t in types], dtype=np.float64)+ w = type_weights(types, heart_weight, edge_weight)+ alloc = largest_remainder(counts * w, n_cells)+ blocks, out_lab, dup = [], [], 0+ for t, n in zip(types, alloc):+ if n <= 0:+ continue+ pool = np.flatnonzero(labels == t)+ replace = pool.size < n+ choice = rng.choice(pool, size=int(n), replace=replace)+ if replace:+ dup += int(n) - int(np.unique(choice).size)+ blocks.append(X[choice])+ out_lab.append(np.full(int(n), t, dtype=object))+ out = sparse.vstack(blocks, format="csr").astype(np.float32)+ out.eliminate_zeros()+ return out, np.concatenate(out_lab) if out_lab else np.array([], dtype=object), dup+++def apply_shift(X, out_lab, deltas, beta, renorm, alpha_deltas=None, alpha=ALPHA):+ """x' = clip(x + beta*delta, 0), then per-cell renormalisation in linear space."""+ groups = {}+ for g, d in deltas.items():+ groups[g] = d+ step = {}+ for g, d in groups.items():+ v = beta * d.astype(np.float32)+ if alpha_deltas is not None and g in alpha_deltas and alpha != 0.0:+ ap = alpha * alpha_deltas[g].astype(np.float32)+ both = np.abs(d) > 0+ cos = float(np.dot(d[both], alpha_deltas[g][both])) if both.any() else 0.0+ if cos < 0: # snapshot axis disagrees with the observed trend: drop it+ v = ap+ else:+ v = v + ap+ step[g] = v+ if not step or all(not np.any(v) for v in step.values()):+ return X+ Xd = np.asarray(X.todense(), dtype=np.float32)+ tot_before = np.expm1(np.clip(Xd, 0, None)).sum(axis=1)+ uniq = [g for g in np.unique(out_lab) if g in step]+ for g in uniq:+ m = out_lab == g+ if not m.any() or not np.any(step[g]):+ continue+ Xd[m] = np.clip(Xd[m] + step[g][None, :], 0.0, None)+ if renorm:+ lin = np.expm1(Xd)+ tot = lin.sum(axis=1)+ scale = np.ones_like(tot)+ ok = tot > 1e-6+ target = np.median(tot_before[tot_before > 1e-6]) if (tot_before > 1e-6).any() else 1.0+ scale[ok] = target / tot[ok]+ Xd = np.log1p(lin * scale[:, None]).astype(np.float32)+ return sparse.csr_matrix(Xd, dtype=np.float32) def main() -> None:@@ -30,14 +120,50 @@ def main() -> None: parser.add_argument("--data", required=True) parser.add_argument("--out", required=True) parser.add_argument("--seed", type=int, default=0)+ parser.add_argument("--beta", type=float, default=0.0)+ parser.add_argument("--alpha", type=float, default=ALPHA)+ parser.add_argument("--renorm", type=int, default=1)+ parser.add_argument("--n-cells", type=int, default=N_CELLS)+ parser.add_argument("--heart-weight", type=float, default=1.6)+ parser.add_argument("--edge-weight", type=float, default=0.25)+ parser.add_argument("--verbose", action="store_true") args = parser.parse_args() manifest = load_manifest(args.data) genes = panel_genes(args.data, manifest)- last = read_stage(args.data, inputs_by_time(manifest)[-1], genes)- n = target_n_cells(manifest, N_CELLS)- X = heart_reweight(last.X, labels_of(last), n_cells=n, seed=args.seed)- write_prediction(X, genes, args.out, seed=args.seed)+ entries = inputs_by_time(manifest)+ last = read_stage(args.data, entries[-1], genes)+ labels = labels_of(last)+ n = target_n_cells(manifest, args.n_cells)++ X, out_lab, dup = resample_with_labels(+ last.X, labels, n, args.heart_weight, args.edge_weight, args.seed+ )+ if args.verbose:+ print(f"[sample] {X.shape} duplicated_rows={dup}")++ if args.beta == 0.0 and (len(entries) < 2 or args.alpha == 0.0):+ write_prediction(X, genes, args.out, seed=args.seed)+ return++ cc_mask = dpt.cell_cycle_genes(args.data, genes)+ if args.verbose:+ print(f"[prior] cell-cycle genes on panel: {int(cc_mask.sum())}")+ deltas = dpt.group_delta(+ last.X, labels, genes, cc_mask, seed=args.seed, verbose=args.verbose+ )++ alpha_deltas = None+ if len(entries) >= 2 and args.alpha != 0.0:+ prev = read_stage(args.data, entries[-2], genes)+ alpha_deltas = dpt.pseudobulk_delta(+ prev.X, labels_of(prev), last.X, labels, genes, seed=args.seed+ )+ if args.verbose:+ print(f"[pb] groups with pseudobulk delta: {len(alpha_deltas)}")++ Xs = apply_shift(X, out_lab, deltas, args.beta, bool(args.renorm), alpha_deltas, args.alpha)+ write_prediction(Xs, genes, args.out, seed=args.seed) if __name__ == "__main__":
调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。
用到的知识库条目
| 编号 | 标题 | 出处 |
|---|---|---|
| k014 | Scorer noise and invariance on our proxy boards | notes/pitfalls/04_scorer_invariance.md |
| k020 | Correcting sampling-scope (dissection) bias in composition | notes/guides/modeling_and_evaluation_guide.html |
| k007 | Interval staging and held-out-window filtering of external data | notes/来件/virtualembryo.ai/rules.md |
分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。
| 改了什么 | 新增 solution/dpt.py(族内单快照成熟方向:GO 细胞周期集 z 分数 → 前 5 PC 定向 → 顶/底三分位对比 → 4 次对半符号收缩 → 99 分位截断)与 final 的伪批量差分支(α=0.35 + 符号门),run.py 加 --beta/--renorm 开关;抽样保真检查发现重复行=0,故未改抽样。提交配置为 --beta 0,与父节点逐位相同。 |
|---|---|
| 各组分数的变化 | cell_state:噪声内:56.90 vs 56.90,+0.00 covariation:噪声内:54.97 vs 54.97,+0.00 de_recovery:噪声内:53.06 vs 53.06,+0.00(提交的是 β=0) direction:噪声内:58.56 vs 58.56,+0.00 榜分:55.97 vs 55.97,+0.00;耗时 0.9s(父 2.3s),内存 1.21GB 持平。变化量表只测到 β=0,Engineer 报的 β=0.15/0.30 分数(55.94/55.80)不在表内,属其自述、未被本次评测证实。 |
| 假设是否成立 | 否 |
| 经验 |
|
| 下一步建议 |
|
对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。
只给统计和最后回答的摘录;完整对话请到原始记录位置里列出的文件看。
分析员
| 角色?调研员写计划、工程师改代码、分析员解读分数、审查员检查作弊。 | 分析员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数?大模型一共回复了几轮(每轮可以调用多个工具)。 | 6 |
| 工具调用?大模型调用读文件、执行命令等工具的次数,按工具名统计。 | 共 6 次:bash 5、write 1 |
| 用时 | 不到 1 分 |
| token 数?输入 / 输出(含思考)的 token 数,来自接口返回的用量。 | 输入 21,579 · 输出 1,580 · 思考 483 |
| 任务(第一行) | 分析节点 n4 的打分结果,写 ANALYSIS.json |
| 最后的回答(摘录) | ANALYSIS.json written |
| 原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。 | /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/analyst.jsonl 15 KB /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/analyst.stderr |
工程师
| 角色 | 工程师 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 25 |
| 工具调用 | 共 30 次:bash 26、write 4 |
| 用时 | 18 分 |
| token 数 | 输入 56,331 · 输出 17,511 · 思考 10,907 |
| 任务(第一行) | 按计划修改程序:重加权锚 + 稳定性收缩的单快照伪时间平移 |
| 最后的回答(摘录) | 做了什么:保留父节点全部组成部件(family 权重、4000 抽样、write_prediction),新增可开关的族内单快照成熟方向平移(GO 细胞周期集 → PC 定向 → 分位对比 → 4 次对半符号收缩),并实现 final 的伪批量分支(α=0.35 + 符号门)。抽样保真检查显示重复行=0,PLAN 步骤 1 不需要。 查分(4 次,剩 16):β=0 逐位复现父节点 55.9694;β=0.15 → 55.94(cell_state +0.5、direction +0.36、covariation −1.34);β=0.30 → 55.80(covariation −2.78);renorm on/off 无差。**de_score 三个 β 完全不变(0.1111)**,即常数式表达对 de_recovery 零作用,且 β 越大 variogram 越差 → 零结果,默认锁回 β=0(1.5 s,vec-check ok)。 下一步最值得试:① 只在 w=1 的族内 top-50 高置信基因上做大步长平移(de_recovery 可能要稀疏、族特异才响应);② 用「族内 rank 混合」这类不离开真实流形的表达扰动救 covariation;③ 用视图内 marker 结构重新标定族占比,替代手调 1.6/0.25。 |
| 原始记录位置 | /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/engineer.jsonl 155 KB /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/engineer.stderr |
调研员
| 角色 | 调研员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 6 |
| 工具调用 | 共 9 次:read 5、bash 3、write 1 |
| 用时 | 6 分 |
| token 数 | 输入 33,668 · 输出 3,163 · 思考 10,258 |
| 任务(第一行) | 为节点 n4(改进,父节点 n2)写改进计划 PLAN.json |
| 最后的回答(摘录) | PLAN.json written |
| 原始记录位置 | /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/researcher.jsonl 74 KB /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/researcher.stderr |
审查员
| 角色 | 审查员 alibaba-token-plan-cn/qwen3.8-max |
|---|---|
| 调用次数 | 1 次 |
| 轮数 | 8 |
| 工具调用 | 共 11 次:bash 7、read 3、write 1 |
| 用时 | 1 分 |
| token 数 | 输入 23,528 · 输出 1,469 · 思考 2,809 |
| 任务(第一行) | 审查节点 n4 的程序是否越界读取、写死目标、钻评分器漏洞 |
| 最后的回答(摘录) | REVIEW.json written |
| 原始记录位置 | /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/reviewer.jsonl 97 KB /home/spark-longxinyang/vec/2026_virtual_embryo/agent/runs/20261001-024931-search-t1-spark-eng-default/nodes/4/reviewer.stderr |