Virtual Embryo Challenge更新于 10-03 19:58(北京时间) / 每 5 分钟更新

总览 · ← 返回运行 20261003-171955-search-t2-embryo-interp-chain-12h

节点 n30

homolog_pair — method and sources

运行?一次完整的自动搜索或 Agent 会话,有自己的锁定配置和证据包。20261003-171955-search-t2-embryo-interp-chain-12h
父节点(种子,没有父节点)
子节点—
操作?种子:人写的起点;改进:在父节点上改;草稿:从头写;修复:修父节点的报错。种子
状态已打分
分数搜索目标分 60.65 · proxy 60.65 · 3 次复测均分 60.31
审查通过 seed (human-written, locked)
用时?从运行开始到结束(或到现在)的挂钟时间。1 分
程序版本5665d23a0da520dd3e134654802d597e2ca0303b (programs.git)

方法说明?节点程序自带的 METHOD.md:这个程序做了什么、为什么。

来自 programs.git 5665d23a0d:solution/METHOD.md

homolog_pair — method and sources

Mechanisms (all computed from the view's inputs at run time; no stage name, size, cell count or type list is written into the program; times only as differences):

  1. Type families: same-name types on both bracket ends; a one-end type joins the other end's type with the highest correlation of stage-centred expression centroids (>= 0.3; Unknown excluded). Connected components = families.
  2. Composition: two-end families linear in logit at t, one-end families at that end's weight; count log-linear in t.
  3. Cells: real cells drawn stratified by type from the nearer end; families the near end lacks (or cannot fill) from the far end.
  4. Within-type expression step on non-zero entries only (sparsity kept): w * (type non-zero-mean difference minus the whole-stage non-zero-mean difference); each cell's own residual kept. Pairing by residual matching (shared PCA + coordinates).
  5. Positions: procrustes3d frame, both ends at the target RMS; blocks of neighbouring output cells translated rigidly by pos_f (0.15) of the block's mean pair displacement (far-end cells by 1 - pos_f).
  6. Scale: log RMS of all inputs vs (time - target): quadratic least squares with >= 3 inputs, else log-linear with damping 0.5. Output rescaled to it.
  7. Canonical frame: centre, PCA rotation, skew signs of axes 1-2, det = +1 (no mirror).

Knowledge used: generic only — cell types persist or are renamed between neighbouring stages (lineage continuity); composition interpolation in logit; McCann displacement interpolation (McCann 1997, Adv. Math. 128:153-179, doi:10.1006/aima.1997.1634) restricted to non-zero entries; coordinate canonicalisation (project knowledge k026). No held-out stage data or statistic. Sources: agent/prompts/ideas_T2__embryo__val_interp.md (T2EI-01, 03, 05, 06); notes/plan/tasks/G39-G43_plan_2026-10-03.md §2 (G40.2).

调研员的计划

名称seed homolog_pair
做法seed homolog_pair: G40.2 人写种子(方向库 T2EI-01 / 03 / 05 / 06)。括号两端(interp_bracket,不写阶段名)按细胞类型同源配对:同名直接配;只在一端出现的类型,用各自阶段内居中的表达质心相关(≥

代码改动?这个节点的程序和父节点程序的逐行差别:绿色是新增,红色是删除。

对比:这个提交的上一版(种子程序:相对空仓库)。改动的文件:solution/EXECUTION.json +1 −0、solution/METHOD.md +24 −0、solution/README.md +11 −0、solution/homolog.py +419 −0、solution/run.py +100 −0

diff --git a/solution/EXECUTION.json b/solution/EXECUTION.jsonnew file mode 100644index 0000000..9d5125c--- /dev/null+++ b/solution/EXECUTION.json@@ -0,0 +1 @@+{"gpu": false}diff --git a/solution/METHOD.md b/solution/METHOD.mdnew file mode 100644index 0000000..9ed978d--- /dev/null+++ b/solution/METHOD.md@@ -0,0 +1,24 @@+# homolog_pair — method and sources++Mechanisms (all computed from the view's inputs at run time; no stage name, size, cell count or type list is+written into the program; times only as differences):++1. Type families: same-name types on both bracket ends; a one-end type joins the other end's type with the highest+   correlation of stage-centred expression centroids (>= 0.3; `Unknown` excluded). Connected components = families.+2. Composition: two-end families linear in logit at t, one-end families at that end's weight; count log-linear in t.+3. Cells: real cells drawn stratified by type from the nearer end; families the near end lacks (or cannot fill)+   from the far end.+4. Within-type expression step on non-zero entries only (sparsity kept): w * (type non-zero-mean difference minus+   the whole-stage non-zero-mean difference); each cell's own residual kept. Pairing by residual matching (shared+   PCA + coordinates).+5. Positions: procrustes3d frame, both ends at the target RMS; blocks of neighbouring output cells translated+   rigidly by pos_f (0.15) of the block's mean pair displacement (far-end cells by 1 - pos_f).+6. Scale: log RMS of all inputs vs (time - target): quadratic least squares with >= 3 inputs, else log-linear with+   damping 0.5. Output rescaled to it.+7. Canonical frame: centre, PCA rotation, skew signs of axes 1-2, det = +1 (no mirror).++Knowledge used: generic only — cell types persist or are renamed between neighbouring stages (lineage continuity);+composition interpolation in logit; McCann displacement interpolation (McCann 1997, Adv. Math. 128:153-179,+doi:10.1006/aima.1997.1634) restricted to non-zero entries; coordinate canonicalisation (project knowledge k026).+No held-out stage data or statistic.+Sources: agent/prompts/ideas_T2__embryo__val_interp.md (T2EI-01, 03, 05, 06); notes/plan/tasks/G39-G43_plan_2026-10-03.md §2 (G40.2).diff --git a/solution/README.md b/solution/README.mdnew file mode 100644index 0000000..7fcce7a--- /dev/null+++ b/solution/README.md@@ -0,0 +1,11 @@+# homolog_pair(T2:embryo:val_interp)++G40.2 人写种子(方向库 T2EI-01 / 03 / 05 / 06)。括号两端(`interp_bracket`,不写阶段名)按细胞类型同源配对:同名直接配;只在一端出现的类型,用各自阶段内居中的表达质心相关(≥ 0.3,`Unknown` 不参与)并入另一端最像的类型,连通分量 = "类型族"。族比例在两端都有时按 logit 线性插值,只在一端的族按该端权重 (1−t)·f_a 或 t·f_b;细胞数 log 线性。细胞主要从近端抽真实细胞(分层),近端不够或没有的族从远端抽。+型内表达:每个细胞按残差匹配(共享 PCA + 坐标)配到对端同名类型(或同族)的一个细胞;表达只改非零元素(零保持零,不稠密化),步长 = w ×(对端非零均值 − 本型非零均值 − 两阶段整体的同一差值),w = 本端到目标的时间比例;每个细胞自己的残差保留(`resid_w` 0)。检出率插值开关 `switch` 默认关:两端检出率差几倍,是技术性的,打开后 proxy 细胞状态组从 56 掉到 31。+坐标:两端 procrustes3d 对齐并缩到目标 RMS;输出细胞按邻域块(种子 + 约 8 个最近邻,Voronoi)为单位刚性平移 pos_f × 块内平均配对位移(近端 pos_f = 0.15,远端细胞走 1 − pos_f 回到近端几何)。尺度:所有输入的 log RMS 对(时间 − 目标时间)拟合,≥ 3 个输入用二次最小二乘(final 走这支),2 个输入用 log 线性 + scale_damp 0.5(proxy 走这支,同 mix)。最后坐标规范化:减质心、PCA 旋转、前两轴按三阶矩定号、第三轴保证 det = +1。+`--ablate pairing|within_type|position`(逗号分隔或重复)关掉对应机制;`--set key=json` 改参数(调试用,harness 不传)。纯 CPU,VM 上 proxy 约 15 s、final 约 7 s。++proxy(VM,2026-10-03,seed 0 / 1 / 2,`scripts/score_proxy.py`):60.22 / 59.67 / 59.87;mix 同 seed 55.96 / 55.96 / 56.26。seed 0 分组:表达 59.5 / **状态 55.8** / 形状 69.7 / **邻域 55.9**(mix:59.1 / 36.3 / 73.6 / 54.9)。+消融(seed 0):关 pairing 61.06(状态 49.1,variogram 掉;形状、邻域升,因为更多远端细胞);关 within_type 60.02(中心化后的型内步长在 proxy 上几乎不动分);关 position 58.29(形状 62.9)。pos_f 敏感性:0.3 → 61.10,0.5 → 61.90(形状 d2 升);默认按计划起点 0.15,没有在 proxy 上调。+final(目标在两个括号输入之间,离较早的一端更近):n = 榜上限,几乎全部来自较早的括号端;尺度走三点二次拟合分支(数值由程序现场算,不在这里给)。伪装视图(时间整体平移一天、输入倒序、键倒序、换路径)输出逐位相同;`vec_check` 通过。+已知弱点:型内表达步长被"减去整体差值"后很小,状态组的提升主要来自组成(族插值 + 近端真实细胞),不是来自表达插值;pos_f 与 scale_damp 都没有在 final 括号上验证过;只用质心相关配对,没有用 Sinkhorn 流出比例。准入记录:`agent/seeds/T2__embryo__val_interp/ADMISSION_2026-10-03_homolog_pair.md`。diff --git a/solution/homolog.py b/solution/homolog.pynew file mode 100644index 0000000..889e77a--- /dev/null+++ b/solution/homolog.py@@ -0,0 +1,419 @@+"""homolog_pair: type-family pairing + logit composition + within-type McCann step + block positions.++Everything here is computed from the two bracketing inputs (and, for the scale, from all inputs); no stage name,+size, cell count or type list is written into the program. Times enter only as differences.++Steps (G40.2, ideas T2EI-01 / 03 / 05 / 06):++1. Families. Types with the same name on both bracket ends are one family. With ``pairing`` on, a type found on one+   end only is joined to the type of the other end whose expression centroid correlates best with it (centroids+   centred by their own stage's mean centroid, so a whole-stage shift does not drive the match), when that+   correlation is >= ``pair_min_corr``. Families are the connected components of these edges.+2. Composition. Family fractions f_a, f_b of families on both ends are interpolated linearly in logit at t;+   a family found on one end only keeps that end's weight ((1 - t) f_a or t f_b); the output count is log-linear in+   t, clipped to the board range.+3. Cells. Each family's quota is drawn (stratified by type) from the end nearer to the target; a family the near end+   lacks, or the part of a quota the near end cannot fill, comes from the far end.+4. Expression (``within_type``). A drawn cell of type T is moved a fraction ``w`` of the way (w = time fraction from+   its own end to the target) towards its reference pool R on the other end (same-name type, else the family). It is+   paired to the R cell nearest to it after removing the type-mean difference (residual matching in a shared PCA,+   plus coordinates); the pair gives the block displacement (step 5) and, with ``resid_w`` > 0, a residual step.+   The expression step changes non-zero entries only (zeros stay zeros: no dense fill): w * (non-zero mean_R -+   non-zero mean_T), minus the same quantity computed over the whole stages (``level_centre``: the stage-wide part,+   e.g. sequencing depth, is not a type's development), floored at a quarter of the entry. The detection-rate+   switch (``switch``: zeros on / non-zeros off at the interpolated detection rate) exists but is off by default:+   the bracket ends differ several-fold in detection rate, which is technical, and switching it collapsed the+   variogram on the proxy (cell state 31 vs 56).+5. Positions (``position``). Both clouds are put in one frame (procrustes3d on shared type centroids) and scaled to+   the target RMS. Output cells are grouped into blocks (a block seed + its nearest drawn neighbours, Voronoi on the+   drawn cells); every cell of a block moves rigidly by pos_f * (mean pair displacement of the block), pos_f = the+   position fraction (near-end cells) or 1 - pos_f (far-end cells, which move back towards the near geometry).+6. Scale. log RMS of every input against (time - target time): with >= 3 inputs a least-squares quadratic, else+   log-linear between the bracket ends with ``scale_damp``. The output is rescaled to that RMS.+7. Canonical frame (T2EI-06): centre, PCA rotation, skew sign of the first two axes, third axis sign so det = +1.+"""++from __future__ import annotations++import numpy as np+from scipy import sparse+from sklearn.neighbors import NearestNeighbors++from src.task2_spatial.frame import align_pair, rms_radius, scale_to_rms+from src.task2_spatial.sample import interp_count++UNKNOWN = "Unknown"+DEFAULTS = {+    "pairing": True,+    "within_type": True,+    "position": True,+    "pair_min_corr": 0.3,+    "pos_f": 0.15,+    "block_k": 8,+    "resid_w": 0.0,+    "coord_w": 1.0,+    "pca_dim": 20,+    "pca_n": 4000,+    "scale_damp": 0.5,+    "count_damp": 1.0,+    "min_type_cells": 5,+    "level": True,+    "switch": False,+    "level_centre": True,+}+++# ------------------------------------------------------------------------------------------------------------------+# helpers+# ------------------------------------------------------------------------------------------------------------------+def _dense(X, idx=None) -> np.ndarray:+    block = X if idx is None else X[idx]+    if sparse.issparse(block):+        return np.asarray(block.toarray(), dtype=np.float32)+    return np.asarray(block, dtype=np.float32)+++def _logit(p):+    p = np.clip(np.asarray(p, dtype=np.float64), 1e-12, 1 - 1e-12)+    return np.log(p) - np.log1p(-p)+++def _sigmoid(z):+    return 1.0 / (1.0 + np.exp(-np.asarray(z, dtype=np.float64)))+++def _type_centroids(stage) -> dict[str, np.ndarray]:+    out = {}+    for lab in np.unique(stage.labels):+        idx = np.flatnonzero(stage.labels == lab)+        out[str(lab)] = np.asarray(stage.X[idx].mean(axis=0)).ravel().astype(np.float64)+    return out+++class _UF:+    def __init__(self, keys):+        self.p = {k: k for k in keys}++    def find(self, k):+        while self.p[k] != k:+            self.p[k] = self.p[self.p[k]]+            k = self.p[k]+        return k++    def union(self, a, b):+        ra, rb = self.find(a), self.find(b)+        if ra != rb:+            self.p[max(ra, rb)] = min(ra, rb)+++# ------------------------------------------------------------------------------------------------------------------+# 1. families+# ------------------------------------------------------------------------------------------------------------------+def families(stage_a, stage_b, pairing: bool, min_corr: float):+    """{('a'|'b', type): family id}, the pairing edges with their correlations."""+    ta = sorted(set(map(str, stage_a.labels)))+    tb = sorted(set(map(str, stage_b.labels)))+    keys = [("a", t) for t in ta] + [("b", t) for t in tb]+    uf = _UF(keys)+    for t in sorted(set(ta) & set(tb)):+        uf.union(("a", t), ("b", t))+    edges = []+    if pairing:+        ca, cb = _type_centroids(stage_a), _type_centroids(stage_b)+        ma = np.mean([ca[t] for t in ta], axis=0)+        mb = np.mean([cb[t] for t in tb], axis=0)+        A = np.stack([ca[t] - ma for t in ta])+        B = np.stack([cb[t] - mb for t in tb])+        An = A / (np.linalg.norm(A, axis=1, keepdims=True) + 1e-12)+        Bn = B / (np.linalg.norm(B, axis=1, keepdims=True) + 1e-12)+        C = An @ Bn.T  # Pearson-like (centred) correlation of type centroids+        # "Unknown" (cells no label could be given, view contract) is never a pairing partner or a paired type+        C[[i for i, x in enumerate(ta) if x == UNKNOWN], :] = -np.inf+        C[:, [j for j, x in enumerate(tb) if x == UNKNOWN]] = -np.inf+        for i, t in enumerate(ta):+            if t in set(tb) or t == UNKNOWN:+                continue+            j = int(np.argmax(C[i]))+            if C[i, j] >= min_corr:+                uf.union(("a", t), ("b", tb[j]))+                edges.append({"a": t, "b": tb[j], "corr": round(float(C[i, j]), 4)})+        for j, t in enumerate(tb):+            if t in set(ta) or t == UNKNOWN:+                continue+            i = int(np.argmax(C[:, j]))+            if C[i, j] >= min_corr:+                uf.union(("a", ta[i]), ("b", t))+                edges.append({"a": ta[i], "b": t, "corr": round(float(C[i, j]), 4)})+    roots = sorted({uf.find(k) for k in keys})+    fid = {r: i for i, r in enumerate(roots)}+    return {k: fid[uf.find(k)] for k in keys}, edges+++# ------------------------------------------------------------------------------------------------------------------+# 2-3. composition and cell draw+# ------------------------------------------------------------------------------------------------------------------+def _stratified(labels: np.ndarray, pool: np.ndarray, k: int, rng) -> np.ndarray:+    """k cells of `pool` (indices), stratified by label, without replacement (k <= len(pool))."""+    if k >= len(pool):+        return np.sort(pool)+    labs = labels[pool]+    types, counts = np.unique(labs, return_counts=True)+    raw = counts / counts.sum() * k+    alloc = np.floor(raw).astype(int)+    for i in np.argsort(-(raw - alloc), kind="stable")[: int(k - alloc.sum())]:+        alloc[i] += 1+    picks = [rng.choice(pool[labs == t], int(a), replace=False) for t, a in zip(types, alloc) if a > 0]+    return np.sort(np.concatenate(picks)) if picks else np.array([], dtype=int)+++def draw_cells(stage_a, stage_b, fam, t: float, n: int, rng):+    """Returns (rows_a, rows_b, family fractions table)."""+    la = np.asarray(stage_a.labels).astype(str)+    lb = np.asarray(stage_b.labels).astype(str)+    fa_id = np.array([fam[("a", x)] for x in la])+    fb_id = np.array([fam[("b", x)] for x in lb])+    fids = sorted(set(fa_id.tolist()) | set(fb_id.tolist()))+    fa = np.array([(fa_id == f).mean() for f in fids])+    fb = np.array([(fb_id == f).mean() for f in fids])+    both = (fa > 0) & (fb > 0)+    # two-end families: linear in logit; one-end families: that end's weight ((1 - t) f_a or t f_b)+    ft = np.where(both, _sigmoid((1 - t) * _logit(np.where(both, fa, 0.5)) + t * _logit(np.where(both, fb, 0.5))),+                  (1 - t) * fa + t * fb)+    ft = ft / ft.sum()+    raw = ft * n+    quota = np.floor(raw).astype(int)+    for i in np.argsort(-(raw - quota), kind="stable")[: int(n - quota.sum())]:+        quota[i] += 1+    near_a = t < 0.5 or (t == 0.5)+    rows_a, rows_b = [], []+    for f, q in zip(fids, quota):+        if q <= 0:+            continue+        pa = np.flatnonzero(fa_id == f)+        pb = np.flatnonzero(fb_id == f)+        first, second = (pa, pb) if near_a else (pb, pa)+        lab1, lab2 = (la, lb) if near_a else (lb, la)+        k1 = min(q, len(first))+        k2 = min(q - k1, len(second))+        s1 = _stratified(lab1, first, k1, rng) if k1 else np.array([], dtype=int)+        s2 = _stratified(lab2, second, k2, rng) if k2 else np.array([], dtype=int)+        rest = q - k1 - k2+        if rest > 0:  # both ends exhausted: repeat cells of the near end's pool (or the far one if empty)+            pool, which = (first, 1) if len(first) else (second, 2)+            extra = rng.choice(pool, rest, replace=True)+            if which == 1:+                s1 = np.concatenate([s1, extra])+            else:+                s2 = np.concatenate([s2, extra])+        (rows_a if near_a else rows_b).append(s1)+        (rows_b if near_a else rows_a).append(s2)+    cat = lambda xs: np.concatenate(xs).astype(int) if xs else np.array([], dtype=int)  # noqa: E731+    table = [{"family": int(f), "f_a": round(float(x), 5), "f_b": round(float(y), 5), "f_t": round(float(w), 5)}+             for f, x, y, w in zip(fids, fa, fb, ft)]+    return cat(rows_a), cat(rows_b), table+++# ------------------------------------------------------------------------------------------------------------------+# 4. expression: hurdle McCann step with residual partner matching+# ------------------------------------------------------------------------------------------------------------------+class SharedPCA:+    def __init__(self, stage_a, stage_b, dim: int, n: int, rng):+        ia = np.sort(rng.choice(stage_a.n, min(n, stage_a.n), replace=False))+        ib = np.sort(rng.choice(stage_b.n, min(n, stage_b.n), replace=False))+        Z = np.vstack([_dense(stage_a.X, ia), _dense(stage_b.X, ib)]).astype(np.float64)+        self.mu = Z.mean(axis=0)+        _, s, vt = np.linalg.svd(Z - self.mu, full_matrices=False)+        k = int(min(dim, vt.shape[0]))+        self.V = vt[:k].T+        self.sd = s[:k] / np.sqrt(max(len(Z) - 1, 1)) + 1e-8++    def project(self, X: np.ndarray) -> np.ndarray:+        return ((X - self.mu) @ self.V) / self.sd+++def _hurdle_stats(X: np.ndarray):+    nz = X > 0+    det = nz.mean(axis=0)+    cnt = nz.sum(axis=0)+    mean_nz = np.where(cnt > 0, X.sum(axis=0) / np.maximum(cnt, 1), np.nan)+    return det, mean_nz+++def move_expression(src, other, rows: np.ndarray, src_coords, other_coords, fam, side: str, w: float,+                    pca: SharedPCA, p: dict, rng):+    """Expression of src[rows] moved a fraction w towards `other` (dense block, sparsity pattern kept) and each+    row's partner coordinates in `other` (None where the cell has no partner pool)."""+    other_side = "b" if side == "a" else "a"+    ls = np.asarray(src.labels).astype(str)+    lo = np.asarray(other.labels).astype(str)+    out = _dense(src.X, rows)+    partner_xyz = np.full((len(rows), 3), np.nan)+    other_types = set(lo.tolist())+    fam_other = np.array([fam[(other_side, x)] for x in lo])+    sub_lab = ls[rows]+    if p["within_type"] and p["level_centre"]:+        _, gN = _hurdle_stats(_dense(src.X, rng.choice(src.n, min(src.n, 4000), replace=False)))+        _, gR = _hurdle_stats(_dense(other.X, rng.choice(other.n, min(other.n, 4000), replace=False)))+        glevel = np.where(np.isfinite(gN) & np.isfinite(gR), gR - gN, 0.0)+    for T in np.unique(sub_lab):+        loc = np.flatnonzero(sub_lab == T)+        if T in other_types:+            R = np.flatnonzero(lo == T)+        else:+            R = np.flatnonzero(fam_other == fam[(side, T)])+        if len(R) < int(p["min_type_cells"]):+            continue  # one-end type: kept as is (expression and position)+        N = np.flatnonzero(ls == T)+        XN = _dense(src.X, N)+        XR = _dense(other.X, R)+        muN, muR = XN.mean(axis=0), XR.mean(axis=0)+        # partner: nearest R cell to (x + muR - muN) in the shared PCA (+ coordinates), i.e. residual matching+        zR = pca.project(XR.astype(np.float64))+        zq = pca.project((out[loc] + (muR - muN)).astype(np.float64))+        cw = float(p["coord_w"])+        if cw > 0:+            scl = np.sqrt(zR.shape[1])+            cR = np.asarray(other_coords[R], dtype=np.float64)+            cq = np.asarray(src_coords[rows[loc]], dtype=np.float64)+            r = rms_radius(np.vstack([cR, cq])) + 1e-8+            zR = np.hstack([zR, cw * scl * cR / r / np.sqrt(3)])+            zq = np.hstack([zq, cw * scl * cq / r / np.sqrt(3)])+        nn = NearestNeighbors(n_neighbors=1).fit(zR)+        part = R[nn.kneighbors(zq)[1][:, 0]]+        partner_xyz[loc] = np.asarray(other_coords[part], dtype=np.float64)+        if not p["within_type"] or w <= 0:+            continue+        dN, mN = _hurdle_stats(XN)+        dR, mR = _hurdle_stats(XR)+        block = out[loc]+        Y = _dense(other.X, part)+        nzb = block > 0+        # non-zero entries: level step of the non-zero means (+ optional residual step towards the partner)+        level = np.where(np.isfinite(mN) & np.isfinite(mR), mR - mN, 0.0)+        if p["level_centre"]:  # remove the stage-wide part of the step (depth / batch), keep the type-specific part+            level = level - glevel+        level = level.astype(np.float32)+        step = np.float32(w if p["level"] else 0.0) * level[None, :]+        if float(p["resid_w"]) > 0:+            res = (Y - np.where(Y > 0, mR, 0)) - (block - np.where(nzb, mN, 0))+            step = step + np.float32(w * float(p["resid_w"])) * np.where((Y > 0) & nzb, res, 0).astype(np.float32)+        moved = np.where(nzb, np.maximum(block + step, 0.25 * block), 0.0)+        # detection: target rate d* = dN + w (dR - dN); switch off / on at the implied rate+        dstar = dN + w * (dR - dN)+        p_off = np.where(dN > 0, np.clip((dN - dstar) / np.maximum(dN, 1e-12), 0, 1), 0.0)+        p_on = np.where(dN < 1, np.clip((dstar - dN) / np.maximum(1 - dN, 1e-12), 0, 1), 0.0)+        if not p["switch"]:+            p_off, p_on = np.zeros_like(p_off), np.zeros_like(p_on)+        u = rng.random(block.shape)+        off = nzb & (u < p_off[None, :])+        on = (~nzb) & (u < p_on[None, :])+        on_val = np.where(Y > 0, Y, np.nan_to_num(mR, nan=0.0)[None, :]).astype(np.float32)+        moved = np.where(off, 0.0, moved)+        moved = np.where(on, on_val, moved)+        out[loc] = moved.astype(np.float32)+    return out, partner_xyz+++# ------------------------------------------------------------------------------------------------------------------+# 5-7. positions, scale, canonical frame+# ------------------------------------------------------------------------------------------------------------------+def block_move(xyz: np.ndarray, partner: np.ndarray, frac: float, k: int, rng) -> np.ndarray:+    """Rigid per-block translation by frac * mean(partner - xyz) of the block (cells without a partner contribute+    nothing; a block without any partner stays)."""+    n = len(xyz)+    if n == 0 or frac == 0:+        return xyz.copy()+    disp = partner - xyz+    has = np.isfinite(disp).all(axis=1)+    n_blocks = max(1, int(np.ceil(n / (k + 1))))+    seeds = np.sort(rng.choice(n, n_blocks, replace=False))+    nn = NearestNeighbors(n_neighbors=1).fit(xyz[seeds])+    blk = nn.kneighbors(xyz)[1][:, 0]+    out = xyz.copy()+    for b in range(n_blocks):+        m = blk == b+        h = m & has+        if h.any():+            out[m] = xyz[m] + frac * disp[h].mean(axis=0)+    return out+++def target_rms(entries_rms: list[tuple[float, float]], rms_a: float, rms_b: float, t: float, damp: float):+    """entries_rms: [(time - target time, rms)] of every input. >= 3 distinct times: quadratic least-squares fit of+    log rms, value at 0. Else log-linear between the bracket ends with damp."""+    dts = np.array([d for d, _ in entries_rms], dtype=np.float64)+    lr = np.log(np.array([r for _, r in entries_rms], dtype=np.float64))+    if len(np.unique(dts)) >= 3:+        coef = np.polyfit(dts, lr, 2)+        return float(np.exp(np.polyval(coef, 0.0))), "quadratic"+    la, lb = np.log(rms_a), np.log(rms_b)+    return float(np.exp(la + damp * t * (lb - la))), "loglinear"+++def canonical(coords: np.ndarray) -> np.ndarray:+    x = coords - coords.mean(axis=0)+    _, _, vt = np.linalg.svd(x, full_matrices=False)+    R = vt.T+    if np.linalg.det(R) < 0:+        R[:, 2] *= -1+    y = x @ R+    for ax in (0, 1):+        if float((y[:, ax] ** 3).mean()) < 0:+            y[:, ax] *= -1+            y[:, 2] *= -1  # keep det = +1 (no mirror image)+    return y+++def jitter_duplicates(coords: np.ndarray, rng) -> np.ndarray:+    _, inv, counts = np.unique(np.round(coords, 5), axis=0, return_inverse=True, return_counts=True)+    dup = counts[np.ravel(inv)] > 1+    if not dup.any():+        return coords+    out = coords.copy()+    out[dup] += rng.normal(0.0, 1e-3 * (rms_radius(coords) + 1e-8), size=(int(dup.sum()), 3))+    return out+++# ------------------------------------------------------------------------------------------------------------------+def interpolate(stage_a, stage_b, t: float, rms_points, params: dict, lo: int, hi: int, seed: int):+    p = dict(DEFAULTS)+    p.update(params)+    rng = np.random.default_rng(seed)+    fam, edges = families(stage_a, stage_b, bool(p["pairing"]), float(p["pair_min_corr"]))+    n = interp_count(stage_a.n, stage_b.n, t, lo, hi, float(p["count_damp"]))+    rows_a, rows_b, comp = draw_cells(stage_a, stage_b, fam, t, n, rng)++    ca, cb, ainfo = align_pair(stage_a.coords, stage_b.coords, stage_a.labels, stage_b.labels, "procrustes3d")+    rms_a, rms_b = rms_radius(stage_a.coords), rms_radius(stage_b.coords)+    trms, scale_rule = target_rms(rms_points, rms_a, rms_b, t, float(p["scale_damp"]))+    ca, cb = scale_to_rms(ca, trms), scale_to_rms(cb, trms)++    pca = SharedPCA(stage_a, stage_b, int(p["pca_dim"]), int(p["pca_n"]), rng)+    ea, pa_partner = move_expression(stage_a, stage_b, rows_a, ca, cb, fam, "a", t, pca, p, rng)+    eb, pb_partner = move_expression(stage_b, stage_a, rows_b, cb, ca, fam, "b", 1.0 - t, pca, p, rng)++    near_a = t <= 0.5+    pos_f = float(p["pos_f"]) if p["position"] else 0.0+    fa = pos_f if near_a else 1.0 - pos_f+    fb = 1.0 - pos_f if near_a else pos_f+    if not p["position"]:+        fa = fb = 0.0+    xa = block_move(ca[rows_a], pa_partner, fa, int(p["block_k"]), rng)+    xb = block_move(cb[rows_b], pb_partner, fb, int(p["block_k"]), rng)++    expr = np.vstack([ea, eb]).astype(np.float32)+    coords = np.vstack([xa, xb])+    coords = jitter_duplicates(coords, rng)+    coords = canonical(scale_to_rms(coords, trms))+    X = sparse.csr_matrix(np.clip(expr, 0.0, None))+    X.eliminate_zeros()+    in_nnz = (stage_a.X.nnz / np.prod(stage_a.X.shape) * len(rows_a) + stage_b.X.nnz / np.prod(stage_b.X.shape) * len(rows_b)) / max(n, 1)+    info = {+        "t": t, "n": int(X.shape[0]), "n_from_a": int(len(rows_a)), "n_from_b": int(len(rows_b)),+        "n_families": len({v for v in fam.values()}), "pair_edges": edges,+        "rms_a": rms_a, "rms_b": rms_b, "target_rms": trms, "scale_rule": scale_rule, "out_rms": rms_radius(coords),+        "pos_f_a": fa, "pos_f_b": fb, "nnz_frac": float(X.nnz / np.prod(X.shape)), "source_nnz_frac": float(in_nnz),+        "n_shared_types": ainfo.get("n_shared_types"), "z_dot": ainfo.get("z_dot"),+        "flags": {k: bool(p[k]) for k in ("pairing", "within_type", "position")}, "composition": comp,+    }+    return X, coords.astype(np.float32), infodiff --git a/solution/run.py b/solution/run.pynew file mode 100644index 0000000..68e74b7--- /dev/null+++ b/solution/run.py@@ -0,0 +1,100 @@+#!/usr/bin/env python3+"""homolog_pair (T2 interpolation): homologous type pairing + logit composition + within-type McCann step.++Brackets the target with the nearest inputs before / after it (view_io.interp_bracket; no stage name in the code),+pairs the two ends' cell types into families (same name, else best centred-centroid correlation), interpolates the+family composition in logit, draws real cells mostly from the nearer end, moves each cell's expression a time-fraction+step towards its paired cell on the other end (hurdle step: sparsity pattern kept, residual kept), moves blocks of+neighbouring cells rigidly a small fraction pos_f along the pair displacement, rescales to an RMS interpolated in time+(three-point quadratic fit when the view has >= 3 inputs) and writes the cloud in a canonical frame. Details:+homolog.py; method and sources: METHOD.md.++  python run.py --data <view> --out <pred.h5ad> --seed <int> [--ablate pairing,within_type,position]++--ablate turns mechanisms off (comma separated or repeated): pairing (only same-name families), within_type (no+expression step), position (no block move; scale and canonical frame stay). With one input / no bracket: copy the+anchor stage (stratified to the board cap), as copy_last.+"""++from __future__ import annotations++import argparse+import json+import sys+from pathlib import Path++import numpy as np++sys.path.insert(0, str(Path(__file__).resolve().parent))++from homolog import interpolate  # noqa: E402+from src.task2_spatial.frame import rms_radius  # noqa: E402+from src.task2_spatial.sample import take  # noqa: E402+from src.task2_spatial.view_io import (  # noqa: E402+    inputs_by_time,+    interp_bracket,+    load_manifest,+    panel_genes,+    read_stage,+    target_time,+    write_t2,+)+from src.task1_temporal.view_io import write_prediction  # noqa: E402++PARAMS = {+    "pair_min_corr": 0.3,   # centred-centroid correlation needed to join a one-end type to a family+    "pos_f": 0.15,          # fraction of the pair displacement a near-end block moves (G40.2 starting point)+    "block_k": 8,           # block = seed cell + ~k nearest drawn cells, moved rigidly+    "resid_w": 0.0,         # 0: each cell keeps its own residual entirely (only the type-level step is applied)+    "coord_w": 1.0,         # weight of coordinates in the partner match (residual PCA + coordinates)+    "scale_damp": 0.5,      # two-input fallback only (log-linear RMS damping, as the mix seed)+    "count_damp": 1.0,+}+ABLATIONS = ("pairing", "within_type", "position")+++def main() -> None:+    parser = argparse.ArgumentParser()+    parser.add_argument("--data", required=True)+    parser.add_argument("--out", required=True)+    parser.add_argument("--seed", type=int, default=0)+    parser.add_argument("--ablate", action="append", default=[])+    parser.add_argument("--set", action="append", default=[], help="key=json_value: override a PARAMS / homolog default")+    args = parser.parse_args()+    off = {x.strip() for a in args.ablate for x in a.split(",") if x.strip()}+    bad = off - set(ABLATIONS)+    if bad:+        raise SystemExit(f"--ablate: unknown {sorted(bad)}; choose from {ABLATIONS}")++    manifest = load_manifest(args.data)+    genes = panel_genes(args.data, manifest)+    lo, hi = int(manifest["min_cells"]), int(manifest["max_cells"])+    a, b, t = interp_bracket(manifest)+    if b is None:+        stage = read_stage(args.data, a, genes)+        n = int(np.clip(stage.n, lo, hi))+        rows = np.sort(take(stage.labels, min(n, stage.n), np.random.default_rng(args.seed)))+        write_t2(args.out, stage.X[rows].toarray(), stage.coords[rows], genes, seed=args.seed)+        return++    tt = target_time(manifest)+    stages, rms_points = {}, []+    for e in inputs_by_time(manifest):+        st = read_stage(args.data, e, genes)+        rms_points.append((float(e["time"]) - tt, rms_radius(st.coords)))+        if e["path"] in (a["path"], b["path"]):+            stages[e["path"]] = st+        else:+            del st+    params = dict(PARAMS)+    params.update({k: (k not in off) for k in ABLATIONS})+    for kv in args.set:+        k, v = kv.split("=", 1)+        params[k] = json.loads(v)+    X, coords, info = interpolate(stages[a["path"]], stages[b["path"]], t, rms_points, params, lo, hi, args.seed)+    print(json.dumps({"bracket": [a["stage"], b["stage"]], **info}, default=float), file=sys.stderr)+    write_prediction(X, genes, args.out, coords=coords, seed=args.seed)+++if __name__ == "__main__":+    main()

调研来源?调研员查到并用到的知识条目和文献检索结果(只列标题和编号)。

没有记录调研来源。

分析结果?分析员写的 ANALYSIS.json:改了什么、各组分数怎么变、假设是否成立、经验和下一步建议。

没有分析结果(ANALYSIS.json)。

对话摘要?每个角色和大模型对话的统计:轮数、工具调用、用时、token 数和最后的回答摘录;原始记录只给路径。

只给统计和最后回答的摘录;完整对话请到原始记录位置?完整对话记录在 Spark 上的文件路径;看板不复制原始内容。里列出的文件看。

这个节点没有大模型对话记录。