Research x AIReferences

参考文献与内部资料清单

这页把我们历史上围绕 Research x AI workflow 参考过的 paper、public data resource、内部报告和项目材料集中列出来。主叙事用中文,模型名、数据库名和 assay 名保留英文。

怎么在分享里用这页

开场

用 public data 组说明:AI 不是从空白开始,而是先把已有知识变成 project context。

流程图

用 Display / NGS 组说明:内部序列池进入平台之前必须 clean、dedupe、validate。

AI 作用

用 Structure / Affinity 组说明:模型不是单一分数,而是 evidence + ranking + uncertainty。

闭环

用 CAR-T / RL 组说明:wet-lab readout 才是 reward signal,目标是提升干湿实验有效比。

迁移性

用 G-protein 组说明:同一套 data → evidence → wet-lab loop 可以迁移到非抗体蛋白。

Public data / 数据源

支撑 target rationale、sequence / structure recall、表位和竞品背景。

UniProt蛋白序列、isoform、domain、物种和功能注释。G-protein design-around 和 antigen sequence 召回的基础。
RCSB PDB解析结构来源。用于 antigen / complex structure、epitope hotspot、界面复核。
AlphaFold DB没有 PDB 时的结构补位。适合先看 domain coverage 和 disordered region 风险。
SAbDab抗体-抗原复合物结构库。用于理解 Ab-Ag docking 的真实分布和 benchmark。
Observed Antibody Space / OAS人源抗体 repertoire 数据。支撑 humanization、sequence prior、clonal-family 建模思路。
PDBBind / affinity datasets亲和力模型训练和校准的常见来源;用于理解 Boltz-2 affinity head、PRODIGY 等模型的边界。
Display / NGS / Semi-de-novo

解释为什么 NGS pool 需要先过滤、验证和结构化,而不是直接按 read count 选。

Lu et al. 2021 — Animal Immunization, In Vitro Display, and Machine Learning for Antibody Discovery抗体发现方法谱系:动物免疫、phage/yeast display、ML-augmented discovery。
Almagro et al. 2023 — Evolution of phage display libraries for therapeutic antibody discovery展示文库演化与 panning 逻辑;对应 semi-de-novo 的 wet-lab pool 来源。
Semi-de-novo validation report内部真实跑批:DLL3 recovery gate + BCMA production run。说明 known binder 必须先能被回收。
BCMA / DLL3 / FcRH5 NGS files巧娥给到的 amino.txt、binding Excel、antigen sequence,是这次流程演示的数据起点。
Structure / Interface AI

支撑 interface AI、PDB evidence package、ChimeraX / Discovery Studio 复核。

Evans et al. 2021 — AlphaFold-Multimer早期 multimer complex prediction 基线;平台保留为 secondary scorer / continuity baseline。
Abramson et al. 2024 — AlphaFold 3统一 protein / ligand / nucleic acid interaction prediction 的新范式;理解结构预测上限。
Yin & Pierce 2024 — Evaluation of AlphaFold antibody-antigen modelingAb-Ag docking 的失败模式和可靠性边界;提醒不要迷信单一 ipTM。
Passaro et al. 2025 — Boltz-2开源 AF3-style complex prediction + affinity head;对应平台 Wave 1 之后的主要 scorer 方向。
Roy Burman et al. 2024 — AlphaRED / AF + Rosetta + replica exchange说明单模型预测后仍需要 hybrid refinement / docking refinement。
AntiConf 2025Ab-Ag complex confidence scorer 思路;对应未来替代裸 ipTM 的 interface confidence layer。
Affinity / Ranking / RL

支撑 candidate ranking、wet-lab reward signal、closed loop 和提升干湿实验有效比。

Boltz-2 affinity head结构预测共享表示上的 affinity scalar;短期优先暴露和校准。
PRODIGY — Xue et al. / Vangone & Bonvincontact-based ΔG / Kd 估计;与 deep model 正交,适合做 cross-check。
AlphaBind — Vasquez et al. 2024 / 2025抗体亲和力优化模型;适合 affinity maturation 和 IASO 数据 fine-tune。
GearBind — Cai et al. 2024几何 GNN 预测 mutation ΔΔG;适合从 parent binder 做 CDR mutation climbing。
Gordon et al. 2024 — Generative HumanizationProduct-of-Experts + antibody language model;humanization 和多目标属性优化参考。
Amin et al. 2024 — CloneBO从 clonal family 学到 BO prior;和 wet-lab feedback / active learning 思路直接相关。
Internal closed-loop RL roadmapReward v1、score calibration、active learning、contextual bandit / RL 阶梯。
CAR-T / Wet-lab validation

支撑从 AI shortlist 到 CAR construct、体外功能筛、体内验证的讲法。

Yu et al. 2024 — FCRL5-directed CAR-T cells exhibit antitumor activity against multiple myeloma亚鸽部分的 case anchor:新靶点 rationale → CAR design → in vitro / xenograft validation。
CT103a / 026 scFv patent lineageBCMA CAR-T 公开背景和竞品/内部对照基线;只做策略参考,不在页面暴露 operational sequence。
CT103a plasmid / CAR cassette reference内部 construct 参考:SP - binder - hinge - TM - costim - CD3ζ;对应 CAR Design 页面。
BCMA BLI panel / AlphaBind fine-tune noteswet-lab Kd 回流后用于校准 affinity model 和 ranking rule。
Wet-lab loop docs表达、binding、杀伤、CD107a、cytokine、exhaustion、in-vivo safety 等 readout 如何回流。
G-protein / Viral envelope

支撑永坤给的 G-protein design-around 逻辑。

Spindler et al. 2025 — Discovery and Validation of Alternatives to VSV-G for Pseudotyping of Lentiviral Vectors for In Vivo替代 viral glycoprotein / pseudotyping 路线的文献锚点。
UniProt viral glycoprotein recall从 Rhabdoviridae / glycoprotein 等 query 召回候选,做 VSV-G similarity 和长度过滤。
VSV-G / Cocal-G reference sequences用于 identity、47/354 equivalent position mapping 和 design-around risk band。
PDB 5OY9 / VSV-G structure reference结构比对锚点;与 Boltz / AlphaFold 预测候选做 TM-score / RMSD 对照。
G-protein smoke report内部 smoke:候选召回、identity band、position mapping、下一轮 structure batch readiness。
共同指向
这些资料共同支撑同一个结论:AI 的作用不是替代 wet-lab,而是把 public data、internal sequence pool、structure evidence、ranking 和 reward signal 串成 closed loop,持续提升干湿实验有效比。