References
Foundational papers, product context, and internal artifacts — the source material behind the Knowledge Map and platform roadmap. Key takeaways from each are inlined below.
Foundational review of antibody discovery covering Generation 1 (animal immunization → hybridoma → humanization), Generation 2 (in vitro display — phage / yeast / mammalian / ribosome), and emerging Generation 3 (ML-augmented sequence priors + structure prediction).
Key takeaways
- Animal immunization takes 1-2 years per target; mouse repertoire is the inherent ceiling on Gen-1 diversity.
- Display tech (phage / yeast / mammalian / ribosome) increases library diversity by orders of magnitude but still requires physical binding selection.
- ML-augmented methods (VAE, transformer, IgLM, AntiBERTy) shift the cost curve but still need an experimental closing loop.
- Read §1-3 for the methodology lineage that the platform's 6-step RFdiffusion workflow sits on top of.
Deep retrospective on Generation-2 display technology. Useful for understanding what the AI workflow inherits from display screening (selection logic, panning rounds, polyclonal-to-monoclonal triage) and what it explicitly displaces (the library-construction step).
Key takeaways
- Library size and naive vs immune vs synthetic library construction trade-offs are still relevant to designing wet-lab follow-up screens after AI candidate generation.
- Display selection — yeast or phage — is the canonical 'closing loop' for AI-designed candidates, mapping onto Step 6 (experimental validation) of the design workflow.
- The 2010s phage-display engineering plays (CDR-H3 length engineering, framework canonicalization) directly inform how to constrain RFdiffusion's design space.
Baker lab. Fine-tuned RFdiffusion + ProteinMPNN + fine-tuned RoseTTAFold2 pipeline that designs epitope-specific VHHs, scFvs, and full antibodies entirely in silico. Cryo-EM confirms atomic-level CDR accuracy. Initial Kd in the tens-to-hundreds of nM range; OrthoRep affinity maturation brings the best designs to single-digit nM while preserving epitope selectivity. The paper that the platform's Stage 2-5 backbone is built on.
Key takeaways
- Three-stage pipeline: RFdiffusion samples backbones conditioned on (target structure, target sequence, hotspot residues, internal framework structure) → ProteinMPNN designs CDR sequences only (framework fixed) → fine-tuned RoseTTAFold2 filters by self-consistency + low PAE.
- Crucial detail: fine-tuning RFdiffusion on antibody complex structures, with the antibody framework provided as a 2D template (pairwise distances + dihedrals, no 3D coords) — lets the model sample diverse rigid-body docks while keeping framework structure fixed. Standard 'vanilla' RFdiffusion can't do this.
- AlphaFold2 fails on antibody-antigen complexes; AlphaFold3 was not available when this work was done. Fine-tuned RF2 is the practical filter. The platform's Wave 1 Boltz-2 / Protenix upgrade addresses this gap with modern open-weights replacements.
- Validation on 4 disease-relevant epitopes including influenza HA + Clostridium difficile TcdB + PHOX2B peptide-MHC. Cryo-EM resolves the actual binding mode of designed VHHs.
- OrthoRep is used for affinity maturation post-design — yeast-display continuous evolution. The platform's W6 step (experimental validation handoff) is the same loop. Display-construct exporter (V1 in the W/U/C/V matrix) is the planned platform piece.
BigHat Biosciences. Re-frames humanization as conditional generative sequence modeling: sample mutations from a masked language model trained on human antibody sequences (OAS / Observed Antibody Space), guided by a Product-of-Experts that combines the MLM with oracle scorers (binding affinity, thermostability, humanness). Outputs many humanized candidates per starting antibody, with therapeutic attributes preserved or improved.
Key takeaways
- Canonical humanization is a manual / semi-manual 'cut-and-paste' that typically yields a single candidate per starting antibody, with no guarantee that affinity or developability are retained.
- This method generates MANY humanized candidates per round, suitable for iterative dry-wet loops: design batch → measure → re-condition the prior → re-sample. The platform's wet-lab feedback loop (/docs/wet-lab-loop) provides exactly the ingest path this method needs.
- Three sampling strategies: Unmasked Sampling (no masking), Autoregressive Denoising Sampling (mask all, infill), Gibbs-like Sampling (mask one residue at a time). Iterative Masking Argmax compares against the Sapiens baseline.
- Product-of-Experts (PoE) formulation: log p(r_i) = MLM(r_i) + Σ oracle_k(r_i) / τ_k. Sampling at each location combines humanness (MLM) with attribute scores (binding, stability, etc.). This is the architecture the platform's Wave 3 'IASO-internal CAR-fit predictor' could be wrapped into.
- Lab-validated on real therapeutic programs: humanized antibodies showed improved binding to target antigens vs Sapiens-style argmax baselines. Direct relevance to the VHH-design page's `humanization` mode.
NYU + BigHat Biosciences. Introduces CloneBO: a Bayesian optimization procedure where the prior is learned from clonal families — sets of related antibody sequences from the human immune system that evolved to bind the same antigen. The generative model (CloneLM) is trained on hundreds of thousands of clonal families. CloneBO outperforms LaMBO and naive greedy in both in silico and wet-lab antibody optimization. Code: github.com/AlanNawzadAmin/CloneBO.
Key takeaways
- Core insight: the human immune system already 'solves' the antibody optimization problem via affinity maturation — it explores sequence space along a manifold of clonal families. Mimicking that exploration via a generative model gives a much better BO prior than starting from random or generic protein priors.
- CloneLM is a large language model that learns to generate sets of related sequences (clonal families) — not single sequences. This is the architectural difference vs IgLM / AntiBERTy which model individual antibodies.
- Bayesian update on previous wet-lab measurements via twisted sequential Monte Carlo. The platform's wet_lab_measurements table → score_calibrations flow is the right substrate to feed this loop.
- Validated wet-lab: stronger and more stable binders found within a limited experimental budget vs prior methods. Directly applicable to the BCMA top-20 → 2nd round → 3rd round affinity-maturation cycle.
- Comparison frame: structure-based de novo design (RFdiffusion-style) generates starting candidates but doesn't use prior measurements; CloneBO sits downstream and optimizes from there. The two are complementary, not competitive. Platform-wise: RFdiffusion is Stage 2-4; CloneBO would slot into Stage 6 (Iterative Optimization).
The patent disclosing the 026 scFv → CT103a CAR construct → equecabtagene autoleucel (福可苏) product lineage. This is a marketed BCMA CAR-T in China (approved 2023). The patent provides public sequence disclosure for the binder.
Key takeaways
- CT103a is a scFv-based, 4-1BB-CD3ζ-signaling lentivirally-delivered CAR-T.
- The 026 scFv (VL-linker-VH or VH-linker-VL composition specified in the patent) is the recognition end of CT103a.
- The patent's epitope claims around BCMA ECD are the comparison baseline for any new BCMA binder — including wet-lab maturation workflows.
- Per [[iaso-sequence-secrets-policy]], even though the sequence is partially patent-disclosed, treat the full operational sequence as internal — env-resolved, never committed.
The CT103a / 福可苏 construct as a SnapGene file. The lentiviral plasmid encoding the CAR cassette: signal peptide + 026 scFv + hinge + transmembrane + 4-1BB + CD3ζ, packaged in a pLKO backbone with kanR selection.
Key takeaways
- Construct architecture: SP — 026 scFv — hinge — TM — 4-1BB (costim) — CD3ζ (signaling). Same architecture as the platform's /car-design cassette diagram.
- Delivery: lentiviral (ex-vivo T-cell transduction). The platform's forward direction is the same cassette delivered in vivo via mRNA-LNP.
- The plasmid file is the ground-truth artifact for cross-referencing what the platform's 'CAR cassette assembly' UI should produce.
- Per [[iaso-sequence-secrets-policy]]: this construct file is internal-only. The .dna file lives in the private repo's reference/ directory; not served from the platform.
Structured 11-section summary of an internal review covering AI capability assessment for BCMA / CAR programs.
Key takeaways
- §1 — Two AI project tracks surfaced: Track 1 (evaluate existing BCMA CAR/scFv/antibody sequences — structure / epitope / affinity / mutation prediction / competitor comparison) and Track 2 (de novo, RFdiffusion, structure-based, dual-target). Both are real platform needs.
- §2 — Mutation-resistance prediction is the highest-leverage clinical narrative: BCMA position-27 (and other) mutations post-treatment; predicting whether CT103a still binds when competitors don't = product-differentiation gold.
- §7 — Interface predictor / scoring auditability is the highest-priority ask, repeated five times. Every score needs biological meaning + calculation method + reliability boundary + relative-ranking guidance.
- §9 — Commercial validation: the platform UI compares favorably to external tools priced up to ¥13M; no external-tool purchases anticipated for at least a year.
- §10 — Naming bug: the page formerly at /mrna is structurally a CAR cassette assembler, not pure mRNA design. Renamed to /car-design (DONE 2026-05-21).
- §11 — Open questions flagged: scoring benchmark dataset, mutation list specifics, competitor sequence sources, validation pathway, variant data ingestion.
Whiteboard sketch of the antibody-therapeutic format landscape from an internal working session. Source for the Layer-0 (format universe) section of the Knowledge Map.

Key takeaways
- Top center: mAb / CAR / TCE / TCR-mimetic as parallel format options, all sourced from antibody recognition.
- Center: CAR → Ab → sequence dependency chain — every format reduces to a binder sequence as the common substrate.
- Right: MHC-II / neoantigen / TCR-related modalities sketched as adjacent branches.
- Bottom: candidate-library clusters with mutation marks (X marks) — visual shorthand for the 'screen candidates against antigen + variants' workflow.
- Verbal framing from the same session: 'plan to use RF diffusion to regenerate antibody for BCMA using wet-lab-validated sequence' — the warm-start regenerate mode, Spec 1 Layer 3.