Platform Walkthrough: Models, Methods, and What Is Actually Running
A technical tour of the discovery platform as it exists today. Every capability below is labelled with its real status. Nothing here is aspirational.
1. The problem, stated as engineering
Therapeutic antibody discovery is a search problem with a brutal economics. The space of possible binders is effectively infinite, the assay that gives a trustworthy answer costs weeks and real money, and the number of candidates a wet lab can carry to a binding measurement in one cycle is on the order of tens.
So the entire job of a computational platform is narrow and specific:
Spend compute to decide which small number of molecules deserve an experiment.
Not "design a drug". Not "replace the wet lab". Rank and triage, so that the scarce resource — assay slots — lands on candidates with the best prior odds. Everything in this document exists to serve that one sentence.
This framing matters because it fixes what "good" means. A model that is beautifully accurate but cannot be run on ten thousand candidates is useless here. A crude score that reliably puts the true binders in the top decile is valuable. The metric is enrichment at the top of a ranked list, not global accuracy.
2. From a wet-lab decision to a computable question
The mapping from biology to computation is the core intellectual content of the platform. It runs in four steps.
Step 1 — Name the decision. Not "is this a good antibody" but something a number can answer. In practice the decisions are: Does this binder engage the target at all? Does it engage the intended epitope? How tightly? Will it survive manufacturing and formulation?
Step 2 — Find the physical quantity behind the decision. Binding is a free-energy difference. Epitope identity is a set of contacting residues. Developability is a set of chemical liabilities in specific sequence positions. Each decision reduces to a structural or sequence observable.
Step 3 — Choose a model that predicts that observable. This is where the model inventory in §4 comes from. Crucially, different decisions need different model families — there is no single model that answers all four questions, and pretending otherwise is the most common way these platforms go wrong.
Step 4 — Convert the observable into a ranking, with the uncertainty attached. A predicted structure is not a decision. Somebody has to say "buried surface area above this threshold, interface confidence above that one". That conversion is a rubric, it is a human judgement, and it should live in versioned code with tests — not in a spreadsheet. The platform now holds its triage rubrics in code for exactly this reason.
The honest summary of step 4: the models produce physical estimates; the rubric turns estimates into decisions; and the rubric is usually the weakest, least examined link in the chain.
3. What kind of model this is — and why it is not a language model
This is the question that most often causes confusion, so it is worth being precise.
A large language model is a next-token predictor over text. It learns the conditional distribution of the next symbol given the preceding symbols, trained on text written by people. Its competence is linguistic and, by extension, whatever reasoning is recoverable from text. It has no representation of physical law.
The structural models used here are learned solvers. Their input is a set of amino-acid chains; their output is a set of three-dimensional coordinates plus a calibrated confidence. What they have learned, from the Protein Data Bank, is the mapping from sequence to folded geometry — effectively an amortised approximation to the physics that a molecular-dynamics or free-energy calculation would compute far more slowly. The architecture is a deep network, and modern ones borrow the transformer and diffusion machinery that language models also use, but the object being modelled is a physical structure under an energy function, not a text distribution.
Three consequences follow, and they matter for how you read every number on this platform:
- The failure modes are physical, not linguistic. These models do not "hallucinate a fact"; they place a loop in the wrong position, or dock a binder onto a plausible-looking but incorrect surface patch. The error is geometric and is caught by geometric checks — pose repeatability, contact-footprint agreement, interface confidence — not by fact-checking.
- They come with a self-reported confidence, and it is meaningful. ipTM, pTM and pLDDT are trained to predict the model's own accuracy. That is a genuinely different situation from a language model's token probabilities, which are not calibrated to truth. Confidence gating is therefore a real tool here.
- They are interpolators over known fold space. Performance degrades on targets unlike anything in the training set. For antibody–antigen complexes specifically, the interface is driven by hypervariable loops with no evolutionary covariation signal, which is precisely the regime where co-folding is weakest. This is the single most important caveat on the entire structural stack.
A separate model family appears for affinity: protein language models. These are trained like language models — masked-token objectives over sequence databases — but they are used as featurizers, not generators. The embedding is a learned representation of a sequence; a small supervised head then maps it to a binding constant. The language-model part supplies the representation; the physics comes from the measured data it is regressed against.
And there is one genuine LLM in the codebase: a chat agent for querying the platform in natural language. Its status is covered honestly in §7 — it is not currently operational in production.
4. The model inventory
Grouped by the decision each model serves. Compute location is described by capability — self-hosted GPU, external API, or local CPU — rather than by supplier.
4.1 Structure and complex prediction — "does it bind, and where"
| Model | Role | Where it runs | Status |
|---|---|---|---|
| Boltz-2 (v2.2.1) | Antibody–antigen co-folding; produces the complex, ipTM/pTM/pLDDT, and multiple sampled poses | Self-hosted GPU (A100), weights cached in a persistent volume | Primary path. Default backend in production |
AlphaFold-Multimer (multimer_v3) | Same job, independent architecture and training | External API | Live, used as the automatic fallback and for cross-checking |
| OpenDDE | Antibody–antigen interface prediction, preview | Self-hosted GPU (A100-80GB) | Live, opt-in. Has its own entry in the workspace |
| AlphaFold 3 | Co-folding | Local HTTP service | Wired and enabled, but depends on an external service being up; it has failed jobs when that service was unreachable |
| ImmuneBuilder (ABodyBuilder2 / NanoBodyBuilder2) | Fast antibody-only structure, no antigen | External API | Live |
Boltz-2 deserves detail because it carries most of the load. It is a diffusion-based co-folding model: given the chains, it denoises toward a joint structure rather than threading onto a template. The platform runs it with a multiple-sequence-alignment front-end for the antigen, and falls back to single-sequence mode automatically when the alignment service is unavailable — recording that it did so, because the two modes are not equally trustworthy. It also supports structure-guided prediction: when the user uploads an experimental antigen structure, the selected chains are converted to a template that steers the prediction, so a known multimeric assembly is preserved instead of being re-folded from scratch. In that mode the platform deliberately refuses to silently fall back to sequence-only prediction, because that would quietly change the scientific question being asked.
4.2 Generative design — "propose a binder that does not exist yet"
All of these are reached through an external API rather than self-hosted.
| Model | What it generates |
|---|---|
| RFantibody (RFdiffusion → ProteinMPNN → RoseTTAFold2) | De-novo antibody / VHH CDR loops against a specified epitope. Diffusion proposes backbone geometry, inverse folding assigns a sequence to it, the folder checks the result |
| Germinal | De-novo VHH / scFv binders |
| BoltzGen | Generative binder design |
| AntiFold | Antibody-specific inverse folding — redesign CDR sequence on a fixed backbone |
| ProteinMPNN | General fixed-backbone inverse folding |
| IgDesign | Antigen-conditioned CDR design |
The pattern worth noticing is the design–validation loop: a generative model proposes, an inverse-folding model makes the proposal expressible as a real sequence, and an independent structure predictor checks whether that sequence actually folds back to the intended geometry. The check must use a different model from the one that generated the design, otherwise it only confirms the generator's own biases. This self-consistency criterion is the standard of evidence in de-novo protein design, and it is the reason the pipeline chains models rather than trusting any single one.
4.3 Affinity — "how tightly"
This is the part of the stack with the least reliable prior art, and the platform's design reflects that.
- ESM-2 3B embeddings + k-nearest-neighbour regression. The production default. The 3-billion-parameter protein language model produces a mean-pooled embedding per sequence; a k-NN regressor (k = 5, cosine distance, inverse-distance weighting) predicts log₁₀ K_D by interpolating among measured binders for the same target. Runs on self-hosted GPU. Observed Spearman ρ ≈ +0.77 on the internal BCMA panel.
- AlphaBind (in-house re-implementation). A transformer head over paired ESM-2 embeddings, fine-tuned per target on measured affinity data. Reaches ρ ≈ +0.88 on the same panel — better than k-NN, but only fires when a per-target checkpoint has been trained and staged, so k-NN remains the headline number.
- PRODIGY. Contact-based empirical free energy from a predicted pose — a physics-flavoured estimate rather than a learned one. Installed and usable.
- Nearest-measured-neighbour estimate. Not a model at all: sequence identity to the closest measured binder, gated at 85% identity, reported as an explicit estimate. It exists so the interface can show something honest for a novel antibody instead of a fabricated number.
Two hard-won lessons are encoded here. First, affinity prediction generalises poorly across targets — every good number above is target-conditional, trained on that target's own measured cohort. Second, a co-folding model's affinity head is not a general answer: the Boltz-2 affinity head only produces meaningful output for small-molecule ligands, not protein–protein interfaces, so it is deliberately excluded from the antibody scoring path.
4.4 Sequence annotation and developability
- ANARCI — IMGT numbering, framework and CDR boundary assignment. Installed and real; a regex-based approximation exists as a fallback but is not the active path in production.
- Germline assignment and VHH hallmark checking — curated reference constants and rules, no model.
- PTM liability scanning — motif rules (N-linked sequon, deamidation, isomerisation, oxidation) weighted by region, with CDR3 weighted most heavily.
- Humanness, aggregation, immunogenicity — external-API tools are wired; several local equivalents fall back to published heuristics when the underlying library is absent. §7 records which is which.
4.5 Algorithms that are not models
An important category, easy to overlook, and often the most reliable part of the stack because it has no training distribution to fall outside of:
- Contact extraction — heavy-atom distance cutoff at 4.5 Å, the SAbDab convention, giving epitope and paratope residue sets.
- Buried surface area — Shrake–Rupley solvent-accessible surface, computed on the complex and on each partner alone.
- pDockQ — a published function of interface pLDDT and contact count that converts confidence into a calibrated probability that the interface is correct.
- Pose-ensemble convergence — across the sampled poses from one run, Cα superposition RMSD plus contact-footprint Jaccard overlap. This measures whether the model is self-consistent, which is a genuinely different question from whether it is confident, and it catches a failure mode that confidence alone misses.
- Contact-fingerprint similarity between candidates — the basis of epitope binning without any new prediction.
5. How models compose into workflows
Individual models do not make decisions. Three composed workflows do.
5.1 Interface analysis — the workhorse
Input: antibody sequence (VHH, scFv, or Fab as separate chains) and an antigen given as sequence, accession, structure identifier, or an uploaded structure. Output: the predicted complex, interface metrics, epitope and paratope residue lists, and a reliability read.
The composition: co-folding model → contact extraction → interface metrics → pose-ensemble QC → confidence banding. The user gets a structure they can inspect and a set of numbers they can rank on. Batch submission runs a panel of candidates against one shared antigen so an entire campaign is one action.
5.2 The NGS triage funnel — the highest-leverage workflow
This is where computation buys the most. Phage or yeast display returns on the order of a million sequences; a wet lab can characterise about a hundred. The funnel closes that six-order-of-magnitude gap in stages, spending progressively more compute on progressively fewer candidates:
- Streaming dedup and liability filtering — CDR3 clustering, disk-backed, no GPU touched.
- Cheap proxy ordering — a transparent sequence heuristic, explicitly not an affinity model and labelled as such in the output, used only to decide what gets folded first.
- Structural scoring — co-folding on the survivors, with a hard budget cap so a web-triggered run cannot silently consume unbounded GPU.
- Panel matrix — epitope coverage against specificity and counter-screen targets, so the shortlist is diverse rather than a hundred variations of the same binder.
- Consensus ranking, validated against known binders — anchor recall as a check on the funnel itself.
The economic logic is explicit: an expensive model runs only where a cheap filter has already concentrated the odds.
5.3 The wet-lab calibration loop
Predictions are ranked, the top candidates go to binding measurement, and the measurements come back into the database. Those measurements then serve two purposes: they retrain the target-conditional affinity models, and they calibrate the mapping from raw score to predicted constant using standard isotonic and log-linear methods. Model selection is by rank correlation against measured values with pre-declared decision bands.
This loop is the only mechanism by which the platform's numbers become trustworthy for a given target. Without it, every affinity number is an extrapolation.
6. The rubric problem — a worked example
A concrete illustration of §2 step 4, because it is the part most teams get wrong.
A recent candidate campaign ranked binders with a rubric that converted interface metrics into a 0–100 score. Reconstructing that rubric from its own published results surfaced three findings:
- The documented formula did not match the arithmetic that actually produced the published numbers. The confidence term was a product of two independent lookups, not one; the sub-scores were three discrete levels rather than an interpolation.
- The rubric thresholded continuous, noisy quantities into discrete buckets. Half the candidates sat close enough to a threshold that a change in buried surface area smaller than the model's own pose-to-pose variation would move them a full bucket — meaning part of the published ranking was not reproducible.
- Its two "independent" dimensions correlated at r ≈ 0.999, because one was a monotone rescaling of the other. A weight intended to capture epitope coverage was in fact spent on a second copy of interface size.
None of these are model failures. The models did their job. The failure was entirely in the conversion layer — and it was invisible while that layer lived in a slide deck. The rubric now lives in versioned code, pinned by tests against the original published tables, so that changing it is a deliberate, reviewable act.
The general lesson: the decision rule deserves the same engineering rigour as the model. It usually does not get it.
7. Honest status
The distinction that matters is between implemented and reachable in production. This section records the difference.
Genuinely running in production:
- Self-hosted GPU co-folding, the default path, with automatic fallback to an independent external predictor.
- The preview interface predictor, as an opt-in backend.
- Self-hosted GPU affinity scoring, embeddings plus k-NN, with the fine-tuned head when a target checkpoint exists.
- The full interface analysis chain: contacts, buried surface area, pDockQ, pose-ensemble QC, confidence banding.
- IMGT numbering via the real library; contact-based free-energy estimation via the real library.
- The NGS triage funnel end to end, with budget caps on GPU spend.
- Codon optimisation for vector work — deterministic tables and simulated annealing, no ML, and the score is documented as a packaging-favourability proxy rather than a titre predictor.
Implemented but not currently operational:
- The natural-language agent. The provider client library is not installed in the production image and no API key is configured, so this feature does not work in production today. Its "run a design" tool never executed a pipeline in any case — it registered a task and wrote a metadata file.
- Local humanness, immunogenicity and aggregation tools. The underlying libraries are absent from the production image, so these endpoints return their documented heuristic approximations rather than the named methods. The output labels the method used, which is the reason this is visible at all.
- Structural superposition and structure-search tools used by the design-around workflow. The binaries are not present in the production image, so that path degrades to its sequence-only branch.
- A self-hosted design runner. The route exists; the runner raises "not implemented".
Present in the repository but not part of the running system:
- A large GPU-backend directory holding runners for a dozen models. None of the application code references it; its entry point is a file of TODOs.
- A set of tool-server integrations pointing at placeholder endpoints.
- Several affinity scorers that are adapters to libraries or services that are not vendored and not installed. Of nine registered scorers, two are live in production, two are real adapters to installed or licensed local tools, and the rest cannot run as shipped.
This section exists because a platform walkthrough that lists capabilities without this distinction is marketing. The value of writing it down is that the gap between the two lists is the actual roadmap.
8. Where the real limits are
Being clear about this is more useful than a feature list.
- Antibody–antigen co-folding is the hardest case for co-folding models. The interface is formed by hypervariable loops that carry no evolutionary covariation signal, which is exactly the information these architectures exploit elsewhere. Confidence metrics are correspondingly less reliable here than for a typical protein complex, and a high ipTM on an antibody complex should be read more sceptically than the same number on an enzyme–substrate pair.
- Affinity does not transfer across targets. Every trustworthy affinity number on the platform is conditional on measured data for that specific target. A new target starts with no such data, and the honest output in that situation is an explicit abstention, not a number.
- Rankings can be dominated by threshold placement rather than by biology, as §6 shows. The mitigation is to report score stability alongside the score.
- Predicted structures are single conformations of flexible objects. Nothing in the current stack models conformational ensembles, induced fit, or the entropic component of binding.
- The wet-lab loop is the rate limiter on trust, and it runs on the wet lab's clock, not on compute's.
9. The core logic, in one paragraph
The platform takes a decision that a wet lab cannot afford to make by brute force, decomposes it into physical observables, predicts those observables with deep-learning solvers that approximate the underlying physics far faster than simulation, converts the predictions into a ranking through an explicit and testable rubric, spends the scarce experimental budget on the top of that ranking, and feeds the resulting measurements back to calibrate the next cycle. The models are the interesting part; the composition, the honesty about confidence, and the calibration loop are the part that determines whether any of it is useful.