Semi-de-novo Triage
Funnel a wet-lab NGS pool down to a ranked shortlist. The lab supplies the diversity; the platform ranks and resolves epitope / specificity.
Workflow
The wet lab supplies the diversity; the platform triages ~10⁶ sequences to a ~10² shortlist.
01Original plan
- 1Library enrichment (phage-display panning + functional sort) → antibody pool
- 2NGS readout → ~10⁶ sequences (with abundance)
- 3Dock each clone against the target and score
- 4Filter by structural domain / sequence
- 5Classify against different antigens (epitope / specificity)
- 6Composite evaluation → best ranking → list usable at the bench
02As implemented
- 1Pool sequences1.2Mthe whole enriched library, streamed
- 2Sequence filter4,000length / liability gates, CDR3 clustering (cheap, full 10⁶)
- 3Fold + affinity400dock survivors only; score pose (iptm + ΔG→Kd)
- 4Epitope × specificitymatrix vs the antigen panel
- 5Composite rank + anchor checkknown positives validate, never train
- 6Shortlist → expression100median predicted Kd ~0.7 nM
03Where the two differ, and why
- Docking all 10⁶ is infeasible → two tiers: a cheap sequence filter first, fold only the survivors.
- Two near-independent signals (interface confidence + predicted ΔG) ranked by consensus — single-metric top-10s shared nothing.
- Abundance is a weak tiebreak, never a gate — rare-but-strong binders still surface.
- Validated on the DLL3 known positives (21/21 recovered) before trusting it on unlabeled pools.
- Predicted, not measured — wet-lab KD is the truth; specificity needs counter-screen constructs.