Overview
Protein design is beginning to look less like an open-ended biological search and more like computer-aided engineering. Chai Discovery is pursuing this transition as a neutral software and modeling platform rather than developing its own drug pipeline. Its progression from Chai 1, a structure-prediction model, to Chai 2 and Chai 3, which generate candidate molecules under structural and therapeutic constraints, reflects the field’s movement from observing proteins to intentionally designing them. The company reports designing antibodies against 50 targets, obtaining binders for roughly half, and achieving an average binding hit rate of about 20% in that study. Its product resembles CAD or Figma more than a chatbot, allowing scientists to select epitopes, specify desired and avoided interactions, inspect structures, and iterate using laboratory results. The largest remaining obstacles include slow experimental validation, scarce labeled biological data, compute infrastructure optimized primarily for language models, and the difficulty of identifying the correct biological target or epitope. Chai expects models to produce increasingly drug-like candidates directly, while its software rises through successive levels of abstraction—from molecule inspection to campaign orchestration and potentially the outer loop of scientific discovery.
Sections
Strategic Insights
Higher-level implications of Chai’s model, product, and partnership strategy.
- The real transition is not simply from slower discovery to faster discovery; it is from probabilistic search toward declarative molecular engineering. That distinction explains why the company emphasizes precise epitopes, selectivity, and new modalities rather than cost savings alone.
- Product interfaces function as part of the scientific control system. A CAD-like environment translates expert intent into spatial constraints, limits invalid requests, and provides the inspection layer required before scientists trust generated molecules.
- A pure platform strategy exchanges ownership of individual drug assets for broader learning across partner portfolios. Its defensibility depends on frontier models, workflow integration, security, and accumulated knowledge of what pharmaceutical teams actually need—not necessarily access to unrestricted partner data.
- As implementation becomes cheaper through AI, attention and experimental capacity become the scarce resources. Chai therefore treats researchers, engineers, and pharmaceutical companies as capital allocators choosing which ideas, compute runs, targets, and validation experiments deserve limited shots on goal.
Key Comparisons
Explicit contrasts between models, discovery methods, workflows, and business strategies.
- Chai 1 predicts the structure produced by a known sequence, whereas Chai 2 generates candidate sequences and structures intended to bind a supplied target.
- Traditional screening searches enormous libraries for rare binders with limited control over binding pose; model-guided design attempts to specify the interaction region, desired targets, and avoided targets before generating candidates.
- A conventional waterfall pipeline separates discovery and optimization into long, gated stages, while the proposed model-driven workflow repeatedly incorporates laboratory feedback into new generations.
- Open or commoditized models may handle simpler modalities, while frontier closed models and specialized products are expected to retain value on harder drug-design tasks.
- Owning a proprietary drug pipeline concentrates economics and data around selected assets; Chai’s partnership model instead seeks reusable improvements across many pharmaceutical portfolios.
Technical Details
Specific architectural, computational, and validation mechanisms described in the interview.
- Chai 1 combines molecular tokenization, a transformer-like trunk, and a diffusion component that emits three-dimensional atomic coordinates.
- Inputs are multimodal: sequence-level tokens represent molecular units, while attached atom attributes include element types, charges, and coordinates.
- Multiple sequence alignments expose conserved positions and correlated mutations, providing evolutionary signals about which residues may be close in three-dimensional space.
- Chai 2 is described as an all-atom diffusion model that can select and place atoms, map them back to amino-acid identities, and jointly refine sequence and structure.
- Computational validation compares a generated sequence against an independent structure-prediction model, then considers prediction confidence and generation diversity to detect self-consistent but collapsed outputs.
- Fold-style pair representations behave roughly like sequences of length L squared; attention over them can produce approximately L-cubed computation and substantial memory-bandwidth costs.
- Chai uses Temporal for durable execution of long-running database, model, and data-pipeline work, including queueing, retries, and orchestration across flaky infrastructure.
- Customer isolation is implemented through aggressive data segmentation and near-single-tenant deployments, reflecting pharmaceutical companies’ sensitivity to intellectual property.
Predictions
Forecasts made or strongly implied by the speakers.
- Models will increasingly generate candidates that are already close to therapeutic grade, reducing the practical distinction between hit discovery and lead optimization.
- Protein-design products will be repeatedly rebuilt at higher levels of abstraction, moving from molecular inspection to campaign orchestration and eventually coordination across biological pathways.
- Simpler protein-design tasks may become commoditized, while frontier systems retain value by addressing harder modalities and offering integrated scientific workflows.
- Faster computational design could bend the worsening economics of pharmaceutical research and allow companies to pursue riskier, more ambitious targets across broader portfolios.
- AI-biology models will become comparable to language models in impact and scale, creating demand for specialized compute hardware, inference optimization, and software stacks.
- Generalist machine-learning researchers and software engineers will play a larger role in biology as computational tools make the field more visual, accessible, and engineerable.