César de la Fuente's lab at the University of Pennsylvania is using OpenAI's Codex and ChatGPT to search living and extinct genomes for new antimicrobial molecules, part of a growing effort to counter drug-resistant infections with computational biology. The lab treats the models as working infrastructure — Codex writes and iterates the bioinformatics pipelines, while ChatGPT reasons over sequences and helps rank candidates worth synthesizing in the wet lab. OpenAI disclosed the collaboration in a research post highlighting how frontier models are being folded into active drug-discovery workflows.
The pitch is straightforward. Antibiotic-resistant infections are one of the fastest-moving public-health problems on the planet, and the pharmaceutical pipeline for new antibiotics has been thin for two decades because the economics of the category are punishing. Computational screens that widen the search space — and cut the time from hypothesis to candidate — are one of the few levers researchers have.
De la Fuente's lab has spent years mining the proteomes of living organisms for encrypted peptides with antimicrobial activity, and more recently extended the search to extinct species, including molecules reconstructed from Neanderthal and mammoth sequences. The addition of large language models changes the tempo. Codex is being used to generate and refine the pipelines that filter billions of candidate sequences down to a tractable shortlist, and ChatGPT handles the reasoning layer on top — explaining why a candidate scored the way it did, flagging structural motifs, and helping researchers decide what to synthesize next.
Key facts
- 01César de la Fuente's lab is applying OpenAI's Codex and ChatGPT to scan genomes for new antimicrobial molecules.
- 02The search covers both living species and extinct organisms, expanding the addressable sequence space beyond conventional screens.
- 03The target is drug-resistant infections, a category the WHO ranks among the top global health threats.
- 04The workflow pairs Codex for pipeline code with ChatGPT for reasoning over biological sequences and candidate ranking.
This is a different use case from the one that dominates coverage of generative models in biology. Rather than training a bespoke protein-design model from scratch, the lab is composing off-the-shelf frontier tools around a domain-specific pipeline. The advantage is speed: a researcher who can describe the filter they want in English gets working code back in minutes, not days, and can iterate the screen as new hypotheses emerge.
The extinct-genome angle is the most distinctive piece of the work. Sequences from species that no longer exist expand the biological search space beyond anything evolution is currently producing, and some of the peptides recovered from ancient DNA have shown activity against modern pathogens in laboratory assays. Language models help make that search tractable by handling the annotation, comparison, and prioritization steps that would otherwise consume most of a graduate student's week.
OpenAI's framing is that Codex and ChatGPT are being used as general-purpose scientific assistants rather than as domain-specific models. That distinction matters for the broader question of how much of frontier research can be accelerated by tools built for software engineers and knowledge workers. If a lab hunting antibiotics can get meaningful lift from the same models a startup uses to ship a web app, the diffusion curve for AI in the sciences is steeper than the specialist-tools narrative suggests.
The lab's earlier work has produced antimicrobial candidates that advanced to preclinical testing, and the group has published on machine-learning-guided peptide discovery in journals including Cell and Nature Biomedical Engineering. Adding Codex and ChatGPT to that stack is an incremental change in method, not a change in target — the endpoint is still a molecule that kills a resistant bacterium in an animal model and, eventually, a patient.
The caveats are the ones that apply to any AI-assisted discovery workflow. Candidates that score well computationally still have to survive synthesis, in-vitro assays, toxicity screens, and animal studies before anyone talks about a clinical program, and the attrition at each stage is steep. Language models can hallucinate biological plausibility as easily as they hallucinate citations, which is why the lab keeps a wet-lab loop tight against the computational one. Nothing about the tooling shortcuts the years of validation that separate a promising sequence from a drug.
For OpenAI, the story is a useful data point in an ongoing argument that Codex and ChatGPT belong inside serious research workflows, not just consumer productivity ones. For the AI-in-biology market, it reinforces a pattern already visible at Isomorphic Labs, Recursion, and a handful of academic groups: the interesting work is happening where general-purpose models meet purpose-built pipelines, and the labs moving fastest are the ones treating frontier LLMs as core infrastructure rather than as a demo. Antibiotics are one of the few categories where a genuine hit would matter to almost everyone alive, which makes it one of the more consequential places to watch that pattern play out.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



