
Image: @OpenAIDevs (X) (video still)
Summary
- OpenAI took its life sciences model GPT-Rosalind out of research preview on September 11 and opened it globally to eligible organizations.
- In a public demo the model ranked three asthma targets and went on to draw a 96-well perturbation assay platemap.
- Who gets to use it is decided by the company's own review, and API access is permitted only for approved internal research.
On September 11 a 56-second screen recording went up on OpenAI's developer account. The window is titled Compare IL33 vs TSLP vs IL1RL1, with asthma program beside it. The work is choosing which of three asthma drug targets to push first. At the bottom right of the screen sit the model name GPT-Rosalind and a reasoning effort set to Extra High.
In that video, which METAL watched in full, the model worked for 9 minutes and 25 seconds before producing a ranking. Drawing on six independent public-evidence lanes from the Life Science: Research plugin and reconciling them against a local package of internal material, it concluded that the overall order put TSLP ahead of IL33, and IL33 ahead of IL1RL1. It attached its reasons down to the line numbers of the evidence. TSLP leads because it is the only pathway here with a record of asthma approval; IL33 has the strongest human-genetics signal but slips back because its clinical precedent is still at phase-2 level. The model itself noted that the one real disagreement among the six lanes was the ordering of IL33 and IL1RL1.
Up to that point this is the work of reading literature and tidying it up. The weight of the video lies in the instruction that follows. The person on screen tells it to design an experiment that tests whether TSLP deserves to stay ahead of IL33. A 96-well perturbation assay co-culturing human airway epithelium with immune cells, several timepoints, anti-TSLP as the primary arm and anti-IL33 as the secondary, controls included: the conditions are thrown out in plain speech. The model wrote one new Python file, then made two more, and drew a 96-well platemap that reads like a seating chart.
The values written on that platemap say what kind of announcement this is. Ninety-six wells, three timepoints at 6, 24 and 48 hours, eight conditions, four technical replicates per condition and timepoint, one plate per donor. The stimulus and two kinds of control are specified as well, and the primary endpoint is IL-13 inhibition between 24 and 48 hours measured against the stimulated control. It is a plan a lab could pin to the bench as it stands, and the narration in the video called this going beyond hypothesis generation to actually building a feedback loop at the bench.
And that narration carries the single most important conditional clause in this announcement. It says that "with the lifted biosafety restrictions, the Life Sciences model is able to generate novel hypotheses, design experiments, and optimize existing protocols for drug discovery research." Restrictions being lifted does not mean lifted for anyone. It means the company decides, by review, for whom they are lifted.
According to OpenAI's announcement, GPT-Rosalind came out of research preview on September 11 and is now open globally to eligible organizations through its trusted-access program. Published pricing takes effect on October 5, and customers who clear the bar can use the model in ChatGPT, Codex and the API. It was first introduced on April 16, and the name comes from Rosalind Franklin, whose work left the evidence that proved decisive in revealing the structure of DNA. The company put up front in its announcement the fact that a new drug in the United States normally takes 10 to 15 years to travel from target discovery to regulatory approval.
A scorecard came with it. By the product page, performance per token rose 53.7% on Genebench, 19.6% on Labworkbench, 18.0% on Medchem Bench and 4.4% on LifeSci Bench. On BixBench, which covers real bioinformatics analysis, it led among models with published scores, and on LABBench2, which measures literature retrieval, database access, sequence manipulation and protocol design, the company said it beat GPT-5.4 on six of eleven tasks. The biggest gain came on the task of designing DNA and enzyme reagents end to end for molecular cloning.
There is also a result from going head to head with people. Together with Dyno Therapeutics, which builds AI-designed gene therapies, OpenAI tested the model on linking RNA sequence to function using unpublished sequences that could not have leaked into training. Compared against 57 scores left behind by experts in the field, the best of ten submissions made in the Codex app landed inside the top 5% on the prediction task and around the top 16% on sequence generation. Being inside the top 5% means performing as well as one in twenty experts in this field.
Read through the lens of institutions, what changed here is not capability but the threshold. OpenAI wrote that participating organizations must be conducting legitimate scientific research with clear public benefit, must maintain governance and compliance controls that prevent misuse, and must restrict access to approved users within secure, well-managed environments. Trusted access starts with qualified enterprise customers in the United States. The help documentation says API access is permitted only for approved internal research, and cannot be used in customer-facing products or external commercial applications. The capability is open worldwide, but eligibility to use that capability is set by the company's own review sheet.
The list of names gives a sense of who has cleared that sheet. Amgen, Moderna, Novo Nordisk, Thermo Fisher Scientific, Oracle Health and Life Sciences, NVIDIA, the Allen Institute, Benchling and the UCSF School of Pharmacy appear among the collaborators. McKinsey, Boston Consulting Group and Bain are attached as advisory partners helping with adoption. Sean Bruich, Senior Vice President of Artificial Intelligence and Data at Amgen, said that "the life sciences field demands precision at every step. The questions are highly complex, the data are highly unique, and the stakes are incredibly high." The company said it is also exploring AI-guided protein and catalyst design with Los Alamos National Laboratory.
The same threshold works in the opposite direction too. OpenAI has set up a separate Rosalind Biodefense program for developers and public-health teams building early detection, preparedness, diagnostics, response and medical countermeasures. It is a design that raises biological capability while allocating that capability to defense first. What becomes clear here is that eligibility review is at once the lock on the door of capability and the mechanism that distributes priority through it. Who earns the right to design an experiment is now a question of review, not of technology.
One piece is open at no cost. The Life Sciences research plugin for Codex is published on GitHub, a bundle of skills connecting more than 50 public multi-omics databases, literature sources and biology tools. Enterprise users who clear the bar use the plugin with GPT-Rosalind, while everyone else can use it with OpenAI's mainline models. The model above the threshold goes to those who pass review; the plumbing below it is left open to all.
What remains is the distance between a plan and a result. METAL has set out reports of AI-designed proteins where prediction and actual outcome diverged. However precise a 96-well platemap is, filling that plate and waiting 48 hours is still a matter of human hands and human time. What OpenAI sold this time is not an answer but a plan for testing one, and the worth of that plan is only ever scored at the bench.
So the real content of this announcement is not model performance. It is that the authority to decide what AI may do in biology has taken its seat on a model company's review sheet ahead of any regulator. The capability is already open worldwide; the only door left is that review sheet.





Comments