


FrontierScience is a new benchmark designed to evaluate AI’s expert-level scientific reasoning across physics, chemistry, and biology. It measures both Olympiad-style problem solving and real research tasks, helping track how well advanced models can support and accelerate scientific work. The benchmark goes beyond theoretical questions by including wet lab experimental reasoning, as demonstrated in a recent study where GPT‑5 optimized a molecular cloning protocol, improving efficiency by 79x through a novel enzymatic mechanism.
FrontierScience spans physics, chemistry, and biology, testing both theoretical problem solving (e.g., Olympiad-style questions) and real research tasks. This ensures the benchmark captures a broad range of expert-level scientific thinking.
Unlike purely theoretical benchmarks, FrontierScience includes evaluations where AI models propose modifications to real laboratory protocols. In the cloning study, GPT‑5 autonomously reasoned about molecular biology steps, suggested enzyme combinations, and incorporated experimental data to iteratively improve outcomes.
The benchmark revealed that GPT‑5 could introduce a previously unreported enzymatic mechanism—RecA-Assisted Pair-and-Finish HiFi Assembly (RAPF) combined with T7 transformation—that boosted cloning efficiency by 79x. This demonstrates the model’s ability to surface non-obvious, experimentally valid solutions.
All wet lab work was conducted in a tightly controlled setting using a benign experimental system. The results feed directly into OpenAI’s Preparedness Framework, helping assess and mitigate risks associated with advanced biological reasoning capabilities.
FrontierScience doesn’t just test what AI knows—it tests whether AI can invent new science in the lab.
Most benchmarks stop at multiple-choice or written answers. FrontierScience goes further by requiring models to propose actionable experimental modifications, then validates those ideas through actual wet lab results. The 79x efficiency gain from a novel enzymatic pathway shows that AI can contribute original, empirically sound insights to biological research—not just summarize existing knowledge.
You’re tracking how close AI is to becoming a genuine research collaborator in the life sciences, or if you need a benchmark that captures both theoretical rigor and practical experimental reasoning. FrontierScience is especially relevant for teams working on AI safety, biosecurity, or the acceleration of drug discovery and protein engineering.
Other tools you might consider
Loading comments…
Maker
async_apple
Visit Website
openai.com/index/accelerating-biological-research-in-the-wet-lab/
Project Info
Product Keywords
Achievement