The GPT-5.4 launch marks OpenAI’s latest step in frontier model development, bringing together reasoning, coding, and agentic workflows into a single model released on February 12, 2026. The model is available in ChatGPT as GPT-5.4 Thinking, in the API, and in Codex. OpenAI also released GPT-5.4 Pro for users requiring maximum performance on complex tasks.
GPT-5.4 Launch: Key Capabilities and Benchmarks
On GDPval, a benchmark testing agents across 44 occupations, GPT-5.4 matches or exceeds industry professionals in 83.0% of comparisons. That figure compares to 70.9% for GPT-5.2. On an internal spreadsheet modeling benchmark, GPT-5.4 achieves a mean score of 87.3%, versus 68.4% for GPT-5.2. Human raters preferred GPT-5.4 presentations 68.0% of the time over those from GPT-5.2.
OpenAI said GPT-5.4 is its most factual model to date. Individual claims are 33% less likely to be false, and full responses are 18% less likely to contain errors, relative to GPT-5.2. Brendan Foody, CEO at Mercor, said GPT-5.4 “excels at creating long-horizon deliverables such as slide decks, financial models, and legal analysis, delivering top performance while running faster and at a lower cost than competitive frontier models.”
Native Computer-Use and Agentic Workflows
GPT-5.4 is OpenAI’s first general-purpose model with native computer-use capabilities, enabling artificial intelligence agents to operate computers and carry out complex workflows across applications. On OSWorld-Verified, the model achieves a 75.0% success rate, surpassing human performance at 72.4% and GPT-5.2’s 47.3%. On WebArena-Verified, it achieves a 67.3% success rate, compared to GPT-5.2’s 65.4%.
The model supports up to 1 million tokens of context, allowing agents to plan, execute, and verify tasks across long horizons. Dod Fraser, CEO at Mainstay, said GPT-5.4 achieved a 95% success rate on the first attempt across approximately 30,000 HOA and property tax portals, completing sessions roughly three times faster while using approximately 70% fewer tokens.
Tool Search and Efficiency Improvements
GPT-5.4 introduces tool search in the API, allowing models to retrieve tool definitions on demand rather than loading all definitions upfront. In testing across 250 tasks from Scale’s MCP Atlas benchmark with all 36 MCP servers enabled, the tool-search configuration reduced total token usage by 47% while maintaining the same accuracy. This approach reduces cost and latency for tool-heavy workflows.
On BrowseComp, which measures persistent web browsing to find hard-to-locate information, GPT-5.4 improves 17 percentage points over GPT-5.2. GPT-5.4 Pro sets a new benchmark result of 89.3% on that test. Wade, CEO at Zapier, said GPT-5.4 “finished the job where previous models gave up,” describing it as the most persistent model tested across hundreds of advanced real-world workflows.
Coding Performance and Codex Integration
GPT-5.4 incorporates the coding capabilities of GPT-5.3-Codex and matches or outperforms it on SWE-Bench Pro while offering lower latency across reasoning efforts. In Codex, a fast mode delivers up to 1.5x faster token velocity using the same model. OpenAI also released an experimental Codex skill called Playwright (Interactive), enabling visual debugging of web and Electron applications.
Lee Robinson, VP of Developer Education at Cursor, said engineers find GPT-5.4 “more natural and assertive than previous models,” noting it works through ambiguous problems without second-guessing itself and parallelizes work proactively. Niko Grupen, Head of Applied Research at Harvey, said the model scored 91% on the BigLaw Bench evaluation for document-heavy legal work.
Safety Measures and Deployment Details
OpenAI classifies GPT-5.4 as High cyber capability under its Preparedness Framework, deploying it with an expanded cybersecurity stack that includes monitoring systems, trusted access controls, and asynchronous blocking for higher-risk requests. The company also introduced a new open-source evaluation called CoT controllability, measuring whether models can obfuscate their reasoning to evade monitoring. OpenAI said GPT-5.4 Thinking’s ability to control its chain-of-thought is low, which it described as a positive safety property.
In ChatGPT, GPT-5.4 Thinking is available to Plus, Team, and Pro users starting today, replacing GPT-5.2 Thinking. GPT-5.2 Thinking remains available for three months under Legacy Models before retirement on June 5, 2026. Enterprise and Edu plan users can enable early access via admin settings. In the API, GPT-5.4 is available as gpt-5.4, and GPT-5.4 Pro as gpt-5.4-pro. Batch and Flex pricing are available at half the standard API rate, while Priority processing is available at twice the standard rate.
“GPT-5.4 sets a new bar for document-heavy legal work. On our BigLaw Bench eval, it scored 91%. Compared to other models, GPT-5.4 is currently better at structuring complex transactional analysis, maintaining accuracy across lengthy contracts, and delivering the high level of detail legal practitioners require.”
Niko Grupen, Head of Applied Research at Harvey

