AI news, models and products, with sources中文
NewsLandmark Models7 Aug 2025

OpenAI releases GPT-5

GPT-5 combines a fast model and a reasoning model with an automatic router, and becomes the default for free ChatGPT users. The launch draws complaints when the router malfunctions and older models are removed; OpenAI restores GPT-4o for paying users.

OpenAI chief executive Sam Altman on stage at TED
OpenAI chief executive Sam Altman at TED in April 2025, four months before GPT-5 launched. File photo. Photo: Steve Jurvetson / Wikimedia Commons, CC BY 2.0

What happened

On Thursday 7 August 2025, in a livestream, OpenAI released GPT-5, the successor to GPT-4 and the model behind its next generation of ChatGPT. It was one of the company’s most anticipated launches since ChatGPT itself. At the time, OpenAI said ChatGPT had more than 700 million weekly users.

The most visible change was that ChatGPT got simpler. Before GPT-5, users had to choose between several models with confusing names, some fast and some slower but more careful. GPT-5 makes that choice for them.

Chief executive Sam Altman called it “the best model in the world” and compared it to having “a team of Ph.D. level experts in your pocket”. He also said it was not yet AGI, because it still cannot keep learning on its own without being retrained.

Background

GPT-5 took longer than expected. In 2023 Altman said OpenAI was not yet training it. According to The Information, the company spent much of the second half of 2024 on a model meant to become GPT-5, code-named Orion, but it did not turn out much better and was released in February 2025 as GPT-4.5 instead.

In the meantime OpenAI’s model menu had grown crowded: GPT-4o, o3, o4-mini and more. Altman had criticised that menu as too complicated. GPT-5 was meant to replace it with a single default. Two days earlier, on 5 August, OpenAI had also released gpt-oss, free models anyone can download and run.

What changed for users

  • Free users got a reasoning model. GPT-5 became the default for all free ChatGPT users. OpenAI’s head of ChatGPT, Nick Turley, said this was the first time free users had access to a reasoning model. Free accounts have lower usage limits, and drop to a smaller “mini” version when they hit them.
  • Paid plans. Plus subscribers ($20 a month) got higher limits. Pro subscribers ($200 a month) got unlimited GPT-5 and GPT-5 Pro, a version that uses extra computing power for harder problems. Team, Edu and Enterprise plans followed the next week.
  • New options. Users could pick one of four preset personalities (Cynic, Robot, Listener and Nerd), choose a colour for chats, and connect ChatGPT to Gmail, Google Calendar and Google Contacts.
  • Beyond ChatGPT. GPT-5 also came to Microsoft Copilot and to developers through OpenAI’s API, in three sizes (gpt-5, gpt-5-mini and gpt-5-nano). The full model cost developers $1.25 per million input tokens and $10 per million output tokens.

OpenAI’s examples of what it could do included building a working app or website from a plain description, which Altman called “software on demand” (often called “vibe coding”), as well as handling calendar tasks and writing research briefs.

How it compared

OpenAI’s headline claim was coding. On SWE-bench Verified, a test built from real bug fixes in public code on GitHub, GPT-5 solved 74.9% of tasks on its first attempt. That was only just ahead of Anthropic’s Claude Opus 4.1, released two days earlier.

SWE-bench Verified: fixing real bugs in public codeShare of tasks solved; higher is better
  • GPT-574.9%
  • Claude Opus 4.174.5%
  • Gemini 2.5 Pro59.6%

Each figure was published by the company that makes the model; test set-ups are not necessarily identical. Source: TechCrunch, “OpenAI’s GPT-5 is here”, 7 Aug 2025

The picture elsewhere was mixed, as TechCrunch noted:

  • On GPQA Diamond, PhD-level science questions, GPT-5 Pro scored 89.4%, ahead of Grok 4 Heavy (88.9%) and Claude Opus 4.1 (80.9%).
  • On Humanity’s Last Exam, very hard questions across many subjects, GPT-5 Pro scored 42% with tools, slightly below xAI’s Grok 4 Heavy at 44.4%.
  • On Tau-bench, which simulates tasks on an airline’s or a shop’s website, GPT-5 scored slightly below o3 on the airline part and slightly below Claude Opus 4.1 on the retail part.

Fewer made-up answers

AI models sometimes state false things confidently, which is called hallucination. The problem had seemed to be getting worse in OpenAI’s recent reasoning models such as o3, which OpenAI had said it did not fully understand. The company said GPT-5 was much better. On a set of real ChatGPT questions, OpenAI said GPT-5 with thinking gave an answer containing a factual error 4.8% of the time.

Responses with factual errors, on real ChatGPT questionsLower is better
  • GPT-5 (thinking)4.8%
  • GPT-4o20.6%
  • o322%

OpenAI's own test and figures. Source: TechCrunch, “OpenAI’s GPT-5 is here”, 7 Aug 2025

On health questions, OpenAI said GPT-5 was more careful to flag possible concerns and help people understand medical results. On its HealthBench Hard Hallucinations test it reported an error rate of 1.6%, against 12.9% for GPT-4o and 15.8% for o3.

Safety changes

OpenAI’s safety research lead, Alex Beutel, said GPT-5 went through 5,000 hours of expert-led safety testing and was less likely to deceive users, for example by claiming to have finished a task it had not finished. OpenAI also changed how the model handles risky questions: instead of simply answering or refusing, it tries to give the most helpful answer that stays safe, an approach OpenAI calls “safe completions”. The aim, Beutel said, was to refuse more genuinely unsafe requests while turning away fewer harmless ones.

Outside testers were less impressed. Within a day, the security firm NeuralTrust said it had got GPT-5 to give detailed instructions for making explosive devices, and another firm, SPLX, reported similar results.

A rocky first week

How the launch unfolded
  1. 17 August: launchGPT-5 replaces the old model menu in ChatGPT. For most users, older models such as GPT-4o disappear at once.
  2. 2Day one: the router breaksUsers report that GPT-5 sometimes seems worse than GPT-4o. Some spot basic slips, such as miscounting letters in a word or misspelling US states on a map.
  3. 38 August: Altman explainsHe says the automatic switcher “broke and was out of commission for a chunk of the day”, so GPT-5 “seemed way dumber”.
  4. 4Backlash over GPT-4oUsers who relied on particular models, or preferred GPT-4o's warmer tone, object to losing them without warning. Altman says Plus users can choose GPT-4o again.
  5. 515 August: a warmer GPT-5After Altman said OpenAI was working to make the model “feel warmer”, a personality update rolls out.

Sources: Wikipedia (GPT-5), The Guardian, 8 Aug 2025

Some users described GPT-5’s tone as “flat” and “uncreative”. Altman wrote on X: “We for sure underestimated how much some of the things that people like in GPT-4o matter to them, even if GPT-5 performs better in most ways.” He said the episode showed that people need better ways to customise the assistant, because no single model suits everyone.

Reactions

Reviews were respectful rather than excited:

  • MIT Technology Review called GPT-5 “above all else, a refined product”. If it was a step towards AGI, it said, it was “a very small step”.
  • The Atlantic described it as “intuitive, fast, and efficient”, and argued that at this stage how a chatbot feels matters more than benchmarks.
  • New York magazine wrote that casual users were unlikely to feel they were using a completely different product, while people using it for software or at work were more likely to notice a change.
  • Ars Technica found GPT-4o tended to give a little more detail and be more personable than the more direct, concise GPT-5.
  • Mashable called its ability to build custom interactive apps from a simple description its “coolest feature by far”.

Altman also drew criticism for raising expectations before launch, for example by comparing GPT-5’s development to the Manhattan Project.

Things to keep in mind

  • The scores are the companies’ own. OpenAI’s numbers come from its launch materials; rival scores were published by Anthropic, Google and xAI under their own test conditions. Leads of a fraction of a point, as on SWE-bench Verified, are within the range where test set-ups matter.
  • “PhD-level” is a marketing phrase. In the same week, users showed GPT-5 miscounting letters in words such as “blueberry”. Part of the problem was the router sending questions to the faster model.
  • Energy use was not disclosed. OpenAI did not publish how much energy GPT-5 uses. Researchers at the University of Rhode Island estimated a medium-length answer at slightly over 18 watt-hours.
  • This is history now. GPT-5 has since been followed by many updates, including GPT-5.6 in July 2026 and the GPT-6 family from September 2026 (see GPT-6 Astra).

What it meant for ordinary users

For most people, GPT-5 changed ChatGPT in two practical ways. Free users got a model that could “think” through harder questions without choosing anything. And everyone lost the ability to pick older models, until protests brought GPT-4o back for paying users. It was also a lesson for OpenAI: many people had grown attached to how a particular model talks, not just to how well it scores.

For today’s ChatGPT plans and models, see our ChatGPT product page and the Models page.

Model summary

Developer
OpenAI
Availability
ChatGPT (including free users), Microsoft Copilot, OpenAI API

Reported benchmark results

SWE-bench Verified74.9%Claude Opus 4.1: 74.5%
AIME 2025 (no tools)94.6%
GPQA Diamond89.4%GPT-5 pro
Hallucination rate on ChatGPT prompts4.8%With thinking; o3: 22%

Data source: OpenAI figures as reported by TechCrunch. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.

Sources