Articles & Research

Timeline

The reasoning race: OpenAI from GPT-4o to GPT-5

Three years in which the frontier moved from faster multimodal chat to models that spend inference time thinking.

Dr. Marisol Vega FACS

Director of Research, American Computer Society

September 2025 · 7 min read

Researcher reviewing neural network visualisations beside GPU server racks
Researcher reviewing neural network visualisations beside GPU server racks

Between May 2024 and August 2025 OpenAI shipped a multimodal flagship, invented a commercial product category around inference-time reasoning, and then folded that capability back into a single default model. For practitioners, the significant change is not capability but cost structure.

From omni to o-series

OpenAI launched GPT-4o on 13 May 2024, an 'omni' model handling text, vision and real-time audio in one network, and made it available to free ChatGPT users. The engineering story was latency: a single model that could take speech in and emit speech out removed the transcription and text-to-speech hops that had made voice assistants feel stilted.

Four months later, on 12 September 2024, the company previewed a different idea. o1-preview and o1-mini were trained to produce long internal chains of reasoning before answering. OpenAI reported that o1 reached roughly the 89th percentile on Codeforces competitive programming problems, performed at the level of a top-500 national qualifier on AIME mathematics, and exceeded a PhD-holder baseline on the GPQA science benchmark. The full o1 model followed on 5 December 2024, alongside a $200-per-month ChatGPT Pro tier — the first mainstream signal that reasoning capacity would be sold by the compute-hour rather than bundled.

“When the provider decides at inference time how much thinking your query deserves, capacity planning stops being a static exercise.”

o3, o4-mini and consolidation

OpenAI introduced o3 and o4-mini on 16 April 2025, extending reasoning models with tool use, and released o3-pro to Pro subscribers and the API on 10 June 2025. On 7 August 2025 the company launched GPT-5, describing it as a significant leap with 'built-in thinking' — a router-style consolidation in which one model decides how much deliberation a query warrants rather than requiring the user to choose a model family.

That consolidation matters more to system designers than any benchmark. When latency and cost per request depend on a routing decision the provider makes at inference time, capacity planning stops being a static exercise. Teams that had budgeted per-token now budget per-outcome, with variance they do not control.

  • 13 May 2024 — GPT-4o launched with real-time audio, vision and text.
  • 12 Sept 2024 — o1-preview and o1-mini introduce inference-time reasoning.
  • 5 Dec 2024 — full o1 ships with the $200/month ChatGPT Pro tier.
  • 16 Apr 2025 — o3 and o4-mini released; o3-pro follows 10 June 2025.
  • 7 Aug 2025 — GPT-5 launched with reasoning built into the default model.

What the Society advises

Members procuring these systems should insist on three things in contract: a documented statement of which model version served a given request, a stable evaluation harness the buyer controls rather than one supplied by the vendor, and a rollback path when a routed upgrade changes behaviour in a regulated workflow.

The reasoning turn is genuinely useful in domains where verification is cheaper than generation — code, proofs, structured extraction. It is least useful, and most expensive, where the ground truth is contested. Practitioners should be able to say which of those two situations they are in before they buy.

Artificial intelligence

Join the professional body behind this work

ACS members receive our research first, free CPD and ethics modules every year, and a route to professional registration assessed by their peers.

Become a member