Articles & Research

Analysis

Anthropic's Claude and the agentic turn in software engineering

Sonnet, Opus and the shift from autocomplete to long-running coding agents.

Elena Marsh MACS

Head of Professional Standards, American Computer Society

June 2025 · 6 min read

Students working on laptops and a single-board computer in a computing classroom
Students working on laptops and a single-board computer in a computing classroom

Anthropic's model line moved from a three-tier naming scheme to a claim that a general-purpose model can hold a software task open for hours. That claim, if it holds, changes what engineering supervision means.

A three-tier line

Claude was first released as a chatbot in March 2023. Since the Claude 3 generation, Anthropic has used a consistent three-tier naming convention — Haiku, Sonnet and Opus — to separate cost-optimised, balanced and maximum-capability models within a single family.

On 21 June 2024 the company launched Claude 3.5 Sonnet, stating that the mid-tier model outperformed the previous top-tier Claude 3 Opus while costing considerably less to run. That inversion, a mid-tier model beating the prior flagship, has become the ordinary rhythm of the field and is worth naming: capability at a fixed price point improves faster than capability at the frontier.

“A pull request produced over four hours of autonomous tool use cannot be reviewed the way a twenty-line suggestion is reviewed.”

Opus 4 and long-running tasks

On 22 May 2025 Anthropic introduced Claude Opus 4 and Claude Sonnet 4, describing Opus 4 as the world's best coding model and emphasising sustained performance on long-running agentic tasks rather than single-turn responses. Sonnet 4 was presented as a direct upgrade to the preceding Sonnet release.

The engineering significance of a long-running agent is not that it writes more code. It is that the unit of review changes. A pull request produced over four hours of autonomous tool use cannot be reviewed the way a twenty-line suggestion is reviewed. The reviewer is no longer checking a diff against an intention they hold in their head; they are auditing a process they did not observe.

Supervision as a competence

The Society's position is that supervision of autonomous coding agents is a distinct professional competence and should be assessed as one. It requires the ability to construct adversarial test cases before the agent runs, to read execution traces rather than only outputs, and to identify the class of failure — specification, tool, or model — when a run goes wrong.

Organisations adopting these tools should record which changes to production systems originated from an agent, who approved them, and on what evidence. This is not bureaucratic caution. It is the minimum record needed to investigate an incident six months later.

  • March 2023 — Claude first released as a chatbot.
  • 21 June 2024 — Claude 3.5 Sonnet outperforms the prior Opus tier at lower cost.
  • 22 May 2025 — Claude Opus 4 and Sonnet 4 launched, emphasising long-running agentic coding.
Artificial intelligenceSkills and education

Join the professional body behind this work

ACS members receive our research first, free CPD and ethics modules every year, and a route to professional registration assessed by their peers.

Become a member