Analysis
Anthropic's Claude and the agentic turn in software engineering
Sonnet, Opus and the shift from autocomplete to long-running coding agents.
Elena Marsh MACS
Head of Professional Standards, American Computer Society
June 2025 · 6 min read

Anthropic's model line moved from a three-tier naming scheme to a claim that a general-purpose model can hold a software task open for hours. That claim, if it holds, changes what engineering supervision means.
A three-tier line
Claude was first released as a chatbot in March 2023. Since the Claude 3 generation, Anthropic has used a consistent three-tier naming convention — Haiku, Sonnet and Opus — to separate cost-optimised, balanced and maximum-capability models within a single family.
On 21 June 2024 the company launched Claude 3.5 Sonnet, stating that the mid-tier model outperformed the previous top-tier Claude 3 Opus while costing considerably less to run. That inversion, a mid-tier model beating the prior flagship, has become the ordinary rhythm of the field and is worth naming: capability at a fixed price point improves faster than capability at the frontier.
“A pull request produced over four hours of autonomous tool use cannot be reviewed the way a twenty-line suggestion is reviewed.”
Opus 4 and long-running tasks
On 22 May 2025 Anthropic introduced Claude Opus 4 and Claude Sonnet 4, describing Opus 4 as the world's best coding model and emphasising sustained performance on long-running agentic tasks rather than single-turn responses. Sonnet 4 was presented as a direct upgrade to the preceding Sonnet release.
The engineering significance of a long-running agent is not that it writes more code. It is that the unit of review changes. A pull request produced over four hours of autonomous tool use cannot be reviewed the way a twenty-line suggestion is reviewed. The reviewer is no longer checking a diff against an intention they hold in their head; they are auditing a process they did not observe.
Supervision as a competence
The Society's position is that supervision of autonomous coding agents is a distinct professional competence and should be assessed as one. It requires the ability to construct adversarial test cases before the agent runs, to read execution traces rather than only outputs, and to identify the class of failure — specification, tool, or model — when a run goes wrong.
Organisations adopting these tools should record which changes to production systems originated from an agent, who approved them, and on what evidence. This is not bureaucratic caution. It is the minimum record needed to investigate an incident six months later.
- March 2023 — Claude first released as a chatbot.
- 21 June 2024 — Claude 3.5 Sonnet outperforms the prior Opus tier at lower cost.
- 22 May 2025 — Claude Opus 4 and Sonnet 4 launched, emphasising long-running agentic coding.
Join the professional body behind this work
ACS members receive our research first, free CPD and ethics modules every year, and a route to professional registration assessed by their peers.
Become a memberMore from ACS Insights
Optimus: a humanoid robot from prototype to production line
Four years from an AI Day slide to a converted Fremont assembly line — and still no commercial sale.
AnalysisGrok, Colossus and the compute arms race
xAI built a 100,000-GPU cluster in 122 days, doubled it, and merged twice. The externalities arrived with the electricity.
ArticleA national consortium to build trust in AI
ACS joins federal partners, universities and industry to strengthen assurance practice for high-impact AI systems.