Claude Code now holds the higher top-line coding-agent score, but Codex gets much closer at a fraction of the measured cost.
Artificial Analysis’ current representative comparison puts Claude Code with Sonnet 5.5 at 68 on its Coding Agent Index, versus 63 for Codex with GPT-6.1 Sol. The composite-score gap disappears at matched xhigh effort, where both agents score 63, but Codex still finishes tasks faster and at lower measured API cost.
Anthropic released Sonnet 5.5 on Sept. 28, saying it improves throughput by more than 30% over Sonnet 5 and can reduce task costs by as much as 30% for most workloads. OpenAI followed Sept. 29 with GPT-6.1 Sol for complex coding and professional work at lower cost than GPT-6 Astra. Those launches fit a broader shift in coding-model pricing and performance, where cost per completed task can matter more than token price alone.
Performance, cost, and speed split the comparison
How we compared Codex and Claude Code
TechRepublic used Artificial Analysis’ Coding Agent Index v1.5, which combines DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA, and reviewed each agent’s current representative configuration alongside matched xhigh settings. We also considered measured API cost, task time, and token use. No overall winner or weighted score was assigned.
| Metric | Claude Code + Sonnet 5.5 (max) | Codex + GPT-6.1 Sol (xhigh) |
|---|---|---|
| Coding Agent Index v1.5 | 68 | 63 |
| DeepSWE v1.1 | 72% | 73% |
| Terminal-Bench 4.0 | 66% | 55% |
| SWE-Atlas-QnA | 67% | 61% |
| Cost per task | $14.19 | $1.04 |
| Average time per task | 1.5 hours | 15.5 min. |
| Tokens per task | 27.7 million | 3.2 million |
Claude Code leads the composite score and two component benchmarks; Codex leads DeepSWE and uses far fewer tokens.
Equalizing effort changes the comparison. Sonnet 5.5 at xhigh falls to 63, matching GPT-6.1 Sol at xhigh, while costing $3.33 per task and averaging 27 minutes. Codex stays at $1.04 and 15.5 minutes.
Public benchmarks are not the same as production outcomes. A July study of command-line AI coding agents linked regular use with roughly 24% more merged pull requests, while also finding that adoption and review capacity shaped the result.
Security and controls can outweigh benchmark gaps
Security flaws have affected both platforms. A VentureBeat investigation detailed a Codex vulnerability that could expose a GitHub OAuth token through a crafted branch name and Claude Code permission bypasses that Anthropic later patched.
Anthropic says Claude Code users approve 93% of permission prompts. In internal testing, its auto-mode classifier cut false positives on benign actions to 0.4%, but produced a 17% false-negative rate across 52 real “overeager” actions.
Both vendors document controls intended to constrain agent access. OpenAI describes sandboxing, approvals, and network controls for Codex, while Anthropic offers managed Claude Code settings for tool permissions, file access, and MCP servers.
Teams that need different interfaces or governance models can also compare other Claude Code alternatives, including GitHub Copilot, Cursor, Google Antigravity CLI, and Kiro.
Claude Code offers the higher current representative benchmark score, while Codex reaches the matched xhigh score with lower measured cost, runtime, and token use. Teams should test both against the repositories they plan to automate and verify credential scope and command-approval policies before granting broader access.
Want to learn more AI tips, tricks, and prompting techniques? Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work.
Learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. eWeek readers get free 7-day access. Check out all the lessons here →