100 lines
3.3 KiB
Markdown
100 lines
3.3 KiB
Markdown
---
|
|
source_url: https://artificialanalysis.ai/agents/coding-agents
|
|
ingested: 2026-07-17
|
|
sha256: bed7d890238a27bedd0ba3b84dbec477c6e755890d1e709dca85af95d313e558
|
|
discovered_from:
|
|
platform: discord
|
|
channel_id: '1477793137064935675'
|
|
channel_name: tw
|
|
message_id: '1527811333935075478'
|
|
author_id: '1477793167486226708'
|
|
posted_at: 2026-07-17T22:56:47.372000000Z
|
|
message_excerpt: "Artificial Analysis coding agent benchmark from tw digest."
|
|
score: 3
|
|
score_reason: "Agent coding benchmark infrastructure extends AI evaluation infrastructure."
|
|
---
|
|
|
|
## Artificial Analysis Coding Agent Benchmarks
|
|
|
|
We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.
|
|
|
|
To compare language models see our [model benchmarks](https://artificialanalysis.ai/models).
|
|
|
|
## Artificial Analysis Coding Agent Index
|
|
|
|
Composite index of 3 benchmarks:
|
|
|
|
- DeepSWE
|
|
Software engineering tasks, 113 tasks
|
|
[By Datacurve](https://deepswe.datacurve.ai/)
|
|
- Terminal-Bench v2
|
|
Agentic terminal use, 84 tasks
|
|
[By Laude Institute](https://www.tbench.ai/benchmarks/terminal-bench-2)
|
|
- SWE-Atlas-QnA
|
|
Technical Q&A, 124 tasks
|
|
[By Scale AI](https://labs.scale.com/leaderboard/sweatlas-qna)
|
|
|
|
Index represents the average pass@1 across 3 runs of each benchmark. Index recently updated to v1.2. [See methodology for details](https://artificialanalysis.ai/methodology/coding-agents-benchmarking)
|
|
|
|
Highlights
|
|
|
|
## Performance
|
|
|
|
Performance across the Artificial Analysis Coding Agent Index.
|
|
|
|
### Artificial Analysis Coding Agent Index
|
|
|
|
Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better
|
|
|
|
## Harness Comparison
|
|
|
|
Artificial Analysis Coding Agent Index by harness for Claude Opus 4.7.
|
|
|
|
### Harness Comparison: Artificial Analysis Coding Agent Index
|
|
|
|
Composite average pass@1 across Claude Code, Cursor CLI, and Opencode for Claude Opus 4.7 · Higher is better
|
|
|
|
## Token Usage
|
|
|
|
Token consumption across the Artificial Analysis Coding Agent Index, including total usage, token mix, efficiency, and per-benchmark breakdowns.
|
|
|
|
### Token Usage per Task
|
|
|
|
Average input, cache, and output tokens per task
|
|
|
|
Prompt cache hit rates can vary significantly by provider routing, which can materially change effective cost.
|
|
|
|
### Artificial Analysis Coding Agent Index vs. Total Tokens
|
|
|
|
Artificial Analysis Coding Agent Index vs. average total tokens per task
|
|
|
|
Most attractive quadrant
|
|
|
|
## Cost
|
|
|
|
Cost across the Artificial Analysis Coding Agent Index based on current per-token API pricing, including cache write pricing and cache discounts where available. Many users will access coding agent harnesses through subscription plan offerings rather than pay-per-token.
|
|
|
|
### Cost per Task
|
|
|
|
Average pay-per-token API cost per task (USD) · Lower is better
|
|
|
|
### Artificial Analysis Coding Agent Index vs. Cost per Task
|
|
|
|
Artificial Analysis Coding Agent Index vs. average pay-per-token API cost per task (USD)
|
|
|
|
Most attractive quadrant
|
|
|
|
## Execution Time
|
|
|
|
Active agent runtime across the Artificial Analysis Coding Agent Index.
|
|
|
|
### Time per Task
|
|
|
|
Average agent wall time per task · Lower is better
|
|
|
|
### Artificial Analysis Coding Agent Index vs. Execution Time
|
|
|
|
Artificial Analysis Coding Agent Index vs. average agent wall time per task
|
|
|
|
Most attractive quadrant
|