add
This commit is contained in:
@@ -0,0 +1,99 @@
|
||||
---
|
||||
source_url: https://artificialanalysis.ai/agents/coding-agents
|
||||
ingested: 2026-07-17
|
||||
sha256: bed7d890238a27bedd0ba3b84dbec477c6e755890d1e709dca85af95d313e558
|
||||
discovered_from:
|
||||
platform: discord
|
||||
channel_id: '1477793137064935675'
|
||||
channel_name: tw
|
||||
message_id: '1527811333935075478'
|
||||
author_id: '1477793167486226708'
|
||||
posted_at: 2026-07-17T22:56:47.372000000Z
|
||||
message_excerpt: "Artificial Analysis coding agent benchmark from tw digest."
|
||||
score: 3
|
||||
score_reason: "Agent coding benchmark infrastructure extends AI evaluation infrastructure."
|
||||
---
|
||||
|
||||
## Artificial Analysis Coding Agent Benchmarks
|
||||
|
||||
We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.
|
||||
|
||||
To compare language models see our [model benchmarks](https://artificialanalysis.ai/models).
|
||||
|
||||
## Artificial Analysis Coding Agent Index
|
||||
|
||||
Composite index of 3 benchmarks:
|
||||
|
||||
- DeepSWE
|
||||
Software engineering tasks, 113 tasks
|
||||
[By Datacurve](https://deepswe.datacurve.ai/)
|
||||
- Terminal-Bench v2
|
||||
Agentic terminal use, 84 tasks
|
||||
[By Laude Institute](https://www.tbench.ai/benchmarks/terminal-bench-2)
|
||||
- SWE-Atlas-QnA
|
||||
Technical Q&A, 124 tasks
|
||||
[By Scale AI](https://labs.scale.com/leaderboard/sweatlas-qna)
|
||||
|
||||
Index represents the average pass@1 across 3 runs of each benchmark. Index recently updated to v1.2. [See methodology for details](https://artificialanalysis.ai/methodology/coding-agents-benchmarking)
|
||||
|
||||
Highlights
|
||||
|
||||
## Performance
|
||||
|
||||
Performance across the Artificial Analysis Coding Agent Index.
|
||||
|
||||
### Artificial Analysis Coding Agent Index
|
||||
|
||||
Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better
|
||||
|
||||
## Harness Comparison
|
||||
|
||||
Artificial Analysis Coding Agent Index by harness for Claude Opus 4.7.
|
||||
|
||||
### Harness Comparison: Artificial Analysis Coding Agent Index
|
||||
|
||||
Composite average pass@1 across Claude Code, Cursor CLI, and Opencode for Claude Opus 4.7 · Higher is better
|
||||
|
||||
## Token Usage
|
||||
|
||||
Token consumption across the Artificial Analysis Coding Agent Index, including total usage, token mix, efficiency, and per-benchmark breakdowns.
|
||||
|
||||
### Token Usage per Task
|
||||
|
||||
Average input, cache, and output tokens per task
|
||||
|
||||
Prompt cache hit rates can vary significantly by provider routing, which can materially change effective cost.
|
||||
|
||||
### Artificial Analysis Coding Agent Index vs. Total Tokens
|
||||
|
||||
Artificial Analysis Coding Agent Index vs. average total tokens per task
|
||||
|
||||
Most attractive quadrant
|
||||
|
||||
## Cost
|
||||
|
||||
Cost across the Artificial Analysis Coding Agent Index based on current per-token API pricing, including cache write pricing and cache discounts where available. Many users will access coding agent harnesses through subscription plan offerings rather than pay-per-token.
|
||||
|
||||
### Cost per Task
|
||||
|
||||
Average pay-per-token API cost per task (USD) · Lower is better
|
||||
|
||||
### Artificial Analysis Coding Agent Index vs. Cost per Task
|
||||
|
||||
Artificial Analysis Coding Agent Index vs. average pay-per-token API cost per task (USD)
|
||||
|
||||
Most attractive quadrant
|
||||
|
||||
## Execution Time
|
||||
|
||||
Active agent runtime across the Artificial Analysis Coding Agent Index.
|
||||
|
||||
### Time per Task
|
||||
|
||||
Average agent wall time per task · Lower is better
|
||||
|
||||
### Artificial Analysis Coding Agent Index vs. Execution Time
|
||||
|
||||
Artificial Analysis Coding Agent Index vs. average agent wall time per task
|
||||
|
||||
Most attractive quadrant
|
||||
Reference in New Issue
Block a user