Files
llm-wiki/raw/articles/artificial-analysis-coding-agent-benchmarks-2026.md
T
2026-07-18 10:09:57 +09:00

3.3 KiB

source_url, ingested, sha256, discovered_from, score, score_reason
source_url ingested sha256 discovered_from score score_reason
https://artificialanalysis.ai/agents/coding-agents 2026-07-17 bed7d890238a27bedd0ba3b84dbec477c6e755890d1e709dca85af95d313e558
platform channel_id channel_name message_id author_id posted_at message_excerpt
discord 1477793137064935675 tw 1527811333935075478 1477793167486226708 2026-07-17T22:56:47.372000000Z Artificial Analysis coding agent benchmark from tw digest.
3 Agent coding benchmark infrastructure extends AI evaluation infrastructure.

Artificial Analysis Coding Agent Benchmarks

We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.

To compare language models see our model benchmarks.

Artificial Analysis Coding Agent Index

Composite index of 3 benchmarks:

Index represents the average pass@1 across 3 runs of each benchmark. Index recently updated to v1.2. See methodology for details

Highlights

Performance

Performance across the Artificial Analysis Coding Agent Index.

Artificial Analysis Coding Agent Index

Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better

Harness Comparison

Artificial Analysis Coding Agent Index by harness for Claude Opus 4.7.

Harness Comparison: Artificial Analysis Coding Agent Index

Composite average pass@1 across Claude Code, Cursor CLI, and Opencode for Claude Opus 4.7 · Higher is better

Token Usage

Token consumption across the Artificial Analysis Coding Agent Index, including total usage, token mix, efficiency, and per-benchmark breakdowns.

Token Usage per Task

Average input, cache, and output tokens per task

Prompt cache hit rates can vary significantly by provider routing, which can materially change effective cost.

Artificial Analysis Coding Agent Index vs. Total Tokens

Artificial Analysis Coding Agent Index vs. average total tokens per task

Most attractive quadrant

Cost

Cost across the Artificial Analysis Coding Agent Index based on current per-token API pricing, including cache write pricing and cache discounts where available. Many users will access coding agent harnesses through subscription plan offerings rather than pay-per-token.

Cost per Task

Average pay-per-token API cost per task (USD) · Lower is better

Artificial Analysis Coding Agent Index vs. Cost per Task

Artificial Analysis Coding Agent Index vs. average pay-per-token API cost per task (USD)

Most attractive quadrant

Execution Time

Active agent runtime across the Artificial Analysis Coding Agent Index.

Time per Task

Average agent wall time per task · Lower is better

Artificial Analysis Coding Agent Index vs. Execution Time

Artificial Analysis Coding Agent Index vs. average agent wall time per task

Most attractive quadrant