add
This commit is contained in:
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: AI Evaluation Infrastructure
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [evaluation, llm, quality, workflow]
|
||||
sources: [raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/harbor-langchain-agent-eval-stack-2026.md, raw/articles/anthropic-claude-sonnet-5-2026.md, raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md, raw/articles/openai-genebench-pro-2026.md, raw/articles/shopify-flow-agent-model-optimization-flywheel-2026.md, raw/articles/vllm-semantic-router-micro-agents-2026.md]
|
||||
sources: [raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/harbor-langchain-agent-eval-stack-2026.md, raw/articles/anthropic-claude-sonnet-5-2026.md, raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md, raw/articles/openai-genebench-pro-2026.md, raw/articles/shopify-flow-agent-model-optimization-flywheel-2026.md, raw/articles/vllm-semantic-router-micro-agents-2026.md, raw/articles/artificial-analysis-coding-agent-benchmarks-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -26,6 +26,8 @@ Shopify の Flow agent fine-tuning 記事は、evaluation infrastructure が pro
|
||||
|
||||
vLLM の Semantic Router / micro-agent 構想は、評価と orchestration が serving layer に入り込む例である。単一の OpenAI-compatible model ID の裏で、router が task に応じて recipe を選び、複数 worker に fan-out し、quorum、disagreement check、output contract repair、synthesis を行う。これは「どの model が強いか」を外から測るだけでなく、router 自体が小さな evaluator / coordinator になり、frontier model 呼び出しの前段で capability と cost/safety policy を組み立てるという設計である。[[loop-engineering]] や [[agent-harness-engineering]] では、application graph だけでなく inference gateway も評価・検証・合議の場になる。^[raw/articles/vllm-semantic-router-micro-agents-2026.md]
|
||||
|
||||
Artificial Analysis の Coding Agent Benchmarks は、agent 評価が「モデル名」だけでなく **harness、benchmark mix、cost、token usage、execution time** を一緒に測る段階へ進んでいることを示す。DeepSWE、Terminal-Bench v2、SWE-Atlas-QnA を合成した Coding Agent Index と、Claude Code / Cursor CLI / Opencode の harness comparison は、同じ model でも実行環境と operator loop によって成績が変わることを可視化する。これは [[agent-harness-engineering]] と [[agent-oriented-cli-design]] に近く、評価基盤が model leaderboard から agent runtime / CLI / workflow 比較へ広がる兆候である。^[raw/articles/artificial-analysis-coding-agent-benchmarks-2026.md]
|
||||
|
||||
## なぜ重要か
|
||||
|
||||
- **Crowdsourced comparison**: 利用者が 2 つのモデル出力を比較する形式は、静的な benchmark では拾いにくい実利用の好みを集められる。
|
||||
@@ -35,6 +37,7 @@ vLLM の Semantic Router / micro-agent 構想は、評価と orchestration が s
|
||||
- **Judgment-heavy scientific evaluation**: GeneBench-Pro のような benchmark は、正解率だけでなく、データ診断、分析方針の変更、因果推論、solver contract の明確さまで評価対象にする。科学 agent を評価するには、clean sandbox と deterministic check だけでなく、trace から判断品質を検査できる問題設計が必要になる。
|
||||
- **Production feedback flywheels**: Shopify Flow の例では、synthetic benchmark、LLM judge、programmatic checker、本番 activation rate、slice analysis、週次 retraining が一つの改善ループになる。評価基盤は「合格判定」ではなく、どのデータを足し、どの形式を変え、どの tool response を削るかを決める運用面になる。
|
||||
- **Serving-layer evaluators**: vLLM Semantic Router のように、model router が quorum、disagreement check、output repair を実行すると、評価は offline benchmark だけでなく、inference request ごとの制御面にも入る。
|
||||
- **Harness-sensitive coding-agent evaluation**: Artificial Analysis のような coding agent leaderboard は、agent が使う CLI/harness、token usage、実行時間、費用を同じ評価面に載せるため、[[agent-harness-engineering]] そのものが比較対象になる。
|
||||
|
||||
## Open Questions
|
||||
|
||||
|
||||
Reference in New Issue
Block a user