Files
llm-wiki/concepts/agent-harness-engineering.md
T
2026-07-18 10:09:57 +09:00

43 lines
13 KiB
Markdown

---
title: Agent Harness Engineering
created: 2026-07-01
updated: 2026-07-17
type: concept
tags: [agent, automation, evaluation, workflow, quality, reliability]
sources: [raw/articles/awesome-harness-engineering-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/github-copilot-vision-ga-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md, raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/vscode-1-110-agent-browser-tools-2026.md, raw/articles/explain-diff-html-agent-skill-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/skamille-respectful-ai-use-guidelines-2026.md, raw/articles/devansh-llm-vulnerability-research-2026.md, raw/articles/slack-agentic-testing-e2e-stack-2026.md]
confidence: medium
---
# Agent Harness Engineering
Agent harness engineering は、AI agent の賢さを model 単体で見ず、周囲の環境・制約・評価・観測性・状態管理を設計して、実務で壊れにくくする考え方。Awesome Harness Engineering は、これを context engineering、evaluation、observability、orchestration、safe autonomy、software architecture の交点として整理し、長時間の coding / research task で agent を dependable にする資料だけを集める方針を明示している。
[[loop-engineering]] が discovery / handoff / verification / persistence / scheduling まで含む「継続ループ」を扱うなら、agent harness engineering は 1 回から数回の agent 実行が正しく進むための足場に近い。context window をどう使うか、失敗をどう残すか、どの tool を許すか、評価をどう再現するか、operator が trace や cost をどう見るかが中心になる。
## 見るべき軸
- **Context / memory / working state**: context window を単なる貼り付け先ではなく作業記憶として扱い、bounded memory、filesystem memory、repo-local instruction、resume artifact を設計する。これは [[llm-wiki-pattern]] のように知識を残す運用とも接続する。
- **Constraints / guardrails / safe autonomy**: sandbox、confirmation mode、tool boundary、prompt-injection mitigation、quality gate で agent の自由度を狭める。ここは [[ai-agent-command-safety]] や [[ai-agent-identity-security]] の権限境界と隣り合う。
- **Specs and workflow design**: AGENTS.md、agent.md、spec-driven development、12 Factor Agents のように、agent が読む仕様と作業手順をプロジェクト側に置く。これは [[agent-oriented-cli-design]] の「道具が agent に使い方を教える」発想の repository 版でもある。
- **Output harnesses for human review**: Geoffrey Litt の `explain-diff-html` skill は、PR / diff / branch の説明を、背景、直感、code walkthrough、interactive quiz 付きの self-contained HTML にまとめる agent instruction である。重要なのは「説明して」で終わらず、初心者向け背景、toy example、diagram family、mobile-readable layout、quiz feedback、code block CSS まで出力要件を固定している点で、agent の成果物を人間が検証しやすい形へ constrained generation する harness として読める。これは [[openwiki]] や [[litho]] の repo documentation loop とも接続する。^[raw/articles/explain-diff-html-agent-skill-2026.md]
- **Respectful handoff / review tax**: Camille Fournier の「respectful AI use」ガイドは、AI policy を security / compliance だけでなく team throughput の問題として扱う。自分が読んでいない AI 生成 code や文書を他人に review させることは、生成者の生産性を同僚の validation tax へ転嫁する。Agent harness は「人間 review を最後に置く」だけでなく、生成者が理解・短縮・分割・説明できる粒度へ落とす制約を持つ必要がある。これは [[agent-oriented-cli-design]] の出力設計や [[ai-agent-command-safety]] の承認境界とも接続する。^[raw/articles/skamille-respectful-ai-use-guidelines-2026.md]
- **Minimal security-research scaffolding**: Devansh の [[llm-assisted-vulnerability-research]] 記事は、脆弱性探索では bloated `AGENT.md` / `SKILLS.md` や広い checklist が context rot を悪化させることがあり、1 ページ程度の threat model、不変条件、thin slice、verifier loop に token を使う方が実用的だとする。これは harness を増やす話ではなく、harness を「注意を散らさず、検証を強制する最小構造」に削る設計として重要である。^[raw/articles/devansh-llm-vulnerability-research-2026.md]
- **Secret-access harnesses**: 1Password Environments MCP Server for Codex は、agent が環境を構成・実行する時に secret value を model context へ入れず、user approval と runtime injection に閉じ込める harness である。agent harness engineering では、tool を増やすだけでなく、credential がどの channel に現れないかを仕様として固定することが安全な自律性の条件になる。^[raw/articles/1password-codex-mcp-secret-access-2026.md]
- **Evals and observability**: skill eval、trace grading、OpenTelemetry、session replay、cost tracking、benchmark を使い、成功/失敗を operator の感覚だけにしない。[[ai-evaluation-infrastructure]] では model / agent を測る市場や基盤が主題だが、harness engineering では eval を個々の workflow の改善 loop に入れる。
- **Agentic test harnesses**: Slack Engineering の E2E 実験は、同じ UI goal を Playwright MCP、Playwright CLI、agent-generated Playwright tests で比較し、model より execution harness の差が reliability / turn count / token cost に効くことを示す。MCP は UI 操作と状態取得をまとめて返すため CLI より少ない turn で済み、複雑な flow でも失敗率が低かった一方、agentic run は snapshot と履歴の再送で高コストになる。したがって agent harness では browser primitive の形、state snapshot の粒度、context compaction、action signature の記録、deterministic test への切り戻しを同時に設計する必要がある。これは [[e2e-coverage-metrics]] と [[safari-mcp-server]] の実践的な接点である。^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
- **Browser harnesses**: GitHub Copilot の VS Code browser tools GA は、agent が live web app を操作し、console error、screenshot、scripted flow を chat へ戻す harness を IDE に組み込む例である。重要なのは browser 操作そのものだけでなく、人間 tab の明示共有、agent tab の session isolation、camera/microphone/geolocation の既定拒否、enterprise allow/deny と workspace trust を同じ harness に入れている点で、これは [[ai-agent-identity-security]] と [[e2e-coverage-metrics]] の接点になる。^[raw/articles/github-copilot-browser-tools-ga-2026.md]
- **Local browser MCP harnesses**: [[safari-mcp-server]] は、Safari Technology Preview の `safaridriver --mcp` を MCP server として公開し、agent が Safari の DOM、network request、console、screenshot、viewport、dialog、tab、page content を直接観測・操作できるようにする。Copilot browser tools が IDE 統合の browser harness なら、Safari MCP は特定ブラウザの実装差、性能、アクセシビリティ、form state を agent loop に入れる local harness である。^[raw/articles/safari-mcp-server-webkit-2026.md]
- **IDE-level agent control surface**: VS Code 1.110 は、agentic browser tools だけでなく、Agent Debug panel、background agent の `/compact` や slash command、session rename、Claude agent の steering / queuing、agent plugins、session memory、chat fork までまとめて入れている。これは browser 操作単体の話ではなく、agent を長時間走らせ、何を読み込んだか・どの tool を呼んだか・どの session へ分岐したかを IDE 側で観測し制御する harness への移行である。auto-approve `/yolo` は便利だが、記事自体も terminal sandboxing と security implication を明示しており、[[ai-agent-command-safety]] と [[ai-agent-identity-security]] の境界設計なしには扱えない。^[raw/articles/vscode-1-110-agent-browser-tools-2026.md]
- **Multimodal context as harness input**: Copilot Vision の一般提供により、VS Code、github.com、Copilot CLI で画像や PDF を prompt に添付できるようになった。agent mode や terminal run が screenshot、設計図、PDF 仕様を同じ context として扱える一方、Business / Enterprise では添付画像・PDF が約 24 時間保持されるため、便利な入力拡張は retention / privacy の設計対象でもある。^[raw/articles/github-copilot-vision-ga-2026.md]
- **Cost guardrails**: Copilot CLI / SDK の AI credit session limit は、model call、subagent、compaction、background work を含む 1 session の消費上限を soft cap として置く。特に無人 automation では、agent が完了まで走り続けるのではなく、上限到達時に wrap up して知らせることが harness の安全機能になる。^[raw/articles/github-copilot-ai-credit-session-limits-2026.md]
- **Workspace harnesses**: [[notion]] Developer Platform は、External Agents API、Workers、CLI、MCP、Markdown API を通じて、agent が業務 workspace 上で data sync、webhook、tool 実行、承認 loop を扱う方向を示している。ここでは chat UI ではなく、workspace そのものが agent harness になり、connection 管理と audit が [[ai-agent-identity-security]] の問題になる。^[raw/articles/notion-developer-platform-agents-workers-2026.md]
- **Operator-facing background agents**: Claude Code 2.1.198 は、背景 agent の完了・入力待ちを Notification hook に出し、worktree 内で終えた code work を commit / push / draft PR まで進め、agent view / task panel / workflow progress の stalled 状態を直す方向へ寄せている。これは model 性能ではなく、長時間 agent を日常運用するための [[loop-engineering]] と [[agent-oriented-cli-design]] の harness 改善である。2.1.196 でも background session survival、auto-resume、streaming idle watchdog、dangerously-skip-permissions の表示修正、MCP OAuth scope 修正が並んでおり、agent harness の価値が「止まらない・見える・勝手に危険側へ倒れない」ことにあると分かる。^[raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md]
- **Agentic deployment harnesses**: AWS Forward Deployed Engineering は、agentic AI を「導入支援込みの運用 harness」として売る動きでもある。FDE は顧客環境に入り、business / engineering / security teams と production AI system を作り、semantic layer・governed/versioned knowledge graph・runbook・architectural documentation・trained internal champion を残して self-sufficiency を目標にする。これは [[llm-wiki-pattern]] 的な知識の残し方と、[[ai-agent-identity-security]] の governance boundary を enterprise deployment に拡張した例として読める。^[raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md]
- **Reference implementations**: SWE-agent、Harbor、Citadel、browser harness、Harness Evolver、skills.sh、Uni-CLI などは、framework そのものより「何を隔離し、何を記録し、何を評価するか」を読む対象になる。
## なぜ重要か
Yuta の関心では、agent harness engineering は「また新しい agent framework が出た」というニュースより重要度が高い。既存の coding agent を複数使い分けるほど、差は model だけでなく、repo-local instruction、sandbox、approval、trace、worktree、eval、cost ledger、session export のような外側の設計に出る。[[pi-coding-agent]] の最小主義や [[abtop]] の operator dashboard も、この harness をどこまで見える形にするかという問題として読める。
Awesome list 形式の資料なので単独の主張は広く浅いが、一次資料・実装・benchmark を横断する地図として価値がある。今後は個別リンクを全部 raw 化するより、実際に使う harness pattern が出たときにこのページから [[loop-engineering]]、[[agent-oriented-cli-design]]、[[ai-evaluation-infrastructure]] へ接続して増補するのがよい。