This commit is contained in:
2026-07-18 10:09:57 +09:00
parent 86cad348b4
commit 8f0e53766a
45 changed files with 6221 additions and 781 deletions
+3 -2
View File
@@ -1,10 +1,10 @@
---
title: Agent Harness Engineering
created: 2026-07-01
updated: 2026-07-02
updated: 2026-07-17
type: concept
tags: [agent, automation, evaluation, workflow, quality, reliability]
sources: [raw/articles/awesome-harness-engineering-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/github-copilot-vision-ga-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md, raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/vscode-1-110-agent-browser-tools-2026.md, raw/articles/explain-diff-html-agent-skill-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/skamille-respectful-ai-use-guidelines-2026.md, raw/articles/devansh-llm-vulnerability-research-2026.md]
sources: [raw/articles/awesome-harness-engineering-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/github-copilot-vision-ga-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md, raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/vscode-1-110-agent-browser-tools-2026.md, raw/articles/explain-diff-html-agent-skill-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/skamille-respectful-ai-use-guidelines-2026.md, raw/articles/devansh-llm-vulnerability-research-2026.md, raw/articles/slack-agentic-testing-e2e-stack-2026.md]
confidence: medium
---
@@ -24,6 +24,7 @@ Agent harness engineering は、AI agent の賢さを model 単体で見ず、
- **Minimal security-research scaffolding**: Devansh の [[llm-assisted-vulnerability-research]] 記事は、脆弱性探索では bloated `AGENT.md` / `SKILLS.md` や広い checklist が context rot を悪化させることがあり、1 ページ程度の threat model、不変条件、thin slice、verifier loop に token を使う方が実用的だとする。これは harness を増やす話ではなく、harness を「注意を散らさず、検証を強制する最小構造」に削る設計として重要である。^[raw/articles/devansh-llm-vulnerability-research-2026.md]
- **Secret-access harnesses**: 1Password Environments MCP Server for Codex は、agent が環境を構成・実行する時に secret value を model context へ入れず、user approval と runtime injection に閉じ込める harness である。agent harness engineering では、tool を増やすだけでなく、credential がどの channel に現れないかを仕様として固定することが安全な自律性の条件になる。^[raw/articles/1password-codex-mcp-secret-access-2026.md]
- **Evals and observability**: skill eval、trace grading、OpenTelemetry、session replay、cost tracking、benchmark を使い、成功/失敗を operator の感覚だけにしない。[[ai-evaluation-infrastructure]] では model / agent を測る市場や基盤が主題だが、harness engineering では eval を個々の workflow の改善 loop に入れる。
- **Agentic test harnesses**: Slack Engineering の E2E 実験は、同じ UI goal を Playwright MCP、Playwright CLI、agent-generated Playwright tests で比較し、model より execution harness の差が reliability / turn count / token cost に効くことを示す。MCP は UI 操作と状態取得をまとめて返すため CLI より少ない turn で済み、複雑な flow でも失敗率が低かった一方、agentic run は snapshot と履歴の再送で高コストになる。したがって agent harness では browser primitive の形、state snapshot の粒度、context compaction、action signature の記録、deterministic test への切り戻しを同時に設計する必要がある。これは [[e2e-coverage-metrics]] と [[safari-mcp-server]] の実践的な接点である。^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
- **Browser harnesses**: GitHub Copilot の VS Code browser tools GA は、agent が live web app を操作し、console error、screenshot、scripted flow を chat へ戻す harness を IDE に組み込む例である。重要なのは browser 操作そのものだけでなく、人間 tab の明示共有、agent tab の session isolation、camera/microphone/geolocation の既定拒否、enterprise allow/deny と workspace trust を同じ harness に入れている点で、これは [[ai-agent-identity-security]] と [[e2e-coverage-metrics]] の接点になる。^[raw/articles/github-copilot-browser-tools-ga-2026.md]
- **Local browser MCP harnesses**: [[safari-mcp-server]] は、Safari Technology Preview の `safaridriver --mcp` を MCP server として公開し、agent が Safari の DOM、network request、console、screenshot、viewport、dialog、tab、page content を直接観測・操作できるようにする。Copilot browser tools が IDE 統合の browser harness なら、Safari MCP は特定ブラウザの実装差、性能、アクセシビリティ、form state を agent loop に入れる local harness である。^[raw/articles/safari-mcp-server-webkit-2026.md]
- **IDE-level agent control surface**: VS Code 1.110 は、agentic browser tools だけでなく、Agent Debug panel、background agent の `/compact` や slash command、session rename、Claude agent の steering / queuing、agent plugins、session memory、chat fork までまとめて入れている。これは browser 操作単体の話ではなく、agent を長時間走らせ、何を読み込んだか・どの tool を呼んだか・どの session へ分岐したかを IDE 側で観測し制御する harness への移行である。auto-approve `/yolo` は便利だが、記事自体も terminal sandboxing と security implication を明示しており、[[ai-agent-command-safety]] と [[ai-agent-identity-security]] の境界設計なしには扱えない。^[raw/articles/vscode-1-110-agent-browser-tools-2026.md]