This commit is contained in:
2026-07-03 00:38:05 +09:00
parent 87eacd39b2
commit 86cad348b4
149 changed files with 24450 additions and 105 deletions
+41
View File
@@ -0,0 +1,41 @@
---
title: Agent Harness Engineering
created: 2026-07-01
updated: 2026-07-02
type: concept
tags: [agent, automation, evaluation, workflow, quality, reliability]
sources: [raw/articles/awesome-harness-engineering-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/github-copilot-vision-ga-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md, raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/vscode-1-110-agent-browser-tools-2026.md, raw/articles/explain-diff-html-agent-skill-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/skamille-respectful-ai-use-guidelines-2026.md, raw/articles/devansh-llm-vulnerability-research-2026.md]
confidence: medium
---
# Agent Harness Engineering
Agent harness engineering は、AI agent の賢さを model 単体で見ず、周囲の環境・制約・評価・観測性・状態管理を設計して、実務で壊れにくくする考え方。Awesome Harness Engineering は、これを context engineering、evaluation、observability、orchestration、safe autonomy、software architecture の交点として整理し、長時間の coding / research task で agent を dependable にする資料だけを集める方針を明示している。
[[loop-engineering]] が discovery / handoff / verification / persistence / scheduling まで含む「継続ループ」を扱うなら、agent harness engineering は 1 回から数回の agent 実行が正しく進むための足場に近い。context window をどう使うか、失敗をどう残すか、どの tool を許すか、評価をどう再現するか、operator が trace や cost をどう見るかが中心になる。
## 見るべき軸
- **Context / memory / working state**: context window を単なる貼り付け先ではなく作業記憶として扱い、bounded memory、filesystem memory、repo-local instruction、resume artifact を設計する。これは [[llm-wiki-pattern]] のように知識を残す運用とも接続する。
- **Constraints / guardrails / safe autonomy**: sandbox、confirmation mode、tool boundary、prompt-injection mitigation、quality gate で agent の自由度を狭める。ここは [[ai-agent-command-safety]] や [[ai-agent-identity-security]] の権限境界と隣り合う。
- **Specs and workflow design**: AGENTS.md、agent.md、spec-driven development、12 Factor Agents のように、agent が読む仕様と作業手順をプロジェクト側に置く。これは [[agent-oriented-cli-design]] の「道具が agent に使い方を教える」発想の repository 版でもある。
- **Output harnesses for human review**: Geoffrey Litt の `explain-diff-html` skill は、PR / diff / branch の説明を、背景、直感、code walkthrough、interactive quiz 付きの self-contained HTML にまとめる agent instruction である。重要なのは「説明して」で終わらず、初心者向け背景、toy example、diagram family、mobile-readable layout、quiz feedback、code block CSS まで出力要件を固定している点で、agent の成果物を人間が検証しやすい形へ constrained generation する harness として読める。これは [[openwiki]] や [[litho]] の repo documentation loop とも接続する。^[raw/articles/explain-diff-html-agent-skill-2026.md]
- **Respectful handoff / review tax**: Camille Fournier の「respectful AI use」ガイドは、AI policy を security / compliance だけでなく team throughput の問題として扱う。自分が読んでいない AI 生成 code や文書を他人に review させることは、生成者の生産性を同僚の validation tax へ転嫁する。Agent harness は「人間 review を最後に置く」だけでなく、生成者が理解・短縮・分割・説明できる粒度へ落とす制約を持つ必要がある。これは [[agent-oriented-cli-design]] の出力設計や [[ai-agent-command-safety]] の承認境界とも接続する。^[raw/articles/skamille-respectful-ai-use-guidelines-2026.md]
- **Minimal security-research scaffolding**: Devansh の [[llm-assisted-vulnerability-research]] 記事は、脆弱性探索では bloated `AGENT.md` / `SKILLS.md` や広い checklist が context rot を悪化させることがあり、1 ページ程度の threat model、不変条件、thin slice、verifier loop に token を使う方が実用的だとする。これは harness を増やす話ではなく、harness を「注意を散らさず、検証を強制する最小構造」に削る設計として重要である。^[raw/articles/devansh-llm-vulnerability-research-2026.md]
- **Secret-access harnesses**: 1Password Environments MCP Server for Codex は、agent が環境を構成・実行する時に secret value を model context へ入れず、user approval と runtime injection に閉じ込める harness である。agent harness engineering では、tool を増やすだけでなく、credential がどの channel に現れないかを仕様として固定することが安全な自律性の条件になる。^[raw/articles/1password-codex-mcp-secret-access-2026.md]
- **Evals and observability**: skill eval、trace grading、OpenTelemetry、session replay、cost tracking、benchmark を使い、成功/失敗を operator の感覚だけにしない。[[ai-evaluation-infrastructure]] では model / agent を測る市場や基盤が主題だが、harness engineering では eval を個々の workflow の改善 loop に入れる。
- **Browser harnesses**: GitHub Copilot の VS Code browser tools GA は、agent が live web app を操作し、console error、screenshot、scripted flow を chat へ戻す harness を IDE に組み込む例である。重要なのは browser 操作そのものだけでなく、人間 tab の明示共有、agent tab の session isolation、camera/microphone/geolocation の既定拒否、enterprise allow/deny と workspace trust を同じ harness に入れている点で、これは [[ai-agent-identity-security]] と [[e2e-coverage-metrics]] の接点になる。^[raw/articles/github-copilot-browser-tools-ga-2026.md]
- **Local browser MCP harnesses**: [[safari-mcp-server]] は、Safari Technology Preview の `safaridriver --mcp` を MCP server として公開し、agent が Safari の DOM、network request、console、screenshot、viewport、dialog、tab、page content を直接観測・操作できるようにする。Copilot browser tools が IDE 統合の browser harness なら、Safari MCP は特定ブラウザの実装差、性能、アクセシビリティ、form state を agent loop に入れる local harness である。^[raw/articles/safari-mcp-server-webkit-2026.md]
- **IDE-level agent control surface**: VS Code 1.110 は、agentic browser tools だけでなく、Agent Debug panel、background agent の `/compact` や slash command、session rename、Claude agent の steering / queuing、agent plugins、session memory、chat fork までまとめて入れている。これは browser 操作単体の話ではなく、agent を長時間走らせ、何を読み込んだか・どの tool を呼んだか・どの session へ分岐したかを IDE 側で観測し制御する harness への移行である。auto-approve `/yolo` は便利だが、記事自体も terminal sandboxing と security implication を明示しており、[[ai-agent-command-safety]] と [[ai-agent-identity-security]] の境界設計なしには扱えない。^[raw/articles/vscode-1-110-agent-browser-tools-2026.md]
- **Multimodal context as harness input**: Copilot Vision の一般提供により、VS Code、github.com、Copilot CLI で画像や PDF を prompt に添付できるようになった。agent mode や terminal run が screenshot、設計図、PDF 仕様を同じ context として扱える一方、Business / Enterprise では添付画像・PDF が約 24 時間保持されるため、便利な入力拡張は retention / privacy の設計対象でもある。^[raw/articles/github-copilot-vision-ga-2026.md]
- **Cost guardrails**: Copilot CLI / SDK の AI credit session limit は、model call、subagent、compaction、background work を含む 1 session の消費上限を soft cap として置く。特に無人 automation では、agent が完了まで走り続けるのではなく、上限到達時に wrap up して知らせることが harness の安全機能になる。^[raw/articles/github-copilot-ai-credit-session-limits-2026.md]
- **Workspace harnesses**: [[notion]] Developer Platform は、External Agents API、Workers、CLI、MCP、Markdown API を通じて、agent が業務 workspace 上で data sync、webhook、tool 実行、承認 loop を扱う方向を示している。ここでは chat UI ではなく、workspace そのものが agent harness になり、connection 管理と audit が [[ai-agent-identity-security]] の問題になる。^[raw/articles/notion-developer-platform-agents-workers-2026.md]
- **Operator-facing background agents**: Claude Code 2.1.198 は、背景 agent の完了・入力待ちを Notification hook に出し、worktree 内で終えた code work を commit / push / draft PR まで進め、agent view / task panel / workflow progress の stalled 状態を直す方向へ寄せている。これは model 性能ではなく、長時間 agent を日常運用するための [[loop-engineering]] と [[agent-oriented-cli-design]] の harness 改善である。2.1.196 でも background session survival、auto-resume、streaming idle watchdog、dangerously-skip-permissions の表示修正、MCP OAuth scope 修正が並んでおり、agent harness の価値が「止まらない・見える・勝手に危険側へ倒れない」ことにあると分かる。^[raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md]
- **Agentic deployment harnesses**: AWS Forward Deployed Engineering は、agentic AI を「導入支援込みの運用 harness」として売る動きでもある。FDE は顧客環境に入り、business / engineering / security teams と production AI system を作り、semantic layer・governed/versioned knowledge graph・runbook・architectural documentation・trained internal champion を残して self-sufficiency を目標にする。これは [[llm-wiki-pattern]] 的な知識の残し方と、[[ai-agent-identity-security]] の governance boundary を enterprise deployment に拡張した例として読める。^[raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md]
- **Reference implementations**: SWE-agent、Harbor、Citadel、browser harness、Harness Evolver、skills.sh、Uni-CLI などは、framework そのものより「何を隔離し、何を記録し、何を評価するか」を読む対象になる。
## なぜ重要か
Yuta の関心では、agent harness engineering は「また新しい agent framework が出た」というニュースより重要度が高い。既存の coding agent を複数使い分けるほど、差は model だけでなく、repo-local instruction、sandbox、approval、trace、worktree、eval、cost ledger、session export のような外側の設計に出る。[[pi-coding-agent]] の最小主義や [[abtop]] の operator dashboard も、この harness をどこまで見える形にするかという問題として読める。
Awesome list 形式の資料なので単独の主張は広く浅いが、一次資料・実装・benchmark を横断する地図として価値がある。今後は個別リンクを全部 raw 化するより、実際に使う harness pattern が出たときにこのページから [[loop-engineering]]、[[agent-oriented-cli-design]]、[[ai-evaluation-infrastructure]] へ接続して増補するのがよい。
+12 -2
View File
@@ -1,10 +1,10 @@
---
title: Agent-Oriented CLI Design
created: 2026-06-30
updated: 2026-06-30
updated: 2026-07-02
type: concept
tags: [agent, cli, dev-tool, workflow, quality]
sources: [raw/articles/agent-oriented-cli-zenn-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md]
sources: [raw/articles/agent-oriented-cli-zenn-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md, raw/articles/stripe-well-known-agent-skills-index-2026.md, raw/articles/comfy-cli-agent-friendly-workflows-2026.md, raw/articles/awesome-openclaw-skills-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/vercel-konsistent-structural-linter-agents-2026.md]
confidence: medium
---
@@ -14,6 +14,8 @@ Agent-oriented CLI design は、人間が目で読んで試行錯誤する端末
重要なのは、エージェントに「推測させない」こと。使い方は wiki や skill 側へ長く写すのではなく、CLI 自体に `skill` や help サブコマンドとして同梱し、スキーマや出力の意味が実装と一緒に更新されるようにする。これは [[wiki-maintenance-loop]] の raw/source と synthesis を分ける考え方にも近く、手順が古くなる場所を減らす設計である。
Stripe の `.well-known/skills/index.json` は、この発想を Web documentation 側へ広げた例として読める。サイトが `stripe-best-practices`、`stripe-projects`、`upgrade-stripe` などの agent skill を機械可読な index として公開し、各 skill が参照ファイルや Stripe MCP / implementation planner へ誘導する。つまり agent-oriented design は CLI の出力だけでなく、サービスの公式ドキュメントが「エージェントがどの手順書を読むべきか」を discovery 可能にする方向へも進んでいる。[[ai-agent-identity-security]] の最小権限や監査と同じく、外部サービスが agent 向け入口を用意するほど、どの guidance を信頼するか・どの権限で実行するかが設計対象になる。
## 設計原則
- **JSON first**: 人間向けの整形テキストではなく、既定で構造化 JSON を返す。結果には `id`、`title`、`snippet`、`source_url`、`synced_at`、`is_stale` など、エージェントが次の判断に使う材料を入れる。
@@ -22,6 +24,14 @@ Agent-oriented CLI design は、人間が目で読んで試行錯誤する端末
- **Defaults over flags**: `--sources` や `--discover` のような細かい選択肢を増やすより、よく使う安全な既定値へ寄せる。フラグが多いほど help が長くなり、エージェントの分岐も増える。
- **Governance hooks**: Microsoft Learn の agent architecture / SDLC module は、agent task を input / output / success criteria で構造化し、PR template、checks、CODEOWNERS、rules、environment gate、observability、tool governance、secret boundary を運用設計へ入れることを強調する。CLI も単独の便利道具ではなく、[[ai-agent-identity-security]] や PR governance に接続される実行面として見るべき。
Comfy CLI shows the same design pressure in media/AI workflow tooling. Its commands expose `--json` envelopes, `error.hint`, `discover`, model schemas, job status/watch/cancel commands, workflow slot editing, and bundled agent skills for Claude Code/Cursor/AGENTS.md-aware tools. That makes a graphical workflow system scriptable by agents without requiring them to scrape UI state or guess command parameters. [[loop-engineering]] benefits because generation jobs, downloads, validation, and workflow edits become inspectable command steps rather than hidden GUI actions.^[raw/articles/comfy-cli-agent-friendly-workflows-2026.md]
OpenClaw Skills shows the ecosystem-level version of the same pattern. A community index sourced from ClawHub lists thousands of installable skills, exposes CLI installation (`openclaw skills install <skill-slug>` / `npx clawhub install <skill-slug>`), groups skills by task domain, and explicitly warns that skills are curated but not audited. This makes skills a distribution mechanism for agent capabilities, not just local documentation; it also raises the same trust questions as [[ai-agent-identity-security]] because an agent-readable capability package can contain prompt injection, tool poisoning, over-broad permissions, or unsafe data handling.^[raw/articles/awesome-openclaw-skills-2026.md]
[[notion]] の Developer Platform は、SaaS 側が「coding agent が使う CLI」を明示している例である。Notion CLI は workspace sign-in、page/database 操作、Workers の build/deploy を担当し、Markdown API や MCP と合わせて、agent が Notion の知識・workflow を machine-readable に扱える入口になる。ただし `curl ... | bash` 型の導入や workspace-scoped OAuth / personal access token は、CLI の使いやすさだけでなく [[ai-agent-identity-security]] の最小権限・監査とセットで見る必要がある。^[raw/articles/notion-developer-platform-agents-workers-2026.md]
Vercel Labs の `konsistent` は、agent-oriented CLI を「出力形式」だけでなく codebase structure の enforcement へ広げる。ESLint / Biome / oxlint が file 内の style を見るのに対し、`konsistent` は package、adapter、provider などが同じ file/export/type 形状を持つかを宣言的に検査する。README は、project-level structural convention が人間の onboarding だけでなく coding agent の予測可能性を上げると説明しており、agent が迷わないための interface は CLI help だけでなく repository layout にも宿る。これは [[agent-harness-engineering]] の specs / workflow design と、[[e2e-coverage-metrics]] 的な implementation-derived denominator の中間にある。^[raw/articles/vercel-konsistent-structural-linter-agents-2026.md]
## なぜ重要か
エージェント向け CLI は、単に「CLI を LLM から呼べるようにする」だけでは足りない。出力が曖昧だったり、エラーが不親切だったり、状態の鮮度が返らなかったりすると、agent loop は誤った仮定のまま進む。逆に、CLI が状態・出典・次アクション・失敗理由を明示すれば、[[loop-engineering]] の verification と persistence が自然に強くなる。
+35
View File
@@ -0,0 +1,35 @@
---
title: Agentic Web Monetization
created: 2026-07-01
updated: 2026-07-01
type: concept
tags: [agent, automation, public-interest, information-integrity]
sources: [raw/articles/cloudflare-monetization-gateway-x402-2026.md, raw/articles/cloudflare-content-independence-day-2025.md, raw/articles/cloudflare-ai-traffic-options-2026.md]
confidence: medium
---
# Agentic Web Monetization
Agentic web monetization は、人間の attention、広告表示、月額 subscription ではなく、AI agent や AI-written software が使う **request / token / outcome** ごとに web 資源へ支払う設計。Cloudflare の Monetization Gateway は、web page、dataset、API、MCP tool など Cloudflare 配下の任意の asset に payment rule と access control をかけ、x402 による HTTP 402 Payment Required flow と stablecoin settlement で、agentic buyer が signup や API key なしに小額決済して資源へアクセスする構想として発表された。^[raw/articles/cloudflare-monetization-gateway-x402-2026.md]
この論点は [[ai-crawler-governance]] の「誰が何の目的で読むか」という access policy を、実際の支払い・価格・settlement まで進める。Cloudflare の Content Independence Day は、AI answer が publisher へ traffic を返さないなら crawl には補償が必要だと主張した。Monetization Gateway は、その対象を crawler content から API、dataset、MCP tool call、developer tooling へ広げ、「agent が使う入力は agent が支払う」という経済層を edge policy の一部にする。
## 何が新しいか
- **支払いが request path に入る**: サーバーは 402 と価格・支払い先を返し、client は proof of payment を付けて再リクエストする。checkout 画面や別 payment API ではなく、HTTP request / response の中で access と支払いが結びつく。
- **payment が credential になる**: x402 では buyer が seller account を持たなくても、支払い証明そのものが一時的な access credential になる。これは [[ai-agent-identity-security]] の identity / authorization と隣接するが、必ずしも事前登録された API key を前提にしない。
- **edge が origin を守る**: Cloudflare は payment verification と enforcement を edge で行い、origin が高頻度の payment / authorization traffic を直接さばかなくてよい設計を強調している。
- **agent が一次的な買い手になる**: 記事は、agent が dataset、API call、tool、compute を人間の逐次承認なしに買う世界を前提にしている。これは [[loop-engineering]] の自律 loop に、予算・支払い・証跡の制御面が必要になることを意味する。
## 評価軸
この領域は便利な micropayment 機能というだけでなく、web の公共性や情報基盤の持続性に関わる。[[information-integrity]] の観点では、報道・専門知識・独立 creator の収益が advertising / referral から agent usage payment へ移る可能性がある一方、edge provider や payment protocol が access norm と価格形成を握る危険もある。
[[human-verified-advertising]] は「人間だけへ広告を出す」方向の対策だが、agentic web monetization は「非人間の利用にも明示的に価格を付ける」方向の対策である。両者は、AI agent によって attention economy が崩れるという同じ問題への別解として扱える。
## Open questions
- Agent が自律的に小額決済するとき、ユーザーの予算、同意、取り消し、監査ログをどこで管理するべきか。
- Payment proof と identity proof を分けるべき場面、結びつけるべき場面は何か。
- 公共性の高い情報や行政情報へ payment gate が広がると、accessibility や情報格差にどんな影響が出るか。
- LLM Wiki のような個人用 source ingestion は、商用 training crawler と違う扱いを受けられるのか。
+36
View File
@@ -0,0 +1,36 @@
---
title: AI Agent Command Safety
created: 2026-06-30
updated: 2026-07-01
type: concept
tags: [agent, security, reliability, automation]
sources: [raw/articles/guardfall-ai-coding-agent-shell-bypass-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/cursor-duneslide-sandbox-escape-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/koi-promptjacking-claude-desktop-rce-2026.md]
confidence: medium
---
# AI Agent Command Safety
AI agent command safety は、AI coding agent や computer-use agent が生成した shell command を、実際に実行される形で検査し、危険な動作を sandbox・承認・最小権限で抑える設計領域。[[ai-agent-identity-security]] が「どの権限で何にアクセスするか」を扱うのに対し、こちらは agent が出した具体的な command が shell や OS に解釈された後に何をするかを扱う。
GuardFall は、この領域が単なる blocklist では足りないことを示す事例である。The Hacker News の要約によると、Adversa AI は GuardFall を、bash が quote や省略表現を展開する前の平文 command だけを検査する guard の弱点として説明している。たとえば text matcher が `rm` を探しても、shell は `r''m` を `rm` として実行できる。つまり「モデルが出した文字列」と「bash が実行する argv」が一致しない限り、安全判定は抜け穴になる。^[raw/articles/guardfall-ai-coding-agent-shell-bypass-2026.md]
## 設計上の含意
- **Shell-aware parsing**: 危険語の文字列検索ではなく、shell と同じ解釈に近い tokenization / AST / argv レベルで検査する。記事では Continue が、bash が見る形で command を分解してから判定する設計で比較的耐えた例として挙げられている。
- **Sandbox before policy**: blocklist は補助であり、既定の network off、workspace 外書き込み制限、throwaway `$HOME`、container / OS sandbox の方が基礎になる。これは Codex の approvals / sandboxing ドキュメントが示す `read-only`、`workspace-write`、network policy の考え方と接続する。
- **No silent auto-exec on untrusted input**: fork PR、booby-trapped repository、package の偽 documentation、config file など、untrusted text が agent の command へ変換される経路では、自動実行や `dangerously-skip-permissions` 型の設定を避ける。
- **Command provenance**: どの file / prompt / tool result が command 生成に影響したかを残さないと、[[ci-cd-runtime-security]] のような実行時証跡や [[loop-engineering]] の verification とつながらない。
Cursor の DuneSlide 事例は、sandbox があるだけでは足りず、「agent が書ける場所」をどう解釈するかがそのまま脱出経路になることを示した。The Hacker News の要約によると、CVE-2026-50548 は `run_terminal_cmd` の `working_directory` を非既定 path にすると Cursor がその path を書き込み許可に追加してしまい、攻撃者が sandbox helper や shell startup file を上書きできる問題だった。CVE-2026-50549 は symlink の実体確認に失敗したとき in-project path を信用する fallback を悪用し、同じく project 外の helper を上書きできた。どちらも MCP や web search のような untrusted source からの prompt injection が、承認なしで local shell control へ進む構図である。^[raw/articles/cursor-duneslide-sandbox-escape-2026.md]
Claude Desktop まわりの事例は、command safety が「生成された shell 文字列」だけではなく、設定同期、MCP/extension、personal preferences、偽 error message まで含む広い実行経路の問題であることを示す。The Register の Pentera Labs 記事では、攻撃者が Claude の account-wide personalization に base64 prompt を入れ、Desktop Commander など command-capable MCP があれば reverse shell、なければ Anthropic 風の偽エラーと install prompt でユーザーに実行させる流れが説明されている。Koi の PromptJacking 報告では、公式 Claude Desktop extensions が unsandboxed MCP server として動き、AppleScript への未 escape URL 補間から web prompt injection → local RCE へ進みうると説明されている。どちらも「agent が command を出す瞬間」より前に、信頼済み assistant の設定・connector・外部 web content が command path へ混ざるため、設定変更監視、extension allowlist、connector sandboxing が command guard と同じ層で必要になる。^[raw/articles/theregister-claude-desktop-double-agent-2026.md] ^[raw/articles/koi-promptjacking-claude-desktop-rce-2026.md]
## なぜ重要か
Yuta の運用では、Hermes の scheduled job、Codex/Claude/OpenCode、local CLI、CI runner が同じ「agent が command を出す」面を共有する。便利な自走 loop ほど、guard を抜けた command が SSH key、cloud credential、wiki、repo、home directory へ届きやすい。したがって agent command safety は、個別 agent の機能ではなく、[[agent-oriented-cli-design]]、[[ai-agent-identity-security]]、[[ci-cd-runtime-security]] を横断する運用品質の条件として扱うべきである。
## Open questions
- shell-aware guard を各 agent が個別実装するのか、共通の command policy engine として切り出すべきか。
- bash 以外の shell、PowerShell、Python one-liner、package manager script、Makefile などをどの粒度で同じ policy にかけるべきか。
- local developer UX を壊さずに、auto-run と human approval の境界をどう観測・調整するか。
+34
View File
@@ -0,0 +1,34 @@
---
title: AI Agent Enabled Cyberattacks
created: 2026-07-02
updated: 2026-07-02
type: concept
tags: [agent, security, automation, reliability]
sources: [raw/articles/jadepuffer-langflow-agentic-ransomware-2026.md]
confidence: medium
---
# AI Agent Enabled Cyberattacks
AI agent enabled cyberattacks は、攻撃者が固定 playbook だけでなく LLM agent の tool-use loop を使い、侵入後の探索、資格情報の収集、横展開、データベース操作、破壊や恐喝までをその場で組み立てる攻撃パターン。[[ai-agent-command-safety]] が「自分の agent が危険な command を実行しない」ための防御なら、こちらは攻撃側も同じ command composition と output-reading loop を使えるという脅威モデルである。
The Hacker News の JADEPUFFER 記事では、Langflow の既知 RCE(CVE-2025-3248)を入口に、AI workflow 基盤上の API key、cloud credential、wallet key、database credential を探し、MinIO の既定認証情報、Nacos の古い認証 bypass と既定 signing key、MySQL root 接続をつないで、設定テーブルの暗号化・削除・身代金要求まで進めた事例として説明されている。個々の手口は新規性の高い 0-day ではなく、既知脆弱性、既定値、広すぎる credential、公開された管理面を agent が組み合わせた点が重要である。^[raw/articles/jadepuffer-langflow-agentic-ransomware-2026.md]
## 防御上の読み方
- **AI workflow server は credential concentrator になる**: Langflow のような agent / workflow builder は、LLM API key、cloud credential、database URL、secret manager token を環境変数や設定として持ちやすい。公開 RCE は単なる shell ではなく、複数サービスへの pivot point になる。これは [[ai-agent-identity-security]] の最小権限・token 分離の問題でもある。
- **古い既知脆弱性の価値が上がる**: agent が探索・試行・失敗修正を安価に回せるほど、未 patch の既知 CVE、既定 password、公開管理 port は「人間が丁寧に狙う対象」から「機械が広く試す対象」へ寄る。
- **runtime evidence が必要になる**: 記事は、攻撃側のコード内コメント、自己修正、600 以上の payload、短時間の pivot を AI 駆動の兆候として扱っている。防御側も endpoint / network / database / CI の実行時 trace を残さないと、何が自動で連鎖したのかを後から説明できない。ここは [[ci-cd-runtime-security]] の runner 監視や [[loop-engineering]] の証跡保存と同じ設計原理である。
- **復旧不能な破壊を前提にする**: 記事の例では暗号鍵が保存・送信されず、支払っても復旧できない可能性があるとされる。したがって ransom negotiation より、隔離、credential 失効、backup 検証、blast radius の縮小が主防御になる。
## Open questions
- agent 駆動らしさを、単なる速いスクリプトや SOAR からどう区別して検知するか。
- AI workflow / notebook / MCP server を、通常の web app より強い credential isolation と outbound policy で扱うべきか。
- defensive agent を使う場合、攻撃 agent と同じ speed/cost advantage をどこまで incident response に持ち込めるか。
## Related
- [[ai-agent-command-safety]]
- [[ai-agent-identity-security]]
- [[ci-cd-runtime-security]]
+13 -2
View File
@@ -1,10 +1,10 @@
---
title: AI Agent Identity Security
created: 2026-06-29
updated: 2026-06-30
updated: 2026-07-02
type: concept
tags: [agent, security, reliability, privacy]
sources: [raw/articles/okta-cross-app-access-partners-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md]
sources: [raw/articles/okta-cross-app-access-partners-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md, raw/articles/ios-ai-chatbot-llm-key-leakage-2026.md, raw/articles/unity-terms-agentic-access-2026.md, raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/cloudflare-monetization-gateway-x402-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/xai-voice-agent-builder-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/email-verification-protocol-draft-2026.md]
confidence: medium
---
@@ -20,8 +20,19 @@ AI agent identity security は、AI エージェントやアプリ間連携が
- **Resource app / MCP server**: Asana、Atlassian、Figma、Linear、Slack、Supabase、Datadog などが、エージェントに文脈や業務データを渡す側になる。
- **Policy and audit**: アクセスが許可される前に企業ポリシーで検査し、操作の監査証跡を残す。これは [[loop-engineering]] の persistence と verification をセキュリティ境界へ移したものでもある。
- **Least privilege for agents**: 常時広い権限を持つ bot token ではなく、必要な範囲に絞った identity-based token を使う。
- **Secret custody outside the model**: 1Password Environments MCP Server for Codex は、coding agent を secret の保管庫ではなく「承認された利用主体」として扱う設計例である。Codex は environment を作成し、変数名を扱い、実行を orchestrate できるが、secret value は MCP channel、model context、local file、terminal へ返さず、1Password が承認済み process の runtime memory にだけ注入する。これにより、agent workflow の速度を保ちながら、credential custody、explicit approval、scope、audit を [[agent-harness-engineering]] 側の実行 loop へ組み込める。^[raw/articles/1password-codex-mcp-secret-access-2026.md]
- **Browser-mediated identity assertions**: Email Verification Protocol draft は、email verification を「メールを送って code を入力させる」方式から、browser が relying party と issuer の間を仲介して signed token を受け渡す方式へ寄せる。issuer は RP identity を直接知る必要がなく、RP は nonce と browser key binding で token を検証するため、friction reduction と privacy separation を同時に狙う標準化案として読める。これは agent 固有ではないが、agent が account creation や delegated workflow を扱う時代には、identity assertion を browser / issuer / RP に分け、過剰な identifier sharing を避ける設計として隣接する。^[raw/articles/email-verification-protocol-draft-2026.md]
- **Local sandbox / approval boundary**: Codex の安全運用ドキュメントは、cloud では隔離 container、CLI/IDE では OS sandbox と approval policy を組み合わせ、既定で network access を切り、workspace 外の編集や network 利用を承認対象にする設計を説明している。`workspace-write`、`read-only`、network proxy、domain allow/deny などの設定は、企業の cross-app 認可だけでなく個人の agent loop でも「どこまで自動実行してよいか」を明示する制御面になる。
- **Browser / device permission boundary**: GitHub Copilot の VS Code browser tools GA は、agent が実ブラウザを開き、click/type/drag、console error、screenshot、scripted flow を使えるようにする一方、人間が開いた tab は `Share with Agent` するまで読めず、agent tab は fresh session で cookie/storage から隔離され、camera/microphone/geolocation は既定拒否になると説明している。browser が agent tool になるほど、tab ownership、session isolation、site allow/deny、workspace trust は identity boundary の一部になる。^[raw/articles/github-copilot-browser-tools-ga-2026.md]
- **Local browser data boundary**: [[safari-mcp-server]] は local に動き、自身では network call せず、AutoFill などの個人情報にはアクセスしないと説明されている。ただし page content、screenshot、console log は接続先 agent へ渡るため、どの agent を信頼するか、どの site/tab を見せるか、captured data が vendor 側でどう扱われるかは運用上の identity / privacy boundary になる。^[raw/articles/safari-mcp-server-webkit-2026.md]
- **Repository governance as identity boundary**: Microsoft Learn の agent architecture / SDLC module は、agent task を input / output / success criteria で定義し、PR template、checks、CODEOWNERS、rules、environment gate を通じて「どの変更が誰の承認で通るか」を設計する。これは [[agent-oriented-cli-design]] の tool-level clarity と同じく、agent の行動を監査可能な境界へ置く方法である。
- **Platform-designated agent access**: Unity の 2026-06-30 Terms of Service は、AI agents、LLM、MCP clients / servers が Unity platform とやり取りする場合、Unity が運用または指定する framework 経由に限ると明記している。これは XAA のような cross-app authorization とは別に、resource platform 側が「どの agent gateway なら許すか」を契約と access policy で決める方向を示す。利用者は account / credential 経由で動く automated caller の責任を負うため、[[agent-harness-engineering]] の tool boundary と契約上の identity boundary が重なる。^[raw/articles/unity-terms-agentic-access-2026.md]
- **Client credential exposure**: iOS の LLM chatbot 調査では、444 本中 282 本が plaintext API key、認証なし backend、再利用可能 token のいずれかで有料 LLM access を露出していた。AI 機能を mobile app に載せるだけでも、key を client に埋め込まない、backend が呼び出し元を検証する、漏れた key を revoke する、といった基本的な identity boundary が実務上の cost / privacy / abuse boundary になる。^[raw/articles/ios-ai-chatbot-llm-key-leakage-2026.md]
- **Telemetry identity leakage**: Claude Code 2.1.196 audit は、source code 本文を送らなくても git remote URL の hash、GitHub Actions の actor/repository ID、account/org UUID、machine/session ID のような識別子が telemetry に載りうると指摘している。[[ai-agent-telemetry-privacy]] では、agent の権限境界だけでなく、agent vendor へ流れる作業文脈の最小化と opt-out の実効性も identity security の一部として扱う。^[raw/articles/claude-code-telemetry-audit-2026.md]
- **Payment as access credential**: Cloudflare Monetization Gateway / x402 は、agentic buyer が request に payment proof を添えて web page、API、dataset、MCP tool にアクセスする設計を示す。支払い証明は一種の credential になるが、記事は同時に Web Bot Auth などで agent identity を求められる余地も残している。[[agentic-web-monetization]] では、誰の agent が、どの予算で、どの resource を買ったかを identity / audit 境界として扱う必要がある。^[raw/articles/cloudflare-monetization-gateway-x402-2026.md]
- **Voice agent as delegated operator**: xAI Voice Agent Builder は、電話番号/SIP、Gmail、Google Calendar、Outlook、Linear、Notion、OneDrive、custom MCP、knowledge base、guardrails、call playback を一体化した no-code voice agent として提示されている。電話応答 agent は単なる chat UI ではなく、顧客本人確認、PII、社内 system 操作、人間への handoff を扱う delegated operator になるため、誰の声・番号・tool 権限で何を実行したかの audit が必要になる。^[raw/articles/xai-voice-agent-builder-2026.md]
- **Synced assistant settings as identity surface**: Claude Desktop の personalization / preferences は account-wide に同期され、The Register の Pentera Labs 記事ではここに攻撃 prompt を入れることで、別端末の Claude Desktop と command-capable MCP connector へ影響を広げられると説明されている。agent identity security では token だけでなく、sync される instruction、skill、extension 設定も「どの actor が変更し、どの端末へ反映されたか」を監査すべき対象になる。^[raw/articles/theregister-claude-desktop-double-agent-2026.md]
- **Command execution boundary**: 権限や token が正しくても、agent が shell command をどう生成・実行するかには別の危険がある。[[ai-agent-command-safety]] は、GuardFall のように text guard と shell interpretation がずれる問題を扱う隣接領域である。
## なぜ重要か
+29
View File
@@ -0,0 +1,29 @@
---
title: AI Agent Telemetry Privacy
created: 2026-07-01
updated: 2026-07-01
type: concept
tags: [agent, privacy, security, data-protection, reliability]
sources: [raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/theregister-claude-code-transcript-retention-2026.md]
confidence: medium
---
# AI Agent Telemetry Privacy
AI agent telemetry privacy は、coding agent や CLI agent が利用状況、エラー、trace、repo 情報、transcript をどこへ送り、どの opt-out が何を止めるのかを扱う論点。[[ai-agent-identity-security]] が「agent が他サービスへ何の権限でアクセスするか」を扱うのに対し、こちらは agent 自体が operator や作業環境について何を観測・送信・保存するかに焦点を置く。
Adnane Khan の Claude Code 2.1.196 audit は、公開文書の「Statsig metrics + Sentry errors」という説明と実装がずれている可能性を示す。bundle 内では Statsig/Sentry SDK ではなく、Anthropic 1P OTLP event logging、Datadog logs、Datadog error tracking、ユーザー設定の 3P OTLP が見つかったとされる。特に、Datadog feature event path が server-side gate に依存し、`DISABLE_TELEMETRY` / `DO_NOT_TRACK` / `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` だけでは完全に止まらない可能性を指摘している。^[raw/articles/claude-code-telemetry-audit-2026.md]
## 見るべき軸
- **送信先の透明性**: agent の telemetry が vendor 直送、Datadog などの third-party、企業内 collector、local file のどれへ流れるかを区別する。公開文書の vendor 名や endpoint が古いと、利用者は実際の data processor を評価できない。
- **Opt-out の実効性**: 環境変数や設定が、metrics、error reporting、feature gate exposure、3P OTLP、update check をそれぞれ止めるかを pipeline ごとに見る。単一の `DISABLE_TELEMETRY` が「全部止まる」とは限らない。
- **Repo / CI identity**: audit は 1P event に git remote URL の 16 文字 SHA-256、GitHub Actions では actor や repository ID が載ると指摘している。source code 本文ではなくても、どの repo で agent を使ったかは強い作業文脈になる。
- **Error stack and local paths**: error tracking は prompt や file content を送らなくても、stack frame に file path や project layout が混ざる可能性がある。これは [[ai-agent-command-safety]] の command provenance と同じく、debuggability と漏えいリスクの境界になる。
- **Local transcript retention**: [[loop-engineering]] では transcript が検証・再開・説明責任の材料になるが、長期保存は credential や source code を抱える危険にもなる。保存期間、削除ログ、backup、export の仕様は privacy 機能でも reliability 機能でもある。
## 運用上の含意
Yuta の agent 運用では、telemetry は単純な「送る/送らない」ではなく、loop の観測性と privacy の交換条件として扱う必要がある。OpenAI Codex の approvals/security docs が示すように、OTel を自分の collector へ送れる設計は [[agent-harness-engineering]] の観測性を高める一方、collector 側の retention と access control を同時に決めなければならない。
この論点は [[data-protection-and-expression]] とも接続する。agent が開発者の作業文脈を観測するほど、利用者の control、説明、削除、第三者提供の透明性が重要になる。特に CLI agent は IDE より権限が広く、shell、repo、CI、browser-use、MCP server へまたがるため、telemetry 設計を product quality と security boundary の一部として読むべきである。
+4 -2
View File
@@ -1,10 +1,10 @@
---
title: AI-Assisted Reverse Engineering
created: 2026-06-29
updated: 2026-06-29
updated: 2026-07-02
type: concept
tags: [security, dev-tool, agent, automation, quality]
sources: [raw/articles/ghidra-mcp-2026.md]
sources: [raw/articles/ghidra-mcp-2026.md, raw/articles/fluxsec-windows-11-kernel-reverse-engineering-etwti-2026.md]
confidence: medium
---
@@ -14,6 +14,8 @@ AI-assisted reverse engineering は、binary 解析、逆コンパイル、型
重要なのは、LLM に「それらしく読ませる」だけでは品質が安定しないこと。逆解析では、関数名、型、構造体、呼び出し関係、根拠コメントが後続作業の足場になるため、一度の推測ミスが広く伝播する。Ghidra MCP の README は、命名規則、型変更の拒否、文書化の完全性得点、batch operation、transaction といった仕組みを通じて、作業のばらつきを道具側で抑えようとしている。
FluxSec の Windows 11 kernel / ETW:TI write-up は、AI 支援の有無にかかわらず reverse engineering の良い検証 loop を示す例として使える。`NtWriteVirtualMemory` から undocumented `MiReadWriteVirtualMemory`、`PsIsProcessLoggingEnabled`、`EtwTiLogReadWriteVm` へ進み、decompiler の読みを KPCR/KTHREAD/KPROCESS/EPROCESS offset、WinDbg breakpoint、bitfield 確認、driver 実装で検証している。AI エージェントに逆解析を手伝わせる場合も、このように「静的な推測 → 動的な観測 → 最小実装で再現」の loop を harness 側で要求しないと、もっともらしい構造体名や flag 解釈が事実として固着しやすい。^[raw/articles/fluxsec-windows-11-kernel-reverse-engineering-etwti-2026.md]
## 見るべき軸
- **読み取りから書き込みへ**: AI が decompile 結果を要約するだけでなく、rename、retype、comment、structure creation まで行うなら、取り消しや検査の境界が必要になる。
+40
View File
@@ -0,0 +1,40 @@
---
title: AI Crawler Governance
created: 2026-07-01
updated: 2026-07-02
type: concept
tags: [agent, automation, information-integrity, media, public-interest]
sources: [raw/articles/cloudflare-ai-traffic-options-2026.md, raw/articles/cloudflare-content-independence-day-ai-options-2026.md, raw/articles/hakuhodo-human-verified-ad-2026.md, raw/articles/cloudflare-content-independence-day-2025.md, raw/articles/cloudflare-monetization-gateway-x402-2026.md]
confidence: medium
---
# AI Crawler Governance
AI crawler governance は、検索、AI agent、model training crawler などの自動アクセスを、サイト運営者・読者・広告・AI 事業者の利害に合わせて分類し制御する設計領域。Cloudflare の 2026-07-01 changelog は、AI traffic を **Search**、**Agent**、**Training** の 3 種類へ分け、顧客がそれぞれ allow / block / 広告表示ページだけ block を選べるようにした。
この分類は、従来の bot 対策より細かい。Search は質問回答や検索で後から referral や補償が期待される crawling、Agent は chat fetch bot や browser-use agent のように人の代理でリアルタイムに動く activity、Training は model の training / fine-tuning のために content を持ち帰る activity とされる。Cloudflare は 2026-09-15 から新規 domain では Training と Agent を広告表示ページ上で既定 block、Search は allow にする予定だとしている。^[raw/articles/cloudflare-ai-traffic-options-2026.md]
Cloudflare の同日 blog は、単に「AI bot」を定義するのではなく、サイト上で何をしているか、何を保存するか、どう再共有するかで分類する姿勢を明確にした。特に、多目的 crawler は Search / Agent / Training を一つの user-agent に混ぜるのではなく目的別に分けるべきだとし、site owner が用途ごとに許可・拒否・広告付きページのみ拒否を選べる状態を透明性の条件としている。これは [[ai-agent-identity-security]] の認可・監査だけでなく、[[agentic-web-monetization]] のような支払い付き access policy の前提にもなる。^[raw/articles/cloudflare-content-independence-day-ai-options-2026.md]
## なぜ重要か
この論点は [[information-integrity]] の「情報基盤の責任」を、AI 時代の web access policy へ移したものとして読める。記事を読む bot がすべて同じではないなら、robots.txt 的な単純な allow/deny だけでは、検索流入を残しつつ training だけ拒否する、あるいは人の代理 agent は許すが広告収益を奪う crawling は止める、といった判断ができない。
また、[[human-verified-advertising]] と同じく、非人間トラフィックが広告やメディアの収益モデルをどう壊すかという問題でもある。広告付きページ上で Agent / Training を既定 block する設計は、AI crawler が content だけでなく広告露出・読者接触・効果測定の前提を迂回しうることを認めている。
Cloudflare の “Content Independence Day” 投稿は、この制御を単なる bot 管理ではなく web の経済モデル再設計として位置づける。旧来の検索は「content をコピーする代わりに traffic を返す」取引だったが、AI answer / AI Overview では derivative answer が消費され、元サイトへの流入が大きく減る。Cloudflare は 2025-07-01 に AI crawler を既定で block し、支払いなしの crawl を拒む方向を示し、将来的には traffic ではなく「AI engine の知識の穴をどれだけ埋めるか」で content value を測る marketplace を構想している。これは Search / Agent / Training の分類に、補償・価値測定・publisher bargaining power の層を足す。^[raw/articles/cloudflare-content-independence-day-2025.md]
2026-07-01 の Cloudflare Monetization Gateway 発表は、この流れを [[agentic-web-monetization]] としてさらに広げる。Pay Per Crawl が crawler に content への支払いを求める段階だったのに対し、Monetization Gateway は API、dataset、MCP tool call など任意の resource に x402 payment rule を置き、agentic buyer が request ごとに支払う構想である。つまり crawler governance は allow/block の policy だけでなく、価格、payment proof、identity、origin protection を含む access economy の設計へ進んでいる。^[raw/articles/cloudflare-monetization-gateway-x402-2026.md]
## 見るべき軸
- **分類の説明可能性**: crawler が Search / Agent / Training のどれに分類されたかを、サイト運営者が後から理解できる必要がある。
- **補償と referral**: Search は許す、Training は拒否するという区別は、content が reader/revenue を返すかどうかを中心にしている。
- **人の代理性**: Agent traffic は bot だが、背後に人間の意図がある場合がある。ここは [[ai-agent-identity-security]] の「誰の権限で行動しているか」と接続する。
- **既定値の政治性**: Cloudflare のような edge provider が新規 domain の default を決めると、個々の publisher だけでなく web 全体の AI access norm を形作る。
## Open questions
- AI crawler の self-identification が信頼できないとき、分類は header、IP reputation、behavior、契約のどれに依存するべきか。
- 個人サイトや OSS docs は、AI agent に読ませたい場合と training を拒否したい場合をどう分けるべきか。
- Search / Agent / Training の区別は、[[llm-wiki-pattern]] のような source ingestion とどう折り合うか。個人の知識管理のための読み取りと、大規模 training の収集は同じ「AI が読む」ではない。
+3 -1
View File
@@ -4,7 +4,7 @@ created: 2026-06-28
updated: 2026-06-30
type: concept
tags: [law, public-interest, privacy, security, data-protection]
sources: [raw/articles/ravi-naik-awo-profile-2026.md, raw/articles/boj-ai-legal-risk-financial-institutions-2026.md]
sources: [raw/articles/ravi-naik-awo-profile-2026.md, raw/articles/boj-ai-legal-risk-financial-institutions-2026.md, raw/articles/ios-ai-chatbot-llm-key-leakage-2026.md]
confidence: medium
---
@@ -31,3 +31,5 @@ AI 開発者の責任は、個別の事故対応にとどまらない。責任
現時点ではこのページは Ravi Naik / AWO profile という単一資料からの入口であり、具体的な法理や裁判上の争点は今後の資料で補う必要がある。
日本銀行金融研究所の「金融機関におけるAI利用に伴う私法上のリスクと管理」は、個人被害や deepfake とは別の角度から、金融機関が AI 開発者・提供者に契約責任を追及する場合、AI を使ったサービスを顧客へ提供する場合、組織内部で取締役が AI ガバナンス体制を構築する場合を整理している。ここでは AI 開発者責任は不法行為だけでなく、契約条項、顧客との説明・合意、内部統制としても現れる。これは [[ai-agent-identity-security]] の権限境界や監査ログが、事故後の説明責任だけでなく契約上の管理義務にも関わることを示す。
iOS の LLM chatbot 444 本を調べた研究では、282 本が plaintext API key、認証なし backend、再利用可能 token のいずれかで有料 LLM access を露出していたとされる。これは「AI 機能を追加した」だけではなく、client に key を置く、backend の認可を省く、漏えい後に revoke しない、といった設計判断が金銭的損害や abuse の責任に直結する例である。[[ai-agent-identity-security]] の最小権限・監査・取り消しは、企業 agent だけでなく消費者向け AI app でも基本線になる。^[raw/articles/ios-ai-chatbot-llm-key-leakage-2026.md]
+17 -2
View File
@@ -1,10 +1,10 @@
---
title: AI Evaluation Infrastructure
created: 2026-06-30
updated: 2026-06-30
updated: 2026-07-02
type: concept
tags: [evaluation, llm, quality, workflow]
sources: [raw/articles/arena-ai-leaderboard-business-2026.md]
sources: [raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/harbor-langchain-agent-eval-stack-2026.md, raw/articles/anthropic-claude-sonnet-5-2026.md, raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md, raw/articles/openai-genebench-pro-2026.md, raw/articles/shopify-flow-agent-model-optimization-flywheel-2026.md, raw/articles/vllm-semantic-router-micro-agents-2026.md]
confidence: medium
---
@@ -14,12 +14,27 @@ AI evaluation infrastructure は、LLM や agent の性能を、単発 benchmark
重要なのは、評価が単なる研究補助ではなく、モデル改善・post-training・企業導入判断の市場そのものになっている点。Arena は text、coding、vision、image generation に加え、Agent Mode のような長時間 workflow も扱う。これは [[loop-engineering]] や [[agentic-hardware-design]] のようなエージェント運用で、最終成果だけでなく、途中の意思決定・失敗・回復をどう測るかという問題に接続する。
LangChain と Harbor の統合記事は、agent 評価では「環境」「指示」「検証スクリプト」を task として束ね、各 trial を clean sandbox で並列実行し、最後に deterministic check を走らせる必要があると整理している。Deep Agents / LangGraph の entrypoint、LangSmith Sandbox、LangSmith tracing を Harbor に接続する構成は、agent 評価を単なる回答採点ではなく、ファイル変更・shell 実行・状態変化まで含む再現可能な実験として扱う具体例である。これは [[ai-agent-command-safety]] や [[ci-cd-runtime-security]] とも接続し、評価環境そのものが安全境界になることを示す。^[raw/articles/harbor-langchain-agent-eval-stack-2026.md]
Anthropic の Claude Sonnet 5 発表は、モデル提供者自身の launch post も評価基盤の一部になっていることを示す。BrowseComp、OSWorld-Verified、agentic safety、prompt-injection resistance、cyber capability、misalignment audit などを、価格・effort level・モデル選択と一緒に提示しており、開発者は「高いモデルを使うか」ではなく、task ごとの cost-performance と安全境界でモデルを選ぶようになる。これは [[loop-engineering]] の evaluator 設計や [[ai-agent-command-safety]] の権限制御と同じ問題系にある。^[raw/articles/anthropic-claude-sonnet-5-2026.md]
Fable 5 / Mythos 5 の再展開記事では、評価基盤が政府協議・業界標準・脆弱性報告 triage にまで広がる。Anthropic は jailbreak の深刻度を capability gain、breadth、weaponization、discoverability で測る枠組みを提案しており、これは [[ai-jailbreak-severity-framework]] として、model launch 前の red-team だけでなく launch 後の 24/7 監視、HackerOne 報告、政府機関による独立評価をつなぐ層になる。^[raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md]
OpenAI の GeneBench-Pro は、評価対象が「正答を出せるか」から、曖昧な研究データを診断し、仮説や分析方針を修正し、下流判断に耐える結論へ閉じる「research taste」へ広がっていることを示す。129 問の合成 computational biology 課題は、因果構造とデータ生成過程を制御することで、もっともらしいが誤った分析が通らないように設計され、外部専門家 review と deterministic grading を組み合わせる。これは [[claude-science]] や [[ai-research-automation]] と同じ科学 agent 領域で、評価が実行環境・データ・分析 trace・判断品質まで含む必要があることを補強する。^[raw/articles/openai-genebench-pro-2026.md]
Shopify の Flow agent fine-tuning 記事は、evaluation infrastructure が production data flywheel と一体化する例である。Sidekick の自然言語→Flow automation 生成では、最初は既存の本番 workflow から synthetic query と tool trajectory を逆算し、hand-crafted benchmark と LLM judge / syntactic checker で評価した。しかし 1% 本番投入では、offline benchmark が同等でも workflow activation rate が 35% 低く、実利用では workflow editing、email configuration、third-party integration、質問だけの会話などが抜けていた。そこで会話を facet 別 LLM judge と tag slice で診断し、高品質な本番会話を training pool へ戻し、低品質な会話を review 隔離し、週次 retraining へ接続している。これは [[loop-engineering]] と [[agent-harness-engineering]] における evaluator が、単発採点ではなく実運用の改善ループになることを示す。^[raw/articles/shopify-flow-agent-model-optimization-flywheel-2026.md]
vLLM の Semantic Router / micro-agent 構想は、評価と orchestration が serving layer に入り込む例である。単一の OpenAI-compatible model ID の裏で、router が task に応じて recipe を選び、複数 worker に fan-out し、quorum、disagreement check、output contract repair、synthesis を行う。これは「どの model が強いか」を外から測るだけでなく、router 自体が小さな evaluator / coordinator になり、frontier model 呼び出しの前段で capability と cost/safety policy を組み立てるという設計である。[[loop-engineering]] や [[agent-harness-engineering]] では、application graph だけでなく inference gateway も評価・検証・合議の場になる。^[raw/articles/vllm-semantic-router-micro-agents-2026.md]
## なぜ重要か
- **Crowdsourced comparison**: 利用者が 2 つのモデル出力を比較する形式は、静的な benchmark では拾いにくい実利用の好みを集められる。
- **Evaluation as business**: 無料 leaderboard の背後で、詳細分析や model lab 向け評価が商用サービスになる。
- **Post-training demand**: Arena は、人間評価やラベリングを提供する Mercor、Surge、Scale AI などと同じ予算を争うと説明されており、評価と訓練改善が近づいている。
- **Agent evaluation**: 長時間 workflow や Agent Mode が評価対象になると、[[wiki-maintenance-loop]] のような自走ジョブでも、単一回答の品質ではなく状態更新・検証・永続化まで測る必要が出る。
- **Judgment-heavy scientific evaluation**: GeneBench-Pro のような benchmark は、正解率だけでなく、データ診断、分析方針の変更、因果推論、solver contract の明確さまで評価対象にする。科学 agent を評価するには、clean sandbox と deterministic check だけでなく、trace から判断品質を検査できる問題設計が必要になる。
- **Production feedback flywheels**: Shopify Flow の例では、synthetic benchmark、LLM judge、programmatic checker、本番 activation rate、slice analysis、週次 retraining が一つの改善ループになる。評価基盤は「合格判定」ではなく、どのデータを足し、どの形式を変え、どの tool response を削るかを決める運用面になる。
- **Serving-layer evaluators**: vLLM Semantic Router のように、model router が quorum、disagreement check、output repair を実行すると、評価は offline benchmark だけでなく、inference request ごとの制御面にも入る。
## Open Questions
@@ -0,0 +1,38 @@
---
title: AI Jailbreak Severity Framework
created: 2026-07-01
updated: 2026-07-01
type: concept
tags: [llm, evaluation, security, reliability]
sources: [raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md]
confidence: medium
---
# AI Jailbreak Severity Framework
AI jailbreak severity framework は、LLM の安全機構を迂回する手法を「成功した/しなかった」だけでなく、どれほど危険で、どれほど急いで直すべきかを共通尺度で扱うための枠組み。[[ai-evaluation-infrastructure]] がモデル能力や agent workflow を測る基盤を扱うのに対し、この概念は安全評価・脆弱性報告・政府や業界との連絡を同じ triage 言語にそろえることを目指す。
Anthropic は Fable 5 / Mythos 5 の輸出管理解除と再展開の記事で、Amazon 研究者の報告を契機に Fable 5 のサイバーセキュリティ安全分類器を強化し、該当の bypass を 99% 超で遮断するようにしたと説明している。一方で、この強化は日常的な coding / debugging の benign request を誤検知しやすくするため、モデル安全は単純な拒否率ではなく、危険行為の取り逃しと正当利用の阻害を同時に測る必要がある。^[raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md]
## 評価軸
Anthropic の提案は、jailbreak を少なくとも次の 4 軸で見る。
| 軸 | 見るもの | 実務上の意味 |
|---|---|---|
| Capability gain | 既存ツールや弱いモデルをどれだけ超える能力を開くか | 既存手段で同じことができるなら緊急度は下がる |
| Breadth of capability gain | 同じ手法が何種類の攻撃や対象に効くか | narrow jailbreak と universal jailbreak を分ける |
| Ease of weaponization | 攻撃へ変えるための人間の手間や再試行回数 | 一発で動くほど対処優先度が上がる |
| Discoverability | 手法の入手しやすさ | 公開済み・拡散済みなら被害化が早い |
この見方は、ソフトウェア脆弱性に CVSS があるように、AI jailbreak にも報告・修正・公開・政府連絡の共通語が必要だという立場に近い。ただし jailbreak の挙動はモデル更新、classifier、prompt、tool 接続、権限境界で変わるため、単一スコアだけで安定的に扱うのは難しい。
## Agent 運用との接続
Agentic coding や自律 job では、jailbreak は chat の不適切回答だけでなく、tool call、shell command、credential、外部 API、永続状態に波及する。したがってこの枠組みは [[ai-agent-command-safety]] や [[ci-cd-runtime-security]] と接続する。たとえば「危険な command を 1 回で出させる」jailbreak は、会話上の失敗よりも実行環境上の影響が大きい。逆に、sandbox・approval・read-only database・network 制限があれば、同じ model-level jailbreak でも運用上の severity は下がる。
## Open questions
- CVSS のような数値化を、モデル・classifier・tool 権限・実行 sandbox が絡む agent system にどう拡張するか。
- 政府や大手 model provider が作る共通枠組みを、個人や小規模 OSS の red-team / disclosure workflow へどう軽量化するか。
- Benign coding/debugging の false positive を増やさずに、危険な capability gain だけを抑える評価データをどう作るか。
+3 -1
View File
@@ -4,7 +4,7 @@ created: 2026-06-28
updated: 2026-06-28
type: concept
tags: [llm, agent, automation, workflow, dev-tool]
sources: [raw/articles/tokium-self-evolving-ai-researcher-2026.md]
sources: [raw/articles/tokium-self-evolving-ai-researcher-2026.md, raw/articles/claude-science-ai-workbench-2026.md]
confidence: medium
---
@@ -18,6 +18,8 @@ AI まわりの変化を追う仕組みは、単に検索結果を集めるだ
運用面では、手順書としての skill とシェルスクリプトを分けている。収集・報告・自動見直しの具体手順を Markdown に置き、スクリプトは実行順序、並列実行、再実行しやすさ、上限ターン数、部分失敗の許容を担当する。この分担は [[wiki-maintenance-loop]] と同じく、人間が毎回判断しなくても続く手入れの形である。
Claude Science pushes the same theme into scientific workbenches: research automation is not only periodic news gathering, but also tool-connected analysis where code, figures, compute environment, citations, and reviewer-agent feedback are preserved as auditable artifacts. Its design suggests that useful research automation needs both connectors to domain sources and a way to keep execution history reproducible enough for later validation. See [[claude-science]] and [[ai-evaluation-infrastructure]].^[raw/articles/claude-science-ai-workbench-2026.md]
## Open Questions
- 報告に採用された回数を、短期の流行と長期の価値のどちらとして扱うか。
+17 -3
View File
@@ -1,16 +1,16 @@
---
title: CI/CD Runtime Security
created: 2026-06-30
updated: 2026-06-30
updated: 2026-07-02
type: concept
tags: [security, supply-chain, quality, reliability, automation]
sources: [raw/articles/cicd-sensor-2026.md, raw/articles/scrutineer-oss-security-workflow-2026.md]
sources: [raw/articles/cicd-sensor-2026.md, raw/articles/scrutineer-oss-security-workflow-2026.md, raw/articles/tangled-spindle-microvm-ci-runners-2026.md, raw/articles/thehackernews-argo-cd-repo-server-unpatched-rce-2026.md, raw/articles/synacktiv-argo-cd-codeql-rce-2026.md, raw/articles/zenn-ai-generated-github-actions-yaml-security-2026.md, raw/articles/flatt-github-actions-credential-leakage-2026.md, raw/articles/github-secret-scanning-public-monitoring-2026.md, raw/articles/microsoft-ghqr-github-quick-review-2026.md, raw/articles/strix-ai-pentesting-agent-2026.md]
confidence: medium
---
# CI/CD Runtime Security
CI/CD runtime security is the practice of observing and constraining what actually runs inside build, test, release, and deployment jobs. The core problem is that CI jobs hold cloud credentials, signing keys, package-registry tokens, and deployment authority, while compromised dependencies or scripts can execute inside short-lived jobs and disappear with the evidence when the job ends. [[ai-agent-identity-security]] covers adjacent authorization and audit concerns for agents; [[loop-engineering]] is relevant because autonomous development loops often depend on these pipelines as their verification and deployment boundary.
CI/CD runtime security is the practice of observing, constraining, and isolating what actually runs inside build, test, release, and deployment jobs. The core problem is that CI jobs hold cloud credentials, signing keys, package-registry tokens, and deployment authority, while compromised dependencies or scripts can execute inside short-lived jobs and disappear with the evidence when the job ends. [[ai-agent-identity-security]] covers adjacent authorization and audit concerns for agents; [[loop-engineering]] is relevant because autonomous development loops often depend on these pipelines as their verification and deployment boundary.
## Why it matters
@@ -20,6 +20,20 @@ Traditional software supply-chain controls often answer where an artifact came f
`cicd-sensor` uses an eBPF-powered sensor for GitHub Actions and GitLab CI/CD. Its baseline detections use process ancestry and correlated signals: for example, credential access by a process descended from `npm install`, or one job reading several credential categories. It can emit per-run logs, graphical job summaries, cloud-routed evidence, and build attestations while keeping data in the operator's own infrastructure rather than sending it to a project-operated SaaS.^[raw/articles/cicd-sensor-2026.md]
Tangled's Spindle microVM engine shows the complementary isolation side of the same problem. Each workflow boots a small QEMU microVM, runs a guest agent over vsock, executes steps as an unprivileged `spindle-workflow` user, and can build a NixOS guest configuration from the workflow file itself. Network access is routed through separate namespaces, slirp layers, DNS filtering, and blackholed special-use ranges so the guest can reach the internet without reaching the host or local private networks. This makes CI runner design part of the trust boundary, not just a scheduling detail.^[raw/articles/tangled-spindle-microvm-ci-runners-2026.md]
Argo CD の repo-server 欠陥は、deployment controller 自体が CI/CD runtime boundary になることを示す。The Hacker News / Synacktiv の報告では、内部 gRPC port に届く attacker が kustomize の `--helm-command` 経由で repo-server 上の code execution を得て、さらに Redis password を読んで deployment cache を poison し、次回 sync で attacker workload を cluster に入れられる。Synacktiv の一次解説は、CodeQL で taint flow を追い、repo-server の request handling から Kubernetes cluster compromise へつながる exploit path と自動化 tool まで示している。Helm chart では network policy が既定で無効なため、「cluster 内部だから安全」という前提が壊れる。CI/CD runtime security では runner だけでなく、repo-server、Redis/cache、GitOps controller、sync loop も最小到達性・署名・監査の対象になる。^[raw/articles/thehackernews-argo-cd-repo-server-unpatched-rce-2026.md] ^[raw/articles/synacktiv-argo-cd-codeql-rce-2026.md]
AI が生成した GitHub Actions YAML は、動作確認だけでなく trigger、checkout 対象、token 権限、cache / artifact / secret の信頼境界を人間が読む必要がある。Zenn の整理では、`pull_request_target` で外部 PR 側の code を checkout して実行しないこと、`permissions` を省略せず read-only から始めること、untrusted trigger からの cache を信用しないこと、`workflow_run` や artifact 経由で権限境界をまたがないことが確認点として挙げられている。これは [[ai-agent-command-safety]] と接続し、AI が作った自動化設定そのものを privileged runtime code としてレビューする必要を示す。^[raw/articles/zenn-ai-generated-github-actions-yaml-security-2026.md]
GMO Flatt Security の GitHub Actions 解説は、OIDC / Trusted Publishing を入れても「認証後に runner 上へ置かれる派生クレデンシャル」は残る、という runtime 視点を強調する。`GITHUB_TOKEN` は `actions/checkout` の credential persistence や `Runner.Worker` のメモリから読まれうるし、AWS/GCP/Azure/Docker などの認証 Action は一時クレデンシャルや設定ファイルを後続 step から到達可能な場所へ置く。Environment 保護、ruleset、claim の数値 ID 検証、job 分離、短い session duration は有効だが、依存関係・Action・正規レビュアー経由で信頼済み経路に悪意ある code が入ると、漏洩を完全には防げない。したがって runner 側の process/network/file trace と cloud 側の異常検知を合わせ、漏洩後の検知・調査・失効手順まで設計する必要がある。^[raw/articles/flatt-github-actions-credential-leakage-2026.md]
GitHub の Secret Protection による public monitoring は、secret leak detection の境界を「自社 repo」から GitHub の公開面全体へ広げる例である。企業メンバーや verified domain の metadata から、個人 fork、OSS repo、issue、pull request、discussion などに漏れた secret を enterprise に帰属させる。これは CI/CD runtime そのものの隔離策ではないが、agent・bot・開発者が組織外の公開面へ token を誤って出す前提で、公開漏洩の発見を incident response loop に入れる実務的な補助線になる。^[raw/articles/github-secret-scanning-public-monitoring-2026.md]
Microsoft の GitHub Quick Review (`ghqr`) は、GitHub Enterprise / org / repo / GHES を横断して security posture を棚卸しする CLI である。Dependabot、secret scanning、code scanning、2FA/SAML、branch protection、CODEOWNERS、Actions workflow permissions、self-hosted runners、audit log、Copilot policy、MCP settings までを Markdown / Excel / JSON に出せるため、CI/CD runtime security を「個別 YAML のレビュー」から「GitHub tenant 全体の定期診断」へ広げる道具として位置づけられる。^[raw/articles/microsoft-ghqr-github-quick-review-2026.md]
[[strix]] は、CI/CD に入る security testing が SAST や設定監査だけでなく、実行中の application へ AI pentest agent を当て、reconnaissance、exploitation、PoC validation、修正案、report まで返す方向へ広がる例である。これは「runner が何をしたかを監視する」cicd-sensor 型の runtime evidence と対になる。Strix のような tool を PR gate に置くなら、検査対象の sandbox、network egress、test credential、false-positive review、auto-fix の human gate まで含めて CI/CD runtime security として設計する必要がある。^[raw/articles/strix-ai-pentesting-agent-2026.md]
## Design implications
For Yuta-style automation, the useful distinction is not just "scan code before it runs" but "record and reason about privileged automation while it runs." Agentic coding systems, scheduled jobs, and deployment workflows should treat CI/CD runtime logs, provenance, and least-privilege boundaries as first-class product requirements. This connects to [[ai-agent-identity-security]] when agents need scoped credentials, and to [[wiki-maintenance-loop]] as an example of recurring automation that should be observable and auditable.
+30
View File
@@ -0,0 +1,30 @@
---
title: Climate Adaptation AI
created: 2026-07-02
updated: 2026-07-02
type: concept
tags: [llm, civic-tech, public-interest, evaluation, reliability]
sources: [raw/articles/jamstec-regional-climate-llm-2026.md]
confidence: medium
---
# Climate Adaptation AI
Climate adaptation AI は、気候予測、地域の行政知識、対策ガイドラインを組み合わせ、自治体や地域事業者が猛暑・豪雨・干ばつ・海面上昇などへの適応策を立案するための AI 支援を指す。単なる一般相談 chatbot ではなく、科学データと地域の意思決定をつなぐ civic-tech 的な道具として見るのが自然で、評価や説明責任の面では [[ai-evaluation-infrastructure]] とも接続する。
JAMSTEC、高知大学、Ridge-i の地域気候特化型 LLM はこの方向の具体例である。Llama 3.3 Swallow 70B Instruct v0.4 をベースに、A-PLAT の気候変動適応論文 338 編や IPCC 評価報告書で気候学に特化させ、RAG を使って地域の適応計画ガイドラインだけでなく d4PDF のアンサンブル気候予測データから数値を検索・抽出できるようにしている。熊谷市の PoC では、RCP8.5 シナリオの将来気温上昇から熱中症患者の増加数を推定し、グリーンカーテンや休息所の増設要件を平均的・楽観的・悲観的ケースごとに提示した。^[raw/articles/jamstec-regional-climate-llm-2026.md]
この事例で重要なのは、LLM が「専門家の代替」ではなく、専門知識と定量データを扱いにくい自治体実務者へ予備的な選択肢を出す点である。科学者、コンサルタント、自治体職員という役割を LLM にシミュレートさせ、効果、コスト、実現可能性を議論させる設計は、[[ai-research-automation]] のような分析支援と、公共サービスの [[inclusive-design]] の中間にある。
## Design implications
- **Domain-specific grounding**: 汎用 LLM だけに任せると気候科学の誤答や hallucination が問題になるため、専門文献、IPCC 報告、地域ガイドライン、数値予測データを明示的に接続する必要がある。
- **Quantitative retrieval**: 気候リスクでは文章の要約だけでなく、予測データベースから数値を取り出し、計算過程を示すことが意思決定の信頼性になる。
- **Local governance**: 自治体ごとの制約、予算、住民合意を扱うため、AI は結論を押しつけるのではなく、複数シナリオと根拠を提示する補助者として設計すべきである。
- **Equity risk**: 専門人材や財源の少ない地域ほど支援価値が大きい一方、データ更新、地域差、説明責任を維持できないと適応能力の格差を逆に固定する危険がある。
## Open questions
- 気候学特化ベンチマークの点数と、自治体が実際に使える計画品質をどう接続して評価するか。
- RAG の参照データを地域ごとに差し替えるとき、古いガイドラインや不完全な地域データをどう検出するか。
- 災害・熱中症対策のような公共性の高い提案で、AI の説明と人間の最終責任をどこで分けるべきか。
+6 -2
View File
@@ -1,10 +1,10 @@
---
title: Data Protection and Expression
created: 2026-06-28
updated: 2026-06-28
updated: 2026-07-02
type: concept
tags: [data-protection, privacy, freedom-expression, law, public-interest, media]
sources: [raw/articles/david-erdos-cambridge-law-profile-2026.md, raw/articles/ravi-naik-awo-profile-2026.md]
sources: [raw/articles/david-erdos-cambridge-law-profile-2026.md, raw/articles/ravi-naik-awo-profile-2026.md, raw/articles/google-zkp-age-assurance-2026.md, raw/articles/longfellow-zk-identity-proofs-2026.md, raw/articles/email-verification-protocol-draft-2026.md]
confidence: medium
---
@@ -24,6 +24,10 @@ Erdos の Cambridge profile は、EU 内でもこの均衡の置き方が大き
[[maria-ressa]] のような報道側の議論は、情報基盤が嘘や暴力を広げる危険を強調する。Erdos の研究は、法制度がその危険に対応するとき、報道・研究・表現の自由まで削りすぎないための地図になる。
Google が公開した age assurance 向け Zero-Knowledge Proof library は、年齢確認のような規制対応を「本人属性を証明するが、それ以外のデータは渡さない」設計へ寄せる例として読める。EU の eIDAS / EUDI Wallet 文脈では、未成年保護や年齢制限サービスの実装が本人確認データの過剰収集になりやすいため、ZKP は [[inclusive-design]] 的な使いやすさと、[[ai-developer-liability]] 的な設計責任の両方に関わる。Longfellow ZK は ISO MDOC、JWT、W3C Verifiable Credentials のような既存 identity standard に対して anonymous credential / zero-knowledge proof を構成する実装で、legacy credential を使いながら disclosure を最小化する方向の実装面を補う。^[raw/articles/google-zkp-age-assurance-2026.md] ^[raw/articles/longfellow-zk-identity-proofs-2026.md]
Email Verification Protocol draft も同じ系譜にある。従来の email one-time code は、ユーザーに mail client への移動を強いるだけでなく、verification email の送受信や relying party / issuer の関係から余計な情報が流れやすい。EVP は browser を仲介者にし、issuer が email control を token 化し、RP 側では nonce と key binding で検証することで、使いやすさと privacy separation を同時に改善しようとしている。^[raw/articles/email-verification-protocol-draft-2026.md]
## この Wiki での扱い
このページは、一般的なプライバシー法のまとめではなく、公共圏で情報を流す行為と、個人の権利を守る制度のせめぎ合いを追うための入口として置く。今後、忘れられる権利、検索エンジンの責任、研究データ、報道例外に関する資料を追加するときは、このページから分岐させる。
+5 -3
View File
@@ -1,10 +1,10 @@
---
title: Digital Gardening CMS
created: 2026-06-28
updated: 2026-06-30
updated: 2026-07-01
type: concept
tags: [wiki, knowledge-base, maintenance, markdown, design]
sources: [raw/articles/principles-for-digital-gardening-2026.md, raw/articles/litho-deepwiki-rs-code-documentation-2026.md]
sources: [raw/articles/principles-for-digital-gardening-2026.md, raw/articles/litho-deepwiki-rs-code-documentation-2026.md, raw/articles/langchain-openwiki-repo-documentation-agent-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md]
confidence: medium
---
@@ -20,7 +20,9 @@ confidence: medium
実装候補として、[[obsidian]] 的な手元優先の編集体験、Cosense 的な共同編集、Nuxt Content の Markdown と拡張構文、WordPress + WPGraphQL、HyperMD や Milkdown などの編集器が挙げられている。Git で全体を管理すると版管理は強くなるが、携帯端末での編集しやすさが弱くなるため、保存形式、同期、編集体験の折り合いが設計上の中心になる。
[[litho]] のようにコードベースから Wiki 風ドキュメントを生成する道具は、この CMS 発想をソフトウェア設計書側に寄せた例。人間が育てる庭とは違い、コード解析と CI/CD によって鮮度を保とうとするが、生成物をどうレビューし、どの情報を手で補うかは同じく設計問題として残る。
[[litho]] や [[openwiki]] のようにコードベースから Wiki 風ドキュメントを生成する道具は、この CMS 発想をソフトウェア設計書側に寄せた例。人間が育てる庭とは違い、コード解析、git diff、CI/CD、scheduled update によって鮮度を保とうとするが、生成物をどうレビューし、どの情報を手で補うかは同じく設計問題として残る。^[raw/articles/langchain-openwiki-repo-documentation-agent-2026.md]
[[notion]] Developer Platform は、garden / CMS を agent が直接使う shared workspace に寄せる方向を示す。Markdown API、MCP、External Agents API、Workers、CLI がそろうと、ページは人間が読む文書であると同時に、agent が同期・変換・承認依頼・外部 tool 呼び出しの状態を置く場所になる。これは [[llm-wiki-pattern]] のような file-first wiki とは逆に、hosted workspace の操作性を優先する設計だが、長期的な可搬性・版管理・権限境界は慎重に見たい。^[raw/articles/notion-developer-platform-agents-workers-2026.md]
## Open Questions
+31
View File
@@ -0,0 +1,31 @@
---
title: E2E Coverage Metrics
created: 2026-06-30
updated: 2026-07-01
type: concept
tags: [quality, reliability, evaluation, workflow]
sources: [raw/articles/knowledgework-e2e-coverage-metrics-2026.md, raw/articles/github-code-coverage-merge-protection-2026.md, raw/articles/realworld-framework-comparison-spec-2026.md]
confidence: medium
---
# E2E Coverage Metrics
E2E coverage metrics are a way to measure whether end-to-end tests actually touch the product surfaces they are supposed to protect. KnowledgeWork's article argues that manually maintained "test cases written / test cases needed" lists drift as features change, so the denominator should be derived from implementation artifacts where possible.
The concrete pattern is to compute **page coverage** from all known product pages versus pages visited during Playwright runs, and **RPC/API coverage** from all service/method definitions versus RPCs observed in test traffic. In their setup, all pages are extracted from Next.js routes, all RPCs from `.proto` definitions, and the test-side observations come from Playwright trace network entries such as page-view and API requests.^[raw/articles/knowledgework-e2e-coverage-metrics-2026.md]
GitHub Code Quality's merge-protection preview shows the same idea being productized as a repository gate: branch rulesets can block pull requests when coverage falls below a minimum percentage, drops too far from the default branch, or both. Its evaluate mode is important operationally because teams can observe the effect of a quality threshold before turning it into a hard merge blocker.^[raw/articles/github-code-coverage-merge-protection-2026.md]
RealWorld adds a benchmark-design angle: many frontend and backend implementations share the same Medium-like app, API specification, backend spec tests, frontend E2E test suite, CSS theme, and hosted demo API. That makes it useful not only as framework learning material, but as a stable surface for comparing generated code, agent-built app variants, and cross-framework regression behavior under one contract. For [[ai-evaluation-infrastructure]], the important part is the common spec/test harness, not the specific app clone.^[raw/articles/realworld-framework-comparison-spec-2026.md]
This is useful for [[ci-cd-runtime-security]] and [[ai-evaluation-infrastructure]] because it treats test execution as observable runtime evidence, not just a green/red result. It also fits [[loop-engineering]]: the loop should store raw observations, compute metrics later, notify people in the place they already work, and keep enough history to change aggregation methods without rerunning old tests.
## Caveat
Implementation coverage is a necessary-condition signal, not a sufficient proof of product quality. Visiting every page or calling every RPC does not guarantee that important user scenarios are asserted. The stronger pattern is to combine implementation-derived coverage with deterministic regression tests for known critical business paths.
## Open Questions
- Which surfaces should define the denominator for non-Next.js or non-RPC products: routes, OpenAPI endpoints, event names, domain actions, or user journeys?
- When should low E2E coverage block a release, and when should it only produce a review item?
- How can AI-generated test additions avoid optimizing for easy-to-cover surfaces while missing high-risk workflows?
+28
View File
@@ -0,0 +1,28 @@
---
title: Human-Verified Advertising
created: 2026-07-01
updated: 2026-07-01
type: concept
tags: [agent, privacy, data-protection, public-interest]
sources: [raw/articles/hakuhodo-human-verified-ad-2026.md, raw/articles/cloudflare-ai-traffic-options-2026.md, raw/articles/cloudflare-content-independence-day-2025.md, raw/articles/cloudflare-monetization-gateway-x402-2026.md]
confidence: medium
---
# Human-Verified Advertising
Human-verified advertising は、AI エージェント、bot、crawler による非人間トラフィックが増える前提で、広告の配信先を「認証された人間」に限定しようとする広告基盤の設計。博報堂DYホールディングスの Ads for Humanity は、World ID を使って個人情報を共有せずに「固有の人間」であることを証明し、Human-Verified Ad Network 上で人間だけに広告を配信する事業として発表された。
この発想は、[[ai-agent-identity-security]] の「誰が何の権限でアクセスしているのか」を、人間と非人間の区別へ広げる。企業向け agent 認可では agent の権限や監査ログが問題になるが、広告ではクリック、表示、フォーム入力、行動データが人間由来かどうかが、そのまま課金、効果測定、配信最適化の前提になる。
## 何が変わるのか
- **広告接触の主体を検証する**: bot や AI エージェントによる不正接触を排除し、広告主が「人間に届いた」ことを検証できるようにする。
- **効果測定の汚染を防ぐ**: 非人間トラフィックがクリックや行動データに混ざると、広告配信アルゴリズムが生活者の実態ではなく bot の振る舞いへ最適化される。
- **記録を改ざん耐性のある証跡にする**: Hakuhodo DY の発表では、LG Electronics のブロックチェーン技術を使い、配信実績を改ざん不能なエビデンスとして保存するとしている。
- **人間認証とプライバシーの緊張を抱える**: World ID は氏名やメールアドレスを共有せずに人間性を証明できると説明されているが、広告のために「人間である証明」を要求する設計は、[[data-protection-and-expression]] や監視広告への反発とも接続する。
## この Wiki での見方
この領域は広告業界の新商品というより、AI エージェント時代に「人間だけを対象にした経済圏」をどう作るかという公共的な設計問題として見る。[[information-integrity]] では数値や可視性が操作対象になるが、human-verified advertising では広告接触と効果測定の数値が、非人間トラフィックによって汚染されることが問題になる。Cloudflare の [[ai-crawler-governance]] も、広告表示ページ上の Agent / Training traffic を既定 block する方向を示しており、広告収益の前提を AI crawler からどう守るかは edge policy と本人性証明の両側から進んでいる。Cloudflare の “Content Independence Day” 投稿は、広告や subscription が traffic に依存していた web の取引そのものが AI answer によって崩れると説明しており、人間向け広告と crawler 補償は同じ revenue-preservation 問題の別解として見られる。Cloudflare Monetization Gateway の [[agentic-web-monetization]] は、非人間トラフィックを排除するのではなく、agent の利用にも request 単位の価格を付ける第三の方向である。^[raw/articles/cloudflare-ai-traffic-options-2026.md] ^[raw/articles/cloudflare-content-independence-day-2025.md] ^[raw/articles/cloudflare-monetization-gateway-x402-2026.md]
一方で、広告主の検証可能性が強まるほど、利用者側には認証負担、排除、追跡可能性への不安が残る。したがって、評価すべき軸は「bot を排除できるか」だけでなく、本人性証明の最小化、同意、撤回、証跡の透明性、広告を見ない自由をどこまで確保するかにある。
+8 -2
View File
@@ -1,10 +1,10 @@
---
title: Inclusive Design
created: 2026-06-28
updated: 2026-06-30
updated: 2026-07-02
type: concept
tags: [accessibility, inclusive-design, public-interest, design]
sources: [raw/articles/arun-japan-symbols-2026.md, raw/articles/aist-avatar-standardization-committee-2026.md, raw/articles/accessibility-conference-chiba-2026.md]
sources: [raw/articles/arun-japan-symbols-2026.md, raw/articles/aist-avatar-standardization-committee-2026.md, raw/articles/accessibility-conference-chiba-2026.md, raw/articles/smashing-accessibility-operational-capability-2026.md, raw/articles/w3c-accessible-names-descriptions-2026.md]
confidence: medium
---
@@ -20,6 +20,12 @@ confidence: medium
アクセシビリティカンファレンスCHIBA 2026 の案内は、包摂性を「講演で語るテーマ」だけでなく、会場設計と体験ブースに落とし込んでいる例として使える。通常版と情報保障版の YouTube 配信、手話通訳と UD トーク、バリアフリートイレやオストメイト設備の明記、平坦な導線・混雑・照明条件の説明は、参加前に必要な情報へ到達できること自体をアクセシビリティとして扱っている。^[raw/articles/accessibility-conference-chiba-2026.md]
Smashing Magazine の「Accessibility Is An Operational Capability」は、AI が UI を高速生成する時代のアクセシビリティを、事後監査や法務チェックではなく運用能力として扱う。問題は「画面上は動くが、意味のある HTML、キーボード操作、focus 管理、状態の露出が欠ける」コードが大量に増えることにあり、品質保証や [[ai-agent-command-safety]] と同じく、生成前の制約と生成後の検証を開発 loop に組み込む必要がある。^[raw/articles/smashing-accessibility-operational-capability-2026.md]
実装パターンとしては、design system の accessible component を再利用し、Definition of Done と PR review にアクセシビリティ確認を入れ、eslint-plugin-jsx-a11y、Pa11y、Storybook addon などを CI や component 開発に置く。これは [[e2e-coverage-metrics]] のような実行証跡を使う品質保証とも近く、アクセシビリティを「覚えていた人が頑張る」ものから、platform が継続的に維持する性質へ変える。^[raw/articles/smashing-accessibility-operational-capability-2026.md]
W3C APG の accessible name / description guidance は、その運用能力を component レベルに落とす具体的な基礎資料である。焦点可能・操作可能な要素には短く区別できる accessible name が必要で、見えるラベルを優先し、HTML の `label` や `caption` のような native technique を使い、`aria-label` / `aria-labelledby` が子要素の内容を隠す場面を理解して testing する必要がある。AI が UI を生成する場合も、見た目のボタンではなく assistive technology が読む名前・役割・状態まで検証しないと、[[agent-harness-engineering]] や [[e2e-coverage-metrics]] の browser harness は本当に使える UI を保証できない。^[raw/articles/w3c-accessible-names-descriptions-2026.md]
同イベントの体験ブースでは、Ontenna が音の特徴を振動と光に変え、エキマトペが駅の音を AI で識別して文字・手話・オノマトペで可視化する。これは [[meaning-making-marks]] のような公共空間の記号設計を、聴覚・身体感覚・文字情報へまたがる multi-modal interface に広げる実装例として読める。^[raw/articles/accessibility-conference-chiba-2026.md]
## 見るべき問い
+6 -2
View File
@@ -1,10 +1,10 @@
---
title: Information Integrity
created: 2026-06-28
updated: 2026-06-30
updated: 2026-07-01
type: concept
tags: [information-integrity, disinformation, public-interest, media, civic-tech, knowledge-base]
sources: [raw/articles/osome-observatory-on-social-media-2026.md, raw/articles/osome-tools-2026.md, raw/articles/maria-ressa-nobel-facts-2026.md, raw/articles/maria-ressa-stop-the-lies-2026.md, raw/articles/tbs-sns-metric-farms-election-disinformation-2026.md, raw/articles/wikipedia-sanger-canvassing-ban-2026.md, raw/articles/wikipedia-fake-russian-history-zhemao-2022.md]
sources: [raw/articles/osome-observatory-on-social-media-2026.md, raw/articles/osome-tools-2026.md, raw/articles/maria-ressa-nobel-facts-2026.md, raw/articles/maria-ressa-stop-the-lies-2026.md, raw/articles/tbs-sns-metric-farms-election-disinformation-2026.md, raw/articles/wikipedia-sanger-canvassing-ban-2026.md, raw/articles/wikipedia-fake-russian-history-zhemao-2022.md, raw/articles/cloudflare-ai-traffic-options-2026.md, raw/articles/cloudflare-content-independence-day-2025.md, raw/articles/cloudflare-monetization-gateway-x402-2026.md]
confidence: medium
---
@@ -40,6 +40,10 @@ Wikipedia の Larry Sanger ban 事例は、情報基盤の integrity が「偽
このケースは [[llm-wiki-pattern]] と [[wiki-maintenance-loop]] にも直接関係する。知識ベースは、リンクや一貫した文体によって信頼感を作れる一方で、出典の存在確認、他言語・一次資料との照合、矛盾検出、編集者権限の分離が弱いと、もっともらしい synthesis が長期間残る。LLM が Wiki を更新する場合も、文章の自然さではなく source provenance と反証可能性を保つことが integrity の中心になる。
## AI crawler と web access policy
[[ai-crawler-governance]] は、情報基盤の integrity を「何が拡散されるか」だけでなく「誰が、何の目的で、どの規則で読みに来るか」へ広げる。Cloudflare の AI traffic 分類は、Search / Agent / Training crawler を分け、検索流入や補償を返す bot と、広告や publisher revenue を迂回して content を持ち帰る bot を別扱いしようとする。さらに “Content Independence Day” 投稿は、AI answer が original source への traffic を返さないと、報道・解説・専門知識の持続性そのものが壊れるという経済的 integrity の論点を出している。Cloudflare Monetization Gateway はこの論点を [[agentic-web-monetization]] へ進め、agent が API、dataset、MCP tool、content を使うたびに支払いと identity / access policy を request path で処理する方向を示す。これは media ecosystem の持続性と AI access norm を edge provider が形作る例として重要である。^[raw/articles/cloudflare-ai-traffic-options-2026.md] ^[raw/articles/cloudflare-content-independence-day-2025.md] ^[raw/articles/cloudflare-monetization-gateway-x402-2026.md]
## 関連領域
この領域は、公共のための技術、報道、偽情報対策、SNS の設計、アクセシビリティ、民主主義の維持とつながる。OSoMe や Ressa のような資料は、流行のニュースとして消費するより、人物・組織・道具・概念に分けて蓄積すると後から参照しやすい。
@@ -0,0 +1,35 @@
---
title: LLM Assisted Vulnerability Research
created: 2026-07-02
updated: 2026-07-02
type: concept
tags: [llm, agent, security, quality, workflow]
sources: [raw/articles/devansh-llm-vulnerability-research-2026.md]
confidence: medium
---
# LLM Assisted Vulnerability Research
LLM assisted vulnerability research は、LLM / coding agent を「全部の脆弱性を探して」と広く投げるのではなく、攻撃面・信頼境界・不変条件を小さく切り、証拠で潰しながら脆弱性を探す作業様式である。Devansh の記事は、Parse Server、HonoJS、ElysiaJS、harden-runner、BullFrog、Better-Hub などで見つけた複数の脆弱性を例に、LLM の価値は巨大な AGENTS.md や長い checklist ではなく、薄い threat model と検証 loop にあると整理している。^[raw/articles/devansh-llm-vulnerability-research-2026.md]
この論点は [[agent-harness-engineering]] の「context をどう狭く保つか」と、[[scrutineer]] / [[strix]] の human-gated security workflow に近い。違いは、Scrutineer や Strix が道具・workflow として外形化しているのに対し、このページの焦点は人間 researcher が Codex / Claude などを使う時の探索単位、prompt frame、verification budget の配分にある。
## 実務上のパターン
- **広すぎる依頼を避ける**: `find all vulnerabilities` は threat model がなく、generic CWE 的な観測や到達不能な理論上の問題を増やしやすい。まず「誰が、どの入口から、何を越えようとするのか」を短く固定する。
- **小さな threat model を作る**: 過去 CVE、security advisory、設計文書、既知の bug class から、その project が過去に失敗した境界を抽出する。Parse Server なら key type / authorization boundary、HonoJS なら JWT / JWKS algorithm handling、harden-runner なら GitHub Actions runner の outbound egress が焦点になる。
- **thin slice に分割する**: auth、session、request parsing、file upload、deserialization、sandbox boundary、plugin boundary など、実際の攻撃面に対応する小さな単位で読む。巨大 context へ全体を詰めるより、slice ごとに entry point、sensitive sink、guard、attacker-controlled input を確認させる。
- **不変条件を破らせる**: `only admins can call X`、`JWT issuer/audience/algorithm must be pinned`、`read-only key must never write`、`egress controls must see every network path` のように、コードが守るべき条件を列挙し、各条件を攻撃者が破れるか調べる。
- **検証に token と時間を使う**: 「モデルが言った」段階で止めず、unit / integration test、PoC request、crash reproduction、sanitizer build、fuzzer、static/invariant check で、成立・不成立を証拠化する。ここは [[ci-cd-runtime-security]] の実行時証跡や [[ai-evaluation-infrastructure]] の評価 loop と同じ発想である。
## Context 設計の教訓
記事は、over-scaffolding、bloated `AGENT.md` / `SKILLS.md`、過剰な事前計画が、脆弱性探索では逆に needle-in-the-haystack 問題を悪化させると主張している。これは長文 context の中央にある重要情報が拾われにくいという context rot / lost-in-the-middle 系の問題と接続する。
実務的には、安定 scaffold は 1 ページ程度の threat model、不変条件、crown-jewel 機能に抑え、残りの budget は focused slice audit と verifier loop に使うのがよい。これは [[agent-oriented-cli-design]] の「道具が少ない文脈で正確に使える」設計ともつながる。
## Prompt frame と安全上の注意
記事には「脆弱性があると仮定する」「exploit を書かせる」「auditor ではなく adversary として考えさせる」など、LLM の探索圧を上げる prompt frame が並ぶ。防御研究の中では有効なことがある一方、攻撃化も容易なので、実行対象、権限、ネットワーク境界、報告先、人間 gate を固定する必要がある。
Yuta の wiki では、この種の知見は無制限な攻撃手順ではなく、[[ai-agent-command-safety]]、[[ai-agent-enabled-cyberattacks]]、[[ai-agent-identity-security]] と並べて、agent を防御研究に使う時の boundary design として扱う。
+10 -3
View File
@@ -1,10 +1,10 @@
---
title: Loop Engineering
created: 2026-06-29
updated: 2026-06-30
updated: 2026-07-01
type: concept
tags: [agent, automation, workflow, quality]
sources: [raw/articles/loop-engineering-anthropic-playbook-2026.md, raw/articles/github-issueops-state-machines-2026.md, raw/articles/horizon-agentic-hardware-design-2026.md, raw/articles/abtop-ai-coding-agent-monitor-2026.md, raw/articles/kiro-ide-1-0-agent-focus-2026.md, raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/pi-coding-agent-2025.md, raw/articles/github-desktop-3-6-worktrees-copilot-2026.md, raw/articles/agent-oriented-cli-zenn-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md]
sources: [raw/articles/loop-engineering-anthropic-playbook-2026.md, raw/articles/github-issueops-state-machines-2026.md, raw/articles/horizon-agentic-hardware-design-2026.md, raw/articles/abtop-ai-coding-agent-monitor-2026.md, raw/articles/kiro-ide-1-0-agent-focus-2026.md, raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/pi-coding-agent-2025.md, raw/articles/github-desktop-3-6-worktrees-copilot-2026.md, raw/articles/agent-oriented-cli-zenn-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md, raw/articles/google-adk-go-2-0-agent-workflows-2026.md, raw/articles/theregister-claude-code-transcript-retention-2026.md, raw/articles/awesome-harness-engineering-2026.md, raw/articles/mastra-typescript-agent-framework-2026.md, raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/github-duplicate-issue-detection-mcp-fields-2026.md, raw/articles/langchain-openwiki-repo-documentation-agent-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md]
confidence: medium
---
@@ -18,14 +18,21 @@ Loop engineering は、LLM やエージェントを「人間が毎回プロン
- **Generator / evaluator separation**: 生成したエージェント自身に採点させると甘くなりやすい。別プロンプト、別モデル、別プロセスの「疑う評価者」を置く方が、[[ai-assisted-reverse-engineering]] の命名・型付け検査や wiki ingest の品質判定にも応用しやすい。
- **Persistence first**: ループの成果はチャットの返答だけでなく、PR、Issue、wiki、state file、log のような再利用可能な場所に残す。残らない自動化は、次回の discovery と評価に使えない。
- **State machine として考える**: GitHub の IssueOps 記事は、Issue、label、comment、approval、Action を状態機械として扱う。これは loop engineering の persistence と scheduling を、監査可能な GitHub timeline に置く方法として読める。
- **State machine として考える**: GitHub の IssueOps 記事は、Issue、label、comment、approval、Action を状態機械として扱う。これは loop engineering の persistence と scheduling を、監査可能な GitHub timeline に置く方法として読める。GitHub の duplicate issue detection と MCP server の issue fields 対応は、この面をさらに agent-friendly にする。重複検出は人間 maintainer の triage 負荷を下げ、MCP 経由の issue field 読み書きは agent が priority、area、date などを持つ構造化された state を作れるようにする。^[raw/articles/github-duplicate-issue-detection-mcp-fields-2026.md]
- **Repository-native loop**: [[agentic-hardware-design]] の HORIZON は、Markdown harness から evaluator / acceptance predicate / git policy を持つ project pack を作り、隔離された worktree 上の diff・commit・log・notes をそのまま探索 trace にする。ループの状態を外部データベースに逃がさず、作業対象の repository 自体に残す設計として重要。
- **Operator observability**: [[abtop]] は Claude Code、Codex CLI、OpenCode の session、token、context、rate limit、child process、open port をローカルで可視化する。複数 agent を同時に回す loop では、成果物だけでなく「いま何が動いているか」「どの資源を占有しているか」も運用対象になる。
- **Agent-native work surface**: [[kiro]] の Agent Focus は、コード編集画面ではなく session、会話、spec、diff を前面に置く。loop を IDE の中に寄せると、人間の仕事は直接編集よりも、仕様・承認・差分確認・権限ルールの管理へ移る。
- **Agent harness minimalism**: [[pi-coding-agent]] は、read / write / edit / bash、tmux、明示的な session file など既存の可視な道具に寄せることで、隠れた sub-agent や巨大な system prompt に頼らない loop を作ろうとする。loop の強さは機能数だけでなく、operator が context、tool result、process、cost をどこまで観測できるかにも依存する。
- **Harness before loop**: [[agent-harness-engineering]] は、context、memory、guardrail、tool boundary、eval、observability を整え、1 回から数回の agent 実行を dependable にする層として読める。loop engineering はその上で discovery、handoff、persistence、scheduling をつなぐので、長期自動化の失敗は loop の設計だけでなく、その下の harness が曖昧なことからも起きる。^[raw/articles/awesome-harness-engineering-2026.md]
- **Worktrees as everyday agent infrastructure**: GitHub Desktop 3.6 の worktree support は、agent が複数 branch / sandbox を使う流れを GUI 側にも取り込む。isolated worktree は [[agentic-hardware-design]] や IssueOps 的な repository-native loop と同じく、並列作業を見える単位に分けるための基礎部品になる。
- **Evaluator market pressure**: [[ai-evaluation-infrastructure]] の Arena 事例は、評価が研究用 leaderboard から商用分析・post-training 改善の基盤へ広がっていることを示す。loop engineering でも、実行する agent だけでなく、それを測る evaluator とデータ収集の設計が競争力になる。
- **Agent-oriented tools**: [[agent-oriented-cli-design]] は、JSON first、actionable error、search/read 分離、鮮度情報、少ないフラグを通じて、agent が推測せずに次の行動へ進める CLI を作る考え方。loop の実行単位である tool が曖昧だと、verification や persistence 以前に誤った状態で進んでしまう。
- **Workflow graph as agent**: [[google-adk]] Go 2.0 は、function / agent / tool / join / dynamic node を edge と route でつなぐ graph そのものを `agent.Agent` として実行する。pause/resume、human-in-the-loop、retry、branch isolation、telemetry を framework primitive にすることで、loop を ad-hoc prompt ではなく、状態を持つ観測可能な実行 graph として扱う方向を示している。^[raw/articles/google-adk-go-2-0-agent-workflows-2026.md]
- **TypeScript workflow framework**: [[mastra]] は、model routing、agent、graph workflow、HITL suspend/resume、MCP server、eval、observability を TypeScript application stack にまとめる。[[google-adk]] が Go/Python の typed graph runtime 寄りなら、Mastra は web application に agent loop を組み込む入口として見られる。^[raw/articles/mastra-typescript-agent-framework-2026.md]
- **Transcript retention is product behavior**: Claude Code の `cleanupPeriodDays` 既定値 30 日をめぐる The Register の報道は、agent loop の会話 transcript が単なる UI 履歴ではなく、設計判断、debugging context、研究上の reasoning trail そのものになりうることを示す。保存しすぎると source code や credential を含む privacy / security risk になる一方、削除が silent で recovery log もないと、operator は永続化されていると思った作業知識を失う。loop engineering では、保存期間、削除ログ、soft delete、backup、state file への要約などを明示的な設計対象にする必要がある。^[raw/articles/theregister-claude-code-transcript-retention-2026.md]
- **Telemetry is part of the loop boundary**: [[ai-agent-telemetry-privacy]] は、agent loop の観測性が product/vendor telemetry とどこで重なるかを問う。trace や error stack は debugging に役立つが、repo hash、CI identity、stack frame、session ID が外部へ出るなら、loop の persistence / observability は retention と opt-out まで含めて設計する必要がある。^[raw/articles/claude-code-telemetry-audit-2026.md]
- **Repo wiki as loop memory**: [[openwiki]] は、コードベースの理解を `openwiki/` に残し、`AGENTS.md` / `CLAUDE.md` からそこへ案内し、GitHub Actions で差分更新する。これは agent loop の working memory を一回の context window から外へ出し、repository に残る保守可能な知識面へ移す例として読める。^[raw/articles/langchain-openwiki-repo-documentation-agent-2026.md]
- **Run budget as loop state**: Copilot CLI / SDK の AI credit session limit は、無人 run の cost を loop state として扱う例である。`--max-ai-credits` のような上限があると、agent は人間が見ていない間に無制限に subagent や compaction を回すのではなく、soft cap 到達時に wrap up して state を返す。これは scheduling と persistence だけでなく、run をいつ止めるかという operator contract でもある。^[raw/articles/github-copilot-ai-credit-session-limits-2026.md]
- **Human judgment is scarce**: 生成は安くなっても、何を通し、何を止め、何を記録するかの判断は希少になる。自動化は人間の判断を消すのではなく、判断すべき点を狭く明確にするべき。
## Failure Modes
@@ -0,0 +1,36 @@
---
title: Open Source Package Supply Chain Attacks
created: 2026-07-01
updated: 2026-07-02
type: concept
tags: [security, supply-chain, dev-tool, reliability]
sources: [raw/articles/checkmarx-operation-navy-ghost-pyrogram-supply-chain-2026.md, raw/articles/ladybird-maintainer-only-development-ai-pr-risk-2026.md]
confidence: medium
---
# Open Source Package Supply Chain Attacks
Open source package supply chain attacks are compromises where the attacker does not need to break into a project or account: they can publish a plausible dependency, fork, plugin, or extension and wait for developers, bots, CI jobs, or production services to install it. This complements [[ci-cd-runtime-security]], which focuses on observing privileged automation while it runs, and [[scrutineer]], which focuses on human-gated vulnerability discovery and disclosure.
## Operation Navy Ghost
Checkmarx's Operation Navy Ghost report describes a PyPI campaign against Telegram bot developers using fake or trojanized `pyrogram` forks. Between November 2025 and June 2026, the attacker published packages such as `vlifegram`, `vlife-gram`, `kelragram`, `pyrogram-navy`, `pyrogram-styled`, `sepgram`, `pyrogram-zeeb`, and `pyrogram-kelra`. The packages looked like legitimate forks but included a hidden `pyrogram/helpers/secret.py` backdoor and modified startup paths.^[raw/articles/checkmarx-operation-navy-ghost-pyrogram-supply-chain-2026.md]
The useful pattern is that the malicious code targeted bot/server environments rather than only local developer machines. It registered Telegram-controlled handlers for Python execution and shell execution, used Telegram itself as command-and-control and exfiltration, and included self-exclusion logic so the attacker's own accounts would not trigger the backdoor. For Yuta-style automation, this is a reminder that package choice, bot tokens, CI runners, and long-lived agent services share the same risk surface: once an installed dependency runs with credentials, network monitoring alone may not show the useful evidence.
Ladybird の開発方針変更は、package registry ではなく open-source contribution path 側の trust model 変化を示す。Ladybird は AI tool によって「大きな patch を出す労力」が善意や長期関与の proxy ではなくなり、browser のように untrusted internet input を実行する project では、一つのよく隠れた脆弱性が深刻な結果を持つとして、public pull request を閉じ、maintainer だけが code を入れる方針へ移った。これは [[ci-cd-runtime-security]] や [[ai-agent-command-safety]] と同じく、AI が生成速度を上げたことで review capacity と responsibility boundary が希少資源になる例である。^[raw/articles/ladybird-maintainer-only-development-ai-pr-risk-2026.md]
## Defensive implications
- Treat dependency names and maintainers as part of the threat model, especially for forks of popular libraries.
- Check installed packages and lockfiles for near-name variants, not only for known CVEs.
- Prefer runtime evidence and containment when automation has credentials: process ancestry, file access, network destinations, and credential reads matter as much as static provenance.
- For bot or agent services, rotate tokens and audit persistence if a malicious package may have run; the package may have accessed environment variables, sessions, files, or cloud credentials.
- Link package-ingest checks with [[ai-agent-command-safety]]: agents can install dependencies, run examples, or execute project scripts, so package manager operations are command-execution boundaries, not just setup steps.
- Treat generated-looking contribution volume as a review-capacity problem, not only a code-quality problem; projects may need narrower trusted committer paths when a disguised vulnerability is high impact.
## Open questions
- How should local agent sandboxes make package installation observable without making day-to-day development too slow?
- Which package-registry trust signals are actually useful to an agent making autonomous install decisions?
- Can CI/CD sensors and local agent logs share enough schema to reconstruct package-originated credential access after the fact?