This commit is contained in:
2026-06-30 22:22:00 +09:00
parent 9ad671521d
commit 87eacd39b2
66 changed files with 10280 additions and 105 deletions
+35
View File
@@ -0,0 +1,35 @@
---
title: Agent-Oriented CLI Design
created: 2026-06-30
updated: 2026-06-30
type: concept
tags: [agent, cli, dev-tool, workflow, quality]
sources: [raw/articles/agent-oriented-cli-zenn-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md]
confidence: medium
---
# Agent-Oriented CLI Design
Agent-oriented CLI design は、人間が目で読んで試行錯誤する端末道具ではなく、Claude Code などの [[loop-engineering|agent loop]] が安全に呼び出し、結果を機械的に判断し、次の行動へ進めるための CLI 設計。Zenn の「AI エージェント向け CLI ツール」記事は、Claude Code 用の横断検索 CLI を Go で作った経験から、人間向け CLI と違う判断基準を整理している。
重要なのは、エージェントに「推測させない」こと。使い方は wiki や skill 側へ長く写すのではなく、CLI 自体に `skill` や help サブコマンドとして同梱し、スキーマや出力の意味が実装と一緒に更新されるようにする。これは [[wiki-maintenance-loop]] の raw/source と synthesis を分ける考え方にも近く、手順が古くなる場所を減らす設計である。
## 設計原則
- **JSON first**: 人間向けの整形テキストではなく、既定で構造化 JSON を返す。結果には `id`、`title`、`snippet`、`source_url`、`synced_at`、`is_stale` など、エージェントが次の判断に使う材料を入れる。
- **Actionable errors**: `index is missing` だけで止めず、`run super-cli-tool sync` のように次のコマンドを直接書く。小さいモデルほど推測の余地を減らす効果が大きい。
- **Search then read**: 重い本文取得と軽い候補検索を分ける。まず `search` で候補を絞り、必要なものだけ `read` する方が、トークン・時間・判断負荷を抑えやすい。
- **Defaults over flags**: `--sources` や `--discover` のような細かい選択肢を増やすより、よく使う安全な既定値へ寄せる。フラグが多いほど help が長くなり、エージェントの分岐も増える。
- **Governance hooks**: Microsoft Learn の agent architecture / SDLC module は、agent task を input / output / success criteria で構造化し、PR template、checks、CODEOWNERS、rules、environment gate、observability、tool governance、secret boundary を運用設計へ入れることを強調する。CLI も単独の便利道具ではなく、[[ai-agent-identity-security]] や PR governance に接続される実行面として見るべき。
## なぜ重要か
エージェント向け CLI は、単に「CLI を LLM から呼べるようにする」だけでは足りない。出力が曖昧だったり、エラーが不親切だったり、状態の鮮度が返らなかったりすると、agent loop は誤った仮定のまま進む。逆に、CLI が状態・出典・次アクション・失敗理由を明示すれば、[[loop-engineering]] の verification と persistence が自然に強くなる。
この設計は Hermes の skill にも当てはまる。skill は長い操作説明を抱え込むより、実際の CLI が `--help` や `agent-guide` を返せるならそこへ誘導し、skill 側は「いつ使うか」と「安全境界」を中心に保つ方が、ツール更新とのずれを減らせる。
## Open Questions
- CLI 側の `agent-guide` は人間向け help と別にすべきか、それとも同じ help を機械可読に拡張すべきか。
- JSON schema、exit code、retryability、rate-limit 情報をどこまで標準化すれば、複数 agent / tool 間で再利用できるか。
- [[abtop]] のような operator UI は、個々の CLI 実行ログや stale 状態をどこまで横断可視化すべきか。
+28
View File
@@ -0,0 +1,28 @@
---
title: Agentic Hardware Design
created: 2026-06-30
updated: 2026-06-30
type: concept
tags: [llm, agent, automation, dev-tool, evaluation, quality]
sources: [raw/articles/horizon-agentic-hardware-design-2026.md]
confidence: medium
---
# Agentic Hardware Design
Agentic hardware design は、RTL や検証資産を一回のコード生成ではなく、実行可能な評価器を持つリポジトリ上でエージェントが反復修正する設計方法として扱う考え方。NVIDIA Research の HORIZON 論文は、ハードウェア設計問題を Markdown harness から project pack に変換し、隔離された git worktree、実行可能 evaluator、acceptance predicate、git/runtime policy を組み合わせて、手放しの agent loop で RTL benchmark を収束させる。
[[loop-engineering]] との違いは、対象が一般の作業ループではなく、RTL・testbench・checker・assertion・debugging といった EDA/ハードウェア設計成果物そのものに寄っている点。HORIZON は git diff、commit、log、notes を状態管理と trace buffer として使い、候補変更を evaluator evidence で受け入れるか拒否する。これは [[ai-research-automation]] のような収集・報告ループよりも、評価器と受理条件が強く組み込まれた self-evolution 型の運用である。
## 見るべき軸
- **Repository as task substrate**: 問題をプロンプトではなく、評価器付きの git worktree として渡す。履歴、diff、commit、replay がそのまま agent の探索記録になる。
- **Markdown harness**: 人間が目的、領域知識、期待成果物、評価基準を Markdown で書き、bootstrap agent が project pack に変換する。これは [[llm-wiki-pattern]] の「Markdown を持続的な知識媒体にする」発想と近いが、出力先は wiki ではなく実行可能な設計タスクである。
- **Executable feedback**: RTL では構文の正しさだけでなく、simulation、coverage、assertion、checker などが収束条件になる。生成物は「それらしい」だけでは足りず、実行証拠で受理される必要がある。
- **Benchmark saturation vs robustness**: HORIZON は複数 RTL benchmark を 100% completion まで進めた一方、論文自身も reward hacking、hidden tests、独立 reference model、formal equivalence などを未解決課題として挙げている。[[wiki-maintenance-loop]] と同じく、見える評価器に過適合しない設計が重要になる。
## Open Questions
- ハードウェア設計 agent で、修復時に見せる診断情報と最終評価に使う hidden / randomized checks をどう分けるべきか。
- PPA 最適化や signoff のように評価が遅い領域で、短い edit-evaluate loop をどう置き換えるか。
- 個人・小規模チームの開発者道具に、HORIZON 的な git-native trace と acceptance gate をどこまで軽量に持ち込めるか。
+36
View File
@@ -0,0 +1,36 @@
---
title: AI Agent Identity Security
created: 2026-06-29
updated: 2026-06-30
type: concept
tags: [agent, security, reliability, privacy]
sources: [raw/articles/okta-cross-app-access-partners-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md]
confidence: medium
---
# AI Agent Identity Security
AI agent identity security は、AI エージェントやアプリ間連携が企業データへアクセスするときに、誰の権限で、何へ、どの範囲で、どの操作をしたのかを追跡・制御する設計領域。Okta の Cross App Access(XAA)発表は、MCP 認証拡張としての Enterprise-Managed Authorization を含め、エージェント時代の認可を OAuth とアイデンティティ管理の延長で標準化しようとする動きとして読める。
問題の出発点は、AI エージェントの接続が静的 API key やユーザーごとの同意画面に依存しがちなこと。これだと、常時特権、見えない同意、監査できないアプリ間移動が起こりやすい。XAA は、すべての接続を中央のアイデンティティポリシーに通し、アクションをログに残し、必要最小限のスコープ付き token を使う方向を示している。
## 見るべき軸
- **Agent as requesting app**: Claude、Cursor、Docker、VS Code、Zoom など、作業を始めるエージェントや開発者道具が、別アプリへのアクセスを要求する主体になる。
- **Resource app / MCP server**: Asana、Atlassian、Figma、Linear、Slack、Supabase、Datadog などが、エージェントに文脈や業務データを渡す側になる。
- **Policy and audit**: アクセスが許可される前に企業ポリシーで検査し、操作の監査証跡を残す。これは [[loop-engineering]] の persistence と verification をセキュリティ境界へ移したものでもある。
- **Least privilege for agents**: 常時広い権限を持つ bot token ではなく、必要な範囲に絞った identity-based token を使う。
- **Local sandbox / approval boundary**: Codex の安全運用ドキュメントは、cloud では隔離 container、CLI/IDE では OS sandbox と approval policy を組み合わせ、既定で network access を切り、workspace 外の編集や network 利用を承認対象にする設計を説明している。`workspace-write`、`read-only`、network proxy、domain allow/deny などの設定は、企業の cross-app 認可だけでなく個人の agent loop でも「どこまで自動実行してよいか」を明示する制御面になる。
- **Repository governance as identity boundary**: Microsoft Learn の agent architecture / SDLC module は、agent task を input / output / success criteria で定義し、PR template、checks、CODEOWNERS、rules、environment gate を通じて「どの変更が誰の承認で通るか」を設計する。これは [[agent-oriented-cli-design]] の tool-level clarity と同じく、agent の行動を監査可能な境界へ置く方法である。
## なぜ重要か
エージェントが社内システム、開発環境、デザイン、会議、監視、データベースを横断し始めると、便利さと同じ速度で攻撃面も広がる。[[ai-assisted-reverse-engineering]] のように専門道具へ書き込み権限を渡す場合や、[[wiki-maintenance-loop]] のように自動で source を保存・分類する場合でも、「どの agent が何をしたか」を後で説明できることが信頼性の条件になる。
この論点は [[ai-developer-liability]] とも接続する。事故や漏えいが起きたとき、単にユーザーの操作やモデル出力だけでなく、開発者がどの権限境界、ログ、承認、取り消し手段を設計していたかが問われる可能性がある。
## Open Questions
- 個人用 Hermes / Discord / wiki 自動化では、企業向け XAA の考え方をどこまで軽量化して使えるか。
- MCP server と agent gateway の認可ログを、開発者が後から読める形でどこに保存するべきか。
- 便利な cross-app agent workflow と、ユーザーが理解できる同意・取り消し UI をどう両立するか。
+4 -2
View File
@@ -1,10 +1,10 @@
---
title: AI Developer Liability
created: 2026-06-28
updated: 2026-06-28
updated: 2026-06-30
type: concept
tags: [law, public-interest, privacy, security, data-protection]
sources: [raw/articles/ravi-naik-awo-profile-2026.md]
sources: [raw/articles/ravi-naik-awo-profile-2026.md, raw/articles/boj-ai-legal-risk-financial-institutions-2026.md]
confidence: medium
---
@@ -29,3 +29,5 @@ AI developer liability は、AI の出力や利用者の行為だけでなく、
AI 開発者の責任は、個別の事故対応にとどまらない。責任の線引きが変わると、AI サービスの安全設計、公開前の検証、記録の保存、通報対応、規制当局との関係が変わる。特に、性的画像、選挙、報道、内部告発、広告技術のように、個人の被害と公共の議論が同時に現れる領域では、[[information-integrity]] と法制度の両方から追う必要がある。
現時点ではこのページは Ravi Naik / AWO profile という単一資料からの入口であり、具体的な法理や裁判上の争点は今後の資料で補う必要がある。
日本銀行金融研究所の「金融機関におけるAI利用に伴う私法上のリスクと管理」は、個人被害や deepfake とは別の角度から、金融機関が AI 開発者・提供者に契約責任を追及する場合、AI を使ったサービスを顧客へ提供する場合、組織内部で取締役が AI ガバナンス体制を構築する場合を整理している。ここでは AI 開発者責任は不法行為だけでなく、契約条項、顧客との説明・合意、内部統制としても現れる。これは [[ai-agent-identity-security]] の権限境界や監査ログが、事故後の説明責任だけでなく契約上の管理義務にも関わることを示す。
+28
View File
@@ -0,0 +1,28 @@
---
title: AI Evaluation Infrastructure
created: 2026-06-30
updated: 2026-06-30
type: concept
tags: [evaluation, llm, quality, workflow]
sources: [raw/articles/arena-ai-leaderboard-business-2026.md]
confidence: medium
---
# AI Evaluation Infrastructure
AI evaluation infrastructure は、LLM や agent の性能を、単発 benchmark ではなく、利用者評価、専門タスク、分析サービス、商用フィードバックループとして継続的に測る層。TechCrunch の Arena 記事では、UC Berkeley の研究プロジェクト由来の AI leaderboard が、1,000 万件超の利用者評価をもとにした公開ランキングから、モデル企業や企業向けの深掘り分析サービスへ広がり、商用開始から 8 カ月で年換算 1 億ドル規模に到達したとされる。
重要なのは、評価が単なる研究補助ではなく、モデル改善・post-training・企業導入判断の市場そのものになっている点。Arena は text、coding、vision、image generation に加え、Agent Mode のような長時間 workflow も扱う。これは [[loop-engineering]] や [[agentic-hardware-design]] のようなエージェント運用で、最終成果だけでなく、途中の意思決定・失敗・回復をどう測るかという問題に接続する。
## なぜ重要か
- **Crowdsourced comparison**: 利用者が 2 つのモデル出力を比較する形式は、静的な benchmark では拾いにくい実利用の好みを集められる。
- **Evaluation as business**: 無料 leaderboard の背後で、詳細分析や model lab 向け評価が商用サービスになる。
- **Post-training demand**: Arena は、人間評価やラベリングを提供する Mercor、Surge、Scale AI などと同じ予算を争うと説明されており、評価と訓練改善が近づいている。
- **Agent evaluation**: 長時間 workflow や Agent Mode が評価対象になると、[[wiki-maintenance-loop]] のような自走ジョブでも、単一回答の品質ではなく状態更新・検証・永続化まで測る必要が出る。
## Open Questions
- 公開 leaderboard の人気と、商用分析の顧客価値はどこまで同じ評価データに依存しているのか。
- Agent Mode のような長時間タスクでは、勝敗や好みだけでなく、再現性、コスト、安全な権限利用、監査証跡をどう評価するべきか。
- 個人用 wiki や Hermes job では、大規模な人間評価を使わずに、どの小さな evaluator を積み重ねれば品質劣化を検出できるか。
+27
View File
@@ -0,0 +1,27 @@
---
title: Avatar Standardization
created: 2026-06-30
updated: 2026-06-30
type: concept
tags: [design, interface, public-interest, inclusive-design]
sources: [raw/articles/aist-avatar-standardization-committee-2026.md]
confidence: medium
---
# Avatar Standardization
アバター標準化は、XR やメタバースで使われる仮想身体を、単なるキャラクター表現ではなくユーザインターフェースとして扱う設計課題。産総研の「アバター国際標準化の国内検討委員会」は、ISO/IEC JTC1/SC35 におけるユーザインターフェースとしてのアバター規格開発に対し、国内の業界やユーザの声を集め、提言や助言を行うための委員会として説明されている。
この論点は [[inclusive-design]] と近い。アバターは「なりたい姿」や文化表現であると同時に、サービス側がユーザに何を許し、どの情報を相手に伝え、どんな身体差・文化差を扱うかを決める interface でもある。公共空間の標識や合図を扱う [[meaning-making-marks]] と同じく、見た目が周囲の行動や解釈を変える。
## Design implications
- アバターの設計情報は、体験品質や安全性に影響するため、利用者と開発者の双方に意味のある規格が必要になる。
- 委員会には VRM、通信、大学、企業、メタバース関連団体、VTuber/有識者などが含まれており、技術仕様だけでなく文化的実践を標準化に接続しようとしている。
- 「アバター」は趣味文化の表現物であるだけでなく、XR サービス上の身体・本人性・可視性・相互行為の設計単位になる。
## Open questions
- 標準化が相互運用性を高める一方で、匿名性、変身、文化的遊びの余地を狭めないようにするにはどうするか。
- アバターが本人性や属性を示すとき、[[information-integrity]] とプライバシーの境界をどう扱うか。
- ユーザ側の声を集める委員会設計は、どの範囲の利用者を代表できるか。
+33
View File
@@ -0,0 +1,33 @@
---
title: CI/CD Runtime Security
created: 2026-06-30
updated: 2026-06-30
type: concept
tags: [security, supply-chain, quality, reliability, automation]
sources: [raw/articles/cicd-sensor-2026.md, raw/articles/scrutineer-oss-security-workflow-2026.md]
confidence: medium
---
# CI/CD Runtime Security
CI/CD runtime security is the practice of observing and constraining what actually runs inside build, test, release, and deployment jobs. The core problem is that CI jobs hold cloud credentials, signing keys, package-registry tokens, and deployment authority, while compromised dependencies or scripts can execute inside short-lived jobs and disappear with the evidence when the job ends. [[ai-agent-identity-security]] covers adjacent authorization and audit concerns for agents; [[loop-engineering]] is relevant because autonomous development loops often depend on these pipelines as their verification and deployment boundary.
## Why it matters
Traditional software supply-chain controls often answer where an artifact came from or how it was built, but they may not preserve enough runtime evidence about what a job process actually did. `cicd-sensor` frames this as an EDR-like gap for CI/CD: open-source defenders exist for many production runtimes, while CI/CD runners have lagged despite holding highly privileged credentials.^[raw/articles/cicd-sensor-2026.md]
## Implementation pattern
`cicd-sensor` uses an eBPF-powered sensor for GitHub Actions and GitLab CI/CD. Its baseline detections use process ancestry and correlated signals: for example, credential access by a process descended from `npm install`, or one job reading several credential categories. It can emit per-run logs, graphical job summaries, cloud-routed evidence, and build attestations while keeping data in the operator's own infrastructure rather than sending it to a project-operated SaaS.^[raw/articles/cicd-sensor-2026.md]
## Design implications
For Yuta-style automation, the useful distinction is not just "scan code before it runs" but "record and reason about privileged automation while it runs." Agentic coding systems, scheduled jobs, and deployment workflows should treat CI/CD runtime logs, provenance, and least-privilege boundaries as first-class product requirements. This connects to [[ai-agent-identity-security]] when agents need scoped credentials, and to [[wiki-maintenance-loop]] as an example of recurring automation that should be observable and auditable.
[[scrutineer]] extends the same supply-chain concern toward open-source vulnerability discovery and disclosure. Instead of watching CI runtime behavior, it uses skill-based AI scans, threat models, maintainer discovery, patch drafting, and release watching to keep unverified model findings inside a human-gated workflow before they reach maintainers.^[raw/articles/scrutineer-oss-security-workflow-2026.md]
## Open questions
- How should teams balance runtime blocking, forensic logging, and false positives in developer-facing pipelines?
- What evidence format is durable enough to connect runtime traces with artifact provenance and code review history?
- Where should CI/CD runtime controls live when agents can run locally, in cloud sandboxes, and inside hosted CI at different stages of one task?
+4 -2
View File
@@ -1,10 +1,10 @@
---
title: Digital Gardening CMS
created: 2026-06-28
updated: 2026-06-28
updated: 2026-06-30
type: concept
tags: [wiki, knowledge-base, maintenance, markdown, design]
sources: [raw/articles/principles-for-digital-gardening-2026.md]
sources: [raw/articles/principles-for-digital-gardening-2026.md, raw/articles/litho-deepwiki-rs-code-documentation-2026.md]
confidence: medium
---
@@ -20,6 +20,8 @@ confidence: medium
実装候補として、[[obsidian]] 的な手元優先の編集体験、Cosense 的な共同編集、Nuxt Content の Markdown と拡張構文、WordPress + WPGraphQL、HyperMD や Milkdown などの編集器が挙げられている。Git で全体を管理すると版管理は強くなるが、携帯端末での編集しやすさが弱くなるため、保存形式、同期、編集体験の折り合いが設計上の中心になる。
[[litho]] のようにコードベースから Wiki 風ドキュメントを生成する道具は、この CMS 発想をソフトウェア設計書側に寄せた例。人間が育てる庭とは違い、コード解析と CI/CD によって鮮度を保とうとするが、生成物をどうレビューし、どの情報を手で補うかは同じく設計問題として残る。
## Open Questions
- ページ型を単一にしたまま、公開状態・到達性・版管理をどう表すか。
+8 -2
View File
@@ -1,10 +1,10 @@
---
title: Inclusive Design
created: 2026-06-28
updated: 2026-06-28
updated: 2026-06-30
type: concept
tags: [accessibility, inclusive-design, public-interest, design]
sources: [raw/articles/arun-japan-symbols-2026.md]
sources: [raw/articles/arun-japan-symbols-2026.md, raw/articles/aist-avatar-standardization-committee-2026.md, raw/articles/accessibility-conference-chiba-2026.md]
confidence: medium
---
@@ -16,6 +16,12 @@ confidence: medium
この資料で面白いのは、包摂性を個別の支援制度だけでなく、公共空間の情報設計として扱っている点。標識や札は、読める人だけに向けた文章ではなく、瞬時に見分けられる形と色で、周囲の行動を少し変える。公共の場で「何をすればよいか」を伝えるという意味では [[meaning-making-marks]] と重なり、社会的な判断材料を壊さず整えるという意味では [[information-integrity]] とも遠くつながる。
[[avatar-standardization]] は、包摂的な設計を XR/メタバース上の仮想身体にも広げる論点。アバターは外見、本人性、相互行為、文化表現を同時に担うため、ユーザ側と開発側の双方に意味のある規格を作る必要がある。
アクセシビリティカンファレンスCHIBA 2026 の案内は、包摂性を「講演で語るテーマ」だけでなく、会場設計と体験ブースに落とし込んでいる例として使える。通常版と情報保障版の YouTube 配信、手話通訳と UD トーク、バリアフリートイレやオストメイト設備の明記、平坦な導線・混雑・照明条件の説明は、参加前に必要な情報へ到達できること自体をアクセシビリティとして扱っている。^[raw/articles/accessibility-conference-chiba-2026.md]
同イベントの体験ブースでは、Ontenna が音の特徴を振動と光に変え、エキマトペが駅の音を AI で識別して文字・手話・オノマトペで可視化する。これは [[meaning-making-marks]] のような公共空間の記号設計を、聴覚・身体感覚・文字情報へまたがる multi-modal interface に広げる実装例として読める。^[raw/articles/accessibility-conference-chiba-2026.md]
## 見るべき問い
- 本人が詳細を説明しなくても、必要な配慮だけが伝わる合図をどう作るか。
+15 -3
View File
@@ -1,10 +1,10 @@
---
title: Information Integrity
created: 2026-06-28
updated: 2026-06-28
updated: 2026-06-30
type: concept
tags: [information-integrity, disinformation, public-interest, media, civic-tech]
sources: [raw/articles/osome-observatory-on-social-media-2026.md, raw/articles/osome-tools-2026.md, raw/articles/maria-ressa-nobel-facts-2026.md, raw/articles/maria-ressa-stop-the-lies-2026.md, raw/articles/tbs-sns-metric-farms-election-disinformation-2026.md]
tags: [information-integrity, disinformation, public-interest, media, civic-tech, knowledge-base]
sources: [raw/articles/osome-observatory-on-social-media-2026.md, raw/articles/osome-tools-2026.md, raw/articles/maria-ressa-nobel-facts-2026.md, raw/articles/maria-ressa-stop-the-lies-2026.md, raw/articles/tbs-sns-metric-farms-election-disinformation-2026.md, raw/articles/wikipedia-sanger-canvassing-ban-2026.md, raw/articles/wikipedia-fake-russian-history-zhemao-2022.md]
confidence: medium
---
@@ -28,6 +28,18 @@ TBS NEWS DIG の「農場」取材は、偽情報の文章や画像だけでな
この点は、情報の完全性を「真偽」だけでなく「何が目立つよう設計され、誰がその見え方を買えるのか」という問題として扱う必要があることを示す。選挙期間中には、実際の多数派ではないものが多数派のように見える危険が高まり、規制・削除要請・透明性・表現の自由の均衡が論点になる。
## 知識基盤の統治と外部動員
Wikipedia の Larry Sanger ban 事例は、情報基盤の integrity が「偽情報を消す」だけではなく、編集・審議の手続きが外部の audience や政治的動員に飲み込まれないようにする統治でもあることを示す。404 Media の記事では、Sanger の WikiProject Intellectual Diversity 自体よりも、X の 91,000 followers を Wikipedia 内の議論へ誘導した off-wiki canvassing が問題視され、Wikipedia 側は consensus building と参加者保護の観点から indefinite ban に至ったと説明されている。
このケースは [[llm-wiki-pattern]] や [[wiki-maintenance-loop]] にも教訓がある。知識ベースは openness を価値にしつつも、編集権限、参加の境界、外部からの圧力、outing risk、AI-generated slop への耐性を設計しなければ、source の信頼性だけでなく synthesis の場そのものが壊れる。公開 wiki や community moderation を扱うときは、content policy と process integrity を分けて見る必要がある。
## 知識基盤で虚構が体系化されるリスク
中国語版 Wikipedia の Zhemao / 折毛事件は、情報の完全性が SNS 上の拡散だけでなく、百科事典型の knowledge-base でも壊れうることを示す。GIGAZINE の要約によれば、1人の編集者が 2010 年ごろから複数アカウントで実在の歴史・人物・国家間対立に架空の鉱山や事件を混ぜ込み、最終的に 206 件の記事と数百万語規模の「架空のロシア史」を中国語版 Wikipedia に作った。問題は単発の嘘ではなく、関連記事・脚注・文体・相互参照がそろったため、読者に体系として見えてしまった点にある。^[raw/articles/wikipedia-fake-russian-history-zhemao-2022.md]
このケースは [[llm-wiki-pattern]] と [[wiki-maintenance-loop]] にも直接関係する。知識ベースは、リンクや一貫した文体によって信頼感を作れる一方で、出典の存在確認、他言語・一次資料との照合、矛盾検出、編集者権限の分離が弱いと、もっともらしい synthesis が長期間残る。LLM が Wiki を更新する場合も、文章の自然さではなく source provenance と反証可能性を保つことが integrity の中心になる。
## 関連領域
この領域は、公共のための技術、報道、偽情報対策、SNS の設計、アクセシビリティ、民主主義の維持とつながる。OSoMe や Ressa のような資料は、流行のニュースとして消費するより、人物・組織・道具・概念に分けて蓄積すると後から参照しやすい。
+43
View File
@@ -0,0 +1,43 @@
---
title: Loop Engineering
created: 2026-06-29
updated: 2026-06-30
type: concept
tags: [agent, automation, workflow, quality]
sources: [raw/articles/loop-engineering-anthropic-playbook-2026.md, raw/articles/github-issueops-state-machines-2026.md, raw/articles/horizon-agentic-hardware-design-2026.md, raw/articles/abtop-ai-coding-agent-monitor-2026.md, raw/articles/kiro-ide-1-0-agent-focus-2026.md, raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/pi-coding-agent-2025.md, raw/articles/github-desktop-3-6-worktrees-copilot-2026.md, raw/articles/agent-oriented-cli-zenn-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md]
confidence: medium
---
# Loop Engineering
Loop engineering は、LLM やエージェントを「人間が毎回プロンプトする道具」としてではなく、自分で次の仕事を見つけ、隔離された実行先へ渡し、別の評価者に検証させ、結果を永続化し、次回実行へつなぐループとして設計する考え方。PDF「Loop Engineering: The Anthropic Playbook for Designing Systems That Prompt Your Agents」は、prompt / context / harness engineering の上に来る第四層としてこの概念を置いている。
中心になる 1 ターンは、discovery、handoff、verification、persistence、scheduling の 5 段階。これは [[wiki-maintenance-loop]] の ingest / query / lint / log / cron にかなり近く、Discord link ingest のような自動キュレーションでは、リンク発見だけでなく評価基準、重複検出、raw 保存、wiki 更新、状態ファイル更新までを同じループの一部として扱う必要がある。
## 設計上の要点
- **Generator / evaluator separation**: 生成したエージェント自身に採点させると甘くなりやすい。別プロンプト、別モデル、別プロセスの「疑う評価者」を置く方が、[[ai-assisted-reverse-engineering]] の命名・型付け検査や wiki ingest の品質判定にも応用しやすい。
- **Persistence first**: ループの成果はチャットの返答だけでなく、PR、Issue、wiki、state file、log のような再利用可能な場所に残す。残らない自動化は、次回の discovery と評価に使えない。
- **State machine として考える**: GitHub の IssueOps 記事は、Issue、label、comment、approval、Action を状態機械として扱う。これは loop engineering の persistence と scheduling を、監査可能な GitHub timeline に置く方法として読める。
- **Repository-native loop**: [[agentic-hardware-design]] の HORIZON は、Markdown harness から evaluator / acceptance predicate / git policy を持つ project pack を作り、隔離された worktree 上の diff・commit・log・notes をそのまま探索 trace にする。ループの状態を外部データベースに逃がさず、作業対象の repository 自体に残す設計として重要。
- **Operator observability**: [[abtop]] は Claude Code、Codex CLI、OpenCode の session、token、context、rate limit、child process、open port をローカルで可視化する。複数 agent を同時に回す loop では、成果物だけでなく「いま何が動いているか」「どの資源を占有しているか」も運用対象になる。
- **Agent-native work surface**: [[kiro]] の Agent Focus は、コード編集画面ではなく session、会話、spec、diff を前面に置く。loop を IDE の中に寄せると、人間の仕事は直接編集よりも、仕様・承認・差分確認・権限ルールの管理へ移る。
- **Agent harness minimalism**: [[pi-coding-agent]] は、read / write / edit / bash、tmux、明示的な session file など既存の可視な道具に寄せることで、隠れた sub-agent や巨大な system prompt に頼らない loop を作ろうとする。loop の強さは機能数だけでなく、operator が context、tool result、process、cost をどこまで観測できるかにも依存する。
- **Worktrees as everyday agent infrastructure**: GitHub Desktop 3.6 の worktree support は、agent が複数 branch / sandbox を使う流れを GUI 側にも取り込む。isolated worktree は [[agentic-hardware-design]] や IssueOps 的な repository-native loop と同じく、並列作業を見える単位に分けるための基礎部品になる。
- **Evaluator market pressure**: [[ai-evaluation-infrastructure]] の Arena 事例は、評価が研究用 leaderboard から商用分析・post-training 改善の基盤へ広がっていることを示す。loop engineering でも、実行する agent だけでなく、それを測る evaluator とデータ収集の設計が競争力になる。
- **Agent-oriented tools**: [[agent-oriented-cli-design]] は、JSON first、actionable error、search/read 分離、鮮度情報、少ないフラグを通じて、agent が推測せずに次の行動へ進める CLI を作る考え方。loop の実行単位である tool が曖昧だと、verification や persistence 以前に誤った状態で進んでしまう。
- **Human judgment is scarce**: 生成は安くなっても、何を通し、何を止め、何を記録するかの判断は希少になる。自動化は人間の判断を消すのではなく、判断すべき点を狭く明確にするべき。
## Failure Modes
- **Nodding loop**: 自己承認しているだけで、独立した検証がない。
- **Amnesiac loop**: 前回の状態、失敗、採用・却下理由が保存されず、毎回同じ判断をやり直す。
- **Manual loop**: 実行自体が人間の思いつきに依存し、定期実行・イベント駆動になっていない。
- **Blind loop**: 何を探すか、何を高く評価するかが固定され、発見結果から rubric が改善されない。
- **Tangled loop**: 並列エージェントや自動 PR が同じ資源を同時に触り、検証や rollback が追いつかない。
## Open Questions
- Hermes の scheduled job では、どの変更を即時実行し、どの変更を review item として止めるべきか。
- [[ai-agent-identity-security]] のような認可・監査の標準を、個人用エージェントのローカル自動化にもどう縮小適用するか。
- IssueOps や [[llm-wiki-pattern]] のような永続メディアを、チャット駆動の一時的な作業とどう接続するか。
+4 -2
View File
@@ -1,10 +1,10 @@
---
title: Wiki Maintenance Loop
created: 2026-06-28
updated: 2026-06-28
updated: 2026-06-29
type: concept
tags: [workflow, maintenance, wiki, agent]
sources: [raw/articles/karpathy-llm-wiki-2026.md, raw/articles/hermes-research-llm-wiki-skill-2026.md, raw/articles/nashsu-llm-wiki-2026.md, raw/articles/tokium-self-evolving-ai-researcher-2026.md]
sources: [raw/articles/karpathy-llm-wiki-2026.md, raw/articles/hermes-research-llm-wiki-skill-2026.md, raw/articles/nashsu-llm-wiki-2026.md, raw/articles/tokium-self-evolving-ai-researcher-2026.md, raw/articles/loop-engineering-anthropic-playbook-2026.md]
confidence: medium
---
@@ -25,3 +25,5 @@ Hermes の `llm-wiki` skill では、毎回 `SCHEMA.md`、`index.md`、recent `l
[[llm-wiki-app]] はこの loop をアプリ側の persistent ingest queue、folder auto-watch、source cleanup、graph/search、MCP/API に拡張している。つまり maintenance loop は単なる checklist ではなく、agent procedure と product feature のどちらにもなりうる。
TOKIUM の [[ai-research-automation]] 事例は、wiki ではなく技術動向の報告作成でも同じ考え方が使えることを示している。検索語や巡回先を固定せず、採用された情報源を評価して候補を昇格・降格させることで、収集対象そのものを手入れの対象にしている。
[[loop-engineering]] の観点では、この maintenance loop は「発見 → handoff → 検証 → 永続化 → scheduling」を持つ小さな自走ループでもある。特に Discord link ingest では、良さそうなリンクを全部入れるのではなく、interest profile、重複検出、raw hash、index/log/state 更新を通じて、次回の判断材料を残すことが品質維持の中心になる。