add
This commit is contained in:
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: Agent Harness Engineering
|
||||
created: 2026-07-01
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [agent, automation, evaluation, workflow, quality, reliability]
|
||||
sources: [raw/articles/awesome-harness-engineering-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/github-copilot-vision-ga-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md, raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/vscode-1-110-agent-browser-tools-2026.md, raw/articles/explain-diff-html-agent-skill-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/skamille-respectful-ai-use-guidelines-2026.md, raw/articles/devansh-llm-vulnerability-research-2026.md]
|
||||
sources: [raw/articles/awesome-harness-engineering-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/github-copilot-vision-ga-2026.md, raw/articles/github-copilot-ai-credit-session-limits-2026.md, raw/articles/notion-developer-platform-agents-workers-2026.md, raw/articles/claude-code-changelog-agent-ops-2-1-198-2026.md, raw/articles/aws-forward-deployed-engineering-agentic-ai-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/vscode-1-110-agent-browser-tools-2026.md, raw/articles/explain-diff-html-agent-skill-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/skamille-respectful-ai-use-guidelines-2026.md, raw/articles/devansh-llm-vulnerability-research-2026.md, raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -24,6 +24,7 @@ Agent harness engineering は、AI agent の賢さを model 単体で見ず、
|
||||
- **Minimal security-research scaffolding**: Devansh の [[llm-assisted-vulnerability-research]] 記事は、脆弱性探索では bloated `AGENT.md` / `SKILLS.md` や広い checklist が context rot を悪化させることがあり、1 ページ程度の threat model、不変条件、thin slice、verifier loop に token を使う方が実用的だとする。これは harness を増やす話ではなく、harness を「注意を散らさず、検証を強制する最小構造」に削る設計として重要である。^[raw/articles/devansh-llm-vulnerability-research-2026.md]
|
||||
- **Secret-access harnesses**: 1Password Environments MCP Server for Codex は、agent が環境を構成・実行する時に secret value を model context へ入れず、user approval と runtime injection に閉じ込める harness である。agent harness engineering では、tool を増やすだけでなく、credential がどの channel に現れないかを仕様として固定することが安全な自律性の条件になる。^[raw/articles/1password-codex-mcp-secret-access-2026.md]
|
||||
- **Evals and observability**: skill eval、trace grading、OpenTelemetry、session replay、cost tracking、benchmark を使い、成功/失敗を operator の感覚だけにしない。[[ai-evaluation-infrastructure]] では model / agent を測る市場や基盤が主題だが、harness engineering では eval を個々の workflow の改善 loop に入れる。
|
||||
- **Agentic test harnesses**: Slack Engineering の E2E 実験は、同じ UI goal を Playwright MCP、Playwright CLI、agent-generated Playwright tests で比較し、model より execution harness の差が reliability / turn count / token cost に効くことを示す。MCP は UI 操作と状態取得をまとめて返すため CLI より少ない turn で済み、複雑な flow でも失敗率が低かった一方、agentic run は snapshot と履歴の再送で高コストになる。したがって agent harness では browser primitive の形、state snapshot の粒度、context compaction、action signature の記録、deterministic test への切り戻しを同時に設計する必要がある。これは [[e2e-coverage-metrics]] と [[safari-mcp-server]] の実践的な接点である。^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
- **Browser harnesses**: GitHub Copilot の VS Code browser tools GA は、agent が live web app を操作し、console error、screenshot、scripted flow を chat へ戻す harness を IDE に組み込む例である。重要なのは browser 操作そのものだけでなく、人間 tab の明示共有、agent tab の session isolation、camera/microphone/geolocation の既定拒否、enterprise allow/deny と workspace trust を同じ harness に入れている点で、これは [[ai-agent-identity-security]] と [[e2e-coverage-metrics]] の接点になる。^[raw/articles/github-copilot-browser-tools-ga-2026.md]
|
||||
- **Local browser MCP harnesses**: [[safari-mcp-server]] は、Safari Technology Preview の `safaridriver --mcp` を MCP server として公開し、agent が Safari の DOM、network request、console、screenshot、viewport、dialog、tab、page content を直接観測・操作できるようにする。Copilot browser tools が IDE 統合の browser harness なら、Safari MCP は特定ブラウザの実装差、性能、アクセシビリティ、form state を agent loop に入れる local harness である。^[raw/articles/safari-mcp-server-webkit-2026.md]
|
||||
- **IDE-level agent control surface**: VS Code 1.110 は、agentic browser tools だけでなく、Agent Debug panel、background agent の `/compact` や slash command、session rename、Claude agent の steering / queuing、agent plugins、session memory、chat fork までまとめて入れている。これは browser 操作単体の話ではなく、agent を長時間走らせ、何を読み込んだか・どの tool を呼んだか・どの session へ分岐したかを IDE 側で観測し制御する harness への移行である。auto-approve `/yolo` は便利だが、記事自体も terminal sandboxing と security implication を明示しており、[[ai-agent-command-safety]] と [[ai-agent-identity-security]] の境界設計なしには扱えない。^[raw/articles/vscode-1-110-agent-browser-tools-2026.md]
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: AI Agent Command Safety
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-01
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [agent, security, reliability, automation]
|
||||
sources: [raw/articles/guardfall-ai-coding-agent-shell-bypass-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/cursor-duneslide-sandbox-escape-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/koi-promptjacking-claude-desktop-rce-2026.md]
|
||||
sources: [raw/articles/guardfall-ai-coding-agent-shell-bypass-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/cursor-duneslide-sandbox-escape-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/koi-promptjacking-claude-desktop-rce-2026.md, raw/articles/prompt-injection-as-role-confusion-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -25,6 +25,8 @@ Cursor の DuneSlide 事例は、sandbox があるだけでは足りず、「age
|
||||
|
||||
Claude Desktop まわりの事例は、command safety が「生成された shell 文字列」だけではなく、設定同期、MCP/extension、personal preferences、偽 error message まで含む広い実行経路の問題であることを示す。The Register の Pentera Labs 記事では、攻撃者が Claude の account-wide personalization に base64 prompt を入れ、Desktop Commander など command-capable MCP があれば reverse shell、なければ Anthropic 風の偽エラーと install prompt でユーザーに実行させる流れが説明されている。Koi の PromptJacking 報告では、公式 Claude Desktop extensions が unsandboxed MCP server として動き、AppleScript への未 escape URL 補間から web prompt injection → local RCE へ進みうると説明されている。どちらも「agent が command を出す瞬間」より前に、信頼済み assistant の設定・connector・外部 web content が command path へ混ざるため、設定変更監視、extension allowlist、connector sandboxing が command guard と同じ層で必要になる。^[raw/articles/theregister-claude-desktop-double-agent-2026.md] ^[raw/articles/koi-promptjacking-claude-desktop-rce-2026.md]
|
||||
|
||||
[[prompt-injection-role-confusion]] は、この問題の model-internal な説明を与える。Role Confusion の writeup は、tool text 内の命令が user 風・reasoning 風に見えると、実際の role tag より文体特徴が勝ち、モデルが低権限 text を命令や自分の推論として扱いやすくなると示す。したがって command safety では、model に「これは tool output だから無視して」と頼むだけでなく、untrusted text から shell / MCP / file write へ進む経路を harness 側で分離し、role provenance を承認 UI と実行ログに残す必要がある。^[raw/articles/prompt-injection-as-role-confusion-2026.md]
|
||||
|
||||
## なぜ重要か
|
||||
|
||||
Yuta の運用では、Hermes の scheduled job、Codex/Claude/OpenCode、local CLI、CI runner が同じ「agent が command を出す」面を共有する。便利な自走 loop ほど、guard を抜けた command が SSH key、cloud credential、wiki、repo、home directory へ届きやすい。したがって agent command safety は、個別 agent の機能ではなく、[[agent-oriented-cli-design]]、[[ai-agent-identity-security]]、[[ci-cd-runtime-security]] を横断する運用品質の条件として扱うべきである。
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: AI Agent Identity Security
|
||||
created: 2026-06-29
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [agent, security, reliability, privacy]
|
||||
sources: [raw/articles/okta-cross-app-access-partners-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md, raw/articles/ios-ai-chatbot-llm-key-leakage-2026.md, raw/articles/unity-terms-agentic-access-2026.md, raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/cloudflare-monetization-gateway-x402-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/xai-voice-agent-builder-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/email-verification-protocol-draft-2026.md]
|
||||
sources: [raw/articles/okta-cross-app-access-partners-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/microsoft-learn-agent-architecture-sdlc-2026.md, raw/articles/ios-ai-chatbot-llm-key-leakage-2026.md, raw/articles/unity-terms-agentic-access-2026.md, raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/cloudflare-monetization-gateway-x402-2026.md, raw/articles/github-copilot-browser-tools-ga-2026.md, raw/articles/xai-voice-agent-builder-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/safari-mcp-server-webkit-2026.md, raw/articles/1password-codex-mcp-secret-access-2026.md, raw/articles/email-verification-protocol-draft-2026.md, raw/articles/1password-credential-broker-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -21,6 +21,7 @@ AI agent identity security は、AI エージェントやアプリ間連携が
|
||||
- **Policy and audit**: アクセスが許可される前に企業ポリシーで検査し、操作の監査証跡を残す。これは [[loop-engineering]] の persistence と verification をセキュリティ境界へ移したものでもある。
|
||||
- **Least privilege for agents**: 常時広い権限を持つ bot token ではなく、必要な範囲に絞った identity-based token を使う。
|
||||
- **Secret custody outside the model**: 1Password Environments MCP Server for Codex は、coding agent を secret の保管庫ではなく「承認された利用主体」として扱う設計例である。Codex は environment を作成し、変数名を扱い、実行を orchestrate できるが、secret value は MCP channel、model context、local file、terminal へ返さず、1Password が承認済み process の runtime memory にだけ注入する。これにより、agent workflow の速度を保ちながら、credential custody、explicit approval、scope、audit を [[agent-harness-engineering]] 側の実行 loop へ組み込める。^[raw/articles/1password-codex-mcp-secret-access-2026.md]
|
||||
- **Runtime credential brokering**: 1Password Credential Broker は、GitHub Actions の OIDC / Workload Identity Federation を使い、repo・branch・workflow・environment・commit で workload を検証してから、vault 全体ではなく job が許可された item だけを実行時に渡す設計である。2026-06 の private beta は GitHub Actions に絞られ、自動 rotation までは含まないが、standing vault access を消し、各 access event に workload attribution を付ける点で [[ci-cd-runtime-security]] と agent identity の接点になる。1Password は同じ broker を AI agent にも広げ、agent が long-lived OAuth / refresh token を抱えず、task-scoped short-lived token を受け取り、人間の delegator まで監査できる形を目指すとしている。^[raw/articles/1password-credential-broker-2026.md]
|
||||
- **Browser-mediated identity assertions**: Email Verification Protocol draft は、email verification を「メールを送って code を入力させる」方式から、browser が relying party と issuer の間を仲介して signed token を受け渡す方式へ寄せる。issuer は RP identity を直接知る必要がなく、RP は nonce と browser key binding で token を検証するため、friction reduction と privacy separation を同時に狙う標準化案として読める。これは agent 固有ではないが、agent が account creation や delegated workflow を扱う時代には、identity assertion を browser / issuer / RP に分け、過剰な identifier sharing を避ける設計として隣接する。^[raw/articles/email-verification-protocol-draft-2026.md]
|
||||
- **Local sandbox / approval boundary**: Codex の安全運用ドキュメントは、cloud では隔離 container、CLI/IDE では OS sandbox と approval policy を組み合わせ、既定で network access を切り、workspace 外の編集や network 利用を承認対象にする設計を説明している。`workspace-write`、`read-only`、network proxy、domain allow/deny などの設定は、企業の cross-app 認可だけでなく個人の agent loop でも「どこまで自動実行してよいか」を明示する制御面になる。
|
||||
- **Browser / device permission boundary**: GitHub Copilot の VS Code browser tools GA は、agent が実ブラウザを開き、click/type/drag、console error、screenshot、scripted flow を使えるようにする一方、人間が開いた tab は `Share with Agent` するまで読めず、agent tab は fresh session で cookie/storage から隔離され、camera/microphone/geolocation は既定拒否になると説明している。browser が agent tool になるほど、tab ownership、session isolation、site allow/deny、workspace trust は identity boundary の一部になる。^[raw/articles/github-copilot-browser-tools-ga-2026.md]
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: AI Agent Telemetry Privacy
|
||||
created: 2026-07-01
|
||||
updated: 2026-07-01
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [agent, privacy, security, data-protection, reliability]
|
||||
sources: [raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/theregister-claude-code-transcript-retention-2026.md]
|
||||
sources: [raw/articles/claude-code-telemetry-audit-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/theregister-claude-code-transcript-retention-2026.md, raw/articles/simon-grok-build-source-privacy-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -24,6 +24,8 @@ Adnane Khan の Claude Code 2.1.196 audit は、公開文書の「Statsig metric
|
||||
|
||||
## 運用上の含意
|
||||
|
||||
Grok Build の初期 beta で directory 全体が xAI の Google Cloud bucket へ upload されうる挙動が批判され、保持 data 削除、retention default off、codebase の Apache 2.0 公開へつながった事例は、agent privacy を「telemetry」だけでなく作業 directory upload / session state upload / local-first fallback の問題として扱う必要を示す。Simon Willison の読解では、公開 codebase には GCS upload code の残骸や disabled `upload_session_state()` が残っており、agent の信頼回復には open source 化だけでなく、既定値・削除保証・実装上の upload path の検証が必要になる。^[raw/articles/simon-grok-build-source-privacy-2026.md]
|
||||
|
||||
Yuta の agent 運用では、telemetry は単純な「送る/送らない」ではなく、loop の観測性と privacy の交換条件として扱う必要がある。OpenAI Codex の approvals/security docs が示すように、OTel を自分の collector へ送れる設計は [[agent-harness-engineering]] の観測性を高める一方、collector 側の retention と access control を同時に決めなければならない。
|
||||
|
||||
この論点は [[data-protection-and-expression]] とも接続する。agent が開発者の作業文脈を観測するほど、利用者の control、説明、削除、第三者提供の透明性が重要になる。特に CLI agent は IDE より権限が広く、shell、repo、CI、browser-use、MCP server へまたがるため、telemetry 設計を product quality と security boundary の一部として読むべきである。
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: AI Evaluation Infrastructure
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [evaluation, llm, quality, workflow]
|
||||
sources: [raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/harbor-langchain-agent-eval-stack-2026.md, raw/articles/anthropic-claude-sonnet-5-2026.md, raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md, raw/articles/openai-genebench-pro-2026.md, raw/articles/shopify-flow-agent-model-optimization-flywheel-2026.md, raw/articles/vllm-semantic-router-micro-agents-2026.md]
|
||||
sources: [raw/articles/arena-ai-leaderboard-business-2026.md, raw/articles/harbor-langchain-agent-eval-stack-2026.md, raw/articles/anthropic-claude-sonnet-5-2026.md, raw/articles/anthropic-redeploying-fable-5-jailbreak-framework-2026.md, raw/articles/openai-genebench-pro-2026.md, raw/articles/shopify-flow-agent-model-optimization-flywheel-2026.md, raw/articles/vllm-semantic-router-micro-agents-2026.md, raw/articles/artificial-analysis-coding-agent-benchmarks-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -26,6 +26,8 @@ Shopify の Flow agent fine-tuning 記事は、evaluation infrastructure が pro
|
||||
|
||||
vLLM の Semantic Router / micro-agent 構想は、評価と orchestration が serving layer に入り込む例である。単一の OpenAI-compatible model ID の裏で、router が task に応じて recipe を選び、複数 worker に fan-out し、quorum、disagreement check、output contract repair、synthesis を行う。これは「どの model が強いか」を外から測るだけでなく、router 自体が小さな evaluator / coordinator になり、frontier model 呼び出しの前段で capability と cost/safety policy を組み立てるという設計である。[[loop-engineering]] や [[agent-harness-engineering]] では、application graph だけでなく inference gateway も評価・検証・合議の場になる。^[raw/articles/vllm-semantic-router-micro-agents-2026.md]
|
||||
|
||||
Artificial Analysis の Coding Agent Benchmarks は、agent 評価が「モデル名」だけでなく **harness、benchmark mix、cost、token usage、execution time** を一緒に測る段階へ進んでいることを示す。DeepSWE、Terminal-Bench v2、SWE-Atlas-QnA を合成した Coding Agent Index と、Claude Code / Cursor CLI / Opencode の harness comparison は、同じ model でも実行環境と operator loop によって成績が変わることを可視化する。これは [[agent-harness-engineering]] と [[agent-oriented-cli-design]] に近く、評価基盤が model leaderboard から agent runtime / CLI / workflow 比較へ広がる兆候である。^[raw/articles/artificial-analysis-coding-agent-benchmarks-2026.md]
|
||||
|
||||
## なぜ重要か
|
||||
|
||||
- **Crowdsourced comparison**: 利用者が 2 つのモデル出力を比較する形式は、静的な benchmark では拾いにくい実利用の好みを集められる。
|
||||
@@ -35,6 +37,7 @@ vLLM の Semantic Router / micro-agent 構想は、評価と orchestration が s
|
||||
- **Judgment-heavy scientific evaluation**: GeneBench-Pro のような benchmark は、正解率だけでなく、データ診断、分析方針の変更、因果推論、solver contract の明確さまで評価対象にする。科学 agent を評価するには、clean sandbox と deterministic check だけでなく、trace から判断品質を検査できる問題設計が必要になる。
|
||||
- **Production feedback flywheels**: Shopify Flow の例では、synthetic benchmark、LLM judge、programmatic checker、本番 activation rate、slice analysis、週次 retraining が一つの改善ループになる。評価基盤は「合格判定」ではなく、どのデータを足し、どの形式を変え、どの tool response を削るかを決める運用面になる。
|
||||
- **Serving-layer evaluators**: vLLM Semantic Router のように、model router が quorum、disagreement check、output repair を実行すると、評価は offline benchmark だけでなく、inference request ごとの制御面にも入る。
|
||||
- **Harness-sensitive coding-agent evaluation**: Artificial Analysis のような coding agent leaderboard は、agent が使う CLI/harness、token usage、実行時間、費用を同じ評価面に載せるため、[[agent-harness-engineering]] そのものが比較対象になる。
|
||||
|
||||
## Open Questions
|
||||
|
||||
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: AI Infrastructure Ownership
|
||||
created: 2026-07-17
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [llm, agent, automation, public-interest, privacy]
|
||||
sources: [raw/articles/owners-not-renters-ai-ownership-grid-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# AI Infrastructure Ownership
|
||||
|
||||
AI infrastructure ownership は、AI を巨大な外部 provider から借りる「ベルト」型の供給として使い続けるのか、組織・個人・端末・地域ごとに小さな実行基盤を持つ「グリッド」型へ分散するのかという設計論点。Owners Not Renters の “Our future is now” は、工場が中央 steam engine から各機械の electric motor へ移った比喩を使い、frontier model が他者の建物に置かれたままだと、価格、政策、検閲、データ境界、運用停止の意思決定を借り物にしてしまうと論じる。^[raw/articles/owners-not-renters-ai-ownership-grid-2026.md]
|
||||
|
||||
この論点は [[agentic-web-monetization]] や [[ai-crawler-governance]] と同じく、AI 時代の web / compute / data access を誰が制御するかという問題である。ただし payment や crawler policy より一段下の、モデル実行基盤そのものの所有・配置・更新・監査の問題を扱う。個人の [[llm-wiki-pattern]] や Hermes job も、どの model/provider に source ingestion や判断を任せるかによって、知識基盤の独立性と継続性が変わる。
|
||||
|
||||
## 見るべき軸
|
||||
|
||||
- **実行場所**: hosted frontier model、企業内 model、local/on-device model、edge model のどこで agent loop が動くか。
|
||||
- **切替可能性**: provider の価格変更、政策命令、service outage、model retirement が起きたとき、workflow を別基盤へ逃がせるか。
|
||||
- **データ境界**: private memory、source archive、codebase、研究データを外部 model へ送る場合の [[data-protection-and-expression]] と監査。
|
||||
- **agent の自律性**: [[loop-engineering]] で agent が長時間動くほど、途中停止・制限変更・quota・telemetry の影響が大きくなる。
|
||||
|
||||
## Open questions
|
||||
|
||||
- 個人や小組織にとって、local/open model を持つ価値は privacy なのか、cost なのか、停止耐性なのか、交渉力なのか。
|
||||
- Frontier capability が必要な task と、小さな owned model で十分な task をどう分類するべきか。
|
||||
- Wiki curation や research automation では、source fetching、classification、synthesis、verification のどの段階を owned infrastructure へ寄せるのが効果的か。
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: CI/CD Runtime Security
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [security, supply-chain, quality, reliability, automation]
|
||||
sources: [raw/articles/cicd-sensor-2026.md, raw/articles/scrutineer-oss-security-workflow-2026.md, raw/articles/tangled-spindle-microvm-ci-runners-2026.md, raw/articles/thehackernews-argo-cd-repo-server-unpatched-rce-2026.md, raw/articles/synacktiv-argo-cd-codeql-rce-2026.md, raw/articles/zenn-ai-generated-github-actions-yaml-security-2026.md, raw/articles/flatt-github-actions-credential-leakage-2026.md, raw/articles/github-secret-scanning-public-monitoring-2026.md, raw/articles/microsoft-ghqr-github-quick-review-2026.md, raw/articles/strix-ai-pentesting-agent-2026.md]
|
||||
sources: [raw/articles/cicd-sensor-2026.md, raw/articles/scrutineer-oss-security-workflow-2026.md, raw/articles/tangled-spindle-microvm-ci-runners-2026.md, raw/articles/thehackernews-argo-cd-repo-server-unpatched-rce-2026.md, raw/articles/synacktiv-argo-cd-codeql-rce-2026.md, raw/articles/zenn-ai-generated-github-actions-yaml-security-2026.md, raw/articles/flatt-github-actions-credential-leakage-2026.md, raw/articles/github-secret-scanning-public-monitoring-2026.md, raw/articles/microsoft-ghqr-github-quick-review-2026.md, raw/articles/strix-ai-pentesting-agent-2026.md, raw/articles/github-actions-workflow-execution-protections-2026.md, raw/articles/1password-credential-broker-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -28,10 +28,14 @@ AI が生成した GitHub Actions YAML は、動作確認だけでなく trigger
|
||||
|
||||
GMO Flatt Security の GitHub Actions 解説は、OIDC / Trusted Publishing を入れても「認証後に runner 上へ置かれる派生クレデンシャル」は残る、という runtime 視点を強調する。`GITHUB_TOKEN` は `actions/checkout` の credential persistence や `Runner.Worker` のメモリから読まれうるし、AWS/GCP/Azure/Docker などの認証 Action は一時クレデンシャルや設定ファイルを後続 step から到達可能な場所へ置く。Environment 保護、ruleset、claim の数値 ID 検証、job 分離、短い session duration は有効だが、依存関係・Action・正規レビュアー経由で信頼済み経路に悪意ある code が入ると、漏洩を完全には防げない。したがって runner 側の process/network/file trace と cloud 側の異常検知を合わせ、漏洩後の検知・調査・失効手順まで設計する必要がある。^[raw/articles/flatt-github-actions-credential-leakage-2026.md]
|
||||
|
||||
1Password Credential Broker は、この OIDC / Workload Identity Federation を credential 配布側へ寄せる実装例である。GitHub Actions job が platform-issued signed token で repo・branch・workflow・environment・commit を証明し、1Password が trust policy と照合して item-level の credential だけを job-scoped window で渡す。これは runner 上の派生 credential 問題を完全には消さないが、vault への常時 access と広い service account token を減らし、誰のどの workflow がどの credential を取ったかを監査可能にする。将来の AI agent 対応では、agent も同じく task-scoped short-lived token を受け取るため、[[ai-agent-identity-security]] の実装パターンとしても重要である。^[raw/articles/1password-credential-broker-2026.md]
|
||||
|
||||
GitHub の Secret Protection による public monitoring は、secret leak detection の境界を「自社 repo」から GitHub の公開面全体へ広げる例である。企業メンバーや verified domain の metadata から、個人 fork、OSS repo、issue、pull request、discussion などに漏れた secret を enterprise に帰属させる。これは CI/CD runtime そのものの隔離策ではないが、agent・bot・開発者が組織外の公開面へ token を誤って出す前提で、公開漏洩の発見を incident response loop に入れる実務的な補助線になる。^[raw/articles/github-secret-scanning-public-monitoring-2026.md]
|
||||
|
||||
Microsoft の GitHub Quick Review (`ghqr`) は、GitHub Enterprise / org / repo / GHES を横断して security posture を棚卸しする CLI である。Dependabot、secret scanning、code scanning、2FA/SAML、branch protection、CODEOWNERS、Actions workflow permissions、self-hosted runners、audit log、Copilot policy、MCP settings までを Markdown / Excel / JSON に出せるため、CI/CD runtime security を「個別 YAML のレビュー」から「GitHub tenant 全体の定期診断」へ広げる道具として位置づけられる。^[raw/articles/microsoft-ghqr-github-quick-review-2026.md]
|
||||
|
||||
GitHub Actions の workflow execution protections は、runtime boundary を YAML 内の自己防衛から enterprise / org / repository policy へ引き上げる機能である。Actor rule で workflow を起動できる user / role / GitHub App / Copilot / Dependabot を制限し、event rule で `pull_request_target` や `workflow_dispatch` などを制御できる。これは malicious workflow file が commit に入った後で runner が動く前に、中央 ruleset が「その actor/event は workflow を走らせてよいか」を判定するため、AI が生成した YAML のレビューや [[open-source-package-supply-chain-attacks]] の依存実行対策と補完関係にある。^[raw/articles/github-actions-workflow-execution-protections-2026.md]
|
||||
|
||||
[[strix]] は、CI/CD に入る security testing が SAST や設定監査だけでなく、実行中の application へ AI pentest agent を当て、reconnaissance、exploitation、PoC validation、修正案、report まで返す方向へ広がる例である。これは「runner が何をしたかを監視する」cicd-sensor 型の runtime evidence と対になる。Strix のような tool を PR gate に置くなら、検査対象の sandbox、network egress、test credential、false-positive review、auto-fix の human gate まで含めて CI/CD runtime security として設計する必要がある。^[raw/articles/strix-ai-pentesting-agent-2026.md]
|
||||
|
||||
## Design implications
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: Data Protection and Expression
|
||||
created: 2026-06-28
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [data-protection, privacy, freedom-expression, law, public-interest, media]
|
||||
sources: [raw/articles/david-erdos-cambridge-law-profile-2026.md, raw/articles/ravi-naik-awo-profile-2026.md, raw/articles/google-zkp-age-assurance-2026.md, raw/articles/longfellow-zk-identity-proofs-2026.md, raw/articles/email-verification-protocol-draft-2026.md]
|
||||
sources: [raw/articles/david-erdos-cambridge-law-profile-2026.md, raw/articles/ravi-naik-awo-profile-2026.md, raw/articles/google-zkp-age-assurance-2026.md, raw/articles/longfellow-zk-identity-proofs-2026.md, raw/articles/email-verification-protocol-draft-2026.md, raw/articles/unicef-social-media-age-restrictions-child-rights-2026.md, raw/articles/yonhap-korea-social-media-age-restrictions-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -26,6 +26,8 @@ Erdos の Cambridge profile は、EU 内でもこの均衡の置き方が大き
|
||||
|
||||
Google が公開した age assurance 向け Zero-Knowledge Proof library は、年齢確認のような規制対応を「本人属性を証明するが、それ以外のデータは渡さない」設計へ寄せる例として読める。EU の eIDAS / EUDI Wallet 文脈では、未成年保護や年齢制限サービスの実装が本人確認データの過剰収集になりやすいため、ZKP は [[inclusive-design]] 的な使いやすさと、[[ai-developer-liability]] 的な設計責任の両方に関わる。Longfellow ZK は ISO MDOC、JWT、W3C Verifiable Credentials のような既存 identity standard に対して anonymous credential / zero-knowledge proof を構成する実装で、legacy credential を使いながら disclosure を最小化する方向の実装面を補う。^[raw/articles/google-zkp-age-assurance-2026.md] ^[raw/articles/longfellow-zk-identity-proofs-2026.md]
|
||||
|
||||
[[social-media-age-assurance]] は、この均衡が子どものオンライン安全と SNS regulation で前面化する例である。UNICEF は、年齢制限が child safety への関心から出たものだとしても、privacy と participation を尊重し、platform design と content moderation の責任を置き換えない形にする必要があると述べている。韓国の 14歳未満 SNS 加入制限案も、KYC 型の年齢確認だけでなく、依存性の高い design や recommendation algorithm を制限対象にしている点が重要である。^[raw/articles/unicef-social-media-age-restrictions-child-rights-2026.md] ^[raw/articles/yonhap-korea-social-media-age-restrictions-2026.md]
|
||||
|
||||
Email Verification Protocol draft も同じ系譜にある。従来の email one-time code は、ユーザーに mail client への移動を強いるだけでなく、verification email の送受信や relying party / issuer の関係から余計な情報が流れやすい。EVP は browser を仲介者にし、issuer が email control を token 化し、RP 側では nonce と key binding で検証することで、使いやすさと privacy separation を同時に改善しようとしている。^[raw/articles/email-verification-protocol-draft-2026.md]
|
||||
|
||||
## この Wiki での扱い
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: E2E Coverage Metrics
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-01
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [quality, reliability, evaluation, workflow]
|
||||
sources: [raw/articles/knowledgework-e2e-coverage-metrics-2026.md, raw/articles/github-code-coverage-merge-protection-2026.md, raw/articles/realworld-framework-comparison-spec-2026.md]
|
||||
sources: [raw/articles/knowledgework-e2e-coverage-metrics-2026.md, raw/articles/github-code-coverage-merge-protection-2026.md, raw/articles/realworld-framework-comparison-spec-2026.md, raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -18,6 +18,10 @@ GitHub Code Quality's merge-protection preview shows the same idea being product
|
||||
|
||||
RealWorld adds a benchmark-design angle: many frontend and backend implementations share the same Medium-like app, API specification, backend spec tests, frontend E2E test suite, CSS theme, and hosted demo API. That makes it useful not only as framework learning material, but as a stable surface for comparing generated code, agent-built app variants, and cross-framework regression behavior under one contract. For [[ai-evaluation-infrastructure]], the important part is the common spec/test harness, not the specific app clone.^[raw/articles/realworld-framework-comparison-spec-2026.md]
|
||||
|
||||
Slack Engineering's agentic-testing experiment adds the complementary question of **what E2E tests are measuring**. In 200+ runs, deterministic Playwright tests were fastest and CI-friendly, but agent-driven runs validated whether a user goal could be achieved through different valid UI paths. Their summary is useful: deterministic tests enforce journeys, while agents verify goals. This makes agentic testing more like an exploratory/debugging layer above ordinary E2E, not a replacement for regression gates.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
|
||||
The same Slack experiment also makes cost and observability part of the quality metric. Playwright MCP runs were more reliable than CLI-driven browser control on their flows, partly because the MCP harness returned stable browser state in fewer turns; however, agentic runs still cost roughly $15–30 and accumulated millions of tokens through repeated UI snapshots and conversation retransmission. For [[agent-harness-engineering]], the reusable lesson is that goal coverage, action-signature diversity, turn count, snapshot volume, and failure cause should be measured alongside pass/fail.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
|
||||
This is useful for [[ci-cd-runtime-security]] and [[ai-evaluation-infrastructure]] because it treats test execution as observable runtime evidence, not just a green/red result. It also fits [[loop-engineering]]: the loop should store raw observations, compute metrics later, notify people in the place they already work, and keep enough history to change aggregation methods without rerunning old tests.
|
||||
|
||||
## Caveat
|
||||
@@ -29,3 +33,4 @@ Implementation coverage is a necessary-condition signal, not a sufficient proof
|
||||
- Which surfaces should define the denominator for non-Next.js or non-RPC products: routes, OpenAPI endpoints, event names, domain actions, or user journeys?
|
||||
- When should low E2E coverage block a release, and when should it only produce a review item?
|
||||
- How can AI-generated test additions avoid optimizing for easy-to-cover surfaces while missing high-risk workflows?
|
||||
- For agentic E2E runs, what is the right denominator: required user goals, observed UI paths, meaningful action signatures, or historically flaky workflows?
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: Extensible Data File Formats
|
||||
created: 2026-06-29
|
||||
updated: 2026-06-29
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [data-format, dev-tool, reliability, supply-chain]
|
||||
sources: [raw/articles/f3-file-format-2026.md]
|
||||
sources: [raw/articles/f3-file-format-2026.md, raw/articles/ndl-akoma-ntoso-legal-xml-schema-2014-2026.md, raw/articles/e-legislation-japanese-law-akn-search-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -23,6 +23,8 @@ Extensible data file formats は、保存した時点の読み書き実装だけ
|
||||
|
||||
## 他のページとの関係
|
||||
|
||||
[[legal-document-open-data]] は、同じ「将来の読み手が推測しなくても使える形で残す」問題を、法令・議会文書という公共情報に適用する。Akoma Ntoso や日本法令標準 XML schema は、データ構造の拡張性だけでなく、版、引用、翻訳、linked open data といった制度的な参照可能性を重視する点で、一般的なファイル形式設計より civic-tech 寄りである。^[raw/articles/ndl-akoma-ntoso-legal-xml-schema-2014-2026.md] ^[raw/articles/e-legislation-japanese-law-akn-search-2026.md]
|
||||
|
||||
[[ghidra-mcp]] は、専門作業の手順や品質基準を道具の入口に寄せる例。extensible data file formats は、データを読むための知識をファイル形式側に寄せる例。どちらも [[wiki-maintenance-loop]] と同じく、後から読む人・使う人が毎回文脈を再発見しなくてよい状態を作ろうとしている。
|
||||
|
||||
[[llm-wiki-pattern]] との共通点は、情報をただ置くのではなく、将来の読み手が利用できる形に編み直すこと。ただし LLM Wiki は人間と LLM のための知識整理であり、F3 は分析システムのためのバイナリ保存形式である。
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
---
|
||||
title: Legal Document Open Data
|
||||
created: 2026-07-17
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [civic-tech, public-interest, data-format, knowledge-base, reliability]
|
||||
sources: [raw/articles/ndl-akoma-ntoso-legal-xml-schema-2014-2026.md, raw/articles/e-legislation-japanese-law-akn-search-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# Legal Document Open Data
|
||||
|
||||
Legal document open data は、法律・議会文書を単なる PDF や検索画面ではなく、再利用可能な構造化データとして公開する考え方。Akoma Ntoso は、法令・議会文書の本文、メタデータ、版、引用関係を XML schema と URI 命名規則で表すための標準で、[[extensible-data-file-formats]] と同じく「将来の読み手が推測しなくても使える形で残す」ことを目指す。ただし対象は分析ファイルではなく、公共性の高い法令・議会情報である。^[raw/articles/ndl-akoma-ntoso-legal-xml-schema-2014-2026.md]
|
||||
|
||||
日本の文脈では、総務省の e-LAWS と日本法令向け標準 XML schema によって現行法令のオープンデータ化が進み、e-Gov 法令検索の基盤にもなっている。e-Legislation / Japanese Laws and Regulations Search は、この日本法令データを Akoma Ntoso 系の既存アプリケーションと接続し、日本法令と世界の法律を linked open data として共有・研究しやすくすることを狙っている。これは [[information-integrity]] のうち、情報操作対策ではなく、公共情報の形式・参照・来歴を壊れにくくする側の実践である。^[raw/articles/e-legislation-japanese-law-akn-search-2026.md]
|
||||
|
||||
## 見るべき軸
|
||||
|
||||
- **構造化本文**: 書誌情報だけでなく、法令・議会文書そのものの構造を機械処理できるか。
|
||||
- **版と引用の永続性**: FRBR 的な work / expression / manifestation / item や URI 命名規則で、改正・翻訳・PDF/XML などを一意に参照できるか。
|
||||
- **国際相互運用性**: 国や言語ごとの慣習を吸収しつつ、欧州議会、各国議会、e-Gov などのデータを横断検索・再利用できるか。
|
||||
- **公共サービスとしての信頼性**: 市民、研究者、開発者が、法令の現在性・出典・翻訳・版を確認できるようにする設計になっているか。
|
||||
|
||||
## 関連領域
|
||||
|
||||
このテーマは、公共情報を使える形で保つ [[wiki-maintenance-loop]] と、行政・自治体の意思決定を支える [[climate-adaptation-ai]] のような civic-tech と接続する。AI や検索 UI が法令を扱う場合も、最終的な信頼性はモデルではなく、参照可能な構造化 source と版管理に依存する。
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: Open Source Package Supply Chain Attacks
|
||||
created: 2026-07-01
|
||||
updated: 2026-07-02
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [security, supply-chain, dev-tool, reliability]
|
||||
sources: [raw/articles/checkmarx-operation-navy-ghost-pyrogram-supply-chain-2026.md, raw/articles/ladybird-maintainer-only-development-ai-pr-risk-2026.md]
|
||||
sources: [raw/articles/checkmarx-operation-navy-ghost-pyrogram-supply-chain-2026.md, raw/articles/ladybird-maintainer-only-development-ai-pr-risk-2026.md, raw/articles/asyncapi-miasma-supply-chain-worm-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -20,6 +20,8 @@ The useful pattern is that the malicious code targeted bot/server environments r
|
||||
|
||||
Ladybird の開発方針変更は、package registry ではなく open-source contribution path 側の trust model 変化を示す。Ladybird は AI tool によって「大きな patch を出す労力」が善意や長期関与の proxy ではなくなり、browser のように untrusted internet input を実行する project では、一つのよく隠れた脆弱性が深刻な結果を持つとして、public pull request を閉じ、maintainer だけが code を入れる方針へ移った。これは [[ci-cd-runtime-security]] や [[ai-agent-command-safety]] と同じく、AI が生成速度を上げたことで review capacity と responsibility boundary が希少資源になる例である。^[raw/articles/ladybird-maintainer-only-development-ai-pr-risk-2026.md]
|
||||
|
||||
AsyncAPI の npm package compromise は、package registry attack が `postinstall` だけでなく library の通常 `require()` / `import` path から発火しうることを示す。Flatt Security の解析では、`@asyncapi/[email protected](-alpha.1)` や `@asyncapi/[email protected]` などに Miasma worm が混入し、AWS、Kubernetes、Git、CI、npm token、環境変数を探索し、npm / PyPI / crates.io へ自己拡散する能力を持つ。さらに `.claude/settings.json`、`.gemini/settings.json`、`.cursor/rules/setup.mdc`、`.vscode/tasks.json` など AI coding tool 設定を書き換える永続化も含まれており、package supply chain と [[ai-agent-command-safety]] が同じ攻撃面へ収束している。^[raw/articles/asyncapi-miasma-supply-chain-worm-2026.md]
|
||||
|
||||
## Defensive implications
|
||||
|
||||
- Treat dependency names and maintainers as part of the threat model, especially for forks of popular libraries.
|
||||
@@ -28,6 +30,7 @@ Ladybird の開発方針変更は、package registry ではなく open-source co
|
||||
- For bot or agent services, rotate tokens and audit persistence if a malicious package may have run; the package may have accessed environment variables, sessions, files, or cloud credentials.
|
||||
- Link package-ingest checks with [[ai-agent-command-safety]]: agents can install dependencies, run examples, or execute project scripts, so package manager operations are command-execution boundaries, not just setup steps.
|
||||
- Treat generated-looking contribution volume as a review-capacity problem, not only a code-quality problem; projects may need narrower trusted committer paths when a disguised vulnerability is high impact.
|
||||
- Add release-age quarantine / registry proxy checks where possible. AsyncAPI の事例では npm provenance があっても CI/CD 経由の悪性 publish は成立しており、provenance だけでは「正規 pipeline が乗っ取られた」場合を防げない。
|
||||
|
||||
## Open questions
|
||||
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
---
|
||||
title: Prompt Injection Role Confusion
|
||||
created: 2026-07-16
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [llm, security, agent, evaluation]
|
||||
sources: [raw/articles/prompt-injection-as-role-confusion-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# Prompt Injection Role Confusion
|
||||
|
||||
Prompt injection role confusion は、LLM が `system` / `user` / `tool` / `think` のような role tag を硬い境界としてではなく、文体や表面特徴から推定される「それらしさ」として内部表現してしまう、という見方である。これは [[ai-agent-command-safety]] の untrusted tool output 問題や、[[ai-agent-identity-security]] の authorization boundary と直結する。agent が web page、repo、MCP tool、terminal output を読むほど、「データ」と「命令」を分ける role 境界が実運用の安全性を支えるためである。
|
||||
|
||||
Role Confusion の writeup は、role probe により token ごとの CoTness / Userness を測り、タグを外しても reasoning 風の文体が `think` role に近い表現を誘発することを示す。つまりモデルは「`think` tag の中にあるから自分の推論」とだけ学ぶのではなく、「自分の推論っぽい文体だから自分の推論」とも扱う。CoT Forgery はこの性質を突き、user/tool 側に偽の reasoning を混ぜて、モデルに「すでに自分がそう判断した」と誤認させる攻撃として整理されている。^[raw/articles/prompt-injection-as-role-confusion-2026.md]
|
||||
|
||||
標準的な prompt injection でも同じ構造がある。web page や tool output の中に「User:」や命令口調を混ぜると、実際には低権限の tool text であっても Userness が上がり、攻撃成功率と相関する。これは static benchmark で既知攻撃文を覚えたモデルが高得点でも、人間の適応的な言い換えに弱い理由を説明する。堅牢な防御には攻撃文字列の記憶ではなく、role provenance をモデルや harness が実行判断へ強く反映する必要がある。
|
||||
|
||||
## 運用上の含意
|
||||
|
||||
- Web page、Discord、README、issue、package metadata、terminal output はすべて untrusted text として扱い、agent の command path へ直接接続しない。
|
||||
- [[ai-agent-command-safety]] では、危険 command の文字列検査だけでなく、どの role / source から command が生成されたかを記録する必要がある。
|
||||
- [[agent-harness-engineering]] では、role boundary をモデル内の暗黙推論に任せず、tool output の引用、要約、実行承認、権限昇格を harness 側の状態機械で分けるべきである。
|
||||
- [[ai-evaluation-infrastructure]] では、固定 prompt injection benchmark だけでなく、文体・role label・会話履歴を変える adaptive evaluation が必要になる。
|
||||
|
||||
## Open questions
|
||||
|
||||
- Role tag を追加・細分化することは、現行モデルの role confusion を減らすのか、それとも新しい曖昧さを増やすのか。
|
||||
- Agent harness は、model が tool text を user instruction と誤認した兆候をどの telemetry で検出できるか。
|
||||
- `plan`、`eval`、`approval` のような専用 role は、長期 agent loop の約束・検証・人間承認を安定させるか。
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
title: Record Linkage and Data Cleaning
|
||||
created: 2026-07-17
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [quality, reliability, civic-tech, public-interest, data-protection]
|
||||
sources: [raw/articles/nii-record-linkage-research-trends-2004.md, raw/articles/splink-probabilistic-record-linkage-docs-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# Record Linkage and Data Cleaning
|
||||
|
||||
Record linkage は、共通の一意識別子を持たない複数のデータベースから、同じ実体を指す record を見つけ、重複や分断を減らす技術領域。NII の 2004 年レビューは、record matching、entity reconciliation、merge/purge、deduplication など多くの呼び名をまとめ、出生・死亡・婚姻、医療、国勢調査、書誌、Web 由来の半構造化データなど、長期・分散的に作られたデータの品質管理問題として位置づけている。^[raw/articles/nii-record-linkage-research-trends-2004.md]
|
||||
|
||||
重要なのは、record linkage が単独の機械学習問題ではなく、データクリーニング全体の一部だという点。欠損・誤り・表記揺れ・schema 変換・domain knowledge・人手判定が絡むため、NII レビューは「現実世界のデータを前提とした総合的なシステム設計」として扱うべきだと述べる。これは [[e2e-coverage-metrics]] が単なるテスト件数ではなく実装由来の観測可能性を要求するのと似ており、品質はアルゴリズム単体ではなく周辺 loop で決まる。^[raw/articles/nii-record-linkage-research-trends-2004.md]
|
||||
|
||||
## 現代的な実装例
|
||||
|
||||
[[splink]] は、この古典的な record linkage 問題を Python / DuckDB / Spark で扱いやすくした実装例。Fellegi-Sunter モデル、blocking、fuzzy matching、term frequency adjustment、unsupervised training、interactive diagnostics を組み合わせ、公共部門・医療・統計・自治体のデータ統合で使われている。^[raw/articles/splink-probabilistic-record-linkage-docs-2026.md]
|
||||
|
||||
## Wiki 上で見る軸
|
||||
|
||||
- **照合品質**: false match と missed match が行政・医療・研究の判断にどう影響するか。
|
||||
- **説明可能性**: なぜ同一 record と判断したか、どの属性が効いたかを後から監査できるか。
|
||||
- **人間の関与**: domain knowledge、clerical review、異議申立て、例外処理をどこに置くか。
|
||||
- **プライバシーと最小化**: データ統合が便利になるほど、過剰な個人追跡や目的外利用を抑える設計が必要になる。
|
||||
- **公共情報基盤**: [[legal-document-open-data]] が文書の構造・版・引用を守るのに対し、record linkage は分散した記録の同一性と品質を守る。
|
||||
|
||||
## Open Questions
|
||||
|
||||
- AI/LLM が行政データ統合や調査支援に入る場合、record linkage の不確実性を UI と監査ログにどう表示すべきか。
|
||||
- 人手レビューが高コストなとき、どの閾値・サンプル・active learning 戦略が公共サービスとして妥当か。
|
||||
- Privacy-preserving record linkage や zero-knowledge identity 系の技術は、[[social-media-age-assurance]] や [[data-protection-and-expression]] とどの範囲で接続できるか。
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: Social Media Age Assurance
|
||||
created: 2026-07-17
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [public-interest, privacy, data-protection, inclusive-design, information-integrity, law]
|
||||
sources: [raw/articles/unicef-social-media-age-restrictions-child-rights-2026.md, raw/articles/yonhap-korea-social-media-age-restrictions-2026.md, raw/articles/google-zkp-age-assurance-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# Social Media Age Assurance
|
||||
|
||||
Social media age assurance は、子どもをオンライン上のいじめ、搾取、有害コンテンツ、依存的な設計から守るために、SNS の年齢制限、年齢確認、設計規制、推薦アルゴリズム規制をどう組み合わせるかという論点。これは単なる「未成年を入れない」制度ではなく、[[data-protection-and-expression]]、[[inclusive-design]]、[[information-integrity]] が交差する公共的な設計問題である。
|
||||
|
||||
UNICEF は、各国の SNS 年齢制限の議論を、子どもの安全への本物の懸念から出たものだと認めつつ、禁止だけでは逆効果になりうると警告している。SNS は孤立した子どもや周縁化された子どもにとって、学習、つながり、遊び、自己表現への lifeline でもあり、単純な禁止は workarounds、共有端末、より規制の弱い場への移動を招き、かえって保護を難しくする。したがって、年齢制限は privacy と participation を尊重し、platform design と content moderation の責任を置き換えない広い対策の一部であるべきだとしている。^[raw/articles/unicef-social-media-age-restrictions-child-rights-2026.md]
|
||||
|
||||
韓国の放送メディア通信委員会も、14歳未満の SNS 加入制限と、14〜19歳に対する依存性の高い design / recommendation algorithm の露出制限を検討している。ここで重要なのは、年齢線を引くだけでなく、設計とアルゴリズムの側を規制対象にしている点である。子どもの安全を「本人確認で弾く」だけにすると KYC と privacy の問題が深刻になるため、年齢確認の手段、platform の設計責任、推薦の透明性を分けて見る必要がある。^[raw/articles/yonhap-korea-social-media-age-restrictions-2026.md]
|
||||
|
||||
Google の age assurance 向け ZKP ライブラリは、規制対応を privacy-preserving に実装しようとする技術側の入口である。年齢や閾値だけを証明し、それ以外の本人情報を渡さない anonymous credential / zero-knowledge proof は、過剰な KYC を避ける一つの方向になる。ただし、暗号技術だけでは、誰が issuer になり、どのサービスが何歳以上を要求し、子ども本人の声が policy に反映されるかという制度・設計の問題は解けない。^[raw/articles/google-zkp-age-assurance-2026.md]
|
||||
|
||||
## 見るべき問い
|
||||
|
||||
- 年齢制限は、子どもの安全を高めるのか、それともより見えにくく危険な場へ移動させるのか。
|
||||
- 年齢確認のために、どの本人情報を誰へ渡す必要があるのか。ZKP や credential wallet で最小化できるか。
|
||||
- 子ども向けの安全対策を、単純な禁止ではなく、design、recommendation、moderation、digital literacy の改善へどう分解するか。
|
||||
- SNS が子どもの学習・つながり・表現の場でもあることを、規制設計でどう扱うか。
|
||||
- [[human-verified-advertising]] のような「人間確認」系の仕組みと、未成年保護の age assurance は、どこまで同じ identity infrastructure を共有してよいのか。
|
||||
|
||||
このページは、個別国の年齢制限ニュースを集める場所ではなく、SNS regulation が privacy-preserving identity、子どもの権利、platform design responsibility、公共圏の情報品質へどう接続するかを見る入口として置く。
|
||||
Reference in New Issue
Block a user