add
This commit is contained in:
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: AI Agent Command Safety
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-01
|
||||
updated: 2026-07-16
|
||||
type: concept
|
||||
tags: [agent, security, reliability, automation]
|
||||
sources: [raw/articles/guardfall-ai-coding-agent-shell-bypass-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/cursor-duneslide-sandbox-escape-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/koi-promptjacking-claude-desktop-rce-2026.md]
|
||||
sources: [raw/articles/guardfall-ai-coding-agent-shell-bypass-2026.md, raw/articles/openai-codex-agent-approvals-security-2026.md, raw/articles/cursor-duneslide-sandbox-escape-2026.md, raw/articles/theregister-claude-desktop-double-agent-2026.md, raw/articles/koi-promptjacking-claude-desktop-rce-2026.md, raw/articles/prompt-injection-as-role-confusion-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -25,6 +25,8 @@ Cursor の DuneSlide 事例は、sandbox があるだけでは足りず、「age
|
||||
|
||||
Claude Desktop まわりの事例は、command safety が「生成された shell 文字列」だけではなく、設定同期、MCP/extension、personal preferences、偽 error message まで含む広い実行経路の問題であることを示す。The Register の Pentera Labs 記事では、攻撃者が Claude の account-wide personalization に base64 prompt を入れ、Desktop Commander など command-capable MCP があれば reverse shell、なければ Anthropic 風の偽エラーと install prompt でユーザーに実行させる流れが説明されている。Koi の PromptJacking 報告では、公式 Claude Desktop extensions が unsandboxed MCP server として動き、AppleScript への未 escape URL 補間から web prompt injection → local RCE へ進みうると説明されている。どちらも「agent が command を出す瞬間」より前に、信頼済み assistant の設定・connector・外部 web content が command path へ混ざるため、設定変更監視、extension allowlist、connector sandboxing が command guard と同じ層で必要になる。^[raw/articles/theregister-claude-desktop-double-agent-2026.md] ^[raw/articles/koi-promptjacking-claude-desktop-rce-2026.md]
|
||||
|
||||
[[prompt-injection-role-confusion]] は、この問題の model-internal な説明を与える。Role Confusion の writeup は、tool text 内の命令が user 風・reasoning 風に見えると、実際の role tag より文体特徴が勝ち、モデルが低権限 text を命令や自分の推論として扱いやすくなると示す。したがって command safety では、model に「これは tool output だから無視して」と頼むだけでなく、untrusted text から shell / MCP / file write へ進む経路を harness 側で分離し、role provenance を承認 UI と実行ログに残す必要がある。^[raw/articles/prompt-injection-as-role-confusion-2026.md]
|
||||
|
||||
## なぜ重要か
|
||||
|
||||
Yuta の運用では、Hermes の scheduled job、Codex/Claude/OpenCode、local CLI、CI runner が同じ「agent が command を出す」面を共有する。便利な自走 loop ほど、guard を抜けた command が SSH key、cloud credential、wiki、repo、home directory へ届きやすい。したがって agent command safety は、個別 agent の機能ではなく、[[agent-oriented-cli-design]]、[[ai-agent-identity-security]]、[[ci-cd-runtime-security]] を横断する運用品質の条件として扱うべきである。
|
||||
|
||||
Reference in New Issue
Block a user