--- title: Prompt Injection Role Confusion created: 2026-07-16 updated: 2026-07-16 type: concept tags: [llm, security, agent, evaluation] sources: [raw/articles/prompt-injection-as-role-confusion-2026.md] confidence: medium --- # Prompt Injection Role Confusion Prompt injection role confusion は、LLM が `system` / `user` / `tool` / `think` のような role tag を硬い境界としてではなく、文体や表面特徴から推定される「それらしさ」として内部表現してしまう、という見方である。これは [[ai-agent-command-safety]] の untrusted tool output 問題や、[[ai-agent-identity-security]] の authorization boundary と直結する。agent が web page、repo、MCP tool、terminal output を読むほど、「データ」と「命令」を分ける role 境界が実運用の安全性を支えるためである。 Role Confusion の writeup は、role probe により token ごとの CoTness / Userness を測り、タグを外しても reasoning 風の文体が `think` role に近い表現を誘発することを示す。つまりモデルは「`think` tag の中にあるから自分の推論」とだけ学ぶのではなく、「自分の推論っぽい文体だから自分の推論」とも扱う。CoT Forgery はこの性質を突き、user/tool 側に偽の reasoning を混ぜて、モデルに「すでに自分がそう判断した」と誤認させる攻撃として整理されている。^[raw/articles/prompt-injection-as-role-confusion-2026.md] 標準的な prompt injection でも同じ構造がある。web page や tool output の中に「User:」や命令口調を混ぜると、実際には低権限の tool text であっても Userness が上がり、攻撃成功率と相関する。これは static benchmark で既知攻撃文を覚えたモデルが高得点でも、人間の適応的な言い換えに弱い理由を説明する。堅牢な防御には攻撃文字列の記憶ではなく、role provenance をモデルや harness が実行判断へ強く反映する必要がある。 ## 運用上の含意 - Web page、Discord、README、issue、package metadata、terminal output はすべて untrusted text として扱い、agent の command path へ直接接続しない。 - [[ai-agent-command-safety]] では、危険 command の文字列検査だけでなく、どの role / source から command が生成されたかを記録する必要がある。 - [[agent-harness-engineering]] では、role boundary をモデル内の暗黙推論に任せず、tool output の引用、要約、実行承認、権限昇格を harness 側の状態機械で分けるべきである。 - [[ai-evaluation-infrastructure]] では、固定 prompt injection benchmark だけでなく、文体・role label・会話履歴を変える adaptive evaluation が必要になる。 ## Open questions - Role tag を追加・細分化することは、現行モデルの role confusion を減らすのか、それとも新しい曖昧さを増やすのか。 - Agent harness は、model が tool text を user instruction と誤認した兆候をどの telemetry で検出できるか。 - `plan`、`eval`、`approval` のような専用 role は、長期 agent loop の約束・検証・人間承認を安定させるか。