3.2 KiB
title, created, updated, type, tags, sources, confidence
| title | created | updated | type | tags | sources | confidence | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Prompt Injection Role Confusion | 2026-07-16 | 2026-07-16 | concept |
|
|
medium |
Prompt Injection Role Confusion
Prompt injection role confusion は、LLM が system / user / tool / think のような role tag を硬い境界としてではなく、文体や表面特徴から推定される「それらしさ」として内部表現してしまう、という見方である。これは ai-agent-command-safety の untrusted tool output 問題や、ai-agent-identity-security の authorization boundary と直結する。agent が web page、repo、MCP tool、terminal output を読むほど、「データ」と「命令」を分ける role 境界が実運用の安全性を支えるためである。
Role Confusion の writeup は、role probe により token ごとの CoTness / Userness を測り、タグを外しても reasoning 風の文体が think role に近い表現を誘発することを示す。つまりモデルは「think tag の中にあるから自分の推論」とだけ学ぶのではなく、「自分の推論っぽい文体だから自分の推論」とも扱う。CoT Forgery はこの性質を突き、user/tool 側に偽の reasoning を混ぜて、モデルに「すでに自分がそう判断した」と誤認させる攻撃として整理されている。^[raw/articles/prompt-injection-as-role-confusion-2026.md]
標準的な prompt injection でも同じ構造がある。web page や tool output の中に「User:」や命令口調を混ぜると、実際には低権限の tool text であっても Userness が上がり、攻撃成功率と相関する。これは static benchmark で既知攻撃文を覚えたモデルが高得点でも、人間の適応的な言い換えに弱い理由を説明する。堅牢な防御には攻撃文字列の記憶ではなく、role provenance をモデルや harness が実行判断へ強く反映する必要がある。
運用上の含意
- Web page、Discord、README、issue、package metadata、terminal output はすべて untrusted text として扱い、agent の command path へ直接接続しない。
- ai-agent-command-safety では、危険 command の文字列検査だけでなく、どの role / source から command が生成されたかを記録する必要がある。
- agent-harness-engineering では、role boundary をモデル内の暗黙推論に任せず、tool output の引用、要約、実行承認、権限昇格を harness 側の状態機械で分けるべきである。
- ai-evaluation-infrastructure では、固定 prompt injection benchmark だけでなく、文体・role label・会話履歴を変える adaptive evaluation が必要になる。
Open questions
- Role tag を追加・細分化することは、現行モデルの role confusion を減らすのか、それとも新しい曖昧さを増やすのか。
- Agent harness は、model が tool text を user instruction と誤認した兆候をどの telemetry で検出できるか。
plan、eval、approvalのような専用 role は、長期 agent loop の約束・検証・人間承認を安定させるか。