This commit is contained in:
2026-07-03 00:38:05 +09:00
parent 87eacd39b2
commit 86cad348b4
149 changed files with 24450 additions and 105 deletions
+28
View File
@@ -0,0 +1,28 @@
---
title: Claude Science
created: 2026-06-30
updated: 2026-06-30
type: entity
tags: [llm, agent, automation, tool, evaluation]
sources: [raw/articles/claude-science-ai-workbench-2026.md]
confidence: medium
---
# Claude Science
Claude Science is Anthropic's beta AI workbench for scientific research. It packages Claude as a coordinating agent inside a research environment that connects to scientific databases, packages, local or remote compute, and domain-specific skills. Unlike a general chat assistant, the product emphasizes auditable artifacts: figures, manuscripts, code, environment details, and message history are kept together so results can be validated and reproduced later.^[raw/articles/claude-science-ai-workbench-2026.md]
The product is especially relevant to [[ai-research-automation]] because it turns research work into a persistent, tool-connected loop rather than a one-off answer. Claude Science can run locally on macOS/Linux or against remote machines over SSH/HPC login nodes, ask before reaching new resources, submit jobs, and fork sessions to compare approaches. A reviewer agent checks citations, calculations, untraceable numbers, and figure/code consistency, which connects the product to [[ai-evaluation-infrastructure]] and [[loop-engineering]].
## Design signals
- **Auditable artifacts**: outputs include the code and environment that produced them, plus plain-language explanations and message history.
- **Compute as part of the loop**: local machines, lab infrastructure, HPC, and Modal-style on-demand compute are treated as execution targets behind explicit review/revoke decisions.
- **Domain skills and connectors**: the beta ships with 60+ curated skills/connectors for areas such as genomics, single-cell, proteomics, structural biology, and cheminformatics.
- **Reviewer agents**: separate critic/reviewer agents are used to catch citation and calculation errors, echoing the generator/evaluator split in [[loop-engineering]].
## Open Questions
- How much of the reproducibility guarantee depends on preserving exact execution environments versus preserving narrative provenance?
- Can the reviewer-agent pattern transfer to personal wiki curation, code review, and scheduled research jobs without becoming too expensive?
- What privacy boundary is acceptable when sensitive lab data remains local but context is still sent to a hosted model?
+29
View File
@@ -0,0 +1,29 @@
---
title: Google Agent Development Kit
created: 2026-06-30
updated: 2026-06-30
type: entity
tags: [tool, agent, dev-tool, workflow]
sources: [raw/articles/google-adk-go-2-0-agent-workflows-2026.md]
confidence: medium
---
# Google Agent Development Kit
Google Agent Development Kit(ADK)は、Go や Python で agent application を実装するための code-first framework。Go 2.0 の発表では、単一の LLM 呼び出しではなく、分類、分岐、fan-out / fan-in、人間の承認、retry、pause/resume を含む実運用向けの workflow を、graph of nodes として表現する方向が強調された。
[[loop-engineering]] の観点では、ADK Go 2.0 は「graph がそのまま agent」として同じ runner / launcher / console で動く点が重要である。workflow の状態を session に残し、human-in-the-loop の interrupt を後続ターンや process restart 後に再開できるため、agent loop を一回きりの chat ではなく、状態を持つ実行単位として扱いやすい。
## 重要な設計要素
- **Graph-based workflow engine**: function node、agent node、tool node、join node、dynamic node、sub-workflow、parallel worker を edge と route でつなぎ、sequence、conditional routing、parallel fan-out/fan-in、loop を構成する。
- **Dynamic orchestration in Go**: 実行順序が runtime data や model の判断で変わる場合は、ordinary Go code から child node を `RunNode` する dynamic node で表現できる。これは [[agent-oriented-cli-design]] と同じく、エージェント向けの制御面を暗黙の prompt ではなく実装可能な interface に落とす方向である。
- **Human-in-the-loop as primitive**: 任意の node が `RequestInput` event で人間に承認・修正・追加情報を求め、handoff または re-entry で workflow を再開できる。
- **Resilience and observability**: node ごとの retry policy、timeout、graph-wide concurrency limit、branch history isolation、telemetry span tree が、agent workflow を観測・再実行しやすい単位にする。
- **Unified runtime**: plain LLM agent と full graph が同じ node runtime に寄るため、単体 agent、sub-agent、workflow の境界が薄くなる。これは [[kiro]] のような agent-native work surface や、[[ai-agent-identity-security]] の承認・監査境界とも接続する。
## 見るべき問い
- Durable resume と session history reconstruction は、個人用の scheduled job や Discord link ingest のような小さな [[wiki-maintenance-loop]] にも取り込めるか。
- Go の型付き node / event stream は、agent workflow の検証や replay をどこまで容易にするか。
- Human-in-the-loop が framework primitive になるほど、承認 UI、audit log、権限境界をどの層で標準化すべきか。
+3 -3
View File
@@ -1,10 +1,10 @@
---
title: Litho
created: 2026-06-30
updated: 2026-06-30
updated: 2026-07-01
type: entity
tags: [tool, wiki, knowledge-base, markdown, dev-tool, automation]
sources: [raw/articles/litho-deepwiki-rs-code-documentation-2026.md]
sources: [raw/articles/litho-deepwiki-rs-code-documentation-2026.md, raw/articles/langchain-openwiki-repo-documentation-agent-2026.md]
confidence: medium
---
@@ -12,7 +12,7 @@ confidence: medium
Litho(`deepwiki-rs`)は、ソースコードから C4 model のアーキテクチャ図と Wiki 風ドキュメントを自動生成する Rust 製の AI 文書化エンジン。README では、コードベース解析、依存関係・構造抽出、LLM によるドキュメント生成、Mermaid 図、CI/CD 連携、外部知識の取り込みを特徴としている。
Yuta の Wiki 文脈では、これは [[digital-gardening-cms]] や [[llm-wiki-pattern]] と同じ「読むたびに都度検索する」のではなく、「コードから持続的に読める知識面を作る」方向の道具。ただし Litho は人間の研究 Wiki というより、コードベース理解・オンボーディング・設計書の鮮度維持に寄っている。
Yuta の Wiki 文脈では、これは [[digital-gardening-cms]] や [[llm-wiki-pattern]] と同じ「読むたびに都度検索する」のではなく、「コードから持続的に読める知識面を作る」方向の道具。ただし Litho は人間の研究 Wiki というより、コードベース理解・オンボーディング・設計書の鮮度維持に寄っている。同じ方向の [[openwiki]] は、C4 図よりも agent instruction file から参照される repo Wiki と scheduled update に重点を置く。
## Design implications
+30
View File
@@ -0,0 +1,30 @@
---
title: Mastra
created: 2026-07-01
updated: 2026-07-01
type: entity
tags: [tool, agent, dev-tool, workflow]
sources: [raw/articles/mastra-typescript-agent-framework-2026.md]
confidence: medium
---
# Mastra
Mastra は、TypeScript で AI agent、workflow、MCP server、AI application を作るための framework。README は「prototype から production-ready application まで」を対象にし、React、Next.js、Node.js への組み込み、または standalone server としての配備を想定している。
[[google-adk]] が Go/Python の graph workflow と session resume を強調するのに対し、Mastra は TypeScript / web application 側の既存 stack に寄せて、model routing、agent、workflow、memory、MCP、eval、observability を一つの開発面にまとめる。[[loop-engineering]] の観点では、agent を一回の chat ではなく、状態を持つ workflow、HITL pause/resume、観測・評価される実行単位として扱うための framework と読める。
## 設計要素
- **Model routing**: OpenAI、Anthropic、Gemini など 40+ provider を標準 interface で扱う。provider 切替を application logic から分離する点は、agent 運用の cost / availability control に関わる。
- **Agents and tools**: agent は goal に対して tool を選び、final answer または停止条件まで内部反復する。これは [[agent-harness-engineering]] の「tool boundary と停止条件を harness 側で設計する」論点と接続する。
- **Graph workflows**: `.then()`、`.branch()`、`.parallel()` のような syntax で多段処理を明示的に組み、必要に応じて agent より deterministic な制御面を持てる。
- **Human-in-the-loop**: workflow / agent を suspend し、storage に実行状態を残して、承認や入力を待ってから再開できる。
- **MCP servers**: agent、tool、structured resource を Model Context Protocol server として公開できる。これは [[ai-agent-identity-security]] の認可・監査・最小権限の設計対象にもなる。
- **Evals and observability**: built-in evals と observability を production essentials として掲げる。model 能力だけでなく、実行 trace と評価を改善 loop に入れる点で [[ai-evaluation-infrastructure]] と近い。
## 見るべき問い
- TypeScript application に自然に組み込める一方で、workflow state、MCP exposure、model routing credentials の権限境界をどこで監査するか。
- Google ADK のような typed graph runtime と比べ、Mastra の web/dev UX は Yuta の既存 automation loop にどのくらい低摩擦で入るか。
- Built-in eval / observability が、個人用 scheduled job や wiki ingest のような小さな loop にも過剰でなく使えるか。
+26
View File
@@ -0,0 +1,26 @@
---
title: Notion
created: 2026-07-01
updated: 2026-07-01
type: entity
tags: [tool, agent, automation, knowledge-base, workflow]
sources: [raw/articles/notion-developer-platform-agents-workers-2026.md]
confidence: medium
---
# Notion
Notion は、文書、データベース、ワークフロー、チーム知識を一つの workspace に集める知識作業ツール。2026 年の Developer Platform 発表では、Notion を人間のドキュメント UI だけでなく、coding agent や外部 agent が読む・書く・実行する共有 canvas として位置づけている。これは [[digital-gardening-cms]] の「ページとリンクを育てる場所」が、[[agent-harness-engineering]] の実行面・承認面に近づく例である。^[raw/articles/notion-developer-platform-agents-workers-2026.md]
## Developer Platform の要点
- **External Agents API**: Claude、Codex、Decagon、自作 agent などを Notion に持ち込み、チケットから coding agent を呼び、チーム承認へつなぐ orchestration layer として使う構想。
- **Workers**: Notion 側の hosted runtime で custom code を動かし、database sync、webhook、agent tool、外部 API 操作を deterministic な処理として実装する。LLM reasoning だけに任せず、[[loop-engineering]] の一部を通常のコードに分離する設計として読める。
- **CLI**: Notion に sign in し、読み書きし、Workers を build/deploy するための CLI を提供する。coding agent が使うことを明示しており、[[agent-oriented-cli-design]] の対象になる。
- **Agent SDK / MCP / Markdown API**: Notion Agent を他アプリへ埋め込む計画、Notion MCP の token 効率化、Markdown API などにより、workspace の知識を agent から扱いやすくしようとしている。
## 読みどころ
Notion の発表は、AI agent の作業場所を IDE やチャットから「業務データがある workspace」へ広げる動きとして重要である。agent がチケット、顧客情報、会議メモ、文書を同じ場所で参照・更新できると便利だが、同時に workspace-scoped OAuth、personal access token、内部 connection の管理、誰がどの agent に何を書かせたかの監査が必要になる。この論点は [[ai-agent-identity-security]] と直結する。
一方で、この raw source は公式 release page であり、細かな API 仕様や sandbox boundary は別 documentation を読む必要がある。現時点では「Notion が agent/workflow substrate になろうとしている」方向性の記録として扱う。
+26
View File
@@ -0,0 +1,26 @@
---
title: OpenWiki
created: 2026-07-01
updated: 2026-07-01
type: entity
tags: [tool, wiki, knowledge-base, markdown, agent, dev-tool, automation]
sources: [raw/articles/langchain-openwiki-repo-documentation-agent-2026.md]
confidence: medium
---
# OpenWiki
OpenWiki は LangChain が公開した、コードベース向けの文書生成・保守 CLI / agent。リポジトリ内に `openwiki/` を作り、コードの構造、主要ロジック、ファイル間の関係、慣習を agent が参照しやすい Wiki として残す。[[litho]] と同じく「コードから持続的な知識面を作る」道具だが、OpenWiki は C4 図よりも、coding agent が必要な文脈を巨大な instruction file に詰め込まず発見できるようにする点を前面に出している。
## Design implications
- `AGENTS.md` や `CLAUDE.md` に Wiki 全体を貼るのではなく、生成済み Wiki への参照と使いどころを追加する。これは [[agent-harness-engineering]] の repo-local instruction 設計に近く、instruction file を「すべての知識」ではなく「探し方の入口」として扱う。
- `openwiki --init` で初期文書を作り、`openwiki --update` と GitHub Actions の定期実行で差分を読み、既存 Wiki を更新する。これは [[wiki-maintenance-loop]] をコードベース文書へ寄せた形で、docs drift を手作業ではなく scheduled loop で抑えようとする。
- OpenRouter、Fireworks、Baseten、OpenAI、Anthropic など複数 provider を扱い、DeepAgents と LangSmith tracing を使える。生成結果だけでなく、文書生成 agent が何をしたかを trace できる点は [[loop-engineering]] の observability / persistence と接続する。
- Karpathy の [[llm-wiki-pattern]]、DeepWiki、AutoWiki への明示的な参照があり、coding agent 用の repo wiki は個人研究 Wiki とは別用途ながら、巨大 context を毎回読み込むより「保守された Markdown 知識面を使う」という同じ設計方向にある。
## Open questions
- 生成された repo Wiki の差分を、人間 reviewer がどの粒度で見るべきか。
- agent が Wiki の古い記述を信じて誤った修正をしたとき、どの evaluator / test / trace が検出するか。
- [[digital-gardening-cms]] 的な人間向け文書と、OpenWiki 的な agent 向け文書を同じリポジトリでどう分けるか。
+33
View File
@@ -0,0 +1,33 @@
---
title: Safari MCP Server
created: 2026-07-02
updated: 2026-07-02
type: entity
tags: [agent, dev-tool, workflow, quality, accessibility]
sources: [raw/articles/safari-mcp-server-webkit-2026.md]
confidence: medium
---
# Safari MCP Server
Safari MCP Server は、Safari Technology Preview 247 で導入された、Web 開発者向けの Model Context Protocol server。`safaridriver --mcp` として動き、Claude、Codex などの MCP 対応 agent から Safari の実ブラウザ window を操作・観測できるようにする。[[agent-harness-engineering]] の観点では、agent に DOM、network request、console output、screenshot、viewport、dialog、tab、page content などを渡し、ブラウザ上の実行結果を見ながら debugging させる browser harness である。^[raw/articles/safari-mcp-server-webkit-2026.md]
## 何ができるか
WebKit の記事は、Safari MCP Server を「Browser → Prompt → Agent」の往復を減らす道具として説明している。agent は Safari でページを開き、computed style や layout を調べ、navigation timing や resource load time を見て性能原因を探し、missing label、ARIA、contrast などのアクセシビリティ問題を確認し、form state や checkout flow のようなユーザー状態を検証できる。これは [[e2e-coverage-metrics]] のような実行証跡ベースの品質確認や、[[inclusive-design]] のアクセシビリティ運用に近い。^[raw/articles/safari-mcp-server-webkit-2026.md]
## Tool surface
公開されている tool は、`browser_console_messages`、`list_network_requests`、`get_network_request`、`evaluate_javascript`、`get_page_content`、`screenshot`、`page_interactions`、`set_viewport_size`、`set_emulated_media`、`list_tabs`、`create_tab`、`switch_tab`、`close_tab`、`navigate_to_url`、`wait_for_navigation` など。[[agent-oriented-cli-design]] と同じく、agent が推測ではなく構造化された tool call で状態を読む点が重要になる。^[raw/articles/safari-mcp-server-webkit-2026.md]
## 安全境界
記事は、Safari MCP Server 自体は local machine 上で動き、自身では network call を行わず、AutoFill などの個人情報や他の Safari 活動へアクセスしないと説明している。一方で、page content、screenshot、console log は接続先 agent へ渡るため、その後の扱いは agent/model 側の責任になる。つまり便利さは [[ai-agent-identity-security]] の browser / device permission boundary と一体であり、どの site、どの tab、どの agent に見せるかを運用で決める必要がある。^[raw/articles/safari-mcp-server-webkit-2026.md]
## 関連
- [[agent-harness-engineering]]
- [[agent-oriented-cli-design]]
- [[ai-agent-identity-security]]
- [[e2e-coverage-metrics]]
- [[inclusive-design]]
+21
View File
@@ -0,0 +1,21 @@
---
title: Strix
created: 2026-07-02
updated: 2026-07-02
type: entity
tags: [tool, agent, security, dev-tool]
sources: [raw/articles/strix-ai-pentesting-agent-2026.md]
confidence: medium
---
# Strix
Strix は、アプリケーションを実行しながら脆弱性を探し、PoC で検証し、修正案や pentest report まで返すことを目指す open-source の AI penetration testing tool。README は、reconnaissance、exploitation、validation を multi-agent orchestration で行い、GitHub Actions / CI/CD に入れて pull request ごとに検査できると説明している。静的解析だけではなく「実際に攻撃を試して成立性を確認する」方向を前面に出している点が、[[ci-cd-runtime-security]] や [[ai-agent-enabled-cyberattacks]] と接続する。^[raw/articles/strix-ai-pentesting-agent-2026.md]
Yuta の関心では、Strix は単なる security scanner というより、攻撃側も防御側も agent loop を使う時代の防御 harness として読むべき。[[scrutineer]] が OSS 脆弱性の発見・検証・開示を人間 gate で抑える workflow なら、Strix はアプリケーションに対する探索・PoC・修正を CI や開発者 CLI へ寄せる。自律 pentest agent を導入する場合は、検査対象・ネットワーク範囲・credential・報告先を明確にし、[[ai-agent-command-safety]] と同じく agent が何を実行できるかを制約する必要がある。
## Watch points
- CI/CD で本当に安全に使うには、target sandbox、network egress、test data、secret exposure、false-positive handling の設計が必要。
- 「real exploit validation」は有用だが、検証 payload が本番・共有環境・第三者サービスへ波及しない boundary が重要。
- Auto-fix や report generation は、[[agent-harness-engineering]] の human-review output と同じく、人間が理解・差し戻しできる形に制約されているかを見る。