From bd14321968db7a7e14ff75783d3fec694b1242ab Mon Sep 17 00:00:00 2001 From: yuta Date: Sun, 28 Jun 2026 21:18:12 +0900 Subject: [PATCH] add --- .../discord-link-ingest/interest-profile.md | 60 ++ .automation/discord-link-ingest/state.md | 47 ++ concepts/llm-wiki-pattern.md | 9 +- concepts/wiki-maintenance-loop.md | 8 +- entities/llm-wiki-app.md | 23 + index.md | 3 +- log.md | 16 + raw/articles/academic-research-skills-2026.md | 703 ++++++++++++++++++ ...chinatalk-transfer-station-economy-2026.md | 126 ++++ ...hub-automated-my-job-better-leader-2026.md | 135 ++++ .../hermes-research-llm-wiki-skill-2026.md | 468 ++++++++++++ raw/articles/howm-2026.md | 89 +++ raw/articles/nashsu-llm-wiki-2026.md | 493 ++++++++++++ ...oring-english-effective-design-doc-2026.md | 487 ++++++++++++ raw/articles/rem-cli-macos-reminders-2026.md | 386 ++++++++++ 15 files changed, 3050 insertions(+), 3 deletions(-) create mode 100644 .automation/discord-link-ingest/interest-profile.md create mode 100644 .automation/discord-link-ingest/state.md create mode 100644 entities/llm-wiki-app.md create mode 100644 raw/articles/academic-research-skills-2026.md create mode 100644 raw/articles/chinatalk-transfer-station-economy-2026.md create mode 100644 raw/articles/github-automated-my-job-better-leader-2026.md create mode 100644 raw/articles/hermes-research-llm-wiki-skill-2026.md create mode 100644 raw/articles/howm-2026.md create mode 100644 raw/articles/nashsu-llm-wiki-2026.md create mode 100644 raw/articles/refactoring-english-effective-design-doc-2026.md create mode 100644 raw/articles/rem-cli-macos-reminders-2026.md diff --git a/.automation/discord-link-ingest/interest-profile.md b/.automation/discord-link-ingest/interest-profile.md new file mode 100644 index 0000000..faba69c --- /dev/null +++ b/.automation/discord-link-ingest/interest-profile.md @@ -0,0 +1,60 @@ +# Discord Link Ingest Interest Profile + +Updated: 2026-06-28 +Source: discrawl read-only analysis of Yuta/toymaker Discord messages. + +## Author IDs considered +- `890593691440398456` — yuta, historical messages through 2025-03-24 +- `890908900520505354` — toymaker, current identity through 2026-06-28 + +## Evidence snapshot +- Total messages by these identities in archive: 42,677 +- URL-containing messages: 13,169 +- Top channels by message volume: `💭|yuta`, `chat`, `notes` +- Recent URL domains are heavily technical/source-oriented: `github.com`, `x.com`, `github.blog`, `zenn.dev`, `openai.com`, `blog.cloudflare.com`, `docs.litellm.ai`, `scrapbox.io`, `obsidian.md`, `hermes-agent.nousresearch.com`, `alphaxiv.org`. +- Keyword signals: AI/LLM, X/Twitter, development/code, workflow automation, wiki/knowledge management, agents/Hermes. + +## Current scoring bias +Start strict. The wiki should remain curated; raw sources are acceptable, but wiki-page upgrades should require durable reuse value. + +### Score 4 — create/update substantial wiki pages +Use for sources that are central to one of these durable themes: +- LLM Wiki / compiled knowledge bases / Obsidian / knowledge management systems +- AI agents, Hermes Agent, Codex/Claude/OpenCode-style automation, multi-agent workflows +- Developer tooling that changes Yuta's automation/dev workflow materially +- Quality engineering, security, supply-chain, infra reliability for AI/software systems +- Technical writeups with implementation details likely to be referenced later +- Niche, exciting design/hack/Hacker News-like material, especially when it exposes an unusual technique, tool, interface, or way of thinking +- Public-interest/public-sector technology, civic infrastructure, accessibility (a11y), inclusive design, and systems that make services more usable or equitable +- Papers or research with clear relevance to LLMs, agents, evaluation, knowledge systems, automation, accessibility, or public-interest technology + +### Score 3 — update existing page only +Use when the source adds a concrete fact, method, comparison, or implementation note to an existing page, but does not deserve a new page. + +### Score 2 — raw/articles only +Use for interesting but not-yet-connected sources: +- good one-off articles +- news with possible future relevance +- tools/repos worth remembering but not yet central +- X/Twitter links only when they point to durable technical content or a thread with reusable insight + +### Score 1 — state/link-only +Use for URLs whose title/context is notable but body extraction fails or value is unclear. + +### Score 0 — skip +- ads/campaigns +- memes/ephemeral posts +- shallow news with no later reuse value +- duplicate URLs +- login-only pages where no useful content is extractable +- media-only links unless explicitly tied to a technical/knowledge workflow + +## Per-domain handling hints +- `github.com`: prioritize repos, READMEs, issues/PRs, commits related to automation, agents, LLMs, dev infra, security, QE. Skip personal one-off commits unless message context says they matter. +- `x.com`: do not auto-upgrade by default. Use only for high-signal technical threads or links to durable sources. +- `zenn.dev`, `github.blog`, `openai.com`, `blog.cloudflare.com`, technical docs/blogs: usually score 2+, score 3/4 if aligned with above themes. +- newspapers/general news: usually score 0-2 unless strongly connected to cyber/security/AI policy or a durable thesis. +- YouTube: usually link-only unless transcript/title/context indicates durable technical value. + +## Learning rule +After each run, if repeated keep/skip decisions reveal a stable preference, append one compact note here or to state.md. Prefer explicit evidence from Discord messages over assumptions from short Hermes conversations. diff --git a/.automation/discord-link-ingest/state.md b/.automation/discord-link-ingest/state.md new file mode 100644 index 0000000..0419f94 --- /dev/null +++ b/.automation/discord-link-ingest/state.md @@ -0,0 +1,47 @@ +# Discord Link Ingest State + +last_checked_at: 2026-06-28T11:58:52Z +last_message_created_at: 2026-06-28T07:26:20.859000000Z +lookback_used: 24h_initial_state_missing +channels: + chat: '1028287639918497822' + tw: '1477793137064935675' + +## Last run summary — 2026-06-28 + +- Messages scanned: 110 +- URL-containing messages: 103 +- Normalized unique URLs found: 613 +- Non-X/non-twitter URLs: 73 +- Raw articles saved: 8 +- Wiki pages created: 1 +- Wiki pages updated: 3 (`concepts/llm-wiki-pattern.md`, `concepts/wiki-maintenance-loop.md`, `index.md`) +- Link-only / failed extraction: Obsidian Headless help (`defuddle` returned empty), alphaxiv 2606.25331 (`defuddle` returned empty), many X/t.co/media/news links below strict threshold. + +## Raw articles saved this run + +- https://github.com/nashsu/llm_wiki → raw/articles/nashsu-llm-wiki-2026.md (score 4) +- https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-llm-wiki → raw/articles/hermes-research-llm-wiki-skill-2026.md (score 4) +- https://github.com/emacsmirror/howm → raw/articles/howm-2026.md (score 2) +- https://github.com/Imbad0202/academic-research-skills → raw/articles/academic-research-skills-2026.md (score 2) +- https://refactoringenglish.com/excerpts/write-an-effective-design-doc/ → raw/articles/refactoring-english-effective-design-doc-2026.md (score 2) +- https://github.blog/developer-skills/github/i-automated-my-job-and-it-made-me-a-better-leader/ → raw/articles/github-automated-my-job-better-leader-2026.md (score 2) +- https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in → raw/articles/chinatalk-transfer-station-economy-2026.md (score 2) +- https://github.com/BRO3886/rem → raw/articles/rem-cli-macos-reminders-2026.md (score 2) + +## Processed high-signal URLs + +- https://github.com/nashsu/llm_wiki +- https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-llm-wiki +- https://github.com/emacsmirror/howm +- https://github.com/Imbad0202/academic-research-skills +- https://refactoringenglish.com/excerpts/write-an-effective-design-doc/ +- https://github.blog/developer-skills/github/i-automated-my-job-and-it-made-me-a-better-leader/ +- https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in +- https://github.com/BRO3886/rem +- https://obsidian.md/ja/help/headless +- https://www.alphaxiv.org/abs/2606.25331 + +## Rubric note + +First run confirmed unusually dense interest in LLM Wiki / knowledge-management tooling (`llm_wiki`, Hermes bundled skill docs, howm, Obsidian Headless) from #chat. Keep scoring this cluster high, but continue strict wiki-page updates: create/update pages only for sources that directly change the LLM Wiki operating model; save adjacent workflow/security/dev-tool links as raw-only unless they connect to existing pages. diff --git a/concepts/llm-wiki-pattern.md b/concepts/llm-wiki-pattern.md index cdf1dcd..392e78d 100644 --- a/concepts/llm-wiki-pattern.md +++ b/concepts/llm-wiki-pattern.md @@ -4,7 +4,7 @@ created: 2026-06-28 updated: 2026-06-28 type: concept tags: [wiki, knowledge-base, synthesis, agent, markdown] -sources: [raw/articles/karpathy-llm-wiki-2026.md] +sources: [raw/articles/karpathy-llm-wiki-2026.md, raw/articles/nashsu-llm-wiki-2026.md, raw/articles/hermes-research-llm-wiki-skill-2026.md] confidence: medium --- @@ -16,8 +16,15 @@ LLM Wiki は、LLM が curated raw sources を読み、持続的な Markdown wik 人間の役割は source selection、問いの設計、レビュー、方向づけ。LLM の役割は要約、相互リンク、既存ページ更新、矛盾検出、index/log 更新などの bookkeeping。閲覧・探索には [[obsidian]] のような wikilink 対応エディタが合う。 +## Implementation Surface + +Karpathy の原案は抽象的な運用パターンだが、実装形態は複数ある。Hermes の `llm-wiki` skill は Markdown directory、`SCHEMA.md`、`index.md`、`log.md`、raw source frontmatter、agent の手順を組み合わせて、軽量な file-based workflow として実行する。一方、[[llm-wiki-app]] は同じ発想を desktop app、persistent ingest queue、folder watch、knowledge graph、semantic search、web clipper、local HTTP API/MCP server まで含む product として具体化している。 + +この差は「単純な Markdown vault を agent が保守する」のか、「専用アプリが ingest/search/API を統合する」のかという設計選択。小さく始めるなら skill + [[obsidian]] で十分だが、source 数が増え、画像/PDF、queue、graph、API 連携が必要になると app 型の surface が効いてくる可能性がある。 + ## Open Questions - どの粒度でページを分けると、後から検索・更新しやすいか。 - チーム運用では、LLM 更新をどこまで自動化し、どこを人間レビューにするか。 - 小規模 wiki では index.md だけで十分か、いつ CLI/MCP search が必要になるか。 +- 専用アプリ化した場合、raw source と wiki page の source of truth をどう保つか。 diff --git a/concepts/wiki-maintenance-loop.md b/concepts/wiki-maintenance-loop.md index c51f3cc..6855c09 100644 --- a/concepts/wiki-maintenance-loop.md +++ b/concepts/wiki-maintenance-loop.md @@ -4,7 +4,7 @@ created: 2026-06-28 updated: 2026-06-28 type: concept tags: [workflow, maintenance, wiki, agent] -sources: [raw/articles/karpathy-llm-wiki-2026.md] +sources: [raw/articles/karpathy-llm-wiki-2026.md, raw/articles/hermes-research-llm-wiki-skill-2026.md, raw/articles/nashsu-llm-wiki-2026.md] confidence: medium --- @@ -17,3 +17,9 @@ LLM Wiki の基本運用は ingest / query / lint のループ。 3. **Lint**: broken links、orphan pages、古い主張、矛盾、frontmatter 不備、肥大化ページを点検する。 この loop によって、[[rag-vs-compiled-wiki|query-time retrieval]] だけでは残りにくい知識整理が継続的に蓄積される。[[obsidian]] の graph view や backlinks は、orphan や hub page を見つける補助になる。 + +## Automation Notes + +Hermes の `llm-wiki` skill では、毎回 `SCHEMA.md`、`index.md`、recent `log.md` を読む orientation が必須になっている。これは自動 ingest が duplicate page creation や schema drift を起こさないための guardrail。新規 raw source には `source_url`、`ingested`、body hash を持たせ、wiki page の frontmatter、index、log を同時に更新する。 + +[[llm-wiki-app]] はこの loop をアプリ側の persistent ingest queue、folder auto-watch、source cleanup、graph/search、MCP/API に拡張している。つまり maintenance loop は単なる checklist ではなく、agent procedure と product feature のどちらにもなりうる。 diff --git a/entities/llm-wiki-app.md b/entities/llm-wiki-app.md new file mode 100644 index 0000000..e647a2d --- /dev/null +++ b/entities/llm-wiki-app.md @@ -0,0 +1,23 @@ +--- +title: LLM Wiki App +created: 2026-06-28 +updated: 2026-06-28 +type: entity +tags: [tool, wiki, knowledge-base, agent, markdown] +sources: [raw/articles/nashsu-llm-wiki-2026.md, raw/articles/hermes-research-llm-wiki-skill-2026.md] +confidence: medium +--- + +# LLM Wiki App + +LLM Wiki App は、Karpathy の [[llm-wiki-pattern]] をデスクトップアプリとして実装しようとするプロジェクト。raw sources から相互リンク付き wiki を作り、[[wiki-maintenance-loop|ingest/query/lint]] をアプリ機能として扱う点で、抽象パターンを日常運用の UI・queue・search・graph に落とし込んでいる。 + +README 上の特徴は、two-step ingest、multimodal PDF/image ingestion、knowledge graph、semantic search、persistent ingest queue、folder import/auto-watch、web clipper、local HTTP API、MCP server、agent skill など。これは [[rag-vs-compiled-wiki]] の「query-time retrieval だけでなく ingest-time synthesis を残す」という違いを、検索・MCP・graph traversal まで含む product surface に広げる試み。 + +Hermes の `llm-wiki` skill は同じパターンを agent procedure として定義しており、アプリではなく Markdown directory と agent の作業規律で実現する。両者を比較すると、[[obsidian]] 互換の素朴な file-based workflow と、専用アプリによる queue/search/API 付き workflow のトレードオフが見える。 + +## Open Questions + +- 専用アプリ化で得られる queue、graph、MCP、web clipper の価値は、plain Markdown + agent procedure の単純さを上回るか。 +- 自動 ingest が進むほど、低品質 source や過剰な page creation をどう抑えるか。 +- 既存の Obsidian vault や Hermes skill workflow と併用する場合、どちらを source of truth にするか。 diff --git a/index.md b/index.md index 3f5eece..f183373 100644 --- a/index.md +++ b/index.md @@ -2,10 +2,11 @@ > Content catalog. Every wiki page listed under its type with a one-line summary. > Read this first to find relevant pages for any query. -> Last updated: 2026-06-28 | Total pages: 4 +> Last updated: 2026-06-28 | Total pages: 5 ## Entities +- [[llm-wiki-app]] — Karpathy の LLM Wiki pattern を desktop app、queue、graph/search、MCP/API 付きで具体化する実装。 - [[obsidian]] — LLM Wiki を閲覧・編集するための Markdown/リンク対応ノートアプリ。 ## Concepts diff --git a/log.md b/log.md index 71e52d0..9f187a1 100644 --- a/log.md +++ b/log.md @@ -14,3 +14,19 @@ - Created: concepts/rag-vs-compiled-wiki.md - Created: concepts/wiki-maintenance-loop.md - Created: entities/obsidian.md + +## [2026-06-28] ingest | Discord-discovered LLM Wiki and workflow links +- Scanned #chat and #tw since 2026-06-27T11:58:52Z via discrawl read-only SQL. +- Source saved: raw/articles/nashsu-llm-wiki-2026.md +- Source saved: raw/articles/hermes-research-llm-wiki-skill-2026.md +- Source saved: raw/articles/howm-2026.md +- Source saved: raw/articles/academic-research-skills-2026.md +- Source saved: raw/articles/refactoring-english-effective-design-doc-2026.md +- Source saved: raw/articles/github-automated-my-job-better-leader-2026.md +- Source saved: raw/articles/chinatalk-transfer-station-economy-2026.md +- Source saved: raw/articles/rem-cli-macos-reminders-2026.md +- Created: entities/llm-wiki-app.md +- Updated: concepts/llm-wiki-pattern.md +- Updated: concepts/wiki-maintenance-loop.md +- Updated: index.md +- Skipped wiki-page updates for broader workflow/security links that were useful as raw sources but outside the current wiki's tight LLM Wiki domain. diff --git a/raw/articles/academic-research-skills-2026.md b/raw/articles/academic-research-skills-2026.md new file mode 100644 index 0000000..0b4c6a5 --- /dev/null +++ b/raw/articles/academic-research-skills-2026.md @@ -0,0 +1,703 @@ +--- +source_url: https://github.com/Imbad0202/academic-research-skills +ingested: 2026-06-28 +sha256: d832c91e9fe08f3732df0a1e983f08d14339d45d23a54b8db347b9f539b246c0 +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520478267658731702' + author_id: '890908900520505354' + posted_at: 2026-06-27T17:17:48.130000000Z + message_excerpt: 'https://github.com/Imbad0202/academic-research-skills' +--- + +# Academic Research Skills for Claude Code + +[![Version](https://img.shields.io/badge/version-v3.13.0-blue)](https://github.com/Imbad0202/academic-research-skills/releases/tag/v3.13.0) +[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.20696614.svg)](https://doi.org/10.5281/zenodo.20696614) +[![License: CC BY-NC 4.0](https://img.shields.io/badge/license-CC%20BY--NC%204.0-lightgrey)](https://creativecommons.org/licenses/by-nc/4.0/) +[![Sponsor](https://img.shields.io/badge/sponsor-Buy%20Me%20a%20Coffee-orange?logo=buy-me-a-coffee)](https://buymeacoffee.com/crucify020v) + +[简体中文版](README.zh-CN.md) | [繁體中文版](README.zh-TW.md) | [日本語版](README.ja-JP.md) | [한국어](README.ko-KR.md) + +A comprehensive suite of Claude Code skills for academic research, covering the full pipeline from research to publication. + +**Install in 30 seconds** (Claude Code CLI / VS Code / JetBrains, v3.7.0+): + +```text +/plugin marketplace add Imbad0202/academic-research-skills +/plugin install academic-research-skills +``` + +Then try `/ars-plan` to walk through your paper structure via Socratic dialogue, or jump to [Quick install](#quick-install) for prerequisites and the traditional symlink flow. + +> **AI is your copilot, not the pilot.** This tool won't write your paper for you. It handles the grunt work — hunting down references, formatting citations, verifying data, checking logical consistency — so you can focus on the parts that actually require your brain: defining the question, choosing the method, interpreting what the data means, and writing the sentence after "I argue that." +> +> Unlike a humanizer, this tool doesn't help you hide the fact that you used AI. It helps you write better. Style Calibration learns your voice from past work. Writing Quality Check catches the patterns that make prose feel machine-generated. The goal is quality, not cheating. + +### Why human-in-the-loop, not full automation? + +Lu et al. (2026, *Nature* 651:914-919) built **The AI Scientist** — the first fully autonomous AI research system to publish a paper through blind peer review at a top-tier ML venue (ICLR 2025 workshop, score 6.33/10 vs workshop average 4.87). Their Limitations section enumerates the failure modes that any fully-autonomous AI research pipeline inherits: implementation bugs, hallucinated results, shortcut reliance, bug-as-insight reframing, methodology fabrication, frame-lock, citation hallucinations. + +ARS is built on the premise that **a human researcher augmented by AI avoids these failure modes better than either alone**. Stage 2.5 and Stage 4.5 integrity gates run a 7-mode blocking checklist (see [`academic-pipeline/references/ai_research_failure_modes.md`](academic-pipeline/references/ai_research_failure_modes.md)); the reviewer offers an opt-in calibration mode that measures its own FNR/FPR against a user-supplied gold set. + +[**Zhao et al.**](https://arxiv.org/abs/2605.07723) (2026-05) audited 111M references across 2.5M papers on arXiv, bioRxiv, SSRN, and PMC. Their conservative estimate is 146,932 hallucinated citations for 2025 alone, with an observed mid-2024 inflection; for the bioRxiv-to-PMC pairing they report 85.3% preprint-to-published persistence. The paper describes "real citations deployed to support claims the cited references do not actually make" as an open challenge. ARS v3.7.1 added trust-chain frontmatter for source provenance; v3.7.3 added locator infrastructure (three-layer citation anchors) for future claim-level audits and surfaces advisory risk signals at cite time (ARS labels the claim-faithfulness gap internally as "L3"; this is ARS terminology, not the paper's). v3.7.x is motivated by Zhao et al.'s corpus-scale findings; corpus-scale evaluation of ARS itself remains future work. + +v3.8 closes the second half of the L3 gap. v3.7.3 made every citation carry a locator anchor; v3.8 adds an opt-in audit pass (`ARS_CLAIM_AUDIT=1`) that fetches the cited source against each anchor and judges whether the claim is actually supported. Five new HIGH-WARN classes (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) gate-refuse output through the formatter terminal hard gate. Calibration is shipped as a 20-tuple gold set with FNR<0.15 + FPR<0.10 acceptance thresholds; ramp-on plan is deferred to post-calibration evidence per v3.8 spec §5. + +v3.3 was inspired by [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfister & Yoon, 2026, Google): Semantic Scholar API verification, anti-leakage protocol, VLM figure verification, and score trajectory tracking. + +--- + +## Architecture & pipeline + +**👉 [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md)** — the full pipeline view: flow diagram, stage-by-stage matrix, data-access flow, skill dependency graph, quality gates, and mode list. + +The architecture doc supersedes the sprawling pipeline description that used to live here. Everything about *what runs in which stage* now lives in one place. + +## Quick install + +**Prerequisites** + +- [Claude Code](https://docs.claude.com/en/docs/claude-code/setup) (latest; plugin packaging requires recent versions) +- `ANTHROPIC_API_KEY` exported, or set on first `claude` run +- *Optional:* Pandoc for DOCX, tectonic + Source Han Serif TC for APA 7.0 PDF (Markdown output works without either) +- *Optional (real Python):* The core skills (research / write / review) need no Python — they are prompt-driven. A **real Python interpreter** is needed only for: the `PreToolUse` write-scope guard (optional subagent hardening — if no real Python is found it cleanly no-ops and the guard is simply inactive; core skills are unaffected), plus a few opt-in features that shell out to Python (revision-patch mode, the submission-package verifier, and the `/ars-cache-invalidate` / `/ars-mark-read` / `/ars-unmark-read` commands). On Windows, note that `python3` is often a non-functional Microsoft Store placeholder rather than real Python; install Python from python.org (or via `winget`) so the launcher can find a real interpreter. The guard launcher is a POSIX shell script and `hooks.json` invokes it through `bash`, so on Windows it needs **Git Bash** (bundled with Git for Windows). With Git Bash present, a missing real Python degrades cleanly (the guard no-ops, silently). Without Git Bash, Claude Code falls back to PowerShell, which cannot run the `.sh` launcher at all: the guard is inactive and the `PreToolUse` hook will log an error per call rather than no-op quietly (accepted degradation — the guard is optional and never blocks your writes, but the hook noise is the trade-off until Git Bash is installed). + +**Plugin install (v3.7.0+, recommended):** + +```text +/plugin marketplace add Imbad0202/academic-research-skills +/plugin install academic-research-skills +``` + +**Verify it works:** run `/ars-plan` and describe a paper you're working on — ARS will start a Socratic dialogue to map out chapter structure. For a single-shot test instead, try `/ars-lit-review "your topic"`. + +**👉 [docs/SETUP.md](docs/SETUP.md)** — full guide: install Claude Code, set up API keys, optional Pandoc/tectonic for DOCX/PDF, cross-model verification (`ARS_CROSS_MODEL`), and five installation methods (Plugin, project skills, global skills, claude.ai Project, repo-cloned). + +**Using Codex CLI?** Install the sibling distribution instead: [`Imbad0202/academic-research-skills-codex`](https://github.com/Imbad0202/academic-research-skills-codex) — same workflow content, Codex-native packaging as a single `$academic-research-suite` skill with `ars-*` aliases. + +## Performance & cost + +**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — per-mode token budgets, full-pipeline estimate (~$4–6 for a 15k-word paper), and recommended Claude Code settings (Auto mode; Agent Team optional). + +## Guides & articles + +- [Academic Writing Shouldn't Be a Solo Act](https://open.substack.com/pub/edwardwu223235/p/academic-writing-shouldnt-be-a-solo?r=4dczl&utm_medium=ios) — full pipeline walkthrough (English) +- [學術寫作不該是一個人的事:一套開源 AI 協作工具如何改變研究者的工作流](https://open.substack.com/pub/edwardwu223235/p/ai?r=4dczl&utm_medium=ios) — 完整使用指南(繁體中文) + +--- + +## Features at a glance + +- **Deep Research** — 13-agent research team with Socratic guided mode, PRISMA systematic review, intent detection, dialogue health monitoring, optional cross-model DA, Semantic Scholar API verification. +- **Academic Paper** — 12-agent paper writing with Style Calibration, Writing Quality Check, LaTeX hardening, visualization, revision coaching, citation conversion, anti-leakage protocol, and VLM figure verification. +- **Academic Paper Reviewer** — 7-agent multi-perspective peer review with 0–100 quality rubrics (EIC + 3 dynamic reviewers + Devil's Advocate), concession threshold protocol, attack intensity preservation, optional cross-model DA critique / calibration, R&R traceability matrix, read-only constraint. +- **Academic Pipeline** — 10-stage pipeline orchestrator with adaptive checkpoints, claim verification, Material Passport, optional `repro_lock`, optional cross-model integrity verification, mid-conversation reinforcement, and score trajectory tracking. +- **Data Access Level Metadata** (v3.3.2+) — every skill declares `data_access_level` (`raw` / `redacted` / `verified_only`); enforced by `scripts/check_data_access_level.py`. Pattern adapted from Anthropic's automated-w2s-researcher (2026). See [`shared/ground_truth_isolation_pattern.md`](shared/ground_truth_isolation_pattern.md). +- **Task Type Annotation** (v3.3.2+) — every skill declares `task_type` (`open-ended` or `outcome-gradable`). All current ARS skills are `open-ended`. +- **Benchmark Report Schema** (v3.3.5+) — JSON Schema + lint for honest benchmark comparisons. See [`shared/benchmark_report_pattern.md`](shared/benchmark_report_pattern.md). +- **Artifact Reproducibility Lockfile** (v3.3.5+) — optional `repro_lock` sub-block on Material Passport. **Configuration documentation, not replay guarantee** — LLM outputs are not byte-reproducible. See [`shared/artifact_reproducibility_pattern.md`](shared/artifact_reproducibility_pattern.md). +- **Experiment Provenance Intake** (#260) — optional `experiment_provenance[]` on the Material Passport records experiments the scholar ran **externally** (ARS never runs experiments), and manuscript claims join to them via `claim_intent_manifest.planned_experiment_ids[]`. The integrity gate (Stage 2.5/4.5) audits each experiment-backed claim against declared provenance — `ALIGNED` / `OVERSTATED` / `NOT_SUPPORTED_BY_PROVENANCE` / `PROVENANCE_INSUFFICIENT` — **without judging whether the experiment itself was correct**. A fail-closed `experiment_intake_declaration` makes "did you run experiments?" an explicit Stage 1 decision (even literature-only runs declare `no_experiments_declared`). See [`shared/handoff_schemas.md`](shared/handoff_schemas.md) §"Experiment Provenance Intake (#260)". + +--- + +## Showcase: real pipeline output + +See the complete artifacts from a real 10-stage pipeline run — peer review reports, integrity verification reports, and the final paper: + +**[Browse all pipeline artifacts →](examples/showcase/)** + +| Artifact | Description | +|---|---| +| [Final Paper (EN)](examples/showcase/full_paper_apa7.pdf) | APA 7.0 formatted, LaTeX-compiled | +| [Final Paper (ZH)](examples/showcase/full_paper_zh_apa7.pdf) | Chinese version, APA 7.0 | +| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: caught 15 fabricated refs + 3 statistical errors | +| [Integrity Report — Final](examples/showcase/integrity_report_stage4.5.pdf) | Stage 4.5: zero regressions confirmed | +| [Peer Review Round 1](examples/showcase/stage3_review_report.pdf) | EIC + 3 Reviewers + Devil's Advocate | +| [Re-Review](examples/showcase/stage3prime_rereview_report.pdf) | Verification after revisions | +| [Peer Review Round 2](examples/showcase/stage3_review_report_r2.pdf) | Follow-up review | +| [Response to Reviewers](examples/showcase/response_to_reviewers_r2.pdf) | Point-by-point author response | +| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Independent full-reference audit: found 21/68 issues missed by 3 rounds of integrity checks | + +--- + +## Companion: Experiment Agent + +If your research involves running experiments (code or human studies) before writing, the [Experiment Agent](https://github.com/Imbad0202/experiment-agent) skill fills the gap between ARS Stage 1 (RESEARCH) and Stage 2 (WRITE). + +``` +ARS Stage 1 RESEARCH → RQ Brief + Methodology Blueprint + ↓ + experiment-agent → run/manage experiments → validate results + ↓ +ARS Stage 2 WRITE → write paper with verified experiment results +``` + +**What it does**: executes code experiments (Python, R, etc.) with real-time monitoring, manages human study protocols with IRB ethics checklist, interprets statistics with 11-type fallacy detection, and verifies reproducibility. + +**How to use together**: pause the ARS pipeline after Stage 1, run experiments in a separate experiment-agent session, then bring the results (with Material Passport) back to ARS Stage 2. ARS requires zero modification. See the [experiment-agent README](https://github.com/Imbad0202/experiment-agent) for setup instructions. + +**Stage 1 intake declaration (#260)**: at Stage 1, ARS detects whether the run will carry experiment-backed claims and sets a fail-closed `experiment_intake_declaration` on the Material Passport. If you ran experiments externally, the scholar enters one `experiment_provenance[]` entry per experiment (`experiment_id`, nested `repro_lock`, `planned_vs_executed[]`, `negative_results[]`, `known_limitations[]`) and the declaration is set to `experiments_declared`; if not, it is set to `no_experiments_declared`. The declaration is **required on every post-#260 passport** — a run that touches no experiments still declares `no_experiments_declared`, so the integrity gate can never be silently bypassed by a forgotten provenance block. The `experiment_id`s are frozen at this intake point; the writers later reference them via `planned_experiment_ids[]`. + +**Teaching-side companion**: [Teaching Skills](https://github.com/YujxZJCN/teaching-skills) applies the ARS architecture (skill ensembles, shared contracts, staged gates, a Course Passport) to the teaching side of academic life — course design → lessons → assessment → delivery → reflection; its `sotl` mode hands classroom-inquiry projects off to ARS deep-research / academic-paper for the publication phase. + +--- + +## Usage + +### Quick Start + +``` +# Start a full research pipeline +You: "I want to write a research paper on AI's impact on higher education QA" + +# Start with Socratic guidance +You: "Guide my research on AI in educational evaluation" + +# Write a paper with guided planning +You: "Guide me through writing a paper on demographic decline" + +# Review an existing paper +You: "Review this paper" (then provide the paper) + +# Check pipeline status +You: "status" +``` + +### Individual Skills + +#### Deep Research (8 modes) + +``` +"Research the impact of AI on higher education" → full mode +"Give me a quick brief on X" → quick mode +"Do a systematic review on X with PRISMA" → systematic-review mode +"Guide my research on X" → socratic mode (guided) +"Fact-check these claims" → fact-check mode +"Do a literature review on X" → lit-review mode +"Compare these papers in WHY/HOW/WHAT format" → three-way-scan mode +"Review this paper's research quality" → review mode +``` + +#### Academic Paper (11 modes) + +``` +"Write a paper on X" → full mode +"Guide me through writing a paper" → plan mode (guided) +"Build a paper outline" → outline-only mode +"I have a draft, here are reviewer comments" → revision mode +"Parse these reviewer comments into a roadmap" → revision-coach mode +"Write an abstract for this paper" → abstract-only mode +"Turn this into a literature review paper" → lit-review mode +"Convert to LaTeX" / "Convert citations to IEEE" → format-convert mode +"Check citations" → citation-check mode +"Generate an AI disclosure statement for NeurIPS" → disclosure mode +"Audit my rebuttal draft against the reviews" → rebuttal-audit mode +``` + +#### Academic Paper Reviewer (6 modes) + +``` +"Review this paper" → full mode (EIC + R1/R2/R3 + Devil's Advocate) +"Quick assessment of this paper" → quick mode +"Guide me to improve this paper" → guided mode +"Check the methodology" → methodology-focus mode +"Verify the revisions" → re-review mode +"Calibrate this reviewer against my gold set" → calibration mode +``` + +#### Academic Pipeline (Orchestrator) + +``` +"I want to write a complete research paper" → full pipeline from Stage 1 +"I already have a paper, review it" → mid-entry at Stage 2.5 (integrity first) +"I received reviewer comments" → mid-entry at Stage 4 +``` + +> Pipeline ends with **Stage 6: Process Summary** — auto-generates a paper creation process record with 6-dimension Collaboration Quality Evaluation (1–100 scoring). + +### Supported Languages + +- **Traditional Chinese** (繁體中文) — default when user writes in Chinese +- **English** — default when user writes in English +- Bilingual abstracts (Chinese + English) for academic papers + +> **Using a different language?** Socratic mode (deep-research) and Plan mode (academic-paper) use **intent-based activation** — they detect the meaning of your request, not specific keywords. This means they work in **any language** without modification. +> +> However, the general `Trigger Keywords` section (which determines whether the skill is activated at all) still lists English and Traditional Chinese keywords. If you find the skill isn't activating reliably in your language, you can add your language's keywords to the `### Trigger Keywords` section in each `SKILL.md` file to improve matching confidence. + +### Supported Citation Formats + +- APA 7.0 (default, including Chinese citation rules) +- Chicago (Notes & Author-Date) +- MLA +- IEEE +- Vancouver + +### Supported Paper Structures + +- IMRaD (empirical research) +- Thematic Literature Review +- Theoretical Analysis +- Case Study +- Policy Brief +- Conference Paper + +--- + +## Skill Details + +Per-agent responsibilities and per-stage artifacts now live in [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md). Version numbers are anchored here so release metadata stays in one place. + +### Deep Research (v2.11.0) + +13-agent research team. Modes: full, quick, review, lit-review, three-way-scan, fact-check, socratic, systematic-review. Full agent roster and artifacts: see ARCHITECTURE.md §3. + +### Academic Paper (v3.2.0) + +12-agent paper writing pipeline. Modes: full, plan, outline-only, revision, revision-coach, abstract-only, lit-review, format-convert, citation-check, disclosure, rebuttal-audit. Output: MD + DOCX (via Pandoc when available) + LaTeX (APA 7.0 `apa7` class / IEEE / Chicago) → PDF via tectonic. Full agent roster and per-phase responsibilities: see ARCHITECTURE.md §3. + +### Academic Paper Reviewer (v1.10.0) + +7-agent multi-perspective review with **0-100 quality rubrics**. Modes: full, re-review, quick, methodology-focus, guided, calibration. **Decision mapping:** ≥80 Accept, 65-79 Minor Revision, 50-64 Major Revision, <50 Reject. First-round review team vs. narrow re-review team boundary: see ARCHITECTURE.md §3 Stage 3 / Stage 3'. + +### Academic Pipeline (v3.13.0) + +10-stage orchestrator with integrity verification, two-stage review, Socratic coaching, and collaboration evaluation. Pipeline guarantees: every stage requires user confirmation checkpoint; integrity verification (Stage 2.5 + 4.5) cannot be skipped; R&R Traceability Matrix (Schema 11) independently verifies author revision claims. v3.4 added the Compliance Agent (PRISMA-trAIce + RAISE) at Stage 2.5 / 4.5. v3.5 adds the **Collaboration Depth Observer** (`collaboration_depth_agent`, advisory only — never blocks) at every FULL/SLIM checkpoint and at pipeline completion. MANDATORY integrity gates (2.5 / 4.5) explicitly skip the observer so compliance checks are not diluted. Based on Wang & Zhang (2026), IJETHE 23:11. Stage-by-stage matrix with agents, artifacts, and gates: see ARCHITECTURE.md §3. + +--- + +## v3.0 Optimizations: What We Discovered About AI's Structural Limits + +### What happened + +While using ARS to write a reflection article about AI in higher education, I ran into three structural problems that no amount of prompt engineering could fix: + +1. **Frame-lock**: I asked the AI to run a devil's advocate debate against its own thesis. It did — four rounds, each more refined than the last. But every round stayed inside the frame I'd set. The DA attacked arguments, never premises. It never asked "are we even discussing the right question?" This is the same pattern that caused the 31% citation error rate in v2.7's stress test: the verifying AI and the generating AI share the same cognitive frame. + +2. **Sycophancy under pushback**: Every time I challenged the DA's attacks, it conceded too quickly. It retracted findings faster than it launched them. The model's training rewards conversational harmony — so "the user pushed back" was treated as evidence that the attack was wrong, when often it just meant the user was persistent. + +3. **Intent misdetection**: The Socratic Mentor kept trying to converge and produce deliverables ("Want me to write this up?") when I was still exploring. It couldn't distinguish "the user wants a deep philosophical discussion" from "the user wants an RQ brief." Both look like engagement, but they need opposite AI behaviors. + +### What we changed (v3.0) + +**Devil's Advocate — Concession Threshold Protocol** (`deep-research` + `academic-paper-reviewer`) +- DA must now score every rebuttal on a 1-5 scale before responding +- Concession only allowed at score ≥4 (rebuttal directly addresses core attack with evidence) +- Score ≤3: hold position and restate the original attack +- Anti-sycophancy rules: no consecutive concessions, concession rate tracking, frame-lock detection after each checkpoint + +**Socratic Mentor — Intent Detection Layer** (`deep-research`) +- Classifies user intent as exploratory vs. goal-oriented at dialogue start and every 3 turns +- Exploratory mode: disables auto-convergence, raises max rounds to 60, prohibits "want me to summarize?" prompts +- Goal-oriented mode: standard convergence behavior +- Anti-premature-closure rules: in exploratory mode, the user decides when to stop + +**Socratic Mentor — Dialogue Health Indicator** (`deep-research`) +- Silent self-assessment every 5 turns on three dimensions: persistent agreement, conflict avoidance, premature convergence +- Auto-injects challenging questions when agreement pattern detected +- Invisible to user (to prevent gaming), but log available for post-session review + +### Why this matters + +These optimizations don't solve AI's structural limits — they make the limits visible and manageable. The DA will still eventually concede if pushed hard enough. The Socratic Mentor will still have some convergence bias. But now there are explicit checkpoints that slow down the sycophancy, force the DA to justify concessions, and prevent the Mentor from wrapping up before the user is ready. + +The deeper lesson: AI literacy isn't about learning to use AI as a tool, following ethics rules, or fearing AI risks. It's about engaging AI deeply enough to discover its structural limits yourself — and your own thinking limits in the process. + +--- + +## License + +This work is licensed under [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/). + +**You are free to:** +- Share — copy and redistribute the material +- Adapt — remix, transform, and build upon the material + +**Under the following terms:** +- **Attribution** — You must give appropriate credit +- **NonCommercial** — You may not use the material for commercial purposes + +**Attribution format:** +``` +Based on Academic Research Skills by Cheng-I Wu +https://github.com/Imbad0202/academic-research-skills +``` + +--- + +## Contributors + +**Cheng-I Wu** (吳政宜) — Author and maintainer + +**[aspi6246](https://github.com/aspi6246)** — Contributor. The v3.1 optimization was inspired by patterns from [Claude-Code-Skills-for-Academics](https://github.com/aspi6246/Claude-Code-Skills-for-Academics): read-only constraint pattern, anti-pattern codification as first-class design, cognitive framework approach (teaching "how to think" not just procedures), and lean skill size philosophy. + +**[mchesbro1](https://github.com/mchesbro1)** — Contributor. Originally proposed and drafted the IS Basket of 8 journals for `academic-paper-reviewer/references/top_journals_by_field.md` ([Issue #5](https://github.com/Imbad0202/academic-research-skills/issues/5)). + +**[cloudenochcsis](https://github.com/cloudenochcsis)** — Contributor. Extended the IS section from the *Basket of 8* to the full *Senior Scholars' Basket of 11* — adding *Decision Support Systems*, *Information & Management*, and *Information and Organization* ([Issue #7](https://github.com/Imbad0202/academic-research-skills/issues/7), [PR #8](https://github.com/Imbad0202/academic-research-skills/pull/8)). Sourced from the [AIS Senior Scholars' List of Premier Journals](https://aisnet.org/research/seniorscholarsbasket/). + +**[eltociear](https://github.com/eltociear)** (Ikko Eltociear Ashimine) — Contributor. Translated the Japanese README ([`README.ja-JP.md`](README.ja-JP.md)) ([PR #161](https://github.com/Imbad0202/academic-research-skills/pull/161)). + +**[xpfo-go](https://github.com/xpfo-go)** (xpfo) — Contributor. Translated the Simplified Chinese README ([`README.zh-CN.md`](README.zh-CN.md)) ([PR #181](https://github.com/Imbad0202/academic-research-skills/pull/181)). + +**[devCharlotte](https://github.com/devCharlotte)** — Contributor. Translated the Korean README ([`README.ko-KR.md`](README.ko-KR.md)) ([PR #469](https://github.com/Imbad0202/academic-research-skills/pull/469)). + +**[Yaobin29](https://github.com/Yaobin29)** — Contributor. Proposed reviewer-response tooling in [PR #433](https://github.com/Imbad0202/academic-research-skills/pull/433); the `deep-research three-way-scan` mode and the `academic-paper rebuttal-audit` mode (rescued from the PR's `audit` concept) were integrated from that contribution in v3.12.1. + +--- + +## Changelog + +### v3.13.0 (2026-06-18) — Hook portability, provider-agnostic verification, guard correctness + +> A minor release hardening the install/runtime surface and extending cross-model reach. **Fixes:** the write-scope guard no longer false-denies a user's own `CLAUDE.md` under the git-clone + symlink install layout (#459, closing the residual half of #448/#449 — `CLAUDE.md` is documentation, not a load-bearing enforcement file, so it leaves the infra-protected list while every load-bearing file stays protected); Windows Python hook portability + graceful no-Python degradation via a cross-platform `hooks/run_guard.sh` launcher that rejects the 0-byte Microsoft Store `python3` stub and never spams the hook log (#454); `draft_writer` dual-phase static union documented + POSIX-safe Windows path matching (#451). **Added:** provider-agnostic cross-model verification accepting OpenAI-compatible endpoints (MiMo, DeepSeek, self-hosted) alongside grounded first-party OpenAI, which is never silently downgraded (#455); an opt-in Socratic adjacent-framing probe (STORM-borrowed perspective expansion, `ARS_SOCRATIC_ADJACENT_PROBE=1`, default OFF, prose-layer only — `deep-research` 2.10.0 → 2.11.0) (#461). `academic-pipeline` tracks the suite at v3.13.0; `academic-paper` and `academic-paper-reviewer` are unchanged. See `CHANGELOG.md` for the per-issue detail. + +### v3.12.1 (2026-06-15) — Reviewer-response triage modes (PR #433 integration) + +> A patch release folding the genuinely-novel parts of an external contribution into existing skills as modes, per ARS's mode-based architecture. **New modes:** `deep-research` `three-way-scan` — a lightweight WHY/HOW/WHAT paper-comparison triage between `quick` and `lit-review`, with per-paper shortlists + a cross-paper synthesis (`deep-research` 2.9.4 → 2.10.0); `academic-paper` `rebuttal-audit` — standalone advisory QA of an author's existing rebuttal/response draft against the reviewer comments (per-comment coverage table + gap list + tone/evidence/misread risk flags), which generates nothing and explicitly suppresses Schema 11 / Material Passport writes / `ready_to_submit` when run standalone (enforced by a `check_rebuttal_audit_guard()` lint with mutation coverage); plus a `revision-coach` scope extension to pushback/disagreement posture and non-journal scopes, and `/ars-3w` + `/ars-rebuttal-audit` slash commands. Routed by input shape: reviewer comments AND a draft → `rebuttal-audit`; comments only → `revision-coach`. Integrated from [@Yaobin29](https://github.com/Yaobin29)'s [PR #433](https://github.com/Imbad0202/academic-research-skills/pull/433). Suite mode count 25 → 27 (still 4 skills). See `CHANGELOG.md` for the per-issue detail. + +### v3.12.0 (2026-06-08) — Kong auto-research feature track: experiment provenance, figure fidelity, cross-paper contradiction, partial-evidence decomposition + +> A minor release shipping the Kong et al. (2026, arXiv:2605.18661) auto-research feature track plus the partial-evidence-trap decomposition work, each reviewed and merged independently. **New features:** Experiment Provenance Intake + claim→experiment alignment — a schema-first evidence-ledger layer for experiment-backed claims, intake-and-alignment only (the scholar runs experiments externally; ARS never executes them) (#260); a Figure/Table Fidelity Gate that checks whether a caption's interpretation follows from the data and whether the manuscript cites the artifact for a claim it supports (#261); a structured Cross-Paper Contradiction inventory making assessed paper-pairs enumerable for scholar confirmation (#262); and sub-claim decomposition before judgment in both the citation judge (#213) and the editorial synthesizer (#214), closing the §F.3.2 partial-evidence trap on both layers. **Guidance + interpretive layer:** concise-output + pressure-stable boundary reinforcement across the report-producing reviewers (#274); a same-family / rubric-aware calibration epistemic note (#273); the retrieved-content instruction/data boundary stated as a standing principle (#367). **Negative scope:** the Kong META (#255) closed with a "Rejected mechanisms" section in `POSITIONING.md` enumerating the five autonomous mechanisms ARS does not do, plus two Tier D design-lesson docs. **Release-discipline lint:** version-consistency invariants 5–7 (#357) and ARCHITECTURE component-version policing (#345). Plus correctness fixes across the cross-model grounding guards (#346 / #349 / #351), the citation-gate cache key and rationale bounding (#359 / #360 / #361), the eval gold set (#250), and ACL/EMNLP disclosure regrounding (#242). The new schemas, manifest field, and all invariants are additive and backward-compatible. `academic-pipeline` tracks the suite at v3.12.0; the other three skill versions are unchanged. See `CHANGELOG.md` for the per-issue detail. + +### v3.11.1 (2026-06-06) — Post-ship correctness, hardening & provenance rollup + +> A patch release rolling up the post-ship fixes surfaced after v3.11.0, each reviewed and merged independently: a cross-model consent-gate extension to the integrity-verification + collaboration-depth paths (#322), a per-entry OpenAlex + Crossref backfill parallelization (#138), and seven correctness/hardening fixes across the citation-existence gate, the v3.10 policy layer, the eval harness, the domain evidence profiles, and the #310 security-boundary edge cases (#323 / #327 / #328 / #329 / #331 / #332 / #333) — including two P1 fixes (#327 domain-profile activation on the no-handoff path, #328 the eval-harness per-class threshold gate). No new features and no breaking schema changes. See `CHANGELOG.md` for the per-issue detail. + +### v3.11.0 (2026-06-04) — Deterministic citation verification gate (#182) + +> Adds a **deterministic citation-existence verification gate** that runs independently of LLM peer review. Every cited reference is cross-checked against up to four bibliographic indexes — Semantic Scholar + OpenAlex + Crossref + the new **arXiv resolver** (`scripts/arxiv_client.py`, no API key needed) — and a per-citation `lookup_verified` status (`{true, false, unresolvable}`) is written to a unified summary, so a fabricated citation with a provably-bogus DOI/arXiv ID is caught by lookup rather than by hoping a reviewer agent notices. The gate **inherits the v3.10 `terminal_policies` opt-in model**: detection always runs, but a `lookup_verified == false` row is terminal **only** when a user opts into `terminal_policies.citation_existence == strict` — default behavior is advisory and `/ars-mark-read`-acknowledgeable. `false` is narrowed to **ID-keyed unmatched** (an exact DOI/arXiv lookup that provably fails), so legitimately-unindexed humanities / non-English / regional citations stay `unresolvable` and never block (a documented precision-over-recall tradeoff). Ships a persistent SQLite verification cache (`~/.cache/ars/verification.db`, 90-day TTL) with an `/ars-cache-invalidate` command, a standalone `verification_gate` API + `verify_passport.py` CLI, and a four-index extension (k=0..4) of the v3.9.0 contamination triangulation matrix (all advisory). `academic-pipeline` tracks the suite at v3.11.0; the other three skill versions are unchanged. Spec: `docs/design/2026-05-21-v3.10-182-promote-citation-gate-spec.md` (§0 amendment + C-V6). + +### v3.10.0 (2026-06-01) — Triangulation policy layer, Kong survey adoptions, eval harness, scoped-write guard + +> Minor release bundling: the opt-in contamination-triangulation **terminal policy layer** (#127 — default citation behavior byte-equivalent to v3.9.0); **Kong et al. 2026 survey adoptions** — the Rebuttal Commitment Ledger (#256/#266/#268/#269) and discipline-relative domain evidence profiles (#259); the **v3.10 measurement infrastructure** — a generalized eval gold set + ranking-lift CI gate (#184); the **scoped-write guard MVP** (#134) — a deterministic `PreToolUse` hook that fences the 23 single-phase agents to their own phase directory and denies them Bash (they use the Grep/Glob and structured editing tools instead); the `/ars-mark-read` plugin commands (#190) plus a broken-on-arrival fix (#195); a Simplified-Chinese README (#185); and CI hardening (#156/#155). `academic-paper` → v3.2.0 and `academic-paper-reviewer` → v1.10.0 for the Commitment-Ledger and domain-profile additions; `academic-pipeline` tracks the suite at v3.10.0. Default skill behavior is unchanged unless a strict policy mode is opted into; the one default-on change is the #134 guard, which constrains the fenced subagents, not user-facing outputs. + +### v3.9.4.2 (2026-05-19) — post-ship hotfix for PR #149 CI discipline gates (codex post-ship) + +> Codex post-ship review of PR #149 (7 CI discipline gates) surfaced 4 P2 findings; v3.9.4.2 hardens 3 of 4. F1: `harness-retirement-monthly.yml` adds `GH_REPO` so scheduled runs have repo context for `gh issue create`. F2: `release-cooldown.yml` filters `PREV_TAG` lookup to `v*` tags so non-release tags cannot bypass cooldown. F3: `release-cooldown.yml` also reads annotated tag subject + accepts `hot-fix` spelling (v3.9.2 was previously a false-negative hotfix). PR #157 follow-up: `[skip-cooldown]` override now read from both commit message AND annotated tag message (self-bootstrapping fix — this tag's cooldown bypass demonstrates F2+F3 work end-to-end). F4 (test-count-monotonic harden) reverted because it surfaced pre-existing `scripts/` package issue, tracked as #154 (since fixed by PR #158) + re-attempt #155. Closes #152. Follow-ups: #155, #156. + +### v3.9.4.1 (2026-05-19) — post-ship hotfix for v3.9.4 temporal verification (#135 codex post-ship) + +> Codex post-ship review of v3.9.4 caught 4 real bugs that per-task subagent reviewers missed. Hotfix patches all 4: (1) `audit()` now wires `citation_provenance` through to P2 and P4 — when a ref slug has `confidence: low` or `conflict`, the verifier emits `TEMPORAL-METADATA-MISSING` instead of using timeline dates as ground truth (spec §3.4 first-party safety check was broken). (2) `_date_to_interval` parses all schema-valid date shapes including `YYYY-MM` (Crossref month precision) and `YYYY-MM-DD..YYYY-MM-DD` (interval); v3.9.4 silently `ValueError`'d on these and skipped the check. (3) P4 now binds direct date captures when ref markers are absent — sentences like "The 2026 policy enabled the 2020 rollout" actually trigger now. (4) `citation_provenance.schema.json` `confidence:high` allOf now requires presence (`then.required`) in addition to non-null, closing the absent-property bypass. 1561 passed (+12 new tests vs v3.9.4 baseline, 0 regression). ARCHITECTURE.md aligned to current state (was stale at v3.8.0). + +### v3.9.4 (2026-05-18) — #135 temporal verification layer (advisory) + +> Deterministic advisory verifier at the Phase 4 → 5 boundary covering 5 temporal failure modes (P1 retrospective arithmetic, P2 anachronistic citation, P3 comparator unmaterialized, P4 causal inversion, P5 deictic present). New Phase 2 sibling `timeline_extraction_agent` owns `phase2_investigation/timeline.yaml` + `phase2_investigation/citation_provenance.yaml`. Verifier script `scripts/temporal_integrity_audit.py` runs 5 passes deterministically. M3 Temporal Integrity Iron Rule added to `report_compiler_agent` + `draft_writer_agent`. M6-minimal: Crossref `issued` + pdftotext cover first-party verification. M7-minimal: date provenance + comparator materialization. M5-stub: user-declared `version_family_id` only. Zero modification to `literature_corpus_entry`, `claim_audit_result`, `claim_intent_manifest`. `bibliography_agent` unmodified (F2 invariant). 3 new sidecar schemas. Coverage estimate: 55-70% baseline / 65-75% with M7 minimal. 1549 passed (+44 new, 0 regression). + +### v3.9.3 (2026-05-18) — #128 housekeeping (shared client utilities + dedup resolvers) + +> Pure refactor + one latent-bug fix from the v3.9.0 `/simplify` review backlog. Extracts `scripts/_text_similarity.py` (3-way client dedup: normalize / similarity / threshold / retry constants) + `scripts/_passport_yaml.py` (2-way migration tool dedup: ruamel.yaml round-trip config) + private `_resolve_by_doi_then_title` helper (2-way resolver body dedup, §3.4 / §3.5 API surface preserved). Standardizes throttle measurement on `time.monotonic` across OpenAlex + Crossref (was `time.time`, NTP-unsafe), aligning with Semantic Scholar. Dual-path import infrastructure on all 5 module-level cross-imports (sibling-first, namespace-package fallback) preserves class identity for `SemanticScholarUnavailable` and bonus-fixes 2 latent-broken `import scripts.X` paths. 1505 passed (+23 new, 0 regression). #128 §4 (parallelize OA + CR per-entry) carried to #138. + +### v3.9.2 (2026-05-18) — #133 phase boundary hot-fix + +> #133 closure (hot-fix layer). Long-term architectural fix tracked as v3.10 active conductor in #134. Adds: routing clarification gate in CLAUDE.md (cross-phase materials → clarify with a-d options, not silent dispatch), 22 single-phase agents get prompt hard fence (`## Phase Boundary (v3.9.2)`), 16 multi-phase / phase-orthogonal / cross-phase-meta agents intentionally NOT fenced (honest framing — prose placebo creates false-enforcement illusion), advisory verifier `scripts/check_pipeline_integrity.py` detects #133 pattern post-hoc. Behavioral smoke tests with cross-model spot-check (100% Opus 4.7, ≥75% Sonnet + GPT-5.5). + +### v3.9.1 (2026-05-18) — #129 + #130 client hardening + +> v3.9.0 hot-fix. Wraps OpenAlex / Crossref response-read failures as `*Unavailable` (#129); guards `check_claim_audit_consistency` against non-string `manifest_id` (#130). No spec change. + +### v3.9.0 (2026-05-17) — #102 cross-index triangulation measurement + +> #102 closure. v3.7.3 shipped single-index (Semantic Scholar) contamination detection; v3.9.0 extends to three-index triangulation (S2 + OpenAlex + Crossref) as **advisory evidence only**. Two new optional booleans (`openalex_unmatched`, `crossref_unmatched`) on `contamination_signals`; manual-entry not-rule extended symmetrically. Finalizer adds a 4-tier advisory matrix (k=0/1/2/3 over present `*_unmatched` fields) with v3.7.3 legacy `CONTAMINATED-UNMATCHED` preserved for the k=1/k_max=1 S2-only case. Formatter pass-through allowlist extends 3 → 9 suffixes; refusal rules 1-10 unchanged per R-L3-2-E. The policy layer (strict modes, hard-block tier, `venue_type` / `triangulation_policy`) is deferred to v3.10 per spec §2.3. k=3 marker is `CONTAMINATED-TRIANGULATION-UNMATCHED` (describes observable, not inferred cause). 3 new firm rules: R-L3-2-C (k computed over present fields), R-L3-2-D (no API-inferred classification), R-L3-2-E (refusal list unchanged; pass-through allowlist extends). + +**Migration:** v3.7.3 corpora — run `python scripts/migrate_literature_corpus_to_v3_9_0.py PATH` to backfill the two new fields. Pre-v3.7.3 corpora — run `migrate_literature_corpus_to_v3_7_3.py` FIRST, then v3.9.0 migration (daisy-chained per spec §3.7; the v3.9.0 tool only acts on entries that already carry `contamination_signals.semantic_scholar_unmatched`). + +### v3.8.2 (2026-05-17) — #118 uncited audit_tool_failure surface + +> #118 closure. The `ARS_CLAIM_AUDIT=1` uncited constraint-judging path used to silently substitute `{"judgment": "NOT_VIOLATED"}` on `JudgeInvocationError`, suppressing HIGH-WARN constraint checks on transient judge outage. v3.8.2 routes those failures through a dedicated `uncited_audit_failures[]` aggregate at MED-WARN advisory tier, mirroring the cited path INV-14 row but using a dedicated schema because `claim_audit_result.ref_slug` is required and the uncited path has no ref to bind. The four option-1..4 trade-offs from the #118 issue body landed on option 2 (new aggregate) — option 4 (re-raise and abort) was rejected for the audit-coverage hit on flaky judge endpoints. + +- **New `uncited_audit_failure.schema.json` aggregate** (spec §3.6). One entry per uncited sentence × manifest pair where the constraint judge raised `JudgeInvocationError`. Same fault-class enum as cited-path INV-14 (`judge_timeout` / `judge_api_error` / `judge_parse_error` / `cache_corruption` / `retrieval_api_error` / `retrieval_timeout` / `retrieval_network_error`). `rule_version: D4-c-v1-uaf-v1`. +- **UAF-INV-1..UAF-INV-6 lint** (spec §6 rule 4d). `finding_id` uniqueness, scoped_manifest_id cross-array integrity, (M, C) pair integrity when manifest_claim_id non-null, per-(sentence, manifest) dedup, rationale fault_class prefix, cross-aggregate exclusivity vs `constraint_violations[]`. +- **Finalizer §5 MED-WARN advisory row**: annotation `[CLAIM-AUDIT-TOOL-FAILURE-UNCITED — ]`, gate passes (retry-next-pass remediation). Formatter REFUSE list unchanged — UAF is advisory. +- **Pipeline integration** (`scripts/claim_audit_pipeline.py`): swallow site at line 1211-1224 removed; `JudgeInvocationError` now emits a UAF row + `continue`s to the next (sentence, manifest) pair. No fake NOT_VIOLATED reaches `constraint_violations[]`. +- **Tests**: 18 new (15 schema/lint TSUAFUncitedAuditFailureInvariants + 3 pipeline integration TP23UncitedJudgeOutageEmitsUAF). Baseline 694 → 712 tests, 0 regression. +- **Agent doc** (`academic-pipeline/agents/claim_ref_alignment_audit_agent.md`): Output emission table grows seventh row; Error handling table grows from 3 surfaces to 4 surfaces with the uncited-path UAF row. + +### v3.8.0 (2026-05-16) — L3 Claim-Faithfulness Locator + Audit (paired milestone) + +> v3.7.3 + v3.8 close the L3 (claim-faithfulness) gap end-to-end. v3.7.3 ships the locator infrastructure — every citation carries a three-layer anchor so future audits can fetch the cited passage. v3.8 ships the audit pass that consumes those anchors, judges whether the cited source supports the claim, and gate-refuses HIGH-WARN violations at the formatter terminal hard gate. The release also bundles 5 audit-trail-shipped feature PRs accumulated since v3.7.0 (#104 / #105 / #108 / #111 / #115). + +- **#103 — `claim_ref_alignment_audit_agent`** (v3.8 PR #121). Opt-in (`ARS_CLAIM_AUDIT=1`, default OFF) Stage 4→5 audit agent. Judges every sampled citation against retrieved excerpt; emits `claim_audit_results[]` + `claim_intent_manifests[]` + `claim_drifts[]` + `uncited_assertions[]` + `constraint_violations[]` aggregates. 8-row finalizer matrix routes HIGH-WARN classes (CLAIM-NOT-SUPPORTED / NEGATIVE-CONSTRAINT-VIOLATION / FABRICATED-REFERENCE / ANCHORLESS / CONSTRAINT-VIOLATION-UNCITED) through the formatter REFUSE rules 6-10. Calibration runner ships with 20-tuple gold set (T-C1 FNR<0.15 + FPR<0.10, T-C2 per-class, T-C3 shape integrity). 8 rounds of dual-track review (R1 codex + Gemini-3.1-pro-preview, R2-R8 codex-only after Gemini quota exhausted); trajectory R1 4P1+2P2 → R8 0P1+4P2 ship gate. +- **v3.7.3 — Three-Layer Citation Emission + contamination signals** (PR #98). `synthesis_agent` / `draft_writer_agent` / `report_compiler_agent` gain `## Three-Layer Citation Emission (v3.7.3)` H2. Every `` carries `` with ` ∈ {quote, page, section, paragraph, none}` (quote anchors capped at 25 words, URL-encoded). `pipeline_orchestrator_agent` finalizer becomes 5-cell with precedence-zero NO-LOCATOR check. `formatter_agent` adds explicit hard-gate refusal for `[UNVERIFIED CITATION — NO QUOTE OR PAGE LOCATOR]`. `literature_corpus_entry.schema.json` adds optional `contamination_signals: { preprint_post_llm_inflection, semantic_scholar_unmatched }` object. `bibliography_agent` computes both signals at ingest. 11-round review trajectory (Codex×10 + Gemini cross-model×1) closed 22 findings. Spec: `docs/design/2026-05-12-ars-v3.7.3-claim-faithfulness-and-contaminated-source-spec.md`. External motivation: Zhao et al. arXiv:2605.07723 (2026-05). +- **#108 — AI disclosure policy-anchor renderer** (audit-trail-shipped 2026-05-14). Adds PRISMA-trAIce / ICMJE / Nature / IEEE policy-anchor disclosure paths alongside the existing venue-track renderer. +- **#111 — `slr_lineage` emission on systematic-review → academic-paper handoff** (2026-05-15). Schema 9 optional boolean `slr_lineage` field; producer `pipeline_orchestrator_agent` writes at every handoff transition; consumer `disclosure` mode dispatches `--policy-anchor=prisma-trAIce` per the §4.3 G2 invariant track gate. +- **#104 — README motivation: Zhao et al. corpus-scale evidence anchor** (2026-05-15). README + `README.zh-TW.md` motivation section frames the v3.7.x line against Zhao et al.'s 146,932 hallucinated-citation finding. +- **#105 — v3.7.3 contamination_signals backfill migration tool** (2026-05-15). `scripts/migrate_literature_corpus_to_v3_7_3.py` retro-computes both contamination signals across pre-v3.7.3 passports. +- **#115 — Semantic Scholar client maturity** (2026-05-15). `scripts/semantic_scholar_client.py` adds 1-req/s throttle (drops to 0.1s when `S2_API_KEY` detected), outage latch on URLError, and `reset_outage_latch()` for long-running cross-passport batches. + +### v3.7.0 (2026-05-05) — Claude Code Plugin Packaging + +> Plugin packaging upgrade: ARS now installs in one line on Claude Code CLI / VS Code / JetBrains via `/plugin marketplace add Imbad0202/academic-research-skills` + `/plugin install academic-research-skills`. The traditional `git clone + symlink to ~/.claude/skills/` flow continues to work — both tracks are first-class. + +- **Plugin manifest + marketplace metadata** (Phase 1, PR #68). `.claude-plugin/plugin.json` declares the suite (4 skills auto-discovered from `skills/` directory via relative symlinks). `.claude-plugin/marketplace.json` registers the plugin so a single GitHub-hosted endpoint serves both the marketplace listing and the plugin source. README + `README.zh-TW.md` + `docs/SETUP.md` carry dual-track install instructions. +- **10 slash commands** at `commands/ars-*.md` (Phase 2.1, PR #69) mapping `MODE_REGISTRY.md` entries to `/ars-` triggers. Model routing is pinned in each command's frontmatter — `opus` for `full` and `revision-coach` (architectural / review-interpretation depth), `sonnet` for the other 8. No Haiku per project policy. +- **3 plugin-shipped agents** at `agents/*_agent.md` (Phase 2.1, PR #69) as relative symlinks to the v3.6.7-hardened downstream agents in `deep-research/agents/`: `synthesis_agent`, `research_architect_agent`, `report_compiler_agent`. Underscore filenames preserved to keep `scripts/check_v3_6_7_pattern_protection.py` hard-pinned paths and INV-3 manifest-confined Clause 1 invariant intact. Symlinks (not copies) preserve a single source of truth and prevent the Pattern C3 attack surface that v3.6.7 §6 inversion sweep + INV-1/2/3 lint closes. (Materialized to real byte-identical copies in #413 — relative symlinks break Windows checkouts without `core.symlinks` and zip-download installs; the single-source guarantee moved to the `scripts/check_agents_mirror_sync.py` byte-equality CI lint.) +- **`model: inherit`** added to those three source agent frontmatters. Inherit chosen over pinning `sonnet` so an opus session running ARS full pipeline keeps opus agents (instead of being capped). The user's `~/.claude/hooks/warn-agent-no-model.sh` PreToolUse hook gates Haiku at the dispatching boundary, so `inherit` resolves through an already-Haiku-free model. +- **SessionStart announce hook** at `hooks/hooks.json` + `scripts/announce-ars-loaded.sh` (Phase 2.2, PR #70). When the plugin loads, the hook injects an `additionalContext` listing the 10 slash commands, the 3 plugin agents, and a token-budget pointer into the LLM's first turn. `startup` and `clear` source values get the full announce; `resume` and `compact` get a one-line ack to avoid burning context. Bash 3.2 compatible — runs on macOS stock `/bin/bash` with no `brew install bash` requirement. +- **Phase 2.2 scope reduction**: a `SubagentStop → run_codex_audit.sh` codex audit hook was scoped out for v3.7.0 due to a contract gap (the SubagentStop payload carries no stage/deliverable info, so the wrapper would have to half-infer required arguments) and an invoker-class boundary (`run_codex_audit.sh` lines 4–7 forbid same-session in-LLM invocation; PostToolUse fires inside the producing session). Real audit-hook integration deferred to a future release when ARS gains a stage/deliverable propagation contract. See `docs/design/2026-04-30-ars-v3.7.0-plugin-packaging-roadmap.md` Update note 2026-05-05 (Phase 2.2 scope reduction). +- **`docs/PERFORMANCE.md` + `.zh-TW.md`** gain a "v3.7.0 Plugin agents and model routing" subsection explaining the inherit semantics and current 3-agent scope boundary. +- **Codex review chain across the three PRs**: 8 inline iterative rounds + 3 fresh PR-level rounds, all converging to 0 P0/P1/P2 findings before merge. The Phase 2.2 fresh PR review caught one P2 (unquoted `${CLAUDE_PLUGIN_ROOT}` breaking install paths with spaces) that the inline rounds missed — confirms the value of separating implementation review (inline) from contract review (fresh). +- **What did NOT change**: the four skill directories, all 25 modes, agent prompts, schema files, and lint contracts. Plugin packaging only adds new top-level surface (`commands/`, `agents/`, `hooks/`, `.claude-plugin/`, `skills/` symlink dir, three plugin-agent `model: inherit` frontmatter additions). Existing 4.3k clone-install users see no breaking change. + +### v3.6.8 (2026-05-03) — Generator-Evaluator Contract Gate (v3.6.6 spec ship) + +> Naming note: this release ships the **v3.6.6 generator-evaluator contract** spec +> and implementation. The v3.6.6 work landed after v3.6.7 due to project sequencing; +> the design doc retains the v3.6.6 internal naming for the contract gate version, +> while the suite release is tagged v3.6.8 to keep the CHANGELOG monotonic. + +- **Schema 13.1** (`shared/sprint_contract.schema.json`) extends Schema 13 with two new `mode` enum values (`writer_full` + `evaluator_full`), two new optional top-level fields (`pre_commitment_artifacts` writer-only, `disagreement_handling` evaluator-only), and 12 `allOf` branches enforcing reviewer- / writer- / evaluator-conditional gates. Existing reviewer contracts validate byte-equivalent under Schema 13.1 (§3.6 zero-touch promise). +- **Two new shipped contract templates** under `shared/contracts/writer/full.json` (D1–D7, F1/F4/F2/F3/F0) and `shared/contracts/evaluator/full.json` (D1–D5, F1/F2/F3/F6/F4/F5/F0). Promoted from design-time artefacts on the spec branch to live shipped status atomically with the Schema 13.1 upgrade. +- **Two-phase orchestration** inside `academic-paper full`: Phase 4 splits into Phase 4a (writer paper-blind pre-commitment) + Phase 4b (writer paper-visible drafting + self-scoring); Phase 6 splits into Phase 6a (evaluator paper-blind pre-commitment) + Phase 6b (evaluator paper-visible scoring + decision). Phase-numbered `` / `` data delimiters mirror the v3.6.2 reviewer pattern. Lint count summary: writer 3+4 / evaluator 5+5 / reviewer 5+6 (reviewer remains zero-touch). +- **`academic-paper` SKILL + agent files** gain a verbatim `## v3.6.6 Generator-Evaluator Contract Protocol` block (101 lines in SKILL.md plus 47 lines in `draft_writer_agent.md` + 57 lines in `peer_reviewer_agent.md`). SKILL.md also adds a new `## Known limitations` section carrying graceful-degradation + cross-session resume forward notes for v3.6.7+. +- **Validator extensions**: `scripts/check_sprint_contract.py` SC-* mode-gating audit (SC-5 + SC-11 reviewer-only; SC-9 extended across all three mode families). 17 new tests bring the validator unit-test count from 54 to 71 (positive + 5 schema-branch negative + 2 §3.6 reviewer regression + 6 mode-gating tests). +- **Manifest CI lint**: `scripts/check_v3_6_6_ab_manifest.py` enforces §6.2 manifest schema + §6.5 git-tracked invariants on `tests/fixtures/v3.6.6-ab/manifest.yaml`. `.github/workflows/spec-consistency.yml` extends the sprint contract validation loop to iterate writer + evaluator template directories alongside the existing reviewer loop, plus runs the new manifest CI lint. +- **A/B evidence fixture stub** at `tests/fixtures/v3.6.6-ab/` (30 files): manifest + README + 6 paper-A inputs/baseline + 1 paper-C inputs/baseline + Stage 3 reviewer excerpt + 6 codex-judge baseline placeholders. Real fixture data populates in follow-up commits before the implementation work fully completes. + +### v3.6.7 (2026-04-30) — Downstream-Agent Pattern Protection (Step 1+2) + +- **Three downstream agents hardened against 13 of 17 documented hallucination/drift patterns**: `synthesis_agent` (A1–A5 narrative-side), the survey-designer mode of `research_architect_agent` (B1–B5 instrument-side), and the abstract-only mode of `report_compiler_agent` (C1–C3 publication-side). Each agent prompt now carries a `PATTERN PROTECTION (v3.6.7)` block. +- **Four reference files in `shared/references/`**: `irb_terminology_glossary.md`, `psychometric_terminology_glossary.md`, `protected_hedging_phrases.md`, `word_count_conventions.md`. The reference files carry operational contracts that the agent prompts cite by path. +- **Cross-model audit prompt template** at `shared/templates/codex_audit_multifile_template.md` with seven audit dimensions and a mandatory three-part Section 4(f) check for `report_compiler_agent` bundles. Failure of any sub-check is a P1 finding. +- **Static lint + 29-test mutation suite**: `scripts/check_v3_6_7_pattern_protection.py` enforces protection-clause presence and obligation-phrase shape; `scripts/test_check_v3_6_7_pattern_protection.py` preserves codex review evidence so future checker regressions surface in CI. Both are wired into `.github/workflows/spec-consistency.yml`. +- **Codex review history**: seven rounds of `gpt-5.5` + `xhigh` cross-model review reached SHIP-OK with zero P1+P2 findings. Step 6 (orchestrator runtime hooks) and Step 8 (synthetic eval case) ship in a follow-up PR. + +### v3.6.5 (2026-04-27) — Material Passport `literature_corpus[]` Consumer Integration + +- **Two Phase 1 literature consumers** wired: `deep-research/agents/bibliography_agent.md` and `academic-paper/agents/literature_strategist_agent.md`. Both follow the same five-step **corpus-first, search-fills-gap** flow when the passport carries a non-empty `literature_corpus[]` and the same four Iron Rules (Same criteria / No silent skip / No corpus mutation / Graceful fallback on parse failure). +- **PRE-SCREENED reproducibility block** in Search Strategy reports: enumerates included / excluded / skipped corpus entries, with F3 zero-hit note and F4a–F4f provenance reporting that compose around partial declaration of `obtained_via` / `obtained_at`. `final_included = pre_screened_included[] ∪ external_included[]` stays neutral — no provenance tags on bibliography entries or literature matrix rows. +- **Consumer protocol reference** at `academic-pipeline/references/literature_corpus_consumers.md` with the canonical PRE-SCREENED template, BAD/GOOD examples, four Iron Rules, and per-consumer reading instructions. +- **CI lint** `scripts/check_corpus_consumer_protocol.py` enforcing nine protocol invariants with manifest-driven consumer list (`scripts/corpus_consumer_manifest.json`). +- **Schema 9 caveat retired**: `shared/handoff_schemas.md` retired the v3.6.4 "Consumer-side integration deferred to v3.6.5+" caveat; replaced with backpointer to the consumer protocol. +- Presence-based, no schema change, no new env flag. Parse failures fall back to external-DB-only flow with a `[CORPUS PARSE FAILURE]` surface. `citation_compliance_agent` corpus integration deferred (target version TBD post-v3.8). +- No breaking changes. Existing user adapters work without modification. + +### v3.6.4 (2026-04-25) — Material Passport `literature_corpus[]` Input Port + +- **`literature_corpus[]` field** added to Schema 9 as an optional input port for user-owned literature. Each entry conforms to `shared/contracts/passport/literature_corpus_entry.schema.json` (CSL-JSON authors, year, title, source_pointer + private optional `abstract` / `user_notes`). +- **Language-neutral adapter contract** at `academic-pipeline/references/adapters/overview.md`: any program (any language) reading a user corpus source can produce conformant `passport.yaml` + `rejection_log.yaml`. Fail-soft entry-level errors, fail-loud adapter-level errors, deterministic ordering. +- **Three reference Python adapters** under `scripts/adapters/`: `folder_scan.py` (filesystem of PDFs), `zotero.py` (Better BibTeX JSON export), `obsidian.py` (vault frontmatter). Starting points only; users are expected to write their own adapters for non-reference sources. +- **Rejection log contract** at `shared/contracts/passport/rejection_log.schema.json` with closed enum of categorical reason values; always emitted (empty when no rejections). +- **CI gates**: `scripts/check_literature_corpus_schema.py` validates schemas + adapter examples; `scripts/sync_adapter_docs.py --check` prevents schema→docs drift; new `pytest.yml` workflow runs `scripts/adapters/tests/` on path-filtered triggers. +- **Input-port-only at v3.6.4**: v3.6.4 shipped the schema and adapter contract without consumer integration. `bibliography_agent` and `literature_strategist_agent` were wired in v3.6.5. +- No breaking changes. + +### v3.6.3 (2026-04-23) — Opt-in Passport Reset Boundary + +- **Opt-in passport reset boundary** (`ARS_PASSPORT_RESET=1`). Promotes every FULL checkpoint to a context-reset boundary. New `resume_from_passport=` mode lets users resume in a fresh Claude Code session from the Material Passport ledger alone. `systematic-review` mode with the flag ON makes reset mandatory at every FULL checkpoint; other modes treat reset as the flag-gated default. Flag OFF preserves pre-v3.6.3 behavior byte-for-byte. +- Schema 9 gains an append-only `reset_boundary[]` ledger with two entry kinds (`kind: boundary` + `kind: resume`). Hash uses JSON Canonical Form + SHA-256 with canonical placeholder for self-reference safety. Optional `pending_decision` handles MANDATORY branch choices. +- New `scripts/check_passport_reset_contract.py` CI lint: every mention of the flag must co-locate a pointer to the authoritative protocol doc. +- Protocol doc: `academic-pipeline/references/passport_as_reset_boundary.md`. +- `docs/PERFORMANCE.md` updated with long-running-session guidance. +- No breaking changes. Flag default is OFF. + +### v3.6.2 (2026-04-23) — Reviewer Sprint Contract Hard Gate + +v3.6.2 introduces Schema 13 sprint contracts and a hard-gate orchestration that forces reviewers to pre-commit their scoring plan before reading the paper. Reviewer-only first test case; writer/evaluator deferred to v3.6.4. See CHANGELOG. + +- **Schema 13 sprint contract** with `panel_size`, `acceptance_dimensions`, `failure_conditions` (with `severity` precedence + panel-relative `cross_reviewer_quantifier`), `measurement_procedure`, optional `override_ladder`, bounded `agent_amendments`. Validator: `scripts/check_sprint_contract.py`. +- **Two-call hard gate.** Reviewers run paper-content-blind Phase 1 + paper-visible Phase 2; Phase 1 output is wrapped in `...` data delimiter to narrow the self-injection surface. +- **Synthesizer three-step mechanical protocol.** Build cross-reviewer matrix → evaluate each `failure_condition` with panel-relative quantifier + recognised expression vocabulary → resolve precedence by `severity`. Forbidden-ops list explicit in `editorial_synthesizer_agent`. +- **Two reviewer templates ship** (`shared/contracts/reviewer/full.json` panel 5; `shared/contracts/reviewer/methodology_focus.json` panel 2). `reviewer_re_review`, `reviewer_calibration`, `reviewer_guided` are reserved in the schema enum but ship without contract templates in v3.6.2; they retain pre-v3.6.2 behaviour. `reviewer_quick` is excluded from the enum entirely. +- `academic-paper-reviewer` SKILL version: `1.8.1 → 1.9.0`. `academic-pipeline` SKILL version: `3.5.1 → 3.6.2` (suite-version invariant). Suite version bumped to `3.6.2`. +- See spec [`docs/design/2026-04-23-ars-v3.6.2-sprint-contract-design.md`](docs/design/2026-04-23-ars-v3.6.2-sprint-contract-design.md) and protocol [`academic-paper-reviewer/references/sprint_contract_protocol.md`](academic-paper-reviewer/references/sprint_contract_protocol.md). + +### v3.5.1 (2026-04-22) — Opt-in Socratic Reading-Check Probe + +v3.5.1 adds an opt-in honesty probe to the Socratic Mentor (`ARS_SOCRATIC_READING_PROBE=1`). Default off. See CHANGELOG. + +- **Opt-in reading-check probe**: when `ARS_SOCRATIC_READING_PROBE=1` is set, the Socratic Mentor fires a one-time honesty probe during goal-oriented sessions where the user has cited a specific paper. Decline is logged without penalty. Outcome flows into the Research Plan Summary and Stage 6 AI Self-Reflection Report. No new agent, no schema change. +- `deep-research` SKILL version: `2.9.0 → 2.9.1`. `academic-pipeline` SKILL version: `3.5.0 → 3.5.1`. Suite version bumped to `3.5.1`. + +### v3.5.0 (2026-04-21) — Collaboration Depth Observer + +- **New agent**: `collaboration_depth_agent` in `academic-pipeline` (Agent Team grows from 3 to 4). Invoked at every FULL/SLIM checkpoint and at pipeline completion; scores user-AI collaboration against a 4-dimension rubric. **Advisory only — never blocks progression.** MANDATORY checkpoints (Stages 2.5 / 4.5 integrity gates) do NOT invoke the observer. +- **New rubric**: [`shared/collaboration_depth_rubric.md`](shared/collaboration_depth_rubric.md) v1.0. Dimensions: Delegation Intensity, Cognitive Vigilance, Cognitive Reallocation, Zone Classification (Zone 1 / Zone 2 / Zone 3). Based on Wang, S., & Zhang, H. (2026). "Pedagogical partnerships with generative AI in higher education: how dual cognitive pathways paradoxically enable transformative learning." *International Journal of Educational Technology in Higher Education*, 23:11. DOI [10.1186/s41239-026-00585-x](https://doi.org/10.1186/s41239-026-00585-x). +- **Cross-model divergence flagged, not averaged**: when `ARS_CROSS_MODEL` is set the observer runs on both models; dimension disagreement > 2 points is reported rather than silently smoothed. `ARS_CROSS_MODEL_SAMPLE_INTERVAL` escape hatch for cost trade-off. +- **Short-stage guard**: stages with fewer than 5 user turns inject a static `insufficient_evidence` block instead of dispatching the full-model observer. +- **Anti-sycophancy discipline**: scores ≥ 7 require specific dialogue-turn citations; Zone 3 triggers re-audit; no motivational framing. +- `academic-pipeline` SKILL version: `3.3.0 → 3.4.0`. Suite version bumped to `3.5.0`. New lint `scripts/check_collaboration_depth_rubric.py` + 10 tests. + +### v3.4.0 (2026-04-20) — Compliance Agent + Schema 12 + +- **Compliance Agent** (shared): single mode-aware agent running PRISMA-trAIce 17 items (SR mode only) + RAISE 4 principles + 8-role matrix. Hooks existing Stage 2.5 / 4.5 Integrity Gates; tier-based block (Mandatory → block, HR → warn, R/O → info). Non-SR entries run principles-only, warn-only. +- **Schema 12 compliance_report** appended to Material Passport via `compliance_history[]` (append-only). +- **3-round user-override ladder** auto-injects `disclosure_addendum` into manuscript. No detection evasion possible. +- **Calibration with transparent reporting**, no hard FNR/FPR gate — self-consistent with `task_type: open-ended`. +- **Upstream freshness CI** warns on PRISMA-trAIce drift (non-blocking). +- **Long-running session docs**: Material Passport as cross-session resume mechanism. + +### v3.3.6 (2026-04-15) — README Streamlining + ARCHITECTURE doc + +- Added `docs/ARCHITECTURE.md` as the single source of truth for pipeline structure (flow, matrix, data-access, dependency graph, quality gates, modes). Merged into main via PR #18. +- Added `docs/SETUP.md` (prerequisites, API keys, Pandoc/tectonic, cross-model verification, installation methods) and `docs/PERFORMANCE.md` (token budgets, recommended Claude Code settings). README links to both instead of inlining them. +- Streamlined README: removed the ASCII pipeline diagram and 16-point key-feature list (superseded by ARCHITECTURE.md); Skill Details section now anchors version numbers and points readers to ARCHITECTURE.md §3 for per-agent rosters. +- Note: no functional change to any skill. Pure documentation reorganization. Suite version bumped to `3.3.6`. + +### v3.3.5 (2026-04-15) +- Added `benchmark_report.schema.json` + `repro_lock` optional block on Material Passport. Both ship with pattern docs, lints, and examples. First formal Python dev dep manifest (`requirements-dev.txt`). + +### v3.3.4 (2026-04-15) — README Changelog Sync Patch + +- Synced the embedded changelog sections in `README.md` and `README.zh-TW.md` so they include the missing `v3.3.3` and `v3.3.2` release summaries. +- Extended `scripts/check_spec_consistency.py` so future README changelog drift fails CI. +### v3.3.3 (2026-04-15) — Release Prep + Lint Hardening + +- Hardened SKILL frontmatter linting: missing closing `---` fences now fail cleanly instead of being parsed as valid YAML. +- Frontmatter that parses as valid YAML but not as a mapping now reports a readable error instead of crashing. +- Fixed the broken showcase link for the post-publication audit report in both READMEs. +- Added README relative-link validation to the spec consistency check so dead links fail CI. +- Aligned the DOCX output contract across the docs: direct `.docx` generation is Pandoc-dependent, with Markdown + conversion instructions as fallback. +- Prepared the `v3.3.3` release: suite version bump, `academic-paper` -> v3.0.2, `academic-pipeline` -> v3.2.2. + +### v3.3.2 (2026-04-15) — Data Access Levels + Task Type Metadata + +- Added `metadata.data_access_level` to all top-level `SKILL.md` files with enforced vocabulary: `raw`, `redacted`, `verified_only`. +- Added `metadata.task_type` to all top-level `SKILL.md` files with enforced vocabulary: `open-ended`, `outcome-gradable`. +- Added lint scripts and unit tests for both metadata fields, wired into the GitHub Actions spec consistency workflow. +- Added `shared/ground_truth_isolation_pattern.md` and linked the new vocabulary from `shared/handoff_schemas.md`. + +### v3.3.1 (2026-04-14) — Spec Consistency Patch + +- Synced README, `.claude/CLAUDE.md`, `MODE_REGISTRY.md`, and `SKILL.md` files to the current mode counts and published skill versions. +- Corrected cross-model wording: integrity sample checks and independent DA critique are implemented today; sixth-reviewer peer review remains planned. +- Clarified adaptive checkpoint semantics so SLIM checkpoints still wait for explicit user confirmation. +- Reaffirmed that Stage 2.5 and Stage 4.5 integrity gates cannot be skipped. +- Added a lightweight spec consistency check and GitHub Actions workflow to catch future drift. + +### v3.3 (2026-04-09) — PaperOrchestra-Inspired Enhancements + +Integrates techniques from [PaperOrchestra](https://arxiv.org/abs/2604.05018) (Song, Song, Pfister & Yoon, 2026, Google). + +- **Semantic Scholar API Verification** — Tier 0 programmatic reference existence check via S2 API. Levenshtein >= 0.70 title matching, DOI mismatch detection, bibliography deduplication via S2 IDs. Graceful degradation if API unavailable. +- **Anti-Leakage Protocol** — Knowledge Isolation Directive prioritizes session materials over LLM parametric memory. Flags `[MATERIAL GAP]` for missing content instead of filling from memory. Reduces Mode 5/6 failure risk. +- **VLM Figure Verification** (optional) — Closed-loop verification of rendered figures using vision-capable LLM. 10-point checklist, max 2 refinement iterations. +- **Score Trajectory Protocol** — Per-dimension rubric score delta tracking across revision rounds (7 dimensions). Detects regressions (delta < -3) and triggers mandatory checkpoint. +- **Stage 2 Parallelization** — Visualization and argument building can run in parallel after outline completion. +- New versions: deep-research v2.8, academic-paper v3.0, academic-pipeline v3.2 + +### v3.2 (2026-04-09) — Lu 2026 Nature Integration + +Integrates insights from Lu et al. (2026, *Nature* 651:914-919) — the first end-to-end autonomous AI research system to pass blind peer review. + +- **7-mode AI Research Failure Mode Checklist** — blocks pipeline at Stage 2.5/4.5 on suspected implementation bugs, hallucinated results, shortcut reliance, bug-as-insight, methodology fabrication, frame-lock. Extends existing 5-type citation hallucination taxonomy. +- **Reviewer Calibration Mode** (academic-paper-reviewer v1.8) — opt-in FNR/FPR/balanced-accuracy measurement against user-supplied gold set. 5× ensembling, cross-model default-on, session-scoped confidence disclosure. +- **Disclosure Mode** (academic-paper v2.9) — venue-specific AI-usage statement generator. v1 covers ICLR, NeurIPS, Nature, Science, ACL, EMNLP. +- **Early-Stopping Criterion** (academic-pipeline v3.1) — convergence check + budget transparency at pipeline start. +- **Fidelity-Originality Mode Spectrum** — classifies all modes across 3 skills per Lu 2026 Fig 1c. +- New versions: academic-paper v2.9, academic-paper-reviewer v1.8, academic-pipeline v3.1 + +### v3.1.1 (2026-04-09) — IS Senior Scholars' Basket of 11 + +External contributions: [@mchesbro1](https://github.com/mchesbro1) originally proposed and drafted the IS Basket of 8 journals ([Issue #5](https://github.com/Imbad0202/academic-research-skills/issues/5)); [@cloudenochcsis](https://github.com/cloudenochcsis) extended it to the full Senior Scholars' Basket of 11 ([Issue #7](https://github.com/Imbad0202/academic-research-skills/issues/7), [PR #8](https://github.com/Imbad0202/academic-research-skills/pull/8)). Updated `academic-paper-reviewer/references/top_journals_by_field.md` Section 7, adding *Decision Support Systems*, *Information & Management*, and *Information and Organization*. Source: [AIS Senior Scholars' List of Premier Journals](https://aisnet.org/research/seniorscholarsbasket/). + +### v3.1 (2026-04-06) — Anti-Context-Rot + Cognitive Frameworks + Lean Size + +Inspired by patterns from [aspi6246/Claude-Code-Skills-for-Academics](https://github.com/aspi6246/Claude-Code-Skills-for-Academics). + +**Wave 1: Anti-Context-Rot Anchors** +- 29 explicit Anti-Patterns across all 4 skills (7-8 per skill, tabular format with "Why It Fails" + "Correct Behavior") +- 22 IRON RULE markers on critical rules that must not be violated even in long conversations +- Read-only constraint on academic-paper-reviewer (reviewers cannot modify the manuscript) + +**Wave 2: Traceability + Cognitive Frameworks + Reinforcement** +- R&R Traceability Matrix (Schema 11): adds "Author's Claim" and "Verified?" columns to re-review output, enabling independent verification of revision claims +- 3 cognitive framework reference files teaching agents "how to think" not just "what to do": + - `argumentation_reasoning_framework.md` — Toulmin model, Bradford Hill causal reasoning, inference to best explanation, epistemic status classification + - `review_quality_thinking.md` — three lenses (internal validity, external validity, contribution), common reviewer traps, calibration questions + - `writing_judgment_framework.md` — clarity test, reader's journey, discipline-specific voice, revision decision matrix +- Mid-conversation reinforcement protocol: stage-specific IRON RULE + Anti-Pattern reminders at every pipeline transition +- Self-check questions at every FULL checkpoint (citation integrity, sycophantic concession, quality trajectory, scope discipline, completeness) + +**Wave 3: Lean Skill Size** +- SKILL.md total size reduced from 142KB to 85KB (−40%) by extracting detailed protocols to `references/` files +- ~15 new reference files created (re-review protocol, guided mode, systematic review, process summary, external review, etc.) +- All IRON RULE markers preserved in SKILL.md; detailed content loaded on demand +- New versions: deep-research v2.7, academic-paper v2.8, academic-paper-reviewer v1.7, academic-pipeline v3.0 + +### v3.0 (2026-04-03) — Anti-Sycophancy + Intent Detection + Dialogue Health +- **Devil's Advocate Concession Threshold** (deep-research + academic-paper-reviewer): DA must score rebuttals 1-5 before responding. Concession only at ≥4. No consecutive concessions. Concession rate tracking. Frame-lock detection after each checkpoint. +- **Attack Intensity Preservation** (academic-paper-reviewer): DA does not soften under pushback. Rebuttal assessment protocol with explicit deflection detection. Anti-sycophancy rules prevent persistent pushback from being treated as valid evidence. +- **Intent Detection Layer** (deep-research socratic): Classifies user intent as exploratory vs. goal-oriented. Exploratory mode disables auto-convergence, raises max rounds, prohibits premature closure. Re-assesses every 3 turns. +- **Dialogue Health Indicator** (deep-research socratic): Silent self-check every 5 turns for persistent agreement, conflict avoidance, premature convergence. Auto-injects challenges when agreement pattern detected. +- **Cross-Model Verification Protocol** (shared, optional): Use GPT-5.4 Pro or Gemini 3.1 Pro for integrity verification sample cross-checks and independent DA critique. Sixth-reviewer peer review remains planned, not yet implemented. Activated by setting `ARS_CROSS_MODEL` env var — without it, everything works as before. See `shared/cross_model_verification.md` for full setup guide, API patterns, and cost estimates. +- **AI Self-Reflection Report** (academic-pipeline Stage 6): Post-pipeline self-assessment of AI behavioral patterns — DA concession rate, checkpoint skip rate, health alerts, sycophancy risk rating (LOW/MEDIUM/HIGH), frame-lock incidents, convergence pattern analysis. Includes irony caveat: "this self-reflection is itself produced by the same AI that may have been sycophantic." +- Origin: Discovered through a 4-round dialectic experiment where the DA conceded too quickly, the Socratic Mentor tried to converge prematurely, and the entire debate stayed locked in a frame the human set. +- Versions: deep-research v2.5, academic-paper-reviewer v1.5, academic-pipeline v2.8 + +### v2.9.1 (2026-04-03) — Skill Metadata +- Added `status: active` and `related_skills` cross-references to all 4 SKILL.md frontmatters. +- Enables skill discovery tools and cross-skill navigation across `deep-research` ↔ `academic-paper` ↔ `academic-paper-reviewer` ↔ `academic-pipeline`. + +### v2.9 (2026-03-27) — Style Calibration + Writing Quality Check +- **Style Calibration** (academic-paper intake Step 10, optional): Provide 3+ past papers and the pipeline learns your writing voice — sentence rhythm, vocabulary preferences, citation integration style. Applied as a soft guide during drafting; discipline conventions always take priority. Priority system: discipline norms (hard) > journal conventions (strong) > personal style (soft). See `shared/style_calibration_protocol.md` +- **Writing Quality Check** (`academic-paper/references/writing_quality_check.md`): Writing quality checklist applied during draft self-review. 5 categories: AI high-frequency term warnings (25 terms), punctuation pattern control (em dash ≤3), throat-clearing opener detection, structural pattern warnings (Rule of Three, uniform paragraphs, synonym cycling), and burstiness checks (sentence length variation). These are good writing rules — not detection evasion +- **Style Profile** carried through academic-pipeline Material Passport (Schema 10 in `shared/handoff_schemas.md`) +- **deep-research** report compiler also consumes both features optionally +- Versions: academic-paper v2.5, deep-research v2.4, academic-pipeline v2.7 + +### v2.8 (2026-03-22) — SCR Loop Phase 1: State-Challenge-Reflect +- **Socratic Mentor Agent** (deep-research + academic-paper): SCR (State-Challenge-Reflect) protocol integration + - **Commitment Gates**: Collect user predictions before presenting evidence at each layer/chapter transition + - **Certainty-Triggered Contradiction**: Detect high-confidence language ("obviously", "clearly") and introduce counterpoints + - **Adaptive Intensity**: Track commitment accuracy, dynamically adjust challenge frequency + - **Self-Calibration Signal (S5)**: New convergence signal tracking user's self-calibration growth across dialogue + - **SCR Switch**: Users can say "skip the predictions" to disable or "turn predictions back on" to re-enable mid-dialogue; Socratic questioning continues normally +- `deep-research/references/socratic_questioning_framework.md`: SCR Overlay Protocol mapping SCR phases to Socratic functions +- Added `CHANGELOG.md` + +### v2.7 (2026-03-09) — Integrity Verification v2.0: Anti-Hallucination Overhaul +- **integrity_verification_agent v2.0**: Anti-Hallucination Mandate (no AI memory verification), eliminated gray-zone classifications (VERIFIED/NOT_FOUND/MISMATCH only), mandatory WebSearch audit trail for every reference, Stage 4.5 fresh independent verification, Gray-Zone Prevention Rule +- **Known Hallucination Patterns**: 5-type taxonomy (TF/PAC/IH/PH/SH) from GPTZero × NeurIPS 2025 study, 5 compound deception patterns, real-world case study, literature statistics +- **Post-publication audit**: Full WebSearch verification of all 68 references found 21 issues (31% error rate) that passed 3 rounds of integrity checks — proving the necessity of external verification +- **Paper corrections**: Removed 4 fabricated references, fixed 6 author errors, corrected 7 metadata errors, fixed 2 format issues + +### v2.6.2 (2026-03-09) — Intent-Based Mode Activation +- **deep-research**: Socratic mode now uses **intent-based activation** instead of keyword matching. Works in any language — detects meaning (e.g., "user wants guided thinking") rather than matching specific strings. +- **academic-paper**: Plan mode now uses **intent-based activation**. Detects intent signals like "user is uncertain how to start" or "user wants step-by-step guidance" in any language. +- Both modes now have a **default rule**: when intent is ambiguous, prefer `socratic`/`plan` over `full` — safer to guide first. +- Two-layer architecture: Layer 1 (skill activation) uses bilingual keywords for matching confidence; Layer 2 (mode routing) uses language-agnostic intent signals. + +### v2.6.1 (2026-03-09) — Bilingual Trigger Keywords +- **deep-research**: Added Traditional Chinese trigger keywords for general activation and Socratic mode. +- **academic-paper**: Added Traditional Chinese trigger keywords and Plan Mode trigger section. +- Both mode selection guides now include bilingual examples and Chinese-specific misselection scenarios. + +### v2.6 / v2.4 / v1.4 (2026-03-08) — 15+ Improvements +- **deep-research v2.3**: New systematic-review / PRISMA mode (7th); 3 new agents (risk_of_bias, meta_analysis, monitoring); PRISMA protocol/report templates; Socratic convergence criteria (4 signals + auto-end); Quick Mode Selection Guide +- **academic-paper v2.4**: 2 new agents (visualization, revision_coach); revision tracking template with 4 status types; citation format conversion (APA↔Chicago↔MLA↔IEEE↔Vancouver); statistical visualization standards; Socratic convergence criteria; revision recovery example; **LaTeX output hardening** — mandatory `apa7` document class, text justification fix (`ragged2e` + `etoolbox`), table column width formula, bilingual abstract centering, standardized font stack (Times New Roman + Source Han Serif TC VF + Courier New), PDF via tectonic only +- **academic-paper-reviewer v1.4**: Quality rubrics with 0-100 scoring and behavioral indicators; decision mapping (≥80 Accept, 65-79 Minor, 50-64 Major, <50 Reject); Quick Mode Selection Guide +- **academic-pipeline v2.6**: Adaptive checkpoint system (FULL/SLIM/MANDATORY); Phase E Claim Verification in integrity checks; Material Passport for mid-entry provenance; cross-skill mode advisor (14 scenarios); team collaboration protocol; enhanced handoff schemas (9 schemas); integrity failure recovery example + +### v2.4 / v1.3 (2026-03-08) +- **academic-pipeline v2.4**: New Stage 6 PROCESS SUMMARY — auto-generates structured paper creation process record (MD → LaTeX → PDF, bilingual); mandatory final chapter: **Collaboration Quality Evaluation** with 6 dimensions scored 1–100 (Direction Setting, Intellectual Contribution, Quality Gatekeeping, Iteration Discipline, Delegation Efficiency, Meta-Learning), honest feedback, and improvement recommendations; pipeline expanded from 9 to 10 stages + +### v2.3 / v1.3 (2026-03-08) +- **academic-pipeline v2.3**: Stage 5 FINALIZE now prompts for formatting style (APA 7.0 / Chicago / IEEE); PDF must compile from LaTeX via `tectonic` (no HTML-to-PDF); APA 7.0 uses `apa7` document class (`man` mode) with XeCJK for bilingual CJK support; font stack: Times New Roman + Source Han Serif TC VF + Courier New + +### v2.2 / v1.3 (2025-03-05) +- **Cross-Agent Quality Alignment**: unified definitions (peer-reviewed, currency rule, CRITICAL severity, source tier) across all agents +- **deep-research v2.2**: synthesis anti-patterns, Socratic auto-end conditions, DOI+WebSearch verification, enhanced ethics integrity check, mode transition matrix +- **academic-paper v2.2**: 4-level argument scoring, plagiarism screening, 2 new failure paths (F11 Desk-Reject Recovery, F12 Conference-to-Journal), Plan→Full mode conversion +- **academic-paper-reviewer v1.3**: DA vs R3 role boundaries, CRITICAL finding criteria, consensus classification (4/3/SPLIT/DA-CRITICAL), confidence score weighting, Asian & Regional Journals reference +- **academic-pipeline v2.2**: checkpoint confirmation semantics, mode switching matrix, failure fallback matrix, state ownership protocol, material version control + +### v2.0.1 (2026-03) +- **Simplify 4 SKILL.md** (-371 lines, -16.5%): remove cross-skill duplication, inline templates → file references, redundant routing tables, duplicate mode selection sections +- Fix revision loop cap contradiction between academic-paper and academic-pipeline + +### v2.0 (2026-02) +- **academic-pipeline v2.0**: 5→9 stages, mandatory integrity verification, two-stage review, Socratic revision coaching, reproducibility guarantees +- **academic-paper-reviewer v1.1**: +Devil's Advocate Reviewer (7th agent), +re-review mode (verification), +post-review Socratic coaching +- New agent: `integrity_verification_agent` — 100% reference/data verification with audit trail +- New agent: `devils_advocate_reviewer_agent` — 8-dimension thesis challenger +- Output order: MD → DOCX via Pandoc when available (else instructions) → ask LaTeX → confirm → PDF + +### v1.0 (2026-02) +- Initial release +- deep-research v2.0 (10 agents, 6 modes including socratic) +- academic-paper v2.0 (10 agents, 8 modes including plan) +- academic-paper-reviewer v1.0 (6 agents, 4 modes including guided) +- academic-pipeline v1.0 (orchestrator) diff --git a/raw/articles/chinatalk-transfer-station-economy-2026.md b/raw/articles/chinatalk-transfer-station-economy-2026.md new file mode 100644 index 0000000..992bd70 --- /dev/null +++ b/raw/articles/chinatalk-transfer-station-economy-2026.md @@ -0,0 +1,126 @@ +--- +source_url: https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in +ingested: 2026-06-28 +sha256: c3cd880788b2b8770ef825c8a396122e874bf9106b2be87bff0983d21af8ac24 +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520677596893806682' + author_id: '890908900520505354' + posted_at: 2026-06-28T06:29:51.923000000Z + message_excerpt: 'https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in' +--- + +### The Transfer Station Economy, Explained + +*[Zilan Qian](https://open.substack.com/users/288758452-zilan-qian?utm_source=mentions) is a research associate at the Oxford China Policy Lab and holds a Master’s degree in Social Science of the Internet from the University of Oxford.* + +On April 23, 2026, the White House [released](https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf) a memo warning that Chinese entities were running “industrial-scale” distillation campaigns against American frontier AI models, leveraging “tens of thousands of proxy accounts” to evade detection. In February 2026, Anthropic similarly [reported](https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks) on Chinese labs’ coordinated distillation attacks using “a single proxy network managed more than 20,000 fraudulent accounts”. Both cases see “proxy” — the middlemen between model users and model providers — as a purposeful design by a selective Chinese frontier labs to systematically extract US AI models. + +Regardless of whether Chinese labs [rely](https://www.interconnects.ai/p/how-much-does-distillation-really) on distillation to “catch up”, both documents misread the proxy economy they’re describing. Underneath the handful of labs sits a much larger market, one that has been operating in public on GitHub, Taobao, Twitter, and Telegram. It is a grey economy of API proxies (commonly called “transfer stations,” 中转站) that lets Chinese developers access Anthropic’s models at as low as 10% of the official price. The participants extend far beyond selective experienced AI researchers, and the motivations are much broader than building a frontier model to catch up. Everyone who wants to use more advanced AI models or tools, be they university professors and students, tech workers, individual developers, or hobbyists, uses API proxies. The logs they generate may have become a commodity, traded for purposes ranging from model training to targeted fraud. + +Meanwhile, every layer of control frontier US AI companies have added (geoblocking, phone verification, credit card requirements, and now live biometric KYC checks) has produced a corresponding layer of evasion infrastructure. These new SMS farms and biometric harvesting operations have implications that extend beyond geopolitics into how frontier AI safety frameworks are designed. + +Building on my [2025 ChinaTalk piece](https://www.chinatalk.media/p/the-grey-market-for-american-llms) on accessing banned American models in China, this update zooms in on the transfer station economy specifically: how it is structured, how it monetizes, and what it reveals about the limits of access blocking and account monitoring as AI governance tools. Unlike 2025’s grey market, however, the 2026 story does not stop at the border between Chinese users and American AI model providers. The transfer station economy exposes blind spots in AI safety frameworks designed to prevent harms that extend beyond the US-China rivalry, from misuse by malicious actors to the erosion of provider traceability, while feeding into criminal markets that exploit ordinary people — many already disadvantaged — caught in the supply chain. + +To illustrate how a transfer station works, let’s take Anthropic, the company with the most rigorous geo-blocking mechanism, and whose models are [very popular](https://recodechinaai.substack.com/p/forget-openai-chinas-ai-labs-are) among Chinese developers, as an example. + +![](https://substackcdn.com/image/fetch/$s_!hqCG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5378825-3d4d-4b23-988d-c5300198963c_1024x1024.jpeg) + +A meme circulated on the Chinese internet: “Do you think you are smarter than Claude?” + +## Geo-blocking and Know-Your-Customer (KYC) + +On the map of Anthropic’s supported countries, China is conspicuously absent, and on the Chinese internet, so is Anthropic – technically speaking. In reality, neither Anthropic’s blockage nor the Great Firewall stops Chinese users from accessing Claude and Claude Code. Claude models [have thrived](https://www.chinatalk.media/p/the-grey-market-for-american-llms) on e-commerce apps like Taobao despite supposed platform and government censorship since 2025, and Singapore, with a population smaller than that of New York City, “surprisingly” [leads](https://www.businesstimes.com.sg/companies-markets/singapore-leads-global-capita-use-anthropics-claude-ai-businesses-ramp-adoption-gic) global per capita use of Anthropic’s Claude in April 2026. + +![](https://substackcdn.com/image/fetch/$s_!vCde!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c7be0b-f286-4b83-a50f-afba15e1ae22_742x764.png) + +Chinese developers joked about the report that Singapore is the top token consumption of Claude on Twitter, implying that this is because the Chinese are routing to Singapore to use the model. “We are all Singaporean from time to time.” “Every day I self-assign my nationality.” “Isn’t it because we all use Singapore’s node?” “Seems that many companies are using Singapore’s node.” + +The Chinese government is not today especially motivated to curb Chinese developers’ access to advanced US models. Anthropic, on the other hand, is serious about it, with its multiple layers of mechanisms to block users in mainland China. At the most basic level, account registration requires phone numbers, overseas credit cards, and matching billing addresses. On September 5, 2025, Anthropic further [prohibited](https://www.anthropic.com/news/updating-restrictions-of-sales-to-unsupported-regions) access from any entity more than 50% owned, directly or indirectly, by companies headquartered in unsupported regions like China, regardless of where that entity operates. This closes the subsidiary loophole that had allowed Chinese-backed firms in foreign countries to retain API access. + +The most recent measure arrived in April 2026. Anthropic [began requiring](https://support.claude.com/en/articles/14328960-identity-verification-on-claude) select users to verify their identity using a government-issued photo ID and a live selfie, making Claude the first major consumer AI platform to implement this level of identity checking. The rollout is selective and triggered by specific use cases or platform integrity flags. For Chinese users accessing Claude through VPN or other intermediaries, the new KYC policy is supposed to make it considerably harder to access Claude–even if Chinese users can fake phone numbers and addresses, they will theoretically have a hard time faking live selfies matched against a physical government document. + +In reality, however, Chinese people not only can access Claude and related tools, but most of the time they can purchase tokens at 10% of the original price. The magic lies in “transfer stations.” + +## What is a “Transfer Station (中转站)”? + +A transfer station (中转站) is what the Chinese developer ecosystem calls an API proxy–an overseas server that sits between a developer and Anthropic’s infrastructure. It accepts API requests, forwards them as if they originated from the transfer station’s location, and passes the response back. The user redirects their software to the proxy’s server instead of Anthropic’s, and pays the API proxy RMB via WeChat or Alipay. This sidesteps both the VPN and the overseas credit card needed for direct access. Prominent transfer stations are catalogued in [community repositories](https://github.com/mn-api/awesome-ai-proxy) and [ranked](https://www.aiapipk.com/) by real-time price and uptime. Below them, a longer tail of small and individual projects comes and goes. + +While this setup sounds functionally identical to legitimate Western API aggregators like OpenRouter, transfer stations operate in an entirely different universe of legality and trust. Legitimate aggregators exist to simplify developer workflows, charging standard rates based on transparent enterprise agreements. Transfer stations, conversely, are built explicitly for evasion, routing data through unaccountable middlemen. + +Just like providing VPN services or selling Claude on Taobao, a transfer station is technically not allowed in China. According to [China’s regulations on the AI services registry](https://oxfordchinapolicylab.org/research/china-s-ai-services-registry-system-a-complete-guide), AI services provided without filing and security assessment are illegal. But just as some small businesses can skip AI registration without punishment, so do most transfer stations. However, the bigger the business, the more unsafe it is to run. + +## The Supply Chain of Transfer Stations + +A transfer station is not a sole entity. It sits in the middle of a layered supply chain, with most participants never interacting with each other directly. + +Upstream are the resource providers: account merchants who bulk-register or acquire Anthropic accounts at scale; SMS verification platforms that supply the foreign phone numbers needed to pass sign-up checks; and, at the more technical end, reverse engineers who analyze Anthropic’s client code to find authentication shortcuts or detect when detection logic has changed. The payment infrastructure with card merchants and proxy networks also enables overseas billing from inside China. + +The upstream also tackles more sophisticated KYC regimes–either by AI or humans. AI services [have demonstrated](https://oecd.ai/en/incidents/2024-02-05-37e8) the ability to generate highly realistic fake IDs capable of bypassing identity verification on major platforms, and deepfake tools now [allow](https://www.kaspersky.com/blog/how-deepfakes-threaten-kyc/51987/) criminals to create digital clones that successfully pass biometric verification remotely.Even if the defender can successfully detect AI faking humans, a more labour-intensive method exists to find real humans. Agents travel to lower-income countries in Africa or Latin America to recruit real individuals willing to complete in-person verification. The Worldcoin black market [offered](https://www.biometricupdate.com/202305/worldcoin-may-have-a-biometric-data-black-market-problem) a documented precedent, with iris scans harvested from KYC merchants in Cambodia and Kenya, sold for under $30. + +![](https://substackcdn.com/image/fetch/$s_!_Bwe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3625c36a-42ec-4728-bc30-b2e5ce18855e_1000x178.png) + +Twitter account advertising KYC verification service. + +In the middle sits the transfer station itself: a software interface that receives users’ requests and forwards them to Anthropic as if they originated from a legitimate account, a payment integration (usually Alipay or WeChat), and the unglamorous operational layer that keeps it running — cycling accounts before they get flagged, balancing load across the pool, and continuously adapting to Anthropic’s abuse-detection updates. + +Downstream are the customers: individual developers using Codex or Claude Code, enterprises routing internal workflows through the proxy, application builders embedding the API in their own products, and secondary resellers who buy wholesale access and repackage it for individual customers on Taobao–as I [documented last year](https://www.chinatalk.media/p/the-grey-market-for-american-llms). + +Almost no one operates the full chain. Most participants own one or two links and monetise those well, resulting in a resilient, modular system. AI model providers can suspend individual operators, but the upstream account pools and downstream customer base remain intact. So long as there are developers who want access to Claude and identity black markets willing to supply the credentials, which are both durable features, a replacement can be stood up quickly. + +![](https://substackcdn.com/image/fetch/$s_!2b1Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae6d85a-cb3f-4ac7-9fa8-a2728fa284f8_875x1520.png) + +A screenshot circulated in a developer WeChat group joking about the supply chain to bypass Anthropic’s KYC; originally in Chinese (up), translation added by the author on the bottom + +## One Fish, Three Meals (一鱼三吃): How to Make Tokens Cheap + +The most curious thing, however, is not how to get access to Claude or Claude Code in China, but how to get it at a ridiculously low price–usually priced at 1 RMB per $1 of tokens — 70–90% below official prices. According to [public](https://x.com/search_ai/status/2029797569141035169) [discussions](https://post.smzdm.com/p/awwg9v8m), there are at least three ways a transfer station makes this possible–often described as “one fish, three meals (一鱼三吃)”. + +**Meal 1: The markup on access.** This is possible because of the upstream resource providers who can stack proxies using at least five relatively “innocent” tactics: + +- bulk-registering API accounts to farm Anthropic’s $5 free credit +- reselling unused quota from others’ accounts +- corporate/educational discount arbitrage +- “APImaxxing” — one $200 Max plan carved up among multiple users via tokens-per-hour quotas, exploiting the gap between Anthropic’s flat subscription price and the far higher cost of equivalent pay-per-token API access + +Beyond these, there is a darker upstream input: accounts purchased using [stolen or fraudulent credit cards](https://www.bbc.co.uk/news/technology-59983950) which can enter the proxy pool at effectively zero cost to the operator. How large this share is relative to the above four “innocent” tactics is difficult to verify, but the two markets likely share some infrastructure and personnel. + +**Meal 2: Swapping models and inflating tokens.** Because users’ inputs and model outputs are mediated through a proxy, users cannot verify which model their request was actually routed to. A user selects Opus 4.7, but the proxy can silently route to Sonnet, Haiku, or, in the worst case, GLM or Qwen, and fraudulently relabel the output. In a [recent paper](https://arxiv.org/pdf/2603.01919) from Germany’s CISPA Helmholtz Center for Information Security (which cited my article last year on [grey market](https://www.chinatalk.media/p/the-grey-market-for-american-llms)!), researchers audited 17 API proxies and found widespread model swapping–API proxy access to “Gemini-2.5” achieved only 37.00% on a medical benchmark, a staggering drop from the 83.82% performance of the official API. On the user end, the tell only comes on complex tasks, when the output feels off (often referred to as 降智, or “dumbed-down”), but there is no clean way to prove it. [Numerous](https://cloud.tencent.com/developer/article/2657436) [public](https://finance.sina.cn/stock/jdts/2026-04-18/detail-inhuwsse1638615.d.html?oid=msc&vt=4) [records](https://zhuanlan.zhihu.com/p/2017542841090483260) highlight concerns that certain API proxies have noticeably compromised model performance. These proxies are suspected of “diluting” (掺水) services by substituting premium frontier models with inferior tiers. + +Besides model swapping, overconsumption of tokens also makes the price per token cheaper, though at the expense of driving up the total cost. Some of it is [structural](https://zhuanlan.zhihu.com/p/2019895121458542180), as proxies that rotate accounts frequently destroy cache continuity as a side effect, forcing users to burn full-price tokens on context that would otherwise be nearly free. Some of it may be deliberate as the proxy providers try to milk more usage. The line between the two is difficult to draw from the outside. + +**Meal 3: The logs are the product.** This is perhaps the most important part as it intersects with data privacy and distillation. Every request that passes through a proxy — full prompt, full response, tool calls, iterations — is sitting on the proxy operator’s server. For AI coding agents, those logs contain long reasoning chains, real engineering decisions, repository context, and human-verified correct outputs. This makes them an ideal dataset for post-training: for supervised fine-tuning on real engineering tasks, and, where full reasoning traces are captured, for distilling Claude’s reasoning patterns into smaller models. Chinese developer communities assert this is happening in at [least](https://blog.csdn.net/m0_68727925/article/details/160189607) [some](https://blog.51cto.com/u_16099269/14568388) [cases](https://x.com/yan5xu/status/2029743983522631698), but whether proxy operators are systematically harvesting and selling these logs, and to whom, remains unverified. However, downstream distillation data does exist on the open web. [Several](https://huggingface.co/datasets/Crownelius/Opus-4.6-Reasoning-3300x) [datasets](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered) of Claude Opus 4.6 reasoning outputs circulate on HuggingFace with no clear source for the outputs. Theoretically, one can clean and sell similar distilled datasets to other model developers in China. + +The first two meals are useful for providing cheaper tokens cheaper than Anthropic officially charges, but to really make prices ridiculously low — at 10%, or even 5%, of the original price — one needs to eat the third meal. And as a Chinese saying goes, there is no free lunch in the world (天下没有免费的午餐). [Several](https://x.com/yan5xu/status/2029743983522631698) [Chinese](https://x.com/howie_serious/status/2031620123413590471) [developers](https://x.com/VincentLogic/status/2046434813125763527) have revealed that the markup business is just customer acquisition, and the log harvest is the actual margin. Users are simultaneously paying customers and unpaid data producers, selling their private data to proxy operators in exchange for a low price. Some also [warn](https://x.com/VincentLogic/status/2046434813125763527) of potential promotion, fraud, and even blackmail based on leaked users’ data from the proxy. To avoid privacy risks, some Chinese developers have also constructed their own Claude Code API proxy and [open-sourced](https://github.com/Wei-Shaw/claude-relay-service) the guidelines. + +## What Know-Your-Customer Cannot Know + +AI usage is gradually shifting from chatbot to tool use. With the rise of agent and token economy, the question of using US models is no longer only about access, but extends to cost-efficiency. This is because the Chinese AI ecosystem, regardless if it is frontier labs, university research groups, individual developers, or hobbyists, is capital-scarce. Meanwhile, the data generated by users through transfer stations demonstrably enters downstream markets, used variably for model training, data brokerage, or fraud. To the extent that distillation is part of that economy, the problem extends far beyond a handful of frontier actors that the government or AI companies in the US might expect. + +History teaches us that access blockage rarely stops determined users. They raise the cost of access, which in turn creates profitable markets for anyone with the expertise to lower it. The Great Firewall made VPN services a thriving cottage industry in China. KYC requirements bred an identification-faking economy, from domestic ID card resellers to biometric harvesting operations in Southeast Asia or Africa. Layered controls by frontier AI companies— geoblocking, phone verification, credit card requirements, and now live biometric checks — have produced the same effect. + +The story, however, goes beyond a “Anthropic/US versus China” framing. This points to an uncomfortable truth about access control, both in terms of geopolitical boundaries and beyond. How a geo-blocked developer walks around the controls is, structurally, the same methods plausibly employed by a terrorist to access a frontier AI model and make destructive bioweapons without being tracked. The access problem is both a unique geopolitical consideration and a shared safety concern. + +Today, [AI](https://www.frontiermodelforum.org/technical-reports/frontier-mitigations/) [safety](https://www.anthropic.com/news/detecting-countering-misuse-aug-2025) research treats system-level access control — in particular, detecting, monitoring, and account suspension for publicly accessible closed-weight models — as an important safeguard. In monitoring, developers control inference infrastructure, including flagging harmful inputs and outputs in real time. Detecting such as KYC requirements assumes that the provider can attribute behaviour to identifiable actors, and account suspension similarly assumes that suspending an account meaningfully denies access. However, US model providers do not control inference for Chinese users routing through a transfer station — the proxy operator does. When a harmful request arrives, rather than seeing the IP of the real user, AI model providers see that of the proxy. And when an account is banned, the upstream supply chain can easily set up a new proxy within hours. + +The problem compounds for more sophisticated monitoring tools. [Anthropic’s Clio system](https://www.anthropic.com/research/clio), designed partly to detect coordinated misuse that is invisible at the individual conversation level, works by identifying patterns across accounts and conversations. It identified, for example, a network of automated accounts using similar prompt structures to generate search engine spam and subsequently banned them. But because requests route through proxies, bans do not meaningfully stop the underlying behaviour. And for deliberately staged attacks — such as distributing a harmful inquiry across multiple stages and proxy accounts, each request individually innocuous — cross-account patterns are far less visible than coordinated spam, where the signal is obvious by design. + +Lastly, the transfer station does not only embody a traditional offence/defence paradigm — whether between US AI companies and Chinese users or between AI safeguards and malicious actors. A black market has a supply chain with its own exploitative logic, and the harms it generates extend well beyond the original question of access. Faces harvested for proxy KYC verification to bypass Anthropic’s system today can be resold to open fraudulent financial accounts, fabricate employment records, or generate deepfakes tomorrow, with the original subject in the Global South bearing the legal and reputational consequences. The same infrastructure that routes Claude requests can be used to defraud users through model substitution, targeted scams based on leaked prompt data, or blackmail. The account-farming operations that keep proxy pools stocked — bulk SMS verification, fraudulent registrations, carded accounts — nurture broader criminal markets for spam calls, phishing texts, fraudulent loan applications, and credit card scams. Many harms have nothing to do with AI or geopolitics. + +But now that every byproduct of the grey market–from the potential danger of terrorists leveraging AI to synthesize the next pandemic to real-life exploitation and crime. As much as the Great Firewall or AI geo-blockage wants to separate who gets access to frontier technology along national lines, as the grey market reveals, the harms are not separable. + +***Acknowledgement:*** + +Zilan is grateful to Alan Chan, Gabriel Wagner, Karuna Nandkumar, and Kayla Blomquist for their helpful feedback. + +The author acknowledges the use of LLMs for preliminary desk research, technical concepts clarification and copy-editing, and is, in fact, very grateful that she can still use VPN to access Claude in mainland China via the Singapore node without triggering the KYC process. + +[^1]: Profiles derived from informal conversations. + +[^2]: An Application Programming Interface, or API, is the channel that lets developers plug their software directly into an AI model — sending requests programmatically to Anthropic's servers and receiving responses back, rather than interacting through a browser. + +[^3]: Specifically, replacing the ANTHROPIC\_BASE\_URL environment variable with the proxy's address. + +[^4]: From informal conversations and desk research. diff --git a/raw/articles/github-automated-my-job-better-leader-2026.md b/raw/articles/github-automated-my-job-better-leader-2026.md new file mode 100644 index 0000000..fd0e565 --- /dev/null +++ b/raw/articles/github-automated-my-job-better-leader-2026.md @@ -0,0 +1,135 @@ +--- +source_url: https://github.blog/developer-skills/github/i-automated-my-job-and-it-made-me-a-better-leader/ +ingested: 2026-06-28 +sha256: b08905453312f75914d6ae524306ed068878212f83fc3f0212cc33fb868e2470 +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520675829493923962' + author_id: '890908900520505354' + posted_at: 2026-06-28T06:22:50.542000000Z + message_excerpt: 'https://github.blog/developer-skills/github/i-automated-my-job-and-it-made-me-a-better-leader/?utm_source=social-x-blog-i-automated-my-job&utm_medium=social&utm_campaign=github-copilot-app-ga-2026' +--- + +Here’s the thing about senior leadership that nobody warns you about: the job isn’t hard because of any single task. It’s hard because your work lives in fifteen different places and your brain is the only system connecting them. + +Meetings bleed into each other. Decisions are made in threads without you. Someone mentioned your name in a planning meeting, and now there’s an action item living in a doc you’ve never seen. You’ll find out about it in two weeks when someone casually asks for an update. Fun. + +Last year, my team almost missed a performance review deadline because it was announced in a channel nobody was watching. One person spent ten minutes searching Slack and couldn’t find it. Another found the date in a random, unrelated channel. I ended up posting “I’ll admit we dropped the ball on following up in Slack, so that’s on me.” That’s the kind of thing that keeps happening when your brain is the only system connecting everything. + +I was spending so much energy on context-switching that I had nothing left for the thinking, connecting, and creating that my role actually requires (and that’s the work I actually like doing). But I started using automations in the [GitHub Copilot app](https://github.com/features/ai/github-app?utm_source=blog-automations&utm_medium=blog&utm_campaign=github-copilot-app-ga-2026), and it changed my entire workflow. Bear with me. + +## What automations actually are + +The [GitHub Copilot app](https://github.com/features/ai/github-app?utm_source=blog-automations&utm_medium=blog&utm_campaign=github-copilot-app-ga-2026) is a standalone desktop app for macOS, Windows, and Linux, built for working *with* agents, not just talking to them. You can run parallel sessions across repositories, each on its own branch and worktree. You can see what agents are doing in real time through canvases, which are bidirectional work surfaces where you and the agent operate on the same plan, terminal, or browser session. Progress is visible and steerable, not buried in chat history. + +Automations are scheduled prompts that run against your real work context: your calendar, your email, your messages, your GitHub repos. They connect through MCP servers and integrations, so they can see what’s happening across all the places your work lives. They tell me what actually needs my attention, which lets me ignore the rest. + +Think of them as agents with a standing brief. You tell them what to care about, how to think, and when to run. Then they just… do it. Every day. Without you remembering to ask. Which is good, because you won’t. + +## What this looks like + +I’m a senior director at GitHub. I lead developer relations. My scope is wide, my calendar is full, and my brain works differently than most people assume. I’m AuDHD, which means I’m good at pattern recognition and deep focus, but genuinely terrible at remembering which thread I promised to follow up on three days ago. + +I didn’t set out to build 40 automations. I was curious about the automations tab, asked the app what it could do, and it suggested things I hadn’t thought of. The first time I set one up, I opened a chat and said something like: “Look across all of my work surfaces, my calendar, my email, my messages, and figure out where I’m dropping balls, where I might need help, and suggest automations that would be useful.” + +It immediately suggested about six. The first drafts weren’t perfect, and that’s okay. You refine them. You give them voice. You teach them how you think. Once I saw what was possible, I kept going. Now I have about 40. (I know. I know.) + +I’m not going to walk through all of them. (You’re welcome.) But here are the categories that matter most, and some highlights from each. + +### The morning brief + +Every day before I open anything, several automations have already run. **Meeting Prep** pulls my calendar and builds context for every meeting, with different formats for one-on-ones vs. large syncs vs. external calls. By the time I sit down, I know what each meeting is about and what I need to bring. **Pre-Meeting Access Check** verifies I actually have access to the docs and links referenced in the invite. No more showing up and realizing the agenda doc is locked. If you’ve never experienced that particular panic, honestly, must be nice. **Daily Triage Digest** sweeps GitHub, email, and messages for anything that needs my attention. + +The cumulative effect is that my mornings went from “frantically opening twelve tabs while pretending I’ve read the agenda” to “reading a few summaries with coffee.” It’s a different life. + +### Staying current + +I cannot be surprised by our own launches. That’s literally the job. + +**Ship Decoder** finds everything GitHub shipped in the last 24 hours and explains it to me in plain language. This is real context I can use in conversations. **Launch Radar** runs weekly and surfaces upcoming launches that touch my team’s space so I’m never blindsided. These two alone probably save me an hour a day of scrolling through channels trying to piece together what happened. I used to spend that hour. I did not enjoy that hour. + +### Career architecture + +This is the category that surprised me most. I built automations that actively work on career development, and if that sounds weird, stay with me. + +**Daily Wins Recap** runs every evening and summarizes what I actually accomplished. This one matters more than it sounds. My default mode is to check something off and immediately move to the next thing. I don’t sit with it. I don’t recognize it. I just keep going. Then performance review season comes around. I have to articulate my impact, and I’m panic-staring at a blank doc trying to remember eight months of work. + +This automation keeps a running record so I don’t have to. Think of it as a gratitude practice backed by real data rather than a task list. It counters the “what did I even do today?” spiral that hits hardest on the busy days. On the days when imposter syndrome is loud, I need something that talks back to it with facts. The robot believes in me even when I don’t. That’s oddly moving? I don’t know. It works. + +### Team and people + +This is where I want to be really honest, because I know you might be thinking: *is she automating the human parts of her job?* + +No. And that distinction matters to me more than anything else in this post. + +**Commitments and Follow-Up Tracker** searches my own messages for things I said I’d do and flags what I haven’t done yet. This one is humbling. And essential. Because when I tell someone “I’ll look into this” and then forget, that’s a trust problem. The automation protects the trust. + +The kudos I write are still mine. The noticing is still mine. The automation just makes sure my brain doesn’t steal it from the people who deserve it. + +These automations don’t replace connection. They *enable* it. They give me back the headspace to actually show up for people. Before this system, I’d walk into conversations distracted or running on fumes, because my brain was full of operational noise. Now when I sit down in a one-on-one, I’m actually present. When I write recognition for my team, it’s specific and real. + +The automations handle the scaffolding. I do the human work. That’s the deal. + +### Maintenance and logistics + +This category covers the boring stuff that quietly eats your week if you let it: **Dependabot PR Triage** finds and merges safe dependency updates across my repos daily. Handled. **Stale Work Finder** surfaces pull requests I opened and forgot, issues that went quiet, branches collecting dust. (We all have those. Don’t lie.) **Travel Logistics Tracker** watches for conference-related threads and consolidates logistics into a single brief. Conference season is chaos. This helps. + +## What my automations look like + +Here’s a real one from my setup, the **Stale Work Finder**, so you can see what these prompts actually look like in practice: + +``` +Find all my stale work across GitHub using the gh CLI. Things that are falling through the cracks. + +Check for: + +- PRs I opened that haven't received a review in 7+ days +- PRs I'm assigned to review that I haven't reviewed yet (older than 3 days) +- Issues assigned to me that have had no activity in 14+ days +- Draft PRs I own that have been drafts for 2+ weeks +- For each item show: repo, title, link, how long it's been stale, and who's involved. + +Format as: + + 1. 🔴 Embarrassingly stale (3+ weeks) + 2. 🟡 Getting dusty (1-3 weeks) + 3. 🟢 Just needs a nudge (under a week) +``` + +That’s it. That’s the whole automation. You write a prompt, set a schedule, and the agent runs it on your behalf. You can get as detailed or as loose as you want. The app fills in context from your connected tools. It runs every Monday for me and the results are… always a little eye-opening. But it’s better to know. + +## The AuDHD part + +I’ll just say it: for me, automations are an accessibility tool. + +AuDHD means my executive function and working memory are wildly inconsistent, and the inconsistency is the hardest part to explain to people. Some days I can hold seventeen threads in my head. Other days I forget I have a meeting in 10 minutes. There is no in between. The gap between those days used to scare me, because my team deserves consistent leadership regardless of what my brain is doing on any given Tuesday. + +These automations narrow that gap. They make me *consistent*. They mean my team gets the same quality of attention whether my executive function showed up today or not. For me, that’s the difference between thriving and slowly burning out. And I’ve done the burning out part. Zero stars, would not recommend. + +## How to start (for real) + +If you’re thinking about building something like this, here’s what I’d say: don’t try to automate everything at once. Start with the one thing that causes you the most friction. + +For me, it was meeting prep. I kept walking into meetings cold because prep required visiting four different tools and synthesizing information I didn’t have bandwidth to synthesize. One automation fixed that. And once I felt that relief, I kept going. And going. And going. + +Could I consolidate some of these? Probably. I have about 40, and I’m sure some of them could be combined. I prefer specificity, but you could easily roll several into one big automation if that’s more your style. + +Here’s the trick that worked for me: open a chat in the GitHub Copilot app and ask it to audit your work surfaces. Where are you dropping balls? Where are the repetitive patterns? What’s the thing you keep meaning to do but never get to? Start there. + +The first draft won’t be perfect. That’s fine. You refine it in conversation. You teach it your voice, your priorities, how you think about “good.” Then you let it run. + +Start with one. See how it feels. + +Then build another. And another. And before you know it, you have 40, and you’re writing a blog post about it. Anyway. + +## The bigger picture + +I think we’re at an interesting moment for how people relate to AI at work. The early conversation about AI at work was mostly about generation. Make me a thing. Write me the code. The reality, at least for me, is more like augmentation of invisible labor. The stuff that burns you out but never shows up in your output. The meta-work nobody acknowledges in performance reviews but everyone is buried in. + +Every leader I know is overwhelmed by context. Every neurodivergent professional I know is spending enormous energy on systems that neurotypical people navigate without thinking about. Automations won’t fix organizational dysfunction or bad management or an unreasonable workload. But they can give you back enough headspace to actually do the work you’re here to do. + +And honestly? That’s enough. That’s a lot. + +And look, this is a GitHub product. It’ll also run your dependency updates, triage your issues, do security sweeps across your repos. The developer workflows are exactly what you’d expect. I just happen to use it for the parts of my job nobody talks about. diff --git a/raw/articles/hermes-research-llm-wiki-skill-2026.md b/raw/articles/hermes-research-llm-wiki-skill-2026.md new file mode 100644 index 0000000..de74d35 --- /dev/null +++ b/raw/articles/hermes-research-llm-wiki-skill-2026.md @@ -0,0 +1,468 @@ +--- +source_url: https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-llm-wiki +ingested: 2026-06-28 +sha256: 1396b232adfa25060f2e3d4985765dd03e7712a86b14dec4bc0c01008cefce40 +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520479163444625478' + author_id: '890908900520505354' + posted_at: 2026-06-27T17:21:21.702000000Z + message_excerpt: 'https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-llm-wiki' +--- + +Karpathy's LLM Wiki: build/query interlinked markdown KB. + +| | | +| --- | --- | +| Source | Bundled (installed by default) | +| Path | `skills/research/llm-wiki` | +| Version | `2.1.0` | +| Author | Hermes Agent | +| License | MIT | +| Platforms | linux, macos, windows | +| Tags | `wiki`, `knowledge-base`, `research`, `notes`, `markdown`, `rag-alternative` | +| Related skills | [`obsidian`](https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/note-taking/note-taking-obsidian), [`arxiv`](https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-arxiv) | + +## Reference: full SKILL.md + +## Karpathy's LLM Wiki + +Build and maintain a persistent, compounding knowledge base as interlinked markdown files. Based on [Andrej Karpathy's LLM Wiki pattern](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f). + +Unlike traditional RAG (which rediscovers knowledge from scratch per query), the wiki compiles knowledge once and keeps it current. Cross-references are already there. Contradictions have already been flagged. Synthesis reflects everything ingested. + +**Division of labor:** The human curates sources and directs analysis. The agent summarizes, cross-references, files, and maintains consistency. + +## When This Skill Activates + +Use this skill when the user: + +- Asks to create, build, or start a wiki or knowledge base +- Asks to ingest, add, or process a source into their wiki +- Asks a question and an existing wiki is present at the configured path +- Asks to lint, audit, or health-check their wiki +- References their wiki, knowledge base, or "notes" in a research context + +## Wiki Location + +**Location:** Set via `WIKI_PATH` environment variable (e.g. in `${HERMES_HOME:-~/.hermes}/.env`). + +If unset, defaults to `~/wiki`. + +```bash +WIKI="${WIKI_PATH:-$HOME/wiki}" +``` + +The wiki is just a directory of markdown files — open it in Obsidian, VS Code, or any editor. No database, no special tooling required. + +## Architecture: Three Layers + +```markdown +wiki/ +├── SCHEMA.md # Conventions, structure rules, domain config +├── index.md # Sectioned content catalog with one-line summaries +├── log.md # Chronological action log (append-only, rotated yearly) +├── raw/ # Layer 1: Immutable source material +│ ├── articles/ # Web articles, clippings +│ ├── papers/ # PDFs, arxiv papers +│ ├── transcripts/ # Meeting notes, interviews +│ └── assets/ # Images, diagrams referenced by sources +├── entities/ # Layer 2: Entity pages (people, orgs, products, models) +├── concepts/ # Layer 2: Concept/topic pages +├── comparisons/ # Layer 2: Side-by-side analyses +└── queries/ # Layer 2: Filed query results worth keeping +``` + +**Layer 1 — Raw Sources:** Immutable. The agent reads but never modifies these.**Layer 2 — The Wiki:** Agent-owned markdown files. Created, updated, and cross-referenced by the agent.**Layer 3 — The Schema:** `SCHEMA.md` defines structure, conventions, and tag taxonomy. + +## Resuming an Existing Wiki (CRITICAL — do this every session) + +When the user has an existing wiki, **always orient yourself before doing anything**: + +① **Read `SCHEMA.md`** — understand the domain, conventions, and tag taxonomy. ② **Read `index.md`** — learn what pages exist and their summaries. ③ **Scan recent `log.md`** — read the last 20-30 entries to understand recent activity. + +```bash +WIKI="${WIKI_PATH:-$HOME/wiki}" +# Orientation reads at session start +read_file "$WIKI/SCHEMA.md" +read_file "$WIKI/index.md" +read_file "$WIKI/log.md" offset= +``` + +Only after orientation should you ingest, query, or lint. This prevents: + +- Creating duplicate pages for entities that already exist +- Missing cross-references to existing content +- Contradicting the schema's conventions +- Repeating work already logged + +For large wikis (100+ pages), also run a quick `search_files` for the topic at hand before creating anything new. + +## Initializing a New Wiki + +When the user asks to create or start a wiki: + +1. Determine the wiki path (from `$WIKI_PATH` env var, or ask the user; default `~/wiki`) +2. Create the directory structure above +3. Ask the user what domain the wiki covers — be specific +4. Write `SCHEMA.md` customized to the domain (see template below) +5. Write initial `index.md` with sectioned header +6. Write initial `log.md` with creation entry +7. Confirm the wiki is ready and suggest first sources to ingest + +### SCHEMA.md Template + +Adapt to the user's domain. The schema constrains agent behavior and ensures consistency: + +```markdown +# Wiki Schema + +## Domain +[What this wiki covers — e.g., "AI/ML research", "personal health", "startup intelligence"] + +## Conventions +\`transformer-architecture.md\` +- Every wiki page starts with YAML frontmatter (see below) +\`[[wikilinks]]\` +\`updated\` +\`index.md\` +\`log.md\` +Provenance markers: + at the end of paragraphs whose claims come from a specific source. This lets a reader trace each + claim back without re-reading the whole raw file. Optional on single-source pages where the +\`sources:\` + +## Frontmatter + \`\`\`yaml + --- + title: Page Title + created: YYYY-MM-DD + updated: YYYY-MM-DD + type: entity | concept | comparison | query | summary + tags: [from taxonomy below] + sources: [raw/articles/source-name.md] + # Optional quality signals: + confidence: high | medium | low # how well-supported the claims are + contested: true # set when the page has unresolved contradictions + contradictions: [other-page-slug] # pages this one conflicts with + --- +``` + +`confidence` and `contested` are optional but recommended for opinion-heavy or fast-moving topics. Lint surfaces `contested: true` and `confidence: low` pages for review so weak claims don't silently harden into accepted wiki fact. + +Raw sources ALSO get a small frontmatter block so re-ingests can detect drift: + +```yaml +--- +source_url: https://example.com/article # original URL, if applicable +ingested: YYYY-MM-DD +sha256: <hex digest of the raw content below the frontmatter> +--- +``` + +The `sha256:` lets a future re-ingest of the same URL skip processing when content is unchanged, and flag drift when it has changed. Compute over the body only (everything after the closing `---`), not the frontmatter itself. + +\[Define 10-20 top-level tags for the domain. Add new tags here BEFORE using them.\] + +Example for AI/ML: + +- Models: model, architecture, benchmark, training +- People/Orgs: person, company, lab, open-source +- Techniques: optimization, fine-tuning, inference, alignment, data +- Meta: comparison, timeline, controversy, prediction + +Rule: every tag on a page must appear in this taxonomy. If a new tag is needed, add it here first, then use it. This prevents tag sprawl. + +## Page Thresholds + +- **Create a page** when an entity/concept appears in 2+ sources OR is central to one source +- **Add to existing page** when a source mentions something already covered +- **DON'T create a page** for passing mentions, minor details, or things outside the domain +- **Split a page** when it exceeds ~200 lines — break into sub-topics with cross-links +- **Archive a page** when its content is fully superseded — move to `_archive/`, remove from index + +## Entity Pages + +One page per notable entity. Include: + +- Overview / what it is +- Key facts and dates +- Relationships to other entities (\[\[wikilinks\]\]) +- Source references + +## Concept Pages + +One page per concept or topic. Include: + +- Definition / explanation +- Current state of knowledge +- Open questions or debates +- Related concepts (\[\[wikilinks\]\]) + +## Comparison Pages + +Side-by-side analyses. Include: + +- What is being compared and why +- Dimensions of comparison (table format preferred) +- Verdict or synthesis +- Sources + +## Update Policy + +When new information conflicts with existing content: + +1. Check the dates — newer sources generally supersede older ones +2. If genuinely contradictory, note both positions with dates and sources +3. Mark the contradiction in frontmatter: `contradictions: [page-name]` +4. Flag for user review in the lint report +```markdown +### index.md Template + +The index is sectioned by type. Each entry is one line: wikilink + summary. + +\`\`\`markdown +# Wiki Index + +> Content catalog. Every wiki page listed under its type with a one-line summary. +> Read this first to find relevant pages for any query. +> Last updated: YYYY-MM-DD | Total pages: N + +## Entities + + +## Concepts + +## Comparisons + +## Queries +``` + +**Scaling rule:** When any section exceeds 50 entries, split it into sub-sections by first letter or sub-domain. When the index exceeds 200 entries total, create a `_meta/topic-map.md` that groups pages by theme for faster navigation. + +### log.md Template + +```markdown +# Wiki Log + +> Chronological record of all wiki actions. Append-only. +\`## [YYYY-MM-DD] action | subject\` +> Actions: ingest, update, query, lint, create, archive, delete +> When this file exceeds 500 entries, rotate: rename to log-YYYY.md, start fresh. + +## [YYYY-MM-DD] create | Wiki initialized +- Domain: [domain] +- Structure created with SCHEMA.md, index.md, log.md +``` + +## Core Operations + +### 1\. Ingest + +When the user provides a source (URL, file, paste), integrate it into the wiki: + +① **Capture the raw source:** + +- URL → use `web_extract` to get markdown, save to `raw/articles/` +- PDF → use `web_extract` (handles PDFs), save to `raw/papers/` +- Pasted text → save to appropriate `raw/` subdirectory +- Name the file descriptively: `raw/articles/karpathy-llm-wiki-2026.md` +- **Add raw frontmatter** (`source_url`, `ingested`, `sha256` of the body). On re-ingest of the same URL: recompute the sha256, compare to the stored value — skip if identical, flag drift and update if different. This is cheap enough to do on every re-ingest and catches silent source changes. + +② **Discuss takeaways** with the user — what's interesting, what matters for the domain. (Skip this in automated/cron contexts — proceed directly.) + +③ **Check what already exists** — search index.md and use `search_files` to find existing pages for mentioned entities/concepts. This is the difference between a growing wiki and a pile of duplicates. + +④ **Write or update wiki pages:** + +- **New entities/concepts:** Create pages only if they meet the Page Thresholds in SCHEMA.md (2+ source mentions, or central to one source) +- **Existing pages:** Add new information, update facts, bump `updated` date. When new info contradicts existing content, follow the Update Policy. +- **Cross-reference:** Every new or updated page must link to at least 2 other pages via `[[wikilinks]]`. Check that existing pages link back. +- **Tags:** Only use tags from the taxonomy in SCHEMA.md +- **Provenance:** On pages synthesizing 3+ sources, append `^[raw/articles/source.md]` markers to paragraphs whose claims trace to a specific source. +- **Confidence:** For opinion-heavy, fast-moving, or single-source claims, set `confidence: medium` or `low` in frontmatter. Don't mark `high` unless the claim is well-supported across multiple sources. + +⑤ **Update navigation:** + +- Add new pages to `index.md` under the correct section, alphabetically +- Update the "Total pages" count and "Last updated" date in index header +- Append to `log.md`: `## [YYYY-MM-DD] ingest | Source Title` +- List every file created or updated in the log entry + +⑥ **Report what changed** — list every file created or updated to the user. + +A single source can trigger updates across 5-15 wiki pages. This is normal and desired — it's the compounding effect. + +### 2\. Query + +When the user asks a question about the wiki's domain: + +① **Read `index.md`** to identify relevant pages. ② **For wikis with 100+ pages**, also `search_files` across all `.md` files for key terms — the index alone may miss relevant content. ③ **Read the relevant pages** using `read_file`. ④ **Synthesize an answer** from the compiled knowledge. Cite the wiki pages you drew from: "Based on \[\[page-a\]\] and \[\[page-b\]\]..." ⑤ **File valuable answers back** — if the answer is a substantial comparison, deep dive, or novel synthesis, create a page in `queries/` or `comparisons/`. Don't file trivial lookups — only answers that would be painful to re-derive. ⑥ **Update log.md** with the query and whether it was filed. + +### 3\. Lint + +When the user asks to lint, health-check, or audit the wiki: + +① **Orphan pages:** Find pages with no inbound `[[wikilinks]]` from other pages. + +```python +# Use execute_code for this — programmatic scan across all wiki pages +import os, re +from collections import defaultdict +wiki = "" +# Scan all .md files in entities/, concepts/, comparisons/, queries/ +# Extract all [[wikilinks]] — build inbound link map +# Pages with zero inbound links are orphans +``` + +② **Broken wikilinks:** Find `[[links]]` that point to pages that don't exist. + +③ **Index completeness:** Every wiki page should appear in `index.md`. Compare the filesystem against index entries. + +④ **Frontmatter validation:** Every wiki page must have all required fields (title, created, updated, type, tags, sources). Tags must be in the taxonomy. + +⑤ **Stale content:** Pages whose `updated` date is >90 days older than the most recent source that mentions the same entities. + +⑥ **Contradictions:** Pages on the same topic with conflicting claims. Look for pages that share tags/entities but state different facts. Surface all pages with `contested: true` or `contradictions:` frontmatter for user review. + +⑦ **Quality signals:** List pages with `confidence: low` and any page that cites only a single source but has no confidence field set — these are candidates for either finding corroboration or demoting to `confidence: medium`. + +⑧ **Source drift:** For each file in `raw/` with a `sha256:` frontmatter, recompute the hash and flag mismatches. Mismatches indicate the raw file was edited (shouldn't happen — raw/ is immutable) or ingested from a URL that has since changed. Not a hard error, but worth reporting. + +⑨ **Page size:** Flag pages over 200 lines — candidates for splitting. + +⑩ **Tag audit:** List all tags in use, flag any not in the SCHEMA.md taxonomy. + +⑪ **Log rotation:** If log.md exceeds 500 entries, rotate it. + +⑫ **Report findings** with specific file paths and suggested actions, grouped by severity (broken links > orphans > source drift > contested pages > stale content > style issues). + +⑬ **Append to log.md:** `## [YYYY-MM-DD] lint | N issues found` + +## Working with the Wiki + +### Searching + +```bash +# Find pages by content +search_files "transformer" path="$WIKI" file_glob="*.md" + +# Find pages by filename +search_files "*.md" target="files" path="$WIKI" + +# Find pages by tag +search_files "tags:.*alignment" path="$WIKI" file_glob="*.md" + +# Recent activity +read_file "$WIKI/log.md" offset= +``` + +### Bulk Ingest + +When ingesting multiple sources at once, batch the updates: + +1. Read all sources first +2. Identify all entities and concepts across all sources +3. Check existing pages for all of them (one search pass, not N) +4. Create/update pages in one pass (avoids redundant updates) +5. Update index.md once at the end +6. Write a single log entry covering the batch + +### Archiving + +When content is fully superseded or the domain scope changes: + +1. Create `_archive/` directory if it doesn't exist +2. Move the page to `_archive/` with its original path (e.g., `_archive/entities/old-page.md`) +3. Remove from `index.md` +4. Update any pages that linked to it — replace wikilink with plain text + "(archived)" +5. Log the archive action + +### Obsidian Integration + +The wiki directory works as an Obsidian vault out of the box: + +- `[[wikilinks]]` render as clickable links +- Graph View visualizes the knowledge network +- YAML frontmatter powers Dataview queries +- The `raw/assets/` folder holds images referenced via `![[image.png]]` + +For best results: + +- Set Obsidian's attachment folder to `raw/assets/` +- Enable "Wikilinks" in Obsidian settings (usually on by default) +- Install Dataview plugin for queries like `TABLE tags FROM "entities" WHERE contains(tags, "company")` + +If using the Obsidian skill alongside this one, set `OBSIDIAN_VAULT_PATH` to the same directory as the wiki path. + +### Obsidian Headless (servers and headless machines) + +On machines without a display, use `obsidian-headless` instead of the desktop app. It syncs vaults via Obsidian Sync without a GUI — perfect for agents running on servers that write to the wiki while Obsidian desktop reads it on another device. + +**Setup:** + +```bash +# Requires Node.js 22+ +npm install -g obsidian-headless + +# Login (requires Obsidian account with Sync subscription) +ob login --email --password '' + +# Create a remote vault for the wiki +ob sync-create-remote --name "LLM Wiki" + +# Connect the wiki directory to the vault +cd ~/wiki +ob sync-setup --vault "" + +# Initial sync +ob sync + +# Continuous sync (foreground — use systemd for background) +ob sync --continuous +``` + +**Continuous background sync via systemd:** + +```markdown +# ~/.config/systemd/user/obsidian-wiki-sync.service +[Unit] +Description=Obsidian LLM Wiki Sync +After=network-online.target +Wants=network-online.target + +[Service] +ExecStart=/path/to/ob sync --continuous +WorkingDirectory=/home/user/wiki +Restart=on-failure +RestartSec=10 + +[Install] +WantedBy=default.target +``` +```bash +systemctl --user daemon-reload +systemctl --user enable --now obsidian-wiki-sync +# Enable linger so sync survives logout: +sudo loginctl enable-linger $USER +``` + +This lets the agent write to `~/wiki` on a server while you browse the same vault in Obsidian on your laptop/phone — changes appear within seconds. + +## Pitfalls + +- **Never modify files in `raw/`** — sources are immutable. Corrections go in wiki pages. +- **Always orient first** — read SCHEMA + index + recent log before any operation in a new session. Skipping this causes duplicates and missed cross-references. +- **Always update index.md and log.md** — skipping this makes the wiki degrade. These are the navigational backbone. +- **Don't create pages for passing mentions** — follow the Page Thresholds in SCHEMA.md. A name appearing once in a footnote doesn't warrant an entity page. +- **Don't create pages without cross-references** — isolated pages are invisible. Every page must link to at least 2 other pages. +- **Frontmatter is required** — it enables search, filtering, and staleness detection. +- **Tags must come from the taxonomy** — freeform tags decay into noise. Add new tags to SCHEMA.md first, then use them. +- **Keep pages scannable** — a wiki page should be readable in 30 seconds. Split pages over 200 lines. Move detailed analysis to dedicated deep-dive pages. +- **Ask before mass-updating** — if an ingest would touch 10+ existing pages, confirm the scope with the user first. +- **Rotate the log** — when log.md exceeds 500 entries, rename it `log-YYYY.md` and start fresh. The agent should check log size during lint. +- **Handle contradictions explicitly** — don't silently overwrite. Note both claims with dates, mark in frontmatter, flag for user review. + +[llm-wiki-compiler](https://github.com/atomicmemory/llm-wiki-compiler) is a Node.js CLI that compiles sources into a concept wiki with the same Karpathy inspiration. It's Obsidian-compatible, so users who want a scheduled/CLI-driven compile pipeline can point it at the same vault this skill maintains. Trade-offs: it owns page generation (replaces the agent's judgment on page creation) and is tuned for small corpora. Use this skill when you want agent-in-the-loop curation; use llmwiki when you want batch compile of a source directory. diff --git a/raw/articles/howm-2026.md b/raw/articles/howm-2026.md new file mode 100644 index 0000000..eef5e2d --- /dev/null +++ b/raw/articles/howm-2026.md @@ -0,0 +1,89 @@ +--- +source_url: https://github.com/emacsmirror/howm +ingested: 2026-06-28 +sha256: 505b37553342b608daa7030f42ebccf50e4f87e87597b677a9c3969f738dd274 +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520477343527731381' + author_id: '890908900520505354' + posted_at: 2026-06-27T17:14:07.800000000Z + message_excerpt: 'https://github.com/emacsmirror/howm' +--- + +# Howm: Write fragmentarily and read collectively. + +Howm is a note-taking tool on Emacs. It is similar to emacs-wiki.el; you can enjoy hyperlinks and full-text search easily. It is not similar to emacs-wiki.el; it can be combined with any format. + +* [Home](https://kaorahi.github.io/howm/) +* [Introduction by Leah Neukirchen](https://leahneukirchen.org/blog/archive/2022/03/note-taking-in-emacs-with-howm.html) (thx!) +* [Detailed tutorial by Andrei Sukhovskii](https://emacs101.github.io/howm.html) (thx!) +* [1-minute introduction on YouTube by Raoul Comninos](https://www.youtube.com/watch?v=RHxYMF1wsmk) (thx!) + +The following screenshot illustrates the Howm linking system: +![screenshot](doc/screenshot.png) + +(Colorscheme: [Modus themes](https://protesilaos.com/emacs/modus-themes).) + +## Quick start + +If you're using a recent version of Emacs and have enabled the [MELPA](https://melpa.org/) community package repository, you can simply place the following in your `~/.config/emacs/init.el` configuration file and restart Emacs: + +```emacs-lisp +(use-package howm + :ensure t) +``` + +After that, you can press e.g. `C-c , ,` to open the main menu, `C-c , a` to see a list of all your notes, or `C-c , c` to capture a new note from anywhere. See the documentation links above for more detailed instructions on how to use Howm. + +Alternatively, here is a configurable snippet. If you want to change the note format, delete the menu file `0000-00-00-000000.txt` beforehand to regenerate a new menu in that format. + +```emacs-lisp +(use-package howm + :ensure t + :init + ;; + ;; Options: Remove the leading ";" in the following lines if you like. + ;; + ;; Format + ;(require 'howm-markdown) ;; Write notes in markdown-mode. (*1) + ;(require 'howm-org) ;; Write notes in Org-mode. (*2) + ;; + ;; Preferences + ;(setq howm-directory "~/Documents/Howm") ;; Where to store the files? + ;(setq howm-follow-theme t) ;; Use your Emacs theme colors. (*3) + ;; + ;; Performance + ;(setq howm-menu-expiry-hours 1) ;; Cache menu N hours. (*4) + ;(setq howm-menu-refresh-after-save nil) ;; Speed up note saving. (*5) + ) +``` + +* (*1) [Markdown-mode](https://jblevins.org/projects/markdown-mode/) must be installed separately. Howm's wiki link `[[...]]` is disabled for syntax compatibility. +* (*2) Just replace `C-c ,` with `C-c ;` in Howm's documentation, including that `C-c ; ;` opens the menu. Howm's wiki link `[[...]]` is disabled for syntax compatibility. +* (*3) Technically, it inherits the colors used by Org-mode; this requires that you either (i) use a theme that defines the Org faces or (ii) that your `init.el` loads Org itself. +* (*4) Howm caches the menu state, so that opening the menu is instant. But it regenerates the menu if either (i) it's been more than 1 hour (the "expiry hours") since last menu update, or (ii) you saved a file that belongs to Howm (howm-mode was active in the buffer). +* (*5) Howm doesn't regenerate the menu if you save a file either, only if it's been more than 1 hour since the last menu update. If you want to update it more frequently, you have to manually refresh it (keybinding `R`). + +For a more extensive overview of how you can customize Howm, you can use the Customize system (`M-x customize-group RET howm RET`). See also [Emacs Wiki](https://www.emacswiki.org/emacs/HowmMode) for various tips. + +## FAQ + +### How does Howm compare to Org-roam? + +TL;DR: Balanced simplicity, clever linking system, UI for contextual reading. + +Most things you can do in Howm are probably possible in Org-roam as well. But Howm is more minimalistic and loose in every way. It is designed on the belief that not aiming for perfect organization is the key to maintaining long-term note-taking without becoming unmanageable. Users seem to appreciate its balance between simplicity and functionality, as well as its smooth navigation UI. The UI, along with the search-based link/backlink system, helps with reading many fragmentary notes collectively. + +### I've heard Howm is unbearably slow, actually? + +There are a few performance options. Once enabled, searches are usually instant, even for [20 years of notes](https://github.com/kaorahi/howm/issues/44#issuecomment-2639467999). + +## Project history + +* 2002-05-29 initial release (v0.1) on sourceforge.jp +* ... +* 2023-02-18 v1.5.1-snapshot4 on [OSDN](https://howm.osdn.jp/) (renamed from sourceforge.jp) +* 2023-05-13 moved the repository to [GitHub](https://github.com/kaorahi/howm) +* 2023-07-30 moved the project home to [GitHub](https://kaorahi.github.io/howm/) diff --git a/raw/articles/nashsu-llm-wiki-2026.md b/raw/articles/nashsu-llm-wiki-2026.md new file mode 100644 index 0000000..105a531 --- /dev/null +++ b/raw/articles/nashsu-llm-wiki-2026.md @@ -0,0 +1,493 @@ +--- +source_url: https://github.com/nashsu/llm_wiki +ingested: 2026-06-28 +sha256: 8f5ad888b9060bfaab7c295fbf46202ce63602bc83955fe466ca2ea1de56f1fa +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520475829438648371' + author_id: '890908900520505354' + posted_at: 2026-06-27T17:08:06.813000000Z + message_excerpt: 'https://github.com/nashsu/llm_wiki' +--- + +# LLM Wiki + +

+ LLM Wiki Logo +

+ +

+ A personal knowledge base that builds itself.
+ LLM reads your documents, builds a structured wiki, and keeps it current. +

+ +

+ What is this? • + Features • + Tech Stack • + Installation • + Credits • + License +

+ +

+ English | 中文 | 日本語 | 한국어 +

+ +--- + +

+ Overview +

+ +## Features + +- **Two-Step Chain-of-Thought Ingest** — LLM analyzes first, then generates wiki pages with source traceability and incremental cache +- **Multimodal Image Ingestion** — extract embedded images from PDFs, generate factual captions with a vision LLM, surface them in image-aware search results with lightbox preview and jump-to-source +- **Optional MinerU PDF Parsing** — use MinerU cloud parsing for complex PDFs with tables, formulas, and dense layouts; the built-in local parser remains the default +- **4-Signal Knowledge Graph** — relevance model with direct links, source overlap, Adamic-Adar, and type affinity +- **Louvain Community Detection** — automatic knowledge cluster discovery with cohesion scoring +- **Graph Insights** — surprising connections and knowledge gaps with one-click Deep Research +- **Vector Semantic Search** — optional embedding-based retrieval via LanceDB, supports any OpenAI-compatible endpoint +- **Persistent Ingest Queue** — serial processing with crash recovery, cancel, retry, and progress visualization +- **Folder Import** — recursive folder import preserving directory structure, folder context as LLM classification hint +- **Source Folder Auto-Watch** — detects external changes in `raw/sources/` and keeps ingest/delete cleanup in sync +- **Deep Research** — LLM-optimized search topics, multi-query web search via Tavily, SerpApi, or SearXNG, auto-ingest results into wiki +- **Async Review System** — LLM flags items for human judgment, predefined actions, pre-generated search queries +- **Chrome Web Clipper** — one-click web page capture with auto-ingest into knowledge base +- **Local HTTP API + MCP Server + AI Agent Skill** — built-in `127.0.0.1:19828` JSON API and bundled MCP server for hybrid search, file read, graph traversal, and source rescan; ready-made [agent skill](https://github.com/nashsu/llm_wiki_skill) installs into Claude Code / Codex with one command (`npx skills add …`) + +## What is this? + +LLM Wiki is a cross-platform desktop application that turns your documents into an organized, interlinked knowledge base — automatically. Instead of traditional RAG (retrieve-and-answer from scratch every time), the LLM **incrementally builds and maintains a persistent wiki** from your sources. Knowledge is compiled once and kept current, not re-derived on every query. + +This project is based on [Karpathy's LLM Wiki pattern](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) — a methodology for building personal knowledge bases using LLMs. We implemented the core ideas as a full desktop application with significant enhancements. + +

+ LLM Wiki Architecture +

+ +## Credits + +The foundational methodology comes from **Andrej Karpathy**'s [llm-wiki.md](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f), which describes the pattern of using LLMs to incrementally build and maintain a personal wiki. The original document is an abstract design pattern; this project is a concrete implementation with substantial extensions. + +## What We Kept from the Original + +The core architecture follows Karpathy's design faithfully: + +- **Three-layer architecture**: Raw Sources (immutable) → Wiki (LLM-generated) → Schema (rules & config) +- **Three core operations**: Ingest, Query, Lint +- **index.md** as the content catalog and LLM navigation entry point +- **log.md** as the chronological operation record with parseable format +- **[[wikilink]]** syntax for cross-references +- **YAML frontmatter** on every wiki page +- **Obsidian compatibility** — the wiki directory works as an Obsidian vault +- **Human curates, LLM maintains** — the fundamental role division + +

+ Obsidian Compatibility +

+ +## What We Changed & Added + +### 1. From CLI to Desktop Application + +The original is an abstract pattern document designed to be copy-pasted to an LLM agent. We built it into a **full cross-platform desktop application** with: +- **Three-column layout**: Knowledge Tree / File Tree (left) + Chat (center) + Preview (right) +- **Icon sidebar** for switching between Wiki, Sources, Search, Graph, Lint, Review, Deep Research, Settings +- **Custom resizable panels** — drag-to-resize left and right panels with min/max constraints +- **Activity panel** — real-time processing status showing file-by-file ingest progress +- **All state persisted** — conversations, settings, review items, project config survive restarts +- **Scenario templates** — Research, Reading, Personal Growth, Business, General — each pre-configures purpose.md and schema.md + +### 2. Purpose.md — The Wiki's Soul + +The original has Schema (how the wiki works) but no formal place for **why** the wiki exists. We added `purpose.md`: +- Defines goals, key questions, research scope, evolving thesis +- LLM reads it during every ingest and query for context +- LLM can suggest updates based on usage patterns +- Different from schema — schema is structural rules, purpose is directional intent + +### 3. Two-Step Chain-of-Thought Ingest + +The original describes a single-step ingest where the LLM reads and writes simultaneously. We split it into **two sequential LLM calls** for significantly better quality: + +``` +Step 1 (Analysis): LLM reads source → structured analysis + - Key entities, concepts, arguments + - Connections to existing wiki content + - Contradictions & tensions with existing knowledge + - Recommendations for wiki structure + +Step 2 (Generation): LLM takes analysis → generates wiki files + - Source summary with frontmatter (type, title, sources[]) + - Entity pages, concept pages with cross-references + - Updated index.md, log.md, overview.md + - Review items for human judgment + - Search queries for Deep Research +``` + +Additional ingest enhancements beyond the original: +- **SHA256 incremental cache** — source file content is hashed before ingest; unchanged files are skipped automatically, saving LLM tokens and time +- **Persistent ingest queue** — serial processing prevents concurrent LLM calls; queue persisted to disk, survives app restart; failed tasks auto-retry up to 3 times +- **Folder import** — recursive folder import preserving directory structure; folder path passed to LLM as classification context (e.g., "papers > energy" helps categorize content) +- **Source folder auto-watch** — files added, edited, or deleted in `raw/sources/` outside the app are picked up automatically and reuse the same ingest/delete lifecycle as in-app actions +- **Queue visualization** — Activity Panel shows progress bar, pending/processing/failed tasks with cancel and retry buttons +- **Auto-embedding** — when vector search is enabled, new pages are automatically embedded after ingest +- **Source traceability** — every generated wiki page includes a `sources: []` field in YAML frontmatter, linking back to the raw source files that contributed to it +- **overview.md auto-update** — global summary page regenerated on every ingest to reflect the latest state of the wiki +- **Guaranteed source summary** — fallback ensures a source summary page is always created, even if the LLM omits it +- **Language-aware generation** — LLM responds in the user's configured language (English or Chinese) +- **Progressive Sources view** — large source folders render progressively while scrolling, keeping big source collections responsive + +### 4. Knowledge Graph with Relevance Model + +

+ Knowledge Graph +

+ +The original mentions `[[wikilinks]]` for cross-references but has no graph analysis. We built a **full knowledge graph visualization and relevance engine**: + +**4-Signal Relevance Model:** +| Signal | Weight | Description | +|--------|--------|-------------| +| Direct link | ×3.0 | Pages linked via `[[wikilinks]]` | +| Source overlap | ×4.0 | Pages sharing the same raw source (via frontmatter `sources[]`) | +| Adamic-Adar | ×1.5 | Pages sharing common neighbors (weighted by neighbor degree) | +| Type affinity | ×1.0 | Bonus for same page type (entity↔entity, concept↔concept) | + +**Graph Visualization (sigma.js + graphology + ForceAtlas2):** +- Node colors by page type or community, sizes scaled by link count (√ scaling) +- Edge thickness and color by relevance weight (green=strong, gray=weak) +- Hover interaction: neighbors stay visible, non-neighbors dim, edges highlight with relevance score label +- Zoom controls (ZoomIn, ZoomOut, Fit-to-screen) +- Position caching prevents layout jumps when data updates +- Legend switches between type counts and community info based on coloring mode + +### 5. Louvain Community Detection + +Not in the original. Automatic discovery of knowledge clusters using the **Louvain algorithm** (graphology-communities-louvain): + +- **Auto-clustering** — discovers which pages naturally group together based on link topology, independent of predefined page types +- **Type / Community toggle** — switch between coloring nodes by page type (entity, concept, source...) or by discovered knowledge cluster +- **Cohesion scoring** — each community scored by intra-edge density (actual edges / possible edges); low-cohesion clusters (< 0.15) flagged with warning +- **12-color palette** — distinct visual separation between clusters +- **Community legend** — shows top node label, member count, and cohesion per cluster + +

+ Louvain Community Detection +

+ +### 6. Graph Insights — Surprising Connections & Knowledge Gaps + +Not in the original. The system **automatically analyzes graph structure** to surface actionable insights: + +**Surprising Connections:** +- Detects unexpected relationships: cross-community edges, cross-type links, peripheral↔hub couplings +- Composite surprise score ranks the most noteworthy connections +- Dismissable — mark connections as reviewed so they don't reappear + +**Knowledge Gaps:** +- **Isolated pages** (degree ≤ 1) — pages with few or no connections to the rest of the wiki +- **Sparse communities** (cohesion < 0.15, ≥ 3 pages) — knowledge areas with weak internal cross-references +- **Bridge nodes** (connecting 3+ clusters) — critical junction pages that hold multiple knowledge areas together + +**Interactive:** +- Click any insight card to **highlight** corresponding nodes and edges in the graph; click again to deselect +- Knowledge gaps and bridge nodes have a **Deep Research button** — triggers LLM-optimized research with domain-aware topics (reads overview.md + purpose.md for context) +- Research topic shown in **editable confirmation dialog** before starting — user can refine topic and search queries + +

+ Graph Insights +

+ +### 7. Optimized Query Retrieval Pipeline + +The original describes a simple query where the LLM reads relevant pages. We built a **multi-phase retrieval pipeline** with optional vector search and budget control: + +``` +Phase 1: Tokenized Search + - English: word splitting + stop word removal + - Chinese: CJK bigram tokenization (每个 → [每个, 个…]) + - Title match bonus (+10 score) + - Searches both wiki/ and raw/sources/ + +Phase 1.5: Vector Semantic Search (optional) + - Embedding via any OpenAI-compatible /v1/embeddings endpoint + - Stored in LanceDB (Rust backend) for fast ANN retrieval + - Cosine similarity finds semantically related pages even without keyword overlap + - Results merged into search: boosts existing matches + adds new discoveries + +Phase 2: Graph Expansion + - Top search results used as seed nodes + - 4-signal relevance model finds related pages + - 2-hop traversal with decay for deeper connections + +Phase 3: Budget Control + - Configurable context window: 4K → 1M tokens + - Proportional allocation: 60% wiki pages, 20% chat history, 5% index, 15% system + - Pages prioritized by combined search + graph relevance score + +Phase 4: Context Assembly + - Numbered pages with full content (not just summaries) + - System prompt includes: purpose.md, language rules, citation format, index.md + - LLM instructed to cite pages by number: [1], [2], etc. +``` + +**Vector Search** is fully optional — disabled by default, enabled in Settings with independent endpoint, API key, and model configuration. When disabled, the pipeline falls back to tokenized search + graph expansion. Benchmark: overall recall improved from 58.2% to 71.4% with vector search enabled. + +### 8. Multi-Conversation Chat with Persistence + +The original has a single query interface. We built **full multi-conversation support**: + +- **Independent chat sessions** — create, rename, delete conversations +- **Conversation sidebar** — quick switching between topics +- **Per-conversation persistence** — each conversation saved to `.llm-wiki/chats/{id}.json` +- **Configurable history depth** — limit how many messages are sent as context (default: 10) +- **Cited references panel** — collapsible section on each response showing which wiki pages were used, grouped by type with icons +- **Reference persistence** — cited pages stored directly in message data, stable across restarts +- **Regenerate** — re-generate the last response with one click (removes last assistant + user message pair, re-sends) +- **Save to Wiki** — archive valuable answers to `wiki/queries/`, then auto-ingest to extract entities/concepts into the knowledge network + +### 9. Thinking / Reasoning Display + +Not in the original. For LLMs that emit `` blocks (DeepSeek, QwQ, etc.): + +- **Streaming thinking** — rolling 5-line display with opacity fade during generation +- **Collapsed by default** — thinking blocks hidden after completion, click to expand +- **Visual separation** — thinking content shown in distinct style, separate from the main response + +### 10. KaTeX Math Rendering + +Not in the original. Full LaTeX math support across all views: + +- **KaTeX rendering** — inline `$...$` and block `$$...$$` formulas rendered via remark-math + rehype-katex +- **Milkdown math plugin** — preview editor renders math natively via @milkdown/plugin-math +- **Auto-detection** — bare `\begin{aligned}` and other LaTeX environments automatically wrapped with `$$` delimiters +- **Unicode fallback** — 100+ symbol mappings (α, ∑, →, ≤, etc.) for simple inline notation outside math blocks + +### 11. Review System (Async Human-in-the-Loop) + +The original suggests staying involved during ingest. We added an **asynchronous review queue**: + +- LLM flags items needing human judgment during ingest +- **Predefined action types**: Create Page, Deep Research, Skip — constrained to prevent LLM hallucination of arbitrary actions +- **Search queries generated at ingest time** — LLM pre-generates optimized web search queries for each review item +- User handles reviews at their convenience — doesn't block ingest + +### 12. Deep Research + +

+ Deep Research +

+ +Not in the original. When the LLM identifies knowledge gaps: + +- **Web search** via Tavily, SerpApi, or SearXNG finds relevant sources with full content extraction (no truncation) +- **Provider-specific configuration** — Tavily and SerpApi use independent API keys; SerpApi supports selectable engines, while SearXNG uses a configured instance URL and search categories +- **Multiple search queries** per topic — LLM-generated at ingest time, optimized for search engines +- **LLM-optimized research topics** — when triggered from Graph Insights, LLM reads overview.md + purpose.md to generate domain-specific topics and queries (not generic keywords) +- **User confirmation dialog** — editable topic and search queries shown for review before research starts +- **LLM synthesizes** findings into a wiki research page with cross-references to existing wiki +- **Thinking display** — `` blocks shown as collapsible sections during synthesis, auto-scroll to latest content +- **Auto-ingest** — research results automatically processed to extract entities/concepts into the wiki +- **Task queue** with 3 concurrent tasks +- **Research Panel** — dedicated sidebar panel with dynamic height, real-time streaming progress + +### 13. Browser Extension (Web Clipper) + +

+ Chrome Extension Web Clipper +

+ +The original mentions Obsidian Web Clipper. We built a **dedicated Chrome Extension** (Manifest V3): + +- **Mozilla Readability.js** for accurate article extraction (strips ads, nav, sidebars) +- **Turndown.js** for HTML → Markdown conversion with table support +- **Project picker** — choose which wiki to clip into (supports multi-project) +- **Local HTTP API** (port 19827, tiny_http) — Extension ↔ App communication +- **Auto-ingest** — clipped content automatically triggers the two-step ingest pipeline +- **Clip watcher** — polls every 3 seconds for new clips, processes automatically +- **Offline preview** — shows extracted content even when app is not running + +### 14. Multi-format Document Support + +The original focuses on text/markdown. We support structured extraction preserving document semantics: + +| Format | Method | +|--------|--------| +| PDF | Built-in pdf-extract (Rust) with file caching; optional MinerU cloud parsing for tables, formulas, and complex layouts | +| DOCX | docx-rs — headings, bold/italic, lists, tables → structured Markdown | +| PPTX | ZIP + XML — slide-by-slide extraction with heading/list structure | +| XLSX/XLS/ODS | calamine — proper cell types, multi-sheet support, Markdown tables | +| Images | Native preview (png, jpg, gif, webp, svg, etc.) | +| Video/Audio | Built-in player | +| Web clips | Readability.js + Turndown.js → clean Markdown | + +> MinerU is optional. When enabled, PDF files are uploaded to MinerU cloud for parsing; keep the built-in parser for sensitive documents. If MinerU fails, LLM Wiki falls back to the built-in parser. MinerU usage is subject to its file size, page count, and quota limits. + +### 15. File Deletion with Cascade Cleanup + +The original has no deletion mechanism. We added **intelligent cascade deletion**: + +- Deleting a source file removes its wiki summary page +- **3-method matching** finds related wiki pages: frontmatter `sources[]` field, source summary page name, frontmatter section references +- **Shared entity preservation** — entity/concept pages linked to multiple sources only have the deleted source removed from their `sources[]` array, not deleted entirely +- **Index cleanup** — removed pages are purged from index.md +- **Wikilink cleanup** — dead `[[wikilinks]]` to deleted pages are removed from remaining wiki pages + +### 16. Configurable Context Window + +Not in the original. Users can configure how much context the LLM receives: + +- **Slider from 4K to 1M tokens** — adapts to different LLM capabilities +- **Proportional budget allocation** — larger windows get proportionally more wiki content +- **60/20/5/15 split** — wiki pages / chat history / index / system prompt + +### 17. Cross-Platform Compatibility + +The original is platform-agnostic (abstract pattern). We handle concrete cross-platform concerns: + +- **Path normalization** — unified `normalizePath()` used across 22+ files, backslash → forward slash +- **Unicode-safe string handling** — char-based slicing instead of byte-based (prevents crashes on CJK filenames) +- **macOS close-to-hide** — close button hides window (app stays running in background), click dock icon to restore, Cmd+Q to quit +- **Windows/Linux close confirmation** — confirmation dialog before quitting to prevent accidental data loss +- **Tauri v2** — native desktop on macOS, Windows, Linux +- **GitHub Actions CI/CD** — automated builds for macOS (ARM + Intel), Windows (.msi), Linux (.deb / .AppImage) + +### 18. Other Additions + +- **i18n** — English + Chinese interface (react-i18next) +- **Settings persistence** — LLM provider, API key, model, context size, language saved via Tauri Store +- **Obsidian config** — auto-generated `.obsidian/` directory with recommended settings +- **Markdown rendering** — GFM tables with borders, proper code blocks, wikilink processing in chat and preview +- **Multi-provider LLM support** — OpenAI, Anthropic, Google, Ollama, Custom — each with provider-specific streaming and headers +- **15-minute timeout** — long ingest operations won't fail prematurely +- **dataVersion signaling** — graph and UI automatically refresh when wiki content changes + +## Tech Stack + +| Layer | Technology | +|-------|-----------| +| Desktop | Tauri v2 (Rust backend) | +| Frontend | React 19 + TypeScript + Vite | +| UI | shadcn/ui + Tailwind CSS v4 | +| Editor | Milkdown (ProseMirror-based WYSIWYG) | +| Graph | sigma.js + graphology + ForceAtlas2 | +| Search | Tokenized search + graph relevance + optional vector (LanceDB) | +| Vector DB | LanceDB (Rust, embedded, optional) | +| PDF | pdf-extract + optional MinerU cloud parser | +| Office | docx-rs + calamine | +| i18n | react-i18next | +| State | Zustand | +| LLM | Streaming fetch (OpenAI, Anthropic, Google, Ollama, Custom) | +| Web Search | Tavily, SerpApi, SearXNG JSON API | + +## Installation + +### Pre-built Binaries + +Download from [Releases](https://github.com/nashsu/llm_wiki/releases): +- **macOS**: `.dmg` (Apple Silicon + Intel) +- **Windows**: `.msi` +- **Linux**: `.deb` / `.AppImage` + +### Build from Source + +```bash +# Prerequisites: Node.js 20+, Rust 1.70+ +git clone https://github.com/nashsu/llm_wiki.git +cd llm_wiki +npm install +npm run tauri dev # Development +npm run tauri build # Production build +``` + +### Chrome Extension + +1. Open `chrome://extensions` +2. Enable "Developer mode" +3. Click "Load unpacked" +4. Select the `extension/` directory + +## Quick Start + +1. Launch the app → Create a new project (choose a template) +2. Go to **Settings** → Configure your LLM provider (API key + model) +3. Optional: configure **Web Search** providers and source folder auto-watch in Settings +4. Go to **Sources** → Import documents (PDF, DOCX, MD, etc.) +5. Watch the **Activity Panel** — LLM automatically builds wiki pages +6. Use **Chat** to query your knowledge base +7. Browse the **Knowledge Graph** to see connections +8. Check **Review** for items needing your attention +9. Run **Lint** periodically to maintain wiki health + +## Local HTTP API + MCP Server + AI Agent Skill + +LLM Wiki ships a built-in local HTTP API at `http://127.0.0.1:19828` (token-protected, `127.0.0.1`-only) so external tools — including AI agents like **Claude Code**, **Codex**, or any HTTP-capable script — can query your wiki: + +- `GET /api/v1/health` — server status (no auth) +- `GET /api/v1/projects` — list projects +- `GET /api/v1/projects/{id}/files` / `files/content` — read files and content +- `GET /api/v1/projects/{id}/reviews?status=unresolved` — export Review tab items for wiki maintenance (`status`: `unresolved`, `resolved`, or `all`; optional `type` and `limit`) +- `PATCH /api/v1/projects/{id}/reviews/{reviewId}` — update one Review item (JSON body `{ "resolved": true, "action": "label" }`; `resolved` defaults to true, pass false to reopen) +- `POST /api/v1/projects/{id}/reviews/resolve` — bulk-resolve Review items (JSON body `{ "ids": [...], "action": "label" }`), returns `{ resolved, notFound, count }`; the Review tab's Refresh button re-reads the result from disk +- `POST /api/v1/projects/{id}/search` — **hybrid** retrieval (keyword + vector) returning `mode`, `tokenHits`, `vectorHits`, per-result `vectorScore` +- `GET /api/v1/projects/{id}/graph` — wikilinks graph +- `POST /api/v1/projects/{id}/sources/rescan` — trigger a backend rescan + +Enable the API, generate a token, and choose whether local unauthenticated access is allowed in **Settings → API + MCP**. + +For MCP-compatible clients, LLM Wiki also includes a local MCP server in `mcp-server/`. After building it with `npm run mcp:build`, **Settings → API + MCP** shows a copyable MCP client configuration with the correct local path for your machine. The MCP tools call the same API surface, so agent clients can list projects, read files, export unresolved Review items, run hybrid search, inspect the graph, and trigger source rescans without custom HTTP glue code. + +### Plug your AI agent in with one command + +A ready-made **agent skill** for LLM Wiki lives in its own repo. Install it into Claude Code / Codex / any skills-compatible runtime: + +```bash +npx skills add https://github.com/nashsu/llm_wiki_skill.git --skill llm_wiki_skill +``` + +After install, the agent can answer prompts like "what does my LLM Wiki say about X", "search my 知识库 for Y", "show the neighborhood of node Z in my wiki graph", and "rescan my wiki sources" by talking to your locally-running app — read-only by default, citing wiki page paths so you can verify in-app. + +- **Skill repo**: +- **Trigger discipline**: it intentionally does **not** trigger on generic "search my notes" / "check my Obsidian / Notion / Logseq" — only when you explicitly name LLM Wiki / `my wiki` / `知识库`. + +## Project Structure + +``` +my-wiki/ +├── purpose.md # Goals, key questions, research scope +├── schema.md # Wiki structure rules, page types +├── raw/ +│ ├── sources/ # Uploaded documents (immutable) +│ └── assets/ # Local images +├── wiki/ +│ ├── index.md # Content catalog +│ ├── log.md # Operation history +│ ├── overview.md # Global summary (auto-updated) +│ ├── entities/ # People, organizations, products +│ ├── concepts/ # Theories, methods, techniques +│ ├── sources/ # Source summaries +│ ├── queries/ # Saved chat answers + research +│ ├── synthesis/ # Cross-source analysis +│ └── comparisons/ # Side-by-side comparisons +├── .obsidian/ # Obsidian vault config (auto-generated) +└── .llm-wiki/ # App config, chat history, review items +``` + +## Star History + + + + + + Star History Chart + + + +## License + +This project is licensed under the **GNU General Public License v3.0** — see [LICENSE](LICENSE) for details. diff --git a/raw/articles/refactoring-english-effective-design-doc-2026.md b/raw/articles/refactoring-english-effective-design-doc-2026.md new file mode 100644 index 0000000..10e1f51 --- /dev/null +++ b/raw/articles/refactoring-english-effective-design-doc-2026.md @@ -0,0 +1,487 @@ +--- +source_url: https://refactoringenglish.com/excerpts/write-an-effective-design-doc/ +ingested: 2026-06-28 +sha256: 2a850ade735f83da6a568b1dcd1c9957466ae164552abde47a48a29d2e8963ed +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520425046827597844' + author_id: '890908900520505354' + posted_at: 2026-06-27T13:46:19.295000000Z + message_excerpt: 'https://refactoringenglish.com/excerpts/write-an-effective-design-doc/' +--- + +A good design doc can save you years of development time. Writing a design doc forces you to think through important decisions before you waste time on the wrong implementation or paint yourself into a corner. It’s also the best way to coordinate design decisions among teammates and partner teams. + +I’ve written design docs as a developer at Google, Microsoft, and within [my own companies](https://mtlynch.io/projects/). The specifics vary, but the underlying principles remain the same. A design doc should articulate the hard problems you’re solving and help your teammates give you feedback. + +Below, I share my approach to creating effective design docs and explain what belongs in a design doc and what does not. + +- [An example design doc](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#an-example-design-doc) +- [When should you write a design doc?](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#when-should-you-write-a-design-doc) +- [How much should you invest into your design doc?](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#how-much-should-you-invest-into-your-design-doc) +- [What belongs in a design doc?](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#what-belongs-in-a-design-doc) + - [What’s the cost of getting it wrong?](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#whats-the-cost-of-getting-it-wrong) +- [Components of a design doc](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#components-of-a-design-doc) + - [Title](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#title) + - [Metadata](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#metadata) + - [Objective](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#objective) + - [Background](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#background) + - [Related documents](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#related-documents) + - [Goals](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#goals) + - [Non-goals](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#non-goals) + - [Scenarios](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#scenarios) + - [Diagrams](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#diagrams) + - [Glossary](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#glossary) + - [Constraints](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#constraints) + - [Service level objectives (SLOs)](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#service-level-objectives-slos) + - [Monitoring / alerting](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#monitoring--alerting) + - [Timeline](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#timeline) + - [Interfaces](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#interfaces) + - [Dependencies / infrastructure](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#dependencies--infrastructure) + - [Security](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#security) + - [Privacy](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#privacy) + - [Legal considerations](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#legal-considerations) + - [Logging](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#logging) + - [Open issues](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#open-issues) + - [Resolved issues](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#resolved-issues) + - [Alternatives considered](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#alternatives-considered) +- [Driving Your Design Doc through Review](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#driving-your-design-doc-through-review) + +## An example design doc + +The most common question I get about design docs is where to find a good one. I’ve never seen a public design doc that I consider high-quality. All of mine are hidden away at the companies that paid me to write them. + +So, I wrote [a design doc from scratch](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/little-moments-design-doc/) based on the principles I’m sharing here. It lays out the design for [a real web app I’m building](https://codeberg.org/mtlynch/little-moments). + + + +I created the design doc before writing any code, and I’m adhering to the design [as I implement the app](https://codeberg.org/mtlynch/little-moments#status). + +- [Little Moments Design Doc](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/little-moments-design-doc/) + +The design is more exhaustive than what I’d normally write for a solo hobby project, but this is roughly the length and depth of a design doc I’d create if I were coordinating work with other people on a professional project. + +## When should you write a design doc? + +The more complex or risky the project, the more valuable it is to write a design doc. + +Consider these questions: + +- Will multiple people coordinate work to implement the design? +- Will the project take more than three months of full-time dev work? +- Will the implementation run in production for several years? +- Does the project involve cross-team collaboration? +- Are the goals and requirements of the project ambiguous? +- Are there catastrophic risks you could prevent at design time (e.g., security flaws, legal risks)? + +If you answered “yes” to any of these questions, then it’s likely worth the effort to write a design doc. If you answered “yes” to two or more, a design doc will almost certainly be worth the effort. + +## How much should you invest into your design doc? + +A design doc can be a simple one-pager or a 50-page document that requires signoff from five different teams. You need to decide how much detail makes sense. + +There’s no universal rule that says how long you should spend on a design doc just like there’s no rule that says how much to test your code. The right investment depends on your team’s goals, risks, deadlines, and culture. Sometimes, the right amount to invest in a design doc is zero. + +## What belongs in a design doc? + +If you specify every possible detail in a design doc, you’ve essentially written the implementation during the design phase. That would defeat the whole purpose of a design doc. + +As a rule of thumb, you can ask a simple question to decide whether a decision belongs in your design doc: what’s the penalty for being wrong? + +### What’s the cost of getting it wrong? + +Not all design decisions are equally important. Some choices are radically more flexible than others. + +For example, if you build a web application in C++ and realize 200k lines later that Ruby on Rails was the better choice, you’re stuck. A from-scratch rewrite [would never work](https://www.joelonsoftware.com/2000/04/06/things-you-should-never-do-part-i/), and even if you manage to write new code in Rails, you still suffer the burden of maintaining code in two wildly different languages. + +Other design decisions are trivial. For example, if your app displays a list of 1,000 articles, should they all appear at once? Or should the user see 20 at a time and click “Load more” to see the next 20? + +It doesn’t matter. + +A “load more” button is not a design-level concern. If you pick one solution, and user feedback tells you you’re wrong, you can fix it in a few hours. You don’t need to detail your entire thought process in your design doc, and you definitely shouldn’t waste review cycles arguing about it. + +## Components of a design doc + +Below, I’ve included common sections to include in your design docs. You generally don’t need every single section for every doc. Choose the subset that make sense for you. + +The first thing your project needs is a title. It’s the way people will refer to your project in conversation, so aim for something with these qualities: + +- **Short**: Easy to say aloud. +- **Distinctive**: Makes it clear which project it refers to. +- **Evocative**: Conceptually represents your project. + +For example, if you were adding a caching layer between your application server and your database server, **RecencyBank** would be a good name. It’s easy to say and describes your project’s purpose. A bad name would be “Project Flying Silver Horse” because it’s verbose and nonsensical. + +Boring but useful, metadata helps your reader understand the basic context of your doc: + +- Who is the author? (name + email address) +- When did you create the doc? +- What’s the authoritative URL? + - Especially if your organization uses [shortlink redirects](https://golinks.github.io/golinks/) like `http://go/recency-bank` + +> **Metadata** +> +> - **Author**: Michael Lynch ([michael@refactoringenglish.com](mailto:michael@refactoringenglish.com)) +> - **Status**: Ready for review +> - **Created**: 2026-06-22 +> - **URL**: http://go/recency-bank-design + +### Objective + +The objective is a one-sentence explanation of your project’s purpose. It should appear on the first page of your doc in plain language that any stakeholder understands. + +> **Objective** +> +> Improve application performance by adding a caching layer between the Trogdor web server and the Postgres database. + +### Background + +The background section explains the context and motivation for the project. It should answer these questions: + +- Why is the team taking on this project? +- What problem does this project solve? +- Were there previous attempts to solve this problem? + +> **Background** +> +> When we launched the Trogdor web app in 2023, pages typically loaded in 100ms or less. After three years, median page loads have ballooned to 600ms, which causes users to perceive our app as sluggish. +> +> We investigated the slowdown and discovered that database lookups make up 80% of page load times. As our data store has grown larger, database lookups have gotten slower. +> +> We also discovered that 95% of database lookups are for the same 3% of database rows. This pattern of usage benefits greatly from memory-backed caching. The cache would serve frequently-accessed data faster and reduce database load for all other queries. + +If this project connects to other documents, link to them so it’s easy for the reader to find them. This includes: + +- Documents from your program manager or testing counterparts on this project (e.g., test plans, functional specs) +- Design docs for related systems +- Design docs for previous iterations of this project + +> **Related documents** +> +> - **Testing plan**: http://go/recency-bank-test-plan +> - **Trogdor performance report**: http://go/trogdor-perf-2026 + +### Goals + +The goals section describes your high-level goals for this project. It should connect logically to the background section and explain what the world looks like after you’ve completed implementation. + +Avoid setting goals in terms of implementation details. Your goals should communicate how the project benefits your users, your team, or your company. + +> **Goals** +> +> - Increase user-perceied responsiveness for Trogdor web app. +> - Reduce database server load. + +### Non-goals + +While the goals define what’s within your project’s scope, the non-goals section delineates what’s out of scope. + +Are there goals that readers might mistakenly assume are within scope for your project? If so, add them as explicit non-goals. + +> **Non-goals** +> +> - Create a general-purpose, reusable caching system +> - The caching layer we add to the Trogdor web app will make application-specific optimizations. Re-using this cache on other systems is out of scope. +> - Location-aware caching +> - It may be useful in the future to support caches that sit geographically close to the end user to reduce latency, but that is out of scope of v1. + +### Scenarios + +If your goal is something like “Add a ‘Share as URL’ button to charts,” the reader might not understand what that looks like in practice. + +The scenarios section allows you to paint a picture for your reader of how your completed system works in the real world. + +> **Scenario: Share a report via URL** +> +> 1. Bob creates a custom report in his KeyMetrics dashboard. +> 2. Bob navigates to the menu bar and clicks “Share > as URL.” +> 3. Bob emails the URL to his teammate, Charlie. +> 4. Charlie clicks the link and sees an exact copy of Bob’s report in read-only mode. + +### Diagrams + +Diagrams are tremendously valuable, though they might not seem that way. + +As the design author, you intuitively understand how the pieces of your plan fit together. You can see the architecture in your head. Your reviewers do not have this mental picture, so the fastest way for them to see it is to draw them a picture. + +![Architecture diagram](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/architecture-diagram.svg) + +Example diagram showing the architecture of a simple web application. + +If you’re not sure what belongs in a diagram, think about these questions: + +- How does data flow through your system? +- How do the different components of your system fit together? +- How does your system interact with its dependencies and downstream clients? +- What communication protocols does your system define? + +Choose a diagramming tool that’s flexible to editing. I’ve seen developers create a beautiful diagram on a whiteboard and photograph it for their design doc. The first draft looks amazing, but then they’re stuck with that diagram forever because they can’t edit the photo without recreating the whole thing from scratch. + +[Excalidraw](https://excalidraw.com/), [draw.io](https://www.drawio.com/), and [Google Drawings](https://docs.google.com/drawings/) are popular diagramming tools that facilitate revisions. There are also languages like [Mermaid](https://mermaid.js.org/), [D2](https://d2lang.com/), and [Graphviz](https://graphviz.org/) that allow you to generate diagrams programmatically. I’ve had good experience using an LLM to create diagramming code for me. Remember to link to the source drawing or code so that your teammates have a way to reproduce the diagram as well. + +### Glossary + +The glossary defines terms that your readers might not recognize. + +Think hard about the potential readers of your doc, especially newer team members and people outside of your immediate team. Will those readers understand the names of internal tools or systems your doc references? + +When possible, use terms that your audience recognizes without having to refer to a glossary. Defining a term in a glossary is better than not defining it at all, but the best solution is to use recognizable terms or define them inline so that the reader doesn’t have to jump around your document. + +> **Glossary** +> +> - **Apposaurus**: the team’s internal load testing tool. We use Apposaurus to simulate a surge of visitors to the Trogdor web app so we can verify the app continues functioning under expected workloads. +> - **Baba-o-styley**: an internal code linter that enforces the company’s code style conventions. + +### Constraints + +If there are major constraints imposed on your design by your budget, clients, infrastructure, or dependencies, explain the constraints so the reader understands the context of your design choices. + +> **Constraints** +> +> Our servers are all RISC-V, so all code and dependencies must run on RISC-V architecture. + +### Service level objectives (SLOs) + +SLOs are the measurable goals that a service offers to its clients or users. You’ve probably heard of service level agreements (SLAs). SLAs are just SLOs plus financial penalties for falling short. + +Within a company, you typically don’t financially penalize your co-workers for mistakes (although, wouldn’t that be kind of fun?). So, design docs define SLOs rather than SLAs. + +An SLO creates a measurable, objective metric for your system’s performance. Your manager might tell you that your app must be “performant on mobile,” but that’s vague. You don’t want to wait until code complete to discover that your manager’s definition of “performant” is <2ms of latency. A well-defined SLO prevents ambiguity by expressing goals in concrete, objective terms. + +The typical considerations for your SLO are: + +- **Uptime / availability**: What percentage of time will your system be available? +- **Latency**: How quickly will your service complete requests? +- **Scale**: What volume of work can your system handle? + +> **Service level objectives** +> +> - Trogdor’s 50th percentile latency for user-facing HTTP requests: <=200ms +> - Postgres 50th percentile query latency: <= 80ms + +### Monitoring / alerting + +Once you nail down your SLOs ([above](https://refactoringenglish.com/excerpts/write-an-effective-design-doc/#service-level-objectives-slos)), it’s time to think about how you’ll measure them in production. + +The simplest way to verify that you’ve achieved your SLOs is to test manually. As your organization matures, you should automate monitoring to discover SLO failures immediately. + +When defining your monitoring strategy, ask yourself these questions: + +- If your service goes down, how will you find out? +- If your service’s performance slows by 100x, how will you know? +- What other events should trigger an alert? + - e.g., spikes in CPU usage, authentication failures, system errors + +> **Monitoring** +> +> The following events will trigger a page to the on-call engineer: +> +> - Trogdor’s 95th percentile latency for user-facing HTTP requests: >= 3s +> - Average CPU usage for Postgres servers during trailing 2m window: >= 90% + +### Timeline + +The timeline section breaks your project into milestones and specifies when project stakeholders will receive their deliverables. + +Choose milestones that [create useful artifacts](https://mtlynch.io/tinypilot-redesign/#structure-for-serial-incremental-results) for stakeholders. For example, start with a UI that shows dummy data, and show that to clients first. If it turns out you misunderstood the client’s requirements, fake data lets you find out early rather than after you’ve already implemented all the plumbing to populate the UI with production data. + +If you don’t know how to estimate project timelines, I highly recommend Joel Spolsky’s, [“Painless Software Schedules.”](https://www.joelonsoftware.com/2000/03/29/painless-software-schedules/) The article is 25 years old, but it remains my favorite software estimation strategy. + +> **Timeline** +> +> - **Milestone 1 (2026-07-01)**: RecencyBank is live in the test environment with a hardcoded subset of cached data (doesn’t read from Postgres). +> - **Milestone 2 (2026-07-17)**: RecencyBank is live in the test environment and caches real data from Postgres. +> - **Milestone 3 (2026-08-03)**: RecencyBank is live in the test environment and enforces cache eviction and lifecycle rules. +> - **Milestone 4 (2026-08-22)**: RecencyBank is fully implemented and deployed to production. + +### Interfaces + +Your project exists to serve people or other software systems, so what do those interactions look like? + +- For graphical systems, what is the user interface? + - Just simple sketches; don’t get bogged down in precise UI choices. +- For software interfaces, what are the API or CLI semantics? +- For file-based interfaces, what is the file format? + +> **Interfaces** +> +> The Trogdor `Server` struct currently depends directly on a `PostgresDB` Go `struct` like this: +> +> ```go +> type Server struct { db PostgresDB } +> ``` +> +> `PostgresDB` has the following exported methods: +> +> ```go +> GetUser(id UserID) (User, error) +> ListUsers() ([]User, error) +> ... +> ``` +> +> We will create a Go `interface` type with the same API surface as `PostgresDB`: +> +> ```go +> type Store interface { +> GetUser(id UserID) (User, error) +> ListUsers() ([]User, error) +> } +> ``` +> +> We will implement a RecencyBank caching type that implements the same `interface` and wraps the backend `PostgresDB` struct. The RecencyBank implementation will cache reads from Postgres and forward requests to Postgres when they mutate state or depend on data not in the cache. +> +> The only change to the `Server` implementation will be replacing the type of one member with the new `interface`: +> +> ```go +> type Server struct { db store.Store } +> ``` + +### Dependencies / infrastructure + +The dependencies section should answer questions like: + +- What programming language(s) will you use? +- On what hardware or service does the code run? +- Where will persistent data live? + +It’s easy to overlook this section, but decisions about language, libraries, and infrastructure have a major impact on the complexity and long-term maintenance costs of your system. + +Think deeply about which dependencies will be difficult to change after implementation, and don’t worry so much about the ones that swap out easily. It’s difficult to change languages or storage backends, but if you’re dissatisfied with the third-party service you use to send emails, you can replace it in an afternoon. + +> **Dependencies** +> +> - **Language**: Go +> - We widely use Go already, and it’s a suitable language for serving highly parallel workflows. +> - **Third-party packages** +> - [bbolt](https://pkg.go.dev/go.etcd.io/bbolt): This is a widely used key-value store implementation that implements many of the features we need for RecencyBank. + +### Security + +To build secure software, developers must integrate security into the full software lifecycle, starting at the design stage. + +The security section should answer questions like: + +- What threats did you consider? + - e.g., what happens if an attacker tries every possible password? What if a user uploads a PDF infected with malware? +- What is the [attack surface](https://en.wikipedia.org/wiki/Attack_surface) of this system? + - i.e., where does it process potentially malicious data? +- What are the trust boundaries? + - At what point does data flow from a less privileged system to a more privileged system? + - e.g., in a web app, requests from the user’s browser cross a trust boundary, as the web server shouldn’t assume input from the browser is safe. + +Even if you think security threats are unlikely or irrelevant in your system, it’s still helpful to document your rationale. Your explanation might prompt reviewers to identify threats you overlooked. + +> **Security** +> +> RecencyBank must not accept direct requests from the public Internet, as it does not enforce any access control. +> +> RecencyBank will run on a segregated network where it only accepts inbound requests from the Trogdor web server and can only make outbound requests to the Postgres server pool. + +### Privacy + +The privacy section is an opportunity to think through the sensitive data your system handles and what safeguards you’ll put in place to keep it secure. It should answer these questions: + +- What sensitive data does your system handle? +- How long will you retain it? +- Who will have access to it? +- How will you protect it? + - e.g., will the data be encrypted at rest and in transit? + +> **Privacy** +> +> RecencyBank contains the same sensitive user data as the Postgres database, so it inherits the privacy policy of our Postgres systems. In particular, engineers may only access RecencyBank systems in production with an associated bug number. Engineers must miimize the user data they access to only what is strictly required to investigate a bug. + +### Legal considerations + +If your system operates in a highly-regulated domain like finance or healthcare, the legal section helps you comply with relevant laws. + +Even outside of regulated domains, think about whether your system could break the law if things go awry. Explain how you’ll steer clear of legal violations that could put your company or clients at risk. + +If you’re publishing your code under an open-source license, define which license you’ve chosen and why. + +> **FizzleCorp contractual compliance** +> +> Our contract with FizzleCorp strictly limits our ability to create new copies of their proprietary FizzlePerfect™ user biometrics. +> +> Fortunately, our legal team reviewed the wording of the contract and confirmed that a caching layer fits within the existing definition of “storage layer,” so we may cache FizzlePefect™ data within RecencyBank without contract renegotiation. + +### Logging + +Logs can be tremendously valuable when you’re investigating a bug, performance issue, or security incident. If you design for effective logging, you’ll make it easier to maintain your system long-term. + +As you think about logging, consider these questions: + +- What critical events does the service log? +- Are there different log levels? + - e.g., informational, warning, error, critical +- Where does the system store its logs? +- How long do you retain your logs? +- Who has access to the logs? +- Is there any sensitive data you must keep out of the logs? + +> **Logging** +> +> RecencyBank logs the following events: +> +> - At initialization, logs the parameters used to initialize RecencyBank as well as RAM capacity and usage on the host. +> - Failures to persist a value in memory. +> - Failures to invalidate the cache after mutating an item in Postgres. + +### Open issues + +As you write your design doc, you’ll likely encounter at least one of the following situations: + +- There’s a flaw in your design, but you’re not sure how to solve it. +- You’re torn between multiple solutions. +- There’s a gap in your design because you need to gather more information. + +Create an appendix in your design doc called “Open Issues” that documents your outstanding issues. + +Each entry in the open issues section should explain: + +- What’s the problem that requires more work? +- What options do you see for resolving the issue? +- What is the immediate next step for resolving the issue? + +> **Open Issue: Choosing RAM size for cache** +> +> We need to decide how much RAM to assign to our caching layer. Adding RAM increases performance, but RAM is expensive, and there are diminishing returns to extra RAM. +> +> There is some optimal amount of RAM that minimizes our infrastructure costs between the caching system and our database. We could theoretically discover that optimal value by setting up a test environment and running several simulations, but running those simulations costs us dev time. +> +> I estimate the cost of creating a test environment and running a single simulation to be 3.0 dev days. Once the infrastructure is in place, additional simulations will take about 0.75 dev days each. +> +> **Proposed solution**: Choose 128 GB of RAM without testing. It’s probably close to optimal, and dev time is significantly more expensive than RAM. +> +> **Next step**: Ask our tech lead to weigh in. + +### Resolved issues + +When you resolve an open issue, summarize the decision, and move it from “Open issues” to a “Resolved issues” section in your design doc. Retain the full discussion for posterity. + +> **Resolved Issue: Choosing RAM size for cache** +> +> **Decision**: Provision 128 GB of RAM to the caching layer. If we’re failing to meet our performance goals and we’re RAM-constrained, we can add more RAM at that point. The dev cost of running tests to discover the perfect RAM size far outweighs the cost of additional RAM. +> +> We need to decide… \[rest of original open issue goes here\] + +### Alternatives considered + +If you anticipate readers asking, “Why didn’t you do X?” it’s helpful to answer that proactively in an “alternatives considered” section. This section is also where you can explain options you rejected, especially if they initially seemed appealing or you researched them extensively. + +I know some developers who spend hours meticulously documenting their every rejected design idea, but I think that’s overkill. As both a reader and author, all I need in the alternatives section is a few brief lines describing strong alternatives and why they didn’t work. + +> **Alternatives Considered** +> +> - Google Cloud Firestore (persistent storage) +> - The durability and reliability was appealing, but I disliked the platform lock-in and the difficulty of testing locally. + +## Driving Your Design Doc through Review + +Once you’ve completed your design doc, the next step is to share it with your team and gather feedback. + +The following section covers techniques for eliciting useful design feedback that moves your project forward rather than stalling it with bickering and confusion: + +- [How to Get Meaningful Feedback on Your Design Document](https://refactoringenglish.com/excerpts/useful-feedback-on-design-docs/) diff --git a/raw/articles/rem-cli-macos-reminders-2026.md b/raw/articles/rem-cli-macos-reminders-2026.md new file mode 100644 index 0000000..4853a87 --- /dev/null +++ b/raw/articles/rem-cli-macos-reminders-2026.md @@ -0,0 +1,386 @@ +--- +source_url: https://github.com/BRO3886/rem +ingested: 2026-06-28 +sha256: 003111dfeb577282e9072e7dc14b7ae90c802aa37349aaa27a52c630f1f9972c +discovered_from: + platform: discord + channel_id: '1028287639918497822' + channel_name: chat + message_id: '1520399169221558354' + author_id: '890908900520505354' + posted_at: 2026-06-27T12:03:29.593000000Z + message_excerpt: 'https://github.com/BRO3886/rem' +--- + +# rem + +A blazing fast CLI for macOS Reminders. Sub-200ms reads AND writes via EventKit, natural language dates, and import/export — all in a single binary. + +**[Documentation](https://rem.sidv.dev)** | **[Architecture](https://rem.sidv.dev/docs/architecture/)** | **[go-eventkit](https://github.com/BRO3886/go-eventkit)** + +## Features + +- **Sub-200ms reads AND writes** — EventKit via cgo (go-eventkit), direct memory access, no IPC +- **Single binary** — EventKit compiled in via cgo, no helper processes +- **Natural language dates** — `tomorrow`, `next friday at 2pm`, `in 3 hours`, `eod` +- **20 commands** — full CRUD, search, stats, overdue, upcoming, interactive mode +- **Multiple output formats** — table, JSON, plain text +- **Native tags** — `#hashtag` in titles or `--tags` flag, stored as real Reminders.app tags +- **Location reminders** — geofence triggers via `--location "lat,lng"`, fire on arrival or departure +- **Shared list support** — full CRUD on shared lists, sharing state in `rem lists`, and moves across the shared-list boundary via copy (macOS has no true move there) +- **Import/Export** — JSON and CSV with full property round-trip (including tags and location triggers) +- **Powered by [go-eventkit](https://github.com/BRO3886/go-eventkit)** — use the same library directly for programmatic Go access +- **Shell completions** — bash, zsh, fish + +## Installation + +### Homebrew + +```bash +brew tap BRO3886/tap +brew install rem-cli +``` + +### Quick install (recommended) + +```bash +curl -fsSL https://rem.sidv.dev/install | bash +``` + +Downloads the latest release, extracts, and installs to `~/.local/bin` (override with `INSTALL_DIR=...`). No sudo needed. + +### Via Go + +```bash +go install github.com/BRO3886/rem/cmd/rem@latest +``` + +Requires Go 1.21+ and Xcode Command Line Tools (cgo compiles EventKit bindings). + +### Manual download + +Download from [GitHub Releases](https://github.com/BRO3886/rem/releases): + +```bash +# Apple Silicon +curl -LO https://github.com/BRO3886/rem/releases/latest/download/rem-darwin-arm64.tar.gz +tar xzf rem-darwin-arm64.tar.gz +mkdir -p ~/.local/bin && mv rem ~/.local/bin/rem + +# Intel +curl -LO https://github.com/BRO3886/rem/releases/latest/download/rem-darwin-amd64.tar.gz +tar xzf rem-darwin-amd64.tar.gz +mkdir -p ~/.local/bin && mv rem ~/.local/bin/rem +``` + +### Build from source + +```bash +git clone https://github.com/BRO3886/rem.git +cd rem +make build +# Binary is at ./bin/rem +``` + +## Requirements + +- macOS 13+ (uses EventKit + private ReminderKit bridge for all reads and writes via go-eventkit, AppleScript only for default list name query) +- Xcode Command Line Tools (for building from source — cgo/clang + framework headers) +- First run will prompt for Reminders app access in System Settings > Privacy & Security + +## Quick Start + +```bash +# List all reminder lists +rem lists --count + +# Create a reminder with an alarm +rem add "Buy groceries" --list Personal --due tomorrow --priority high --remind-me 15m + +# Create with tags (parsed from title + --tags flag) +rem add "Review PR #work #urgent" --tags "deploy" + +# Location reminder: fires when arriving (default) or leaving +rem add "Buy milk" --location "37.3318,-122.0312" --radius 200 +rem add "Take out trash" --location "37.3318,-122.0312" --on-leave + +# List incomplete reminders +rem list --list Work --incomplete + +# Search reminders +rem search "meeting" + +# Show reminder details +rem show + +# Complete a reminder +rem complete + +# Show statistics +rem stats +``` + +## Commands + +### Reminders + +```bash +# Create +rem add "Title" [--list LIST] [--due DATE] [--priority high|medium|low] [--notes TEXT] [--url URL] [-F/--flagged] [-t/--tags TAGS] [-r/--remind-me DURATION] [--repeat PATTERN] [--location "LAT,LNG"] [--radius METERS] [--on-arrive|--on-leave] +rem add -i # Interactive creation + +# List +rem list [--list LIST] [--incomplete] [--completed] [--flagged] [--due-before DATE] [--due-after DATE] [-o json|table|plain] +rem ls # Alias + +# Show +rem show # Full or partial ID +rem get -o json + +# Update +rem update [-t/--title TEXT] [--due DATE] [--priority LEVEL] [--notes TEXT] [--url URL] [--add-tags TAGS] [--remove-tags TAGS] [-r/--remind-me DURATION] [--repeat PATTERN] [--list LIST] [--location "LAT,LNG"|none] [--radius METERS] [--on-arrive|--on-leave] + +# Complete / Uncomplete (support multiple IDs) +rem complete [id2 id3...] +rem done # Alias +rem uncomplete [id2 id3...] + +# Flag / Unflag (support multiple IDs) +rem flag [id2 id3...] +rem unflag [id2 id3...] + +# Delete (supports multiple IDs) +rem delete [id2 id3...] # Asks for confirmation +rem rm --force # Skip confirmation (-f / --yes / -y also work) + +# Today — due and overdue reminders +rem today +``` + +### Lists + +```bash +# View all lists +rem lists +rem lists --count # Show reminder counts + +# Create a list +rem list-mgmt create "My List" +rem lm new "Shopping" # Alias + +# Rename a list +rem list-mgmt rename "Old Name" "New Name" + +# Delete a list +rem list-mgmt delete "Name" # Asks for confirmation +rem lm rm "Name" --force +``` + +### Search & Analytics + +```bash +rem search "query" [--list LIST] [--incomplete] +rem stats # Overall statistics +rem overdue # Overdue reminders +rem upcoming [--days 7] # Upcoming due dates +``` + +### Import / Export + +```bash +# Export +rem export --list Work --format json > work.json +rem export --format csv --output-file reminders.csv +rem export --incomplete --format json + +# Import +rem import work.json +rem import reminders.csv --list "Imported" +rem import --dry-run data.json # Preview without creating +``` + +### Interactive Mode + +```bash +rem interactive # Full interactive menu +rem i # Alias +rem add -i # Interactive add +``` + +### Output Formats + +All list/show commands support `--output` (`-o`): + +```bash +rem list -o table # Default, formatted table +rem list -o json # Machine-readable JSON +rem list -o plain # Simple text +rem list -o json | jq '.[].name' # Pipe to jq +``` + +Color output respects `NO_COLOR`: +```bash +NO_COLOR=1 rem list +rem list --no-color +``` + +### AI Agent Skills + +```bash +rem skills install # Interactive picker (shows confirmation prompt) +rem skills install --agent claude # Claude Code only +rem skills install --agent all # All supported agents +rem skills install --dry-run # Preview files without writing +rem skills status # Check installation status +rem skills uninstall # Remove the skill +``` + +### Shell Completions + +```bash +# Bash +rem completion bash > /usr/local/etc/bash_completion.d/rem + +# Zsh +rem completion zsh > "${fpath[1]}/_rem" + +# Fish +rem completion fish > ~/.config/fish/completions/rem.fish +``` + +## Date Parsing + +Date parsing is powered by [`go-eventkit/dateparser`](https://github.com/BRO3886/go-eventkit): + +| Input | Meaning | +|-------|---------| +| `now` | Current date and time | +| `today` | Today at 9:00 AM | +| `tomorrow` | Tomorrow at 9:00 AM | +| `next monday` | Next Monday at 9:00 AM | +| `monday 2pm` | Next Monday at 2:00 PM | +| `next friday at 2pm` | Next Friday at 2:00 PM | +| `in 2 days` | 2 days from now | +| `in 3 hours` | 3 hours from now | +| `5 days ago` | 5 days before now | +| `eod` / `end of day` | Today at 5:00 PM | +| `this week` | End of current week | +| `next week` | Next Monday at 9:00 AM | +| `next month` | 1st of next month at 9:00 AM | +| `mar 15` | March 15 at 9:00 AM | +| `5pm` | Today (or tomorrow) at 5:00 PM | +| `today 5pm` | Today at 5:00 PM | +| `2026-02-15` | February 15, 2026 | +| `2026-02-15 14:30` | February 15, 2026 at 2:30 PM | + +## Go API + +rem is powered by [**go-eventkit**](https://github.com/BRO3886/go-eventkit) — use it directly for programmatic access to macOS Reminders in your own Go programs: + +```bash +go get github.com/BRO3886/go-eventkit +``` + +```go +package main + +import ( + "fmt" + "time" + + "github.com/BRO3886/go-eventkit/reminders" +) + +func main() { + client, err := reminders.New() + if err != nil { + panic(err) + } + + // Create a reminder + due := time.Now().Add(24 * time.Hour) + r, err := client.CreateReminder(reminders.CreateReminderInput{ + Title: "Buy groceries", + ListName: "Personal", + DueDate: &due, + Priority: reminders.PriorityHigh, + }) + if err != nil { + panic(err) + } + fmt.Println("Created:", r.ID) + + // List incomplete reminders + items, _ := client.Reminders( + reminders.WithList("Personal"), + reminders.WithCompleted(false), + ) + for _, item := range items { + fmt.Printf("- %s (due: %v)\n", item.Title, item.DueDate) + } + + // Complete a reminder + client.CompleteReminder(r.ID) + + // Get all lists + lists, _ := client.Lists() + for _, l := range lists { + fmt.Printf("%s (%d reminders)\n", l.Title, l.Count) + } +} +``` + +See the [go-eventkit README](https://github.com/BRO3886/go-eventkit) for the full API reference. + +## Architecture + +``` +rem/ +├── cmd/rem/ # CLI entry point +│ ├── main.go +│ └── commands/ # Cobra command definitions +├── internal/ +│ ├── service/ # Service layer wrapping go-eventkit (AppleScript only for default list name) +│ ├── reminder/ # Domain models (Reminder, List, Priority) +│ ├── export/ # JSON & CSV import/export +│ ├── skills/ # Agent skill install/uninstall/status +│ ├── update/ # Background update check (GitHub releases) +│ └── ui/ # Table formatting, colored output +├── skills/rem-cli/ # Embedded agent skill files +├── website/ # Hugo documentation site +├── Makefile +├── LICENSE +└── README.md +``` + +**All reads and writes** — including reminder CRUD and list CRUD — go through `go-eventkit` (`github.com/BRO3886/go-eventkit`) — an Objective-C EventKit bridge compiled into the binary via cgo. Direct in-process access to the Reminders store, no IPC. All operations complete in under 200ms. + +**Flagged, tag, and list-sharing operations** use the private ReminderKit bridge in go-eventkit — EventKit doesn't expose these properties, but `REMReminder.flagged`, `REMReminder.hashtags`, and `REMList.isShared` do. Tags degrade gracefully if the private API becomes unavailable. AppleScript is only used for the default list name query. + +**Shared lists** work like any other list for creates, reads, updates, flags, tags, and deletes. Moving a reminder across a shared-list boundary is the one exception: macOS refuses a true move there (even Apple's own apps copy and delete behind the scenes), so rem does the same — the reminder is copied to the target with all fields intact, the original is deleted, and a warning with the new ID is printed to stderr. + +## Performance + +Tested with 224 reminders across 12 lists: + +| Command | Time | +|---------|------| +| `rem lists` | 0.12s | +| `rem list` (all 224) | 0.13s | +| `rem show` (by prefix) | 0.11s | +| `rem search` | 0.11s | +| `rem stats` | 0.17s | + +See [Performance docs](https://rem.sidv.dev/docs/performance/) for the full optimization story (JXA at 60s → EventKit at 0.13s). + +## Known Limitations + +- **macOS only** — requires EventKit framework and osascript +- **No subtasks** — not exposed via EventKit +- **Tags, flagged, and list-sharing state use private API** — reads/writes go through Apple's private ReminderKit framework since EventKit doesn't expose these properties. If Apple changes the private API in a future macOS version, tags and flagged degrade gracefully (the reminder is still created/updated, a warning is printed — see [#44](https://github.com/BRO3886/rem/issues/44)) and lists simply report as unshared +- **No true move across shared-list boundaries** — macOS refuses it at the account level, so rem moves to/from shared lists by copy + delete; the reminder gets a new ID (warned on stderr). See [#50](https://github.com/BRO3886/rem/issues/50) +- **Immutable lists** cannot be renamed or deleted (system lists like Siri suggestions) + +## License + +MIT