--- title: Claude Science created: 2026-06-30 updated: 2026-06-30 type: entity tags: [llm, agent, automation, tool, evaluation] sources: [raw/articles/claude-science-ai-workbench-2026.md] confidence: medium --- # Claude Science Claude Science is Anthropic's beta AI workbench for scientific research. It packages Claude as a coordinating agent inside a research environment that connects to scientific databases, packages, local or remote compute, and domain-specific skills. Unlike a general chat assistant, the product emphasizes auditable artifacts: figures, manuscripts, code, environment details, and message history are kept together so results can be validated and reproduced later.^[raw/articles/claude-science-ai-workbench-2026.md] The product is especially relevant to [[ai-research-automation]] because it turns research work into a persistent, tool-connected loop rather than a one-off answer. Claude Science can run locally on macOS/Linux or against remote machines over SSH/HPC login nodes, ask before reaching new resources, submit jobs, and fork sessions to compare approaches. A reviewer agent checks citations, calculations, untraceable numbers, and figure/code consistency, which connects the product to [[ai-evaluation-infrastructure]] and [[loop-engineering]]. ## Design signals - **Auditable artifacts**: outputs include the code and environment that produced them, plus plain-language explanations and message history. - **Compute as part of the loop**: local machines, lab infrastructure, HPC, and Modal-style on-demand compute are treated as execution targets behind explicit review/revoke decisions. - **Domain skills and connectors**: the beta ships with 60+ curated skills/connectors for areas such as genomics, single-cell, proteomics, structural biology, and cheminformatics. - **Reviewer agents**: separate critic/reviewer agents are used to catch citation and calculation errors, echoing the generator/evaluator split in [[loop-engineering]]. ## Open Questions - How much of the reproducibility guarantee depends on preserving exact execution environments versus preserving narrative provenance? - Can the reviewer-agent pattern transfer to personal wiki curation, code review, and scheduled research jobs without becoming too expensive? - What privacy boundary is acceptable when sensitive lab data remains local but context is still sent to a hosted model?