Files
2026-07-03 00:38:05 +09:00

2.3 KiB

title, created, updated, type, tags, sources, confidence
title created updated type tags sources confidence
Claude Science 2026-06-30 2026-06-30 entity
llm
agent
automation
tool
evaluation
raw/articles/claude-science-ai-workbench-2026.md
medium

Claude Science

Claude Science is Anthropic's beta AI workbench for scientific research. It packages Claude as a coordinating agent inside a research environment that connects to scientific databases, packages, local or remote compute, and domain-specific skills. Unlike a general chat assistant, the product emphasizes auditable artifacts: figures, manuscripts, code, environment details, and message history are kept together so results can be validated and reproduced later.^[raw/articles/claude-science-ai-workbench-2026.md]

The product is especially relevant to ai-research-automation because it turns research work into a persistent, tool-connected loop rather than a one-off answer. Claude Science can run locally on macOS/Linux or against remote machines over SSH/HPC login nodes, ask before reaching new resources, submit jobs, and fork sessions to compare approaches. A reviewer agent checks citations, calculations, untraceable numbers, and figure/code consistency, which connects the product to ai-evaluation-infrastructure and loop-engineering.

Design signals

  • Auditable artifacts: outputs include the code and environment that produced them, plus plain-language explanations and message history.
  • Compute as part of the loop: local machines, lab infrastructure, HPC, and Modal-style on-demand compute are treated as execution targets behind explicit review/revoke decisions.
  • Domain skills and connectors: the beta ships with 60+ curated skills/connectors for areas such as genomics, single-cell, proteomics, structural biology, and cheminformatics.
  • Reviewer agents: separate critic/reviewer agents are used to catch citation and calculation errors, echoing the generator/evaluator split in loop-engineering.

Open Questions

  • How much of the reproducibility guarantee depends on preserving exact execution environments versus preserving narrative provenance?
  • Can the reviewer-agent pattern transfer to personal wiki curation, code review, and scheduled research jobs without becoming too expensive?
  • What privacy boundary is acceptable when sensitive lab data remains local but context is still sent to a hosted model?