flake: add coding-agent validation workflow

This commit is contained in:
2026-08-06 20:16:17 +09:00
parent 5ebcbb4abf
commit 483d343f48
24 changed files with 1399 additions and 96 deletions
+21 -14
View File
@@ -95,26 +95,33 @@ points cannot express the requirement.
### 4. Prove the change
Run the validation matrix in
[references/review-checklist.md](references/review-checklist.md). At minimum:
Invoke the `validate-nix-change` skill and use the task-owned files as its
explicit path set. At minimum:
1. Format the task-owned files with the repository formatter and run
`git diff --check`.
2. Inspect the complete task diff for accidental files, duplication, leaked
1. Inspect `nix run .#check -- plan --paths <task-path>... --json` and confirm
the reported units, host classes, and real hosts are correct.
2. Run `nix run .#check -- fast --paths <task-path>...` during the edit loop.
3. Inspect the complete task diff for accidental files, duplication, leaked
secrets, forced values, direct enable assignments, and unrelated rewrites.
3. Run `nix flake check`.
4. Run `pre-commit run --all-files`.
5. Evaluate every affected real host without switching it. For a
cross-platform unit or profile, evaluate both NixOS and nix-darwin even if
only one class changed. Build an affected configuration with `--no-link`
when the current platform can build it.
6. Verify selection as well as syntax: confirm that the expected package,
4. Run `nix run .#check -- all --paths <task-path>...` after the structure is
complete. This evaluates every flake system and builds affected targets for
the current platform without activation.
5. Verify selection as well as syntax: confirm that the expected package,
program, service, group, cask, or external module appears in the resulting
configuration.
6. Add a `pkgs.testers.runNixOSTest` check through the `test-nixos-service`
skill when service startup or another runtime contract cannot be proved by
evaluation and a system build.
Use `nix run .#check -- full` only for CI, scheduled maintenance, or an explicit
repository-wide audit. These commands never activate the live system. Do not
run `nh os switch`, `nixos-rebuild switch`, `darwin-rebuild switch`,
`home-manager switch`, or an equivalent activation command as validation.
If a command is unavailable, blocked by the environment, or fails for a
pre-existing reason, diagnose it and report the exact gap. Never silently skip
a required check or weaken the implementation to make a check pass.
pre-existing reason, invoke the `debug-nix-failure` skill, diagnose it, and
report the exact gap. Never silently skip a required check or weaken the
implementation to make a check pass.
### 5. Audit before completion
@@ -72,56 +72,47 @@ user-owned mutable state unless the requested policy explicitly owns it.
## Validation matrix
Run checks from the repository root and keep the exact results for the handoff.
Do not switch or activate a live system merely to validate a change.
Use the `validate-nix-change` skill and run checks from the repository root.
Keep exact results for the handoff. The validation app never switches or
activates a live system.
### Always
1. Format task-owned files. If the worktree contains unrelated user changes,
pass only task-owned paths to the configured formatter when supported.
2. Run `git diff --check`.
3. Review `git status --short`, `git diff --stat`, and the complete `git diff`.
4. Run `nix flake check`.
5. Run `pre-commit run --all-files`.
1. Run `nix run .#check -- plan --paths <task-path>... --json` and inspect the
affected units and hosts.
2. Run `nix run .#check -- fast --paths <task-path>...` during implementation.
3. Review `git status --short`, `git diff --stat`, and the complete task diff.
4. Run `nix run .#check -- all --paths <task-path>...` before handoff. It runs
all-system evaluation and compatible targeted builds without activation.
5. Reserve `nix run .#check -- full` for CI, scheduled maintenance, or an
explicit repository-wide audit.
Do not run `nh os switch`, `nixos-rebuild switch`, `darwin-rebuild switch`,
`home-manager switch`, or an equivalent activation command.
### NixOS or Home Manager on NixOS
- Evaluate each affected host's system toplevel derivation.
- Build at least one affected NixOS configuration with `--no-link` when the
current system supports it.
- Confirm that the validation plan includes each affected real NixOS host.
- Build affected NixOS configurations with `--no-link` through the validation
app when the current system supports them.
- Inspect the resulting option that proves selection: for example
`environment.systemPackages`, the user's `home.packages`,
`systemd.services`, `users.users.<name>.extraGroups`, or the upstream
`programs`/`services` option.
Typical build shape:
```sh
nix build .#nixosConfigurations.<host>.config.system.build.toplevel --no-link
```
### nix-darwin or Home Manager on Darwin
- Evaluate every affected Darwin host even when running on Linux.
- Confirm that every affected Darwin host is evaluated even when running on
Linux.
- Inspect `homebrew.casks` or `homebrew.brews` for Homebrew-backed additions.
- Evaluate the relevant Home Manager program or package option.
- Build a Darwin configuration only on a compatible Darwin builder; otherwise
report that build as an explicit runtime-validation gap.
Typical evaluation shapes:
```sh
nix eval --raw .#darwinConfigurations.<host>.system.drvPath
nix eval --json .#darwinConfigurations.<host>.config.homebrew.casks
```
Confirm the exact output attribute against the current flake before using a
command; do not paste these shapes blindly.
- Build a Darwin configuration only on a compatible Darwin runner or builder;
otherwise report the build as an explicit platform gap.
### Profiles and cross-platform changes
- Determine transitive selection through `meta.includes`, not only direct
mentions.
- Confirm transitive selection through `meta.includes`, not only direct
mentions. The validation plan computes reverse dependency closure.
- Evaluate every real host selecting the changed profile.
- Evaluate both host classes for a cross-platform profile, even if only one
current fragment changed.
@@ -134,6 +125,8 @@ command; do not paste these shapes blindly.
### Runtime-dependent behavior
Evaluation and builds cannot prove GUI appearance, credentials, network access,
hardware behavior, or successful daemon interaction. State the precise manual
post-activation check needed for those behaviors. Never describe evaluation as
hardware behavior, successful daemon interaction, or reboot state. For
reusable NixOS behavior, use the `test-nixos-service` skill and add a
`pkgs.testers.runNixOSTest` check. State the precise manual post-activation check
needed for physical hardware or external systems. Never describe evaluation as
a runtime test.
+63
View File
@@ -0,0 +1,63 @@
---
name: debug-nix-failure
description: Diagnose failures from parsing, Nix module evaluation, derivation builds, flake checks, NixOS tests, or Home Manager activation logs without changing the live system. Use when `nix run .#check`, `nix flake check`, `nix build`, CI, or a user-provided activation log fails.
---
# Debug a Nix Failure
Classify the failure before changing code. Preserve the original command,
complete error, first causal frame, and affected attribute. Never run a live
switch or activation to reproduce a validation failure.
## Identify the failing layer
- **Parse or format:** syntax location, malformed string, unmatched delimiter,
or formatter-owned rewrite.
- **Static analysis:** dead binding, suspicious expression, ShellCheck finding,
secret scan, or workflow lint.
- **Module evaluation:** missing option, wrong type, assertion, infinite
recursion, conflicting definitions, Registry selection, or unsupported host
class.
- **Derivation instantiation/build:** missing dependency, hash mismatch, patch
failure, compiler/test failure, sandbox violation, or unsupported platform.
- **NixOS test:** failed unit, timeout, command assertion, network readiness, or
reboot state.
- **Activation/runtime:** filesystem conflict, activation script, systemd unit,
hardware, credential, or external-service behavior. Diagnose only from logs
supplied by the user unless they explicitly request a non-switch inspection
command.
## Reproduce the narrowest failing operation
Start with the stage and paths reported by the validation app:
```sh
nix run .#check -- plan --paths <task-path>... --json
nix run .#check -- fast --paths <task-path>...
nix run .#check -- eval --paths <task-path>...
nix run .#check -- build --paths <task-path>...
```
For a single attribute, use `nix eval --show-trace` on its `drvPath` before a
build. For a failed derivation, retain `--print-build-logs` and inspect
`nix log <drv-path>` when the summary omits the causal lines.
## Read traces selectively
1. Find the first repository-owned frame or option path.
2. Separate the immediate failure from wrapper frames in `modules.nix`,
`lib.evalModules`, or flake-parts.
3. Inspect the option declaration and every definition contributing to it.
4. Confirm package and option names against locked inputs, not memory.
5. Check whether the failure reproduces on the base revision before calling it
task-owned.
Do not respond to a type or ownership error with import-order changes,
`lib.mkForce`, global arguments, or an overlay unless repository evidence shows
that those mechanisms are the correct owner.
## Finish with a bounded diagnosis
Report the failing layer, root cause, minimal correction, rerun command, and any
remaining platform or runtime gap. Include enough of the error to identify it,
but do not paste large unrelated logs.
+64
View File
@@ -0,0 +1,64 @@
---
name: test-nixos-service
description: Add or extend a non-activating NixOS VM or container test for service startup, sockets, timers, permissions, firewall behavior, reboot state, and inter-service dependencies. Use when evaluation and a system build cannot prove the requested runtime behavior.
---
# Test NixOS Runtime Behavior
Prefer `pkgs.testers.runNixOSTest` for reusable NixOS behavior that can be
proved without the user's physical machine. Do not activate the host
configuration and do not substitute a live `nh os switch` for a deterministic
test.
## Define the observable contract
List the runtime facts that must hold, such as:
- a systemd unit reaches `active`;
- a socket or port is listening;
- a timer triggers its service;
- a user can or cannot read a file;
- a group membership grants access;
- a firewall permits one path and blocks another;
- state survives a reboot;
- one service waits for another dependency.
Exclude behavior that requires physical GPU, fingerprint, audio, display,
Secure Boot, TPM, private credentials, or an external provider unless the test
can model it explicitly.
## Implement the smallest useful machine
Create a test under `tests/` and expose it through `checks.<system>`. Import the
owning module or Registry selection instead of copying its implementation into
the test. Use only the packages, users, files, and network peers required by the
contract.
Typical shape:
```nix
pkgs.testers.runNixOSTest {
name = "service-name";
nodes.machine = {
# Enable the owning unit or import the module under test.
};
testScript = ''
machine.start()
machine.wait_for_unit("service-name.service")
machine.succeed("systemctl is-active service-name.service")
'';
}
```
Use `wait_for_unit`, `wait_for_open_port`, `succeed`, `fail`, and explicit
reboots to express outcomes. Avoid arbitrary sleeps when a readiness condition
exists.
## Validate and report
Run the targeted test through its flake check, then run the repository
validation app for the task paths. Report the test attribute and assertions
that passed. State clearly which hardware or external behavior remains outside
the VM/container model.
+58
View File
@@ -0,0 +1,58 @@
---
name: update-flake-input
description: Update one or more pinned flake inputs with bounded lock-file changes and non-activating Linux and Darwin validation. Use for dependency refreshes, input-specific updates, automated lock-file pull requests, or diagnosing a regression introduced by flake.lock.
---
# Update a Flake Input
Keep the update scope explicit and treat `flake.lock` as generated dependency
state. Never activate a host as part of this workflow; do not run `nh os switch`,
`nixos-rebuild switch`, `darwin-rebuild switch`, `home-manager switch`, or an
equivalent command.
## Bound the update
1. Read `AGENTS.md`, inspect `git status --short`, and preserve unrelated work.
2. Record the input names and the behavior or version change being requested.
3. Prefer an input-specific update:
```sh
nix flake update <input-name>
```
Use an unrestricted `nix flake update` only when the task explicitly requests a
full refresh. Do not hand-edit lock nodes.
## Audit the lock diff
Inspect the complete `flake.lock` diff. Confirm that changed nodes are the
requested inputs or unavoidable followers and that source owners, repositories,
reference types, and hashes remain expected. Investigate unexpected node
replacement, disappearing followers, or a large transitive graph rewrite before
validation.
## Validate without activation
Run the common validation workflow against the lock file:
```sh
nix run .#check -- plan --paths flake.lock --json
nix run .#check -- fast --paths flake.lock
nix run .#check -- eval --paths flake.lock
nix run .#check -- build --paths flake.lock
```
A lock-file change requires full evaluation and the complete native check set.
Linux and Darwin builds must run on compatible runners. The scheduled update
workflow uploads the candidate lock file, builds Linux and Darwin checks, and
creates a pull request only after both pass.
When a failure appears only after the update, invoke `debug-nix-failure`, compare
the failing derivation or option with the base lock, and narrow the responsible
input before adding an override or patch.
## Report the result
List requested and transitively changed inputs, validation commands and native
platform results, any package or option migration, and remaining manual runtime
checks. Evaluation or a native build is not activation.
+128
View File
@@ -0,0 +1,128 @@
---
name: validate-nix-change
description: Plan and run efficient, non-activating validation for edits to this NixOS, nix-darwin, and Home Manager flake. Use after changing Nix modules, hosts, profiles, overlays, flake outputs, tests, scripts, CI, or agent configuration; before handing off a task; or when deciding which real hosts must be evaluated or built.
---
# Validate a Nix Change
Use the repository validation app as the source of truth for change impact and
validation commands. It derives affected hosts from Registry ownership,
`meta.includes`, host selections, fragment class, and host-local paths.
Never activate a live configuration as part of this workflow. Do not run
`nh os switch`, `nixos-rebuild switch`, `darwin-rebuild switch`,
`home-manager switch`, or an equivalent activation command. The user owns live
activation separately.
## Establish the validation scope
1. Read `AGENTS.md` and run `git status --short` before editing.
2. Preserve unrelated user changes. Track the paths owned by the current task,
including newly created untracked files.
3. Inspect the plan before expensive checks:
```sh
nix run .#check -- plan --paths <task-path>... --json
```
When validating a committed pull-request range, use:
```sh
nix run .#check -- plan --base <base-sha> --json
```
The app automatically uses a `path:` flake reference when task paths are
untracked, so newly created Registry fragments are visible to Nix without
staging them.
## Run checks in increasing cost order
### Fast edit loop
After each coherent edit, parse Nix files, validate project JSON, TOML, and
skill frontmatter, check whitespace, and run the configured hooks only for
task-owned files:
```sh
nix run .#check -- fast --paths <task-path>...
```
Do not replace this with `pre-commit run --all-files` during the edit loop.
Unrelated repository files must not become part of the task merely because an
existing check fails elsewhere.
### Evaluation
After the implementation is structurally complete, evaluate every affected
NixOS and Darwin derivation plus the supporting checks without realizing or
activating them. Flake-wide paths additionally evaluate every flake system:
```sh
nix run .#check -- eval --paths <task-path>...
```
This proves module evaluation, option types, assertions, Registry selection,
and derivation instantiation. It does not prove a successful build or runtime
behavior.
### Compatible builds
Build affected configurations for the current platform with no result link:
```sh
nix run .#check -- build --paths <task-path>...
```
The app reports incompatible targets as evaluated but skipped for native build.
A Darwin target must be built by a compatible Darwin runner or builder; a Linux
evaluation is not a Darwin build.
### Final task validation
Before handoff, run the cumulative task check. It applies file checks only to
task-owned paths, evaluates every flake system, and builds affected native
targets:
```sh
nix run .#check -- all --paths <task-path>...
```
Use `--all-hosts` only when a deliberate audit must report every registered
host as affected. Use the exhaustive command for CI, scheduled maintenance, or
an explicit repository-wide audit:
```sh
nix run .#check -- full
```
`full` runs hooks over every tracked file and builds every check for the current
platform through `nix-fast-build`.
## Add runtime tests when needed
Evaluation and builds do not prove service startup, socket behavior, firewall
rules, users and groups, permissions, reboot behavior, or network interaction.
For reusable NixOS behavior, invoke the `test-nixos-service` skill and add a
`pkgs.testers.runNixOSTest` check. Hardware, credentials, GUI appearance, and
external services may still require a precisely described manual check after
the user activates the configuration.
## Diagnose failures by layer
Invoke the `debug-nix-failure` skill when a stage fails. Fix the first failing
layer before running a more expensive one. Do not hide a pre-existing failure,
weaken an assertion, add `lib.mkForce`, or skip a required host merely to make
the task appear green.
## Report evidence precisely
Conclude with:
- task-owned paths;
- affected units and hosts reported by `plan`;
- each command run and its result;
- which targets were parsed, evaluated, built, or runtime-tested;
- any compatible-platform or manual-runtime gap.
Never describe evaluation as a build, a build as activation, or a VM test as
proof of hardware-specific behavior.