Files
llm-wiki/concepts/e2e-coverage-metrics.md
2026-07-18 10:09:57 +09:00

4.7 KiB
Raw Permalink Blame History

title, created, updated, type, tags, sources, confidence
title created updated type tags sources confidence
E2E Coverage Metrics 2026-06-30 2026-07-17 concept
quality
reliability
evaluation
workflow
raw/articles/knowledgework-e2e-coverage-metrics-2026.md
raw/articles/github-code-coverage-merge-protection-2026.md
raw/articles/realworld-framework-comparison-spec-2026.md
raw/articles/slack-agentic-testing-e2e-stack-2026.md
medium

E2E Coverage Metrics

E2E coverage metrics are a way to measure whether end-to-end tests actually touch the product surfaces they are supposed to protect. KnowledgeWork's article argues that manually maintained "test cases written / test cases needed" lists drift as features change, so the denominator should be derived from implementation artifacts where possible.

The concrete pattern is to compute page coverage from all known product pages versus pages visited during Playwright runs, and RPC/API coverage from all service/method definitions versus RPCs observed in test traffic. In their setup, all pages are extracted from Next.js routes, all RPCs from .proto definitions, and the test-side observations come from Playwright trace network entries such as page-view and API requests.^[raw/articles/knowledgework-e2e-coverage-metrics-2026.md]

GitHub Code Quality's merge-protection preview shows the same idea being productized as a repository gate: branch rulesets can block pull requests when coverage falls below a minimum percentage, drops too far from the default branch, or both. Its evaluate mode is important operationally because teams can observe the effect of a quality threshold before turning it into a hard merge blocker.^[raw/articles/github-code-coverage-merge-protection-2026.md]

RealWorld adds a benchmark-design angle: many frontend and backend implementations share the same Medium-like app, API specification, backend spec tests, frontend E2E test suite, CSS theme, and hosted demo API. That makes it useful not only as framework learning material, but as a stable surface for comparing generated code, agent-built app variants, and cross-framework regression behavior under one contract. For ai-evaluation-infrastructure, the important part is the common spec/test harness, not the specific app clone.^[raw/articles/realworld-framework-comparison-spec-2026.md]

Slack Engineering's agentic-testing experiment adds the complementary question of what E2E tests are measuring. In 200+ runs, deterministic Playwright tests were fastest and CI-friendly, but agent-driven runs validated whether a user goal could be achieved through different valid UI paths. Their summary is useful: deterministic tests enforce journeys, while agents verify goals. This makes agentic testing more like an exploratory/debugging layer above ordinary E2E, not a replacement for regression gates.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]

The same Slack experiment also makes cost and observability part of the quality metric. Playwright MCP runs were more reliable than CLI-driven browser control on their flows, partly because the MCP harness returned stable browser state in fewer turns; however, agentic runs still cost roughly $15–30 and accumulated millions of tokens through repeated UI snapshots and conversation retransmission. For agent-harness-engineering, the reusable lesson is that goal coverage, action-signature diversity, turn count, snapshot volume, and failure cause should be measured alongside pass/fail.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]

This is useful for ci-cd-runtime-security and ai-evaluation-infrastructure because it treats test execution as observable runtime evidence, not just a green/red result. It also fits loop-engineering: the loop should store raw observations, compute metrics later, notify people in the place they already work, and keep enough history to change aggregation methods without rerunning old tests.

Caveat

Implementation coverage is a necessary-condition signal, not a sufficient proof of product quality. Visiting every page or calling every RPC does not guarantee that important user scenarios are asserted. The stronger pattern is to combine implementation-derived coverage with deterministic regression tests for known critical business paths.

Open Questions

  • Which surfaces should define the denominator for non-Next.js or non-RPC products: routes, OpenAPI endpoints, event names, domain actions, or user journeys?
  • When should low E2E coverage block a release, and when should it only produce a review item?
  • How can AI-generated test additions avoid optimizing for easy-to-cover surfaces while missing high-risk workflows?
  • For agentic E2E runs, what is the right denominator: required user goals, observed UI paths, meaningful action signatures, or historically flaky workflows?