37 lines
4.7 KiB
Markdown
37 lines
4.7 KiB
Markdown
---
|
||
title: E2E Coverage Metrics
|
||
created: 2026-06-30
|
||
updated: 2026-07-17
|
||
type: concept
|
||
tags: [quality, reliability, evaluation, workflow]
|
||
sources: [raw/articles/knowledgework-e2e-coverage-metrics-2026.md, raw/articles/github-code-coverage-merge-protection-2026.md, raw/articles/realworld-framework-comparison-spec-2026.md, raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||
confidence: medium
|
||
---
|
||
|
||
# E2E Coverage Metrics
|
||
|
||
E2E coverage metrics are a way to measure whether end-to-end tests actually touch the product surfaces they are supposed to protect. KnowledgeWork's article argues that manually maintained "test cases written / test cases needed" lists drift as features change, so the denominator should be derived from implementation artifacts where possible.
|
||
|
||
The concrete pattern is to compute **page coverage** from all known product pages versus pages visited during Playwright runs, and **RPC/API coverage** from all service/method definitions versus RPCs observed in test traffic. In their setup, all pages are extracted from Next.js routes, all RPCs from `.proto` definitions, and the test-side observations come from Playwright trace network entries such as page-view and API requests.^[raw/articles/knowledgework-e2e-coverage-metrics-2026.md]
|
||
|
||
GitHub Code Quality's merge-protection preview shows the same idea being productized as a repository gate: branch rulesets can block pull requests when coverage falls below a minimum percentage, drops too far from the default branch, or both. Its evaluate mode is important operationally because teams can observe the effect of a quality threshold before turning it into a hard merge blocker.^[raw/articles/github-code-coverage-merge-protection-2026.md]
|
||
|
||
RealWorld adds a benchmark-design angle: many frontend and backend implementations share the same Medium-like app, API specification, backend spec tests, frontend E2E test suite, CSS theme, and hosted demo API. That makes it useful not only as framework learning material, but as a stable surface for comparing generated code, agent-built app variants, and cross-framework regression behavior under one contract. For [[ai-evaluation-infrastructure]], the important part is the common spec/test harness, not the specific app clone.^[raw/articles/realworld-framework-comparison-spec-2026.md]
|
||
|
||
Slack Engineering's agentic-testing experiment adds the complementary question of **what E2E tests are measuring**. In 200+ runs, deterministic Playwright tests were fastest and CI-friendly, but agent-driven runs validated whether a user goal could be achieved through different valid UI paths. Their summary is useful: deterministic tests enforce journeys, while agents verify goals. This makes agentic testing more like an exploratory/debugging layer above ordinary E2E, not a replacement for regression gates.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||
|
||
The same Slack experiment also makes cost and observability part of the quality metric. Playwright MCP runs were more reliable than CLI-driven browser control on their flows, partly because the MCP harness returned stable browser state in fewer turns; however, agentic runs still cost roughly $15–30 and accumulated millions of tokens through repeated UI snapshots and conversation retransmission. For [[agent-harness-engineering]], the reusable lesson is that goal coverage, action-signature diversity, turn count, snapshot volume, and failure cause should be measured alongside pass/fail.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||
|
||
This is useful for [[ci-cd-runtime-security]] and [[ai-evaluation-infrastructure]] because it treats test execution as observable runtime evidence, not just a green/red result. It also fits [[loop-engineering]]: the loop should store raw observations, compute metrics later, notify people in the place they already work, and keep enough history to change aggregation methods without rerunning old tests.
|
||
|
||
## Caveat
|
||
|
||
Implementation coverage is a necessary-condition signal, not a sufficient proof of product quality. Visiting every page or calling every RPC does not guarantee that important user scenarios are asserted. The stronger pattern is to combine implementation-derived coverage with deterministic regression tests for known critical business paths.
|
||
|
||
## Open Questions
|
||
|
||
- Which surfaces should define the denominator for non-Next.js or non-RPC products: routes, OpenAPI endpoints, event names, domain actions, or user journeys?
|
||
- When should low E2E coverage block a release, and when should it only produce a review item?
|
||
- How can AI-generated test additions avoid optimizing for easy-to-cover surfaces while missing high-risk workflows?
|
||
- For agentic E2E runs, what is the right denominator: required user goals, observed UI paths, meaningful action signatures, or historically flaky workflows?
|