add
This commit is contained in:
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: E2E Coverage Metrics
|
||||
created: 2026-06-30
|
||||
updated: 2026-07-01
|
||||
updated: 2026-07-17
|
||||
type: concept
|
||||
tags: [quality, reliability, evaluation, workflow]
|
||||
sources: [raw/articles/knowledgework-e2e-coverage-metrics-2026.md, raw/articles/github-code-coverage-merge-protection-2026.md, raw/articles/realworld-framework-comparison-spec-2026.md]
|
||||
sources: [raw/articles/knowledgework-e2e-coverage-metrics-2026.md, raw/articles/github-code-coverage-merge-protection-2026.md, raw/articles/realworld-framework-comparison-spec-2026.md, raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
@@ -18,6 +18,10 @@ GitHub Code Quality's merge-protection preview shows the same idea being product
|
||||
|
||||
RealWorld adds a benchmark-design angle: many frontend and backend implementations share the same Medium-like app, API specification, backend spec tests, frontend E2E test suite, CSS theme, and hosted demo API. That makes it useful not only as framework learning material, but as a stable surface for comparing generated code, agent-built app variants, and cross-framework regression behavior under one contract. For [[ai-evaluation-infrastructure]], the important part is the common spec/test harness, not the specific app clone.^[raw/articles/realworld-framework-comparison-spec-2026.md]
|
||||
|
||||
Slack Engineering's agentic-testing experiment adds the complementary question of **what E2E tests are measuring**. In 200+ runs, deterministic Playwright tests were fastest and CI-friendly, but agent-driven runs validated whether a user goal could be achieved through different valid UI paths. Their summary is useful: deterministic tests enforce journeys, while agents verify goals. This makes agentic testing more like an exploratory/debugging layer above ordinary E2E, not a replacement for regression gates.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
|
||||
The same Slack experiment also makes cost and observability part of the quality metric. Playwright MCP runs were more reliable than CLI-driven browser control on their flows, partly because the MCP harness returned stable browser state in fewer turns; however, agentic runs still cost roughly $15–30 and accumulated millions of tokens through repeated UI snapshots and conversation retransmission. For [[agent-harness-engineering]], the reusable lesson is that goal coverage, action-signature diversity, turn count, snapshot volume, and failure cause should be measured alongside pass/fail.^[raw/articles/slack-agentic-testing-e2e-stack-2026.md]
|
||||
|
||||
This is useful for [[ci-cd-runtime-security]] and [[ai-evaluation-infrastructure]] because it treats test execution as observable runtime evidence, not just a green/red result. It also fits [[loop-engineering]]: the loop should store raw observations, compute metrics later, notify people in the place they already work, and keep enough history to change aggregation methods without rerunning old tests.
|
||||
|
||||
## Caveat
|
||||
@@ -29,3 +33,4 @@ Implementation coverage is a necessary-condition signal, not a sufficient proof
|
||||
- Which surfaces should define the denominator for non-Next.js or non-RPC products: routes, OpenAPI endpoints, event names, domain actions, or user journeys?
|
||||
- When should low E2E coverage block a release, and when should it only produce a review item?
|
||||
- How can AI-generated test additions avoid optimizing for easy-to-cover surfaces while missing high-risk workflows?
|
||||
- For agentic E2E runs, what is the right denominator: required user goals, observed UI paths, meaningful action signatures, or historically flaky workflows?
|
||||
|
||||
Reference in New Issue
Block a user