Suite health
78.2% of executions passed across the analysed window · 19 Jan 2026 – 7 Feb 2026
- Runs analysed
- 20
- Tests tracked
- 231
- Executions
- 4,620
- Browsers
- 3
Latest (Run 20): 231 tests — 171 passed, 54 failed, 6 skipped (74% pass rate).
Suite summary
Every one of the 231 tracked tests, by how it behaves across 20 runs.
- Passing consistently14462.3%
- Flaky2812.1%
- Consistently failing3414.7%
- Newly failing166.9%
- Recently fixed31.3%
- Skipped62.6%
Flaky tests over time
Tests whose outcome alternates, counted per run.
Show values
| Run | flaky tests |
|---|---|
| Run 1 | 0 |
| Run 2 | 0 |
| Run 3 | 25 |
| Run 4 | 25 |
| Run 5 | 25 |
| Run 6 | 25 |
| Run 7 | 25 |
| Run 8 | 25 |
| Run 9 | 25 |
| Run 10 | 25 |
| Run 11 | 25 |
| Run 12 | 25 |
| Run 13 | 25 |
| Run 14 | 25 |
| Run 15 | 25 |
| Run 16 | 25 |
| Run 17 | 25 |
| Run 18 | 25 |
| Run 19 | 28 |
| Run 20 | 28 |
Up 28 since the first run — instability is spreading to more tests.
Retries per run
How hard the suite worked to recover, run by run.
Show values
| Run | retries |
|---|---|
| Run 1 | 25 |
| Run 2 | 50 |
| Run 3 | 25 |
| Run 4 | 47 |
| Run 5 | 46 |
| Run 6 | 47 |
| Run 7 | 46 |
| Run 8 | 47 |
| Run 9 | 25 |
| Run 10 | 50 |
| Run 11 | 25 |
| Run 12 | 47 |
| Run 13 | 46 |
| Run 14 | 47 |
| Run 15 | 46 |
| Run 16 | 47 |
| Run 17 | 25 |
| Run 18 | 56 |
| Run 19 | 25 |
| Run 20 | 47 |
Swings between 25 and 56 per run — retry load is uneven, so some runs are recovering far more than others.
Pass rate over the run window
Every analysed run, oldest first. Hover for that run's figures.
Show run-by-run figures
| Run | Pass rate | Failed | Flaky | Retries |
|---|---|---|---|---|
| Run 1 | 81.4% | 40 | 0 | 25 |
| Run 2 | 71.9% | 62 | 0 | 50 |
| Run 3 | 82.7% | 37 | 25 | 25 |
| Run 4 | 81% | 41 | 25 | 47 |
| Run 5 | 73.6% | 58 | 25 | 46 |
| Run 6 | 81% | 41 | 25 | 47 |
| Run 7 | 73.6% | 58 | 25 | 46 |
| Run 8 | 81% | 41 | 25 | 47 |
| Run 9 | 82.7% | 37 | 25 | 25 |
| Run 10 | 71.9% | 62 | 25 | 50 |
| Run 11 | 82.7% | 37 | 25 | 25 |
| Run 12 | 81% | 41 | 25 | 47 |
| Run 13 | 73.6% | 58 | 25 | 46 |
| Run 14 | 81% | 41 | 25 | 47 |
| Run 15 | 73.6% | 58 | 25 | 46 |
| Run 16 | 81% | 41 | 25 | 47 |
| Run 17 | 82.7% | 37 | 25 | 25 |
| Run 18 | 70.6% | 65 | 25 | 56 |
| Run 19 | 82.7% | 37 | 28 | 25 |
| Run 20 | 74% | 54 | 28 | 47 |
What the analyzer recommends
Generated by the deterministic engine, in its own words.
Pass rate is below 80%. Review recent changes.
High34 stable failure(s) consistently failing. Prioritize these.
High28 flaky test(s) found. Review retry configurations and test isolation.
Medium
Run highlights
The facts behind the numbers above.
Suite scope
231 tests · 20 runs
144 passed in every run.
Latest run (Run 20)
74% pass rate
171 passed, 54 failed, 6 skipped.
Flaky tests
28 detected
Outcomes alternate between pass and fail across the window.
Newly failing
16 tests
Passed earlier in the window and are failing now.
Consistently failing
34 tests
Failed in every run — reproducible, not intermittent.
Dominant failure category
Network
426 of the classified failures (24.1%) — the largest single concentration.
Retry behaviour
819 retries
3 test(s) recovered on retry; the rest failed again.
Browser spread
webkit worst at 24.4%
chromium is lowest at 18.5% — a 5.9 point spread.
Awaiting human review
18 failures
The engine declined to name a cause and routed these to a person.
Failure intelligence
20 deterministic rules analyze failure patterns before AI investigation.
| Signal | Category | Tests | Confidence | Next step |
|---|---|---|---|---|
Needs manual investigation Could not confidently determine a specific root cause — this failure needs a manual look chromium, firefox, webkit | Unknown | 9 | 27% | Needs human review |
Test is consistently broken — not flaky, genuinely failing chromium, firefox, webkit | Stability | 6 | 99% | Needs human review |
Network request failed at the protocol level — unreachable host, DNS failure, or TLS error chromium, firefox, webkit | Network | 6 | 98% | Evidence supports this signal |
Authentication required — the request was not authenticated chromium, firefox, webkit | Authentication | 6 | 86% | Evidence supports this signal |
A network request failed — API unreachable, timeout, or DNS failure chromium, firefox, webkit | Network | 6 | 78% | Evidence supports this signal |
The locator did not match any element in the DOM chromium, firefox, webkit | Locator | 4 | 91% | Evidence supports this signal |
Page or API response too slow chromium, firefox, webkit | Timeout | 4 | 88% | Evidence supports this signal |
Element exists in the DOM but is not visible — timing or rendering issue chromium, firefox, webkit | Locator | 4 | 78% | Evidence supports this signal |
The locator matched multiple elements — Playwright's strict mode requires a single match chromium, firefox, webkit | Locator | 3 | 99% | Evidence supports this signal |
Authorization denied — the authenticated user lacks required permissions chromium, firefox, webkit | Authentication | 3 | 94% | Evidence supports this signal |
The requested endpoint or page was not found chromium, firefox, webkit | Network | 3 | 94% | Evidence supports this signal |
Internal server error — the backend encountered an unexpected condition chromium, firefox, webkit | Network | 3 | 94% | Evidence supports this signal |
Expected title does not match actual page title chromium, firefox, webkit | Assertion | 3 | 92% | Evidence supports this signal |
The browser or page was unexpectedly closed during test execution chromium, firefox, webkit | Environment | 3 | 89% | Evidence supports this signal |
Expected value did not match actual value — test assertion failed chromium, firefox, webkit | Assertion | 3 | 87% | Evidence supports this signal |
Element was present but not ready for interaction when clicked chromium, firefox, webkit | Locator | 3 | 82% | Evidence supports this signal |
The network connection was abruptly terminated by the server or network chromium, firefox, webkit | Network | 3 | 77% | Evidence supports this signal |
Race condition or async timing issue — the test outcome depends on execution order chromium, firefox, webkit | Stability | 3 | 67% | Needs human review |
The element was removed from the DOM and re-rendered between location and interaction webkit | Locator | 2 | 81% | Evidence supports this signal |
Expected text does not match actual text on the page firefox | Assertion | 1 | 85% | Evidence supports this signal |
3 of 20 rules need human review because the available evidence isn't strong enough to support a conclusion.
Featured investigations
Three failures showing three different investigation outcomes.
Checkout — declined card retry
A payment confirmation test that fails, then passes when retried.
Billing — monthly invoice generation
An invoice test that fails on every run with a server error.
Settings — theme preference
A test that failed recently, with very little evidence captured.
Deterministic analysis
Runs before the AI, on every failure.
- Distinct failure signatures28
- Escalated to human review18
- Most common categoryNetwork
Failure categories
How failures classify.
- Element / selector2221.8%
- Network2120.8%
- Unclassified1615.8%
- Backend / API1514.9%
- Authentication98.9%
- Assertion76.9%
Failure rate by browser
Share of executions that failed.
- webkit27.3%21 of 77
- firefox22.1%17 of 77
- chromium20.8%16 of 77