62/ 100
At risk

Suite health

78.2% of executions passed across the analysed window · 19 Jan 2026 – 7 Feb 2026

Runs analysed
20
Tests tracked
231
Executions
4,620
Browsers
3

Latest (Run 20): 231 tests — 171 passed, 54 failed, 6 skipped (74% pass rate).

Suite summary

Every one of the 231 tracked tests, by how it behaves across 20 runs.

Total tests
231
Passing
144
Passing on retry
3
Flaky
28
Newly failing
16
Consistently failing
34
Skipped
6
  • Passing consistently14462.3%
  • Flaky2812.1%
  • Consistently failing3414.7%
  • Newly failing166.9%
  • Recently fixed31.3%
  • Skipped62.6%

Flaky tests over time

Tests whose outcome alternates, counted per run.

28now · ranged 0–28 across the window
01530Run 1Run 20
Show values
Runflaky tests
Run 10
Run 20
Run 325
Run 425
Run 525
Run 625
Run 725
Run 825
Run 925
Run 1025
Run 1125
Run 1225
Run 1325
Run 1425
Run 1525
Run 1625
Run 1725
Run 1825
Run 1928
Run 2028

Up 28 since the first run — instability is spreading to more tests.

Retries per run

How hard the suite worked to recover, run by run.

47in the latest run · 41 average · 819 total
03060Run 1Run 20
Show values
Runretries
Run 125
Run 250
Run 325
Run 447
Run 546
Run 647
Run 746
Run 847
Run 925
Run 1050
Run 1125
Run 1247
Run 1346
Run 1447
Run 1546
Run 1647
Run 1725
Run 1856
Run 1925
Run 2047

Swings between 25 and 56 per run — retry load is uneven, so some runs are recovering far more than others.

Pass rate over the run window

Every analysed run, oldest first. Hover for that run's figures.

65%78%90%Run 1Run 20
Show run-by-run figures
RunPass rateFailedFlakyRetries
Run 181.4%40025
Run 271.9%62050
Run 382.7%372525
Run 481%412547
Run 573.6%582546
Run 681%412547
Run 773.6%582546
Run 881%412547
Run 982.7%372525
Run 1071.9%622550
Run 1182.7%372525
Run 1281%412547
Run 1373.6%582546
Run 1481%412547
Run 1573.6%582546
Run 1681%412547
Run 1782.7%372525
Run 1870.6%652556
Run 1982.7%372825
Run 2074%542847

What the analyzer recommends

Generated by the deterministic engine, in its own words.

  • Pass rate is below 80%. Review recent changes.

    High
  • 34 stable failure(s) consistently failing. Prioritize these.

    High
  • 28 flaky test(s) found. Review retry configurations and test isolation.

    Medium

Run highlights

The facts behind the numbers above.

  • Suite scope

    231 tests · 20 runs

    144 passed in every run.

  • Latest run (Run 20)

    74% pass rate

    171 passed, 54 failed, 6 skipped.

  • Flaky tests

    28 detected

    Outcomes alternate between pass and fail across the window.

  • Newly failing

    16 tests

    Passed earlier in the window and are failing now.

  • Consistently failing

    34 tests

    Failed in every run — reproducible, not intermittent.

  • Dominant failure category

    Network

    426 of the classified failures (24.1%) — the largest single concentration.

  • Retry behaviour

    819 retries

    3 test(s) recovered on retry; the rest failed again.

  • Browser spread

    webkit worst at 24.4%

    chromium is lowest at 18.5% — a 5.9 point spread.

  • Awaiting human review

    18 failures

    The engine declined to name a cause and routed these to a person.

Failure intelligence

20 deterministic rules analyze failure patterns before AI investigation.

SignalCategoryTestsConfidenceNext step

Needs manual investigation

Could not confidently determine a specific root cause — this failure needs a manual look

chromium, firefox, webkit

Unknown9
27%
Needs human review

Test is consistently broken — not flaky, genuinely failing

chromium, firefox, webkit

Stability6
99%
Needs human review

Network request failed at the protocol level — unreachable host, DNS failure, or TLS error

chromium, firefox, webkit

Network6
98%
Evidence supports this signal

Authentication required — the request was not authenticated

chromium, firefox, webkit

Authentication6
86%
Evidence supports this signal

A network request failed — API unreachable, timeout, or DNS failure

chromium, firefox, webkit

Network6
78%
Evidence supports this signal

The locator did not match any element in the DOM

chromium, firefox, webkit

Locator4
91%
Evidence supports this signal

Page or API response too slow

chromium, firefox, webkit

Timeout4
88%
Evidence supports this signal

Element exists in the DOM but is not visible — timing or rendering issue

chromium, firefox, webkit

Locator4
78%
Evidence supports this signal

The locator matched multiple elements — Playwright's strict mode requires a single match

chromium, firefox, webkit

Locator3
99%
Evidence supports this signal

Authorization denied — the authenticated user lacks required permissions

chromium, firefox, webkit

Authentication3
94%
Evidence supports this signal

The requested endpoint or page was not found

chromium, firefox, webkit

Network3
94%
Evidence supports this signal

Internal server error — the backend encountered an unexpected condition

chromium, firefox, webkit

Network3
94%
Evidence supports this signal

Expected title does not match actual page title

chromium, firefox, webkit

Assertion3
92%
Evidence supports this signal

The browser or page was unexpectedly closed during test execution

chromium, firefox, webkit

Environment3
89%
Evidence supports this signal

Expected value did not match actual value — test assertion failed

chromium, firefox, webkit

Assertion3
87%
Evidence supports this signal

Element was present but not ready for interaction when clicked

chromium, firefox, webkit

Locator3
82%
Evidence supports this signal

The network connection was abruptly terminated by the server or network

chromium, firefox, webkit

Network3
77%
Evidence supports this signal

Race condition or async timing issue — the test outcome depends on execution order

chromium, firefox, webkit

Stability3
67%
Needs human review

The element was removed from the DOM and re-rendered between location and interaction

webkit

Locator2
81%
Evidence supports this signal

Expected text does not match actual text on the page

firefox

Assertion1
85%
Evidence supports this signal

3 of 20 rules need human review because the available evidence isn't strong enough to support a conclusion.

Featured investigations

Three failures showing three different investigation outcomes.

Deterministic analysis

Runs before the AI, on every failure.

Rules fired
20/20
Failures analysed
78
  • Distinct failure signatures28
  • Escalated to human review18
  • Most common categoryNetwork
See all 20 detection rules →

Failure categories

How failures classify.

  • Element / selector2221.8%
  • Network2120.8%
  • Unclassified1615.8%
  • Backend / API1514.9%
  • Authentication98.9%
  • Assertion76.9%

Failure rate by browser

Share of executions that failed.

  • webkit27.3%21 of 77
  • firefox22.1%17 of 77
  • chromium20.8%16 of 77
Retries recorded819
Recovered on retry3 tests