Skip to content

E2E: make flaky tests pass on the first attempt instead of relying on retries #301

Description

@anjuchamantha

Problem

The E2E suite retries failed tests: retries: process.env.CI ? 1 : 2 in
dpdp-integration-test-suite/playwright.config.ts. A test that fails and then passes on retry is
reported as flaky, and the run still goes green. Retries hide intermittent failures instead of
surfacing them, and some of those failures are real product behaviour that users can hit too.

Goal

Every test passes on its first attempt:

  1. Measure which tests are flaky.
  2. Find the root cause of each and fix it where it lives, whether that's the test, the portal, or
    the server config.
  3. Then stop relying on retries to go green. Playwright 1.62 (the version in use) supports
    failOnFlakyTests in config and --fail-on-flaky-tests on the CLI, so CI could fail a run that
    only passed because of a retry.

Known cases

Measured across 44 recent CI database runs (2026-09-22 to 2026-09-28): 22 MySQL, 11 PostgreSQL and
11 h2, from the upstream PR E2E runs and the fork's nightlies and release dry runs.

Case Seen in the 44 runs Status
1. 03.03.01 silent sign-in fails 0 Diagnosed, rare
2. Deep link lands on /dashboard 0 Undiagnosed
3. EN-5001 on event publish (09.07.05, 09.08.06) 3, all MySQL Tracked in #333
4. 08.09.01 test timeout 1, MySQL Seen once

1. 03.03.01 — purpose search by partial name (diagnosed, rare)

Observed: once, in a super-tenant run against a fresh IS 7.3.0 + PostgreSQL install
(172 passed, 1 flaky, 3 skipped). It passed on retry. It hasn't recurred in the 44 runs above.

Cause, from the Playwright trace and the IS logs:

  1. The test signs in, creates a purpose, then does a full page load of /purposes, twice, about a
    second apart. Each full page load makes the portal sign in again silently with
    GET /oauth2/authorize, reusing the IS login session.
  2. The first two silent sign-ins succeeded. The third redirected back with
    error=invalid_request&error_description=Invalid authorization request. At that moment IS
    logged:
    WARN {org.wso2.carbon.identity.oauth.endpoint.util.AuthzUtil} - Cannot find AuthenticationResult from the cache
    
  3. The portal showed "Unable to load your session." (authorization.loadFailed, with a Try
    again
    button). The purpose list never rendered.
  4. The test waited for the Search by purpose name field until the 30-second test timeout. No
    search request was ever sent; every API call that did go out returned 200 in 16–115 ms.

Not database-related: the warning appeared exactly once, at the failure. It is absent from a
full 356-test PostgreSQL run and from a MySQL-backed server's log.

Two things to check:

  • Why IS misses its cached sign-in result on back-to-back silent authorize requests. This part
    is not verified; it needs a look at IS's authorize flow.
  • The portal doesn't retry a failed silent sign-in. App.tsx goes straight to the
    loadFailed screen, and useCurrentUserQuery is retry: false. A real user who reloads twice
    quickly could hit the same screen.

Possible fixes:

  • Test-side: reach the list through the portal's own navigation instead of a second
    page.goto('purposes'), so there's no extra full sign-in. Small, but it only hides the glitch.
  • Portal-side: retry a silent sign-in once on error=invalid_request before showing
    loadFailed. This also protects real users.

2. Deep-linked goto() sometimes lands on /dashboard (undiagnosed)

Recorded under "Known flakiness → Still open" in dpdp-integration-test-suite/TEST-SCENARIOS.md:
a deep-linked goto() occasionally ends up on /dashboard, and the test times out on a page that
never rendered. It hasn't recurred in the 44 runs above.

It goes through the same step as case 1: a full page load followed by the portal's silent
sign-in. It may be the same root cause, or a sibling; worth diagnosing both together.

A lead from the code, not verified: the portal saves the requested route
(rememberReturnPath in utils/authClient.ts, kept in sessionStorage) only on the path where no
sign-in code is already waiting. If a code left over from an earlier, interrupted sign-in is still
parked when a deep link loads, the route is never saved, and the app opens its home page. Confirming
it needs a new occurrence with its trace.

3. Event publish returns 500 EN-5001 on MySQL (tracked in #333)

09.07.05 and 09.08.06 failed their first attempt while seeding an event: POST /events returned
500 {"code":"EN-5001","message":"Event publish failed"} after 0.1s, and the retry passed. 3 times in
22 MySQL runs, never on PostgreSQL or h2. It's a server-side bug, likely a MySQL lock conflict in the
publish transaction. See #333 for the analysis and job links.

4. 08.09.01 — one test timeout on MySQL

08.09.01 hit its 60s test timeout once, in the fork's nightly
36308707757
(MySQL), and passed on retry. Seen once, so not diagnosed yet.

Suggested approach

  1. Measure. Run the full suite several times with --retries=0 --workers=2 (the configured
    local worker count), on MySQL and on PostgreSQL. Use the JSON reporter or HTML report to list
    first-attempt failures, and keep each failed test's trace (trace: 'retain-on-failure' is
    already set).
  2. Diagnose each case from its trace, as done for case 1: the action timeline, the page's
    network requests, the last screencast frame, and the IS log at that timestamp.
  3. Fix at the source. Record each outcome in TEST-SCENARIOS.md "Known flakiness", as the
    04.05.06 entry does.
  4. Gate it. Once a few consecutive runs show 0 flaky, enable failOnFlakyTests for CI. Keep
    retries so that a run still produces a trace on failure, but no longer passes because of a
    retry.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type/ImprovementMarks enhancements or improvements to existing features

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions