Problem
The E2E suite retries failed tests: retries: process.env.CI ? 1 : 2 in
dpdp-integration-test-suite/playwright.config.ts. A test that fails and then passes on retry is
reported as flaky, and the run still goes green. Retries hide intermittent failures instead of
surfacing them, and some of those failures are real product behaviour that users can hit too.
Goal
Every test passes on its first attempt:
- Measure which tests are flaky.
- Find the root cause of each and fix it where it lives, whether that's the test, the portal, or
the server config.
- Then stop relying on retries to go green. Playwright 1.62 (the version in use) supports
failOnFlakyTests in config and --fail-on-flaky-tests on the CLI, so CI could fail a run that
only passed because of a retry.
Known cases
Measured across 44 recent CI database runs (2026-09-22 to 2026-09-28): 22 MySQL, 11 PostgreSQL and
11 h2, from the upstream PR E2E runs and the fork's nightlies and release dry runs.
| Case |
Seen in the 44 runs |
Status |
1. 03.03.01 silent sign-in fails |
0 |
Diagnosed, rare |
2. Deep link lands on /dashboard |
0 |
Undiagnosed |
3. EN-5001 on event publish (09.07.05, 09.08.06) |
3, all MySQL |
Tracked in #333 |
4. 08.09.01 test timeout |
1, MySQL |
Seen once |
1. 03.03.01 — purpose search by partial name (diagnosed, rare)
Observed: once, in a super-tenant run against a fresh IS 7.3.0 + PostgreSQL install
(172 passed, 1 flaky, 3 skipped). It passed on retry. It hasn't recurred in the 44 runs above.
Cause, from the Playwright trace and the IS logs:
- The test signs in, creates a purpose, then does a full page load of
/purposes, twice, about a
second apart. Each full page load makes the portal sign in again silently with
GET /oauth2/authorize, reusing the IS login session.
- The first two silent sign-ins succeeded. The third redirected back with
error=invalid_request&error_description=Invalid authorization request. At that moment IS
logged:
WARN {org.wso2.carbon.identity.oauth.endpoint.util.AuthzUtil} - Cannot find AuthenticationResult from the cache
- The portal showed "Unable to load your session." (
authorization.loadFailed, with a Try
again button). The purpose list never rendered.
- The test waited for the
Search by purpose name field until the 30-second test timeout. No
search request was ever sent; every API call that did go out returned 200 in 16–115 ms.
Not database-related: the warning appeared exactly once, at the failure. It is absent from a
full 356-test PostgreSQL run and from a MySQL-backed server's log.
Two things to check:
- Why IS misses its cached sign-in result on back-to-back silent authorize requests. This part
is not verified; it needs a look at IS's authorize flow.
- The portal doesn't retry a failed silent sign-in.
App.tsx goes straight to the
loadFailed screen, and useCurrentUserQuery is retry: false. A real user who reloads twice
quickly could hit the same screen.
Possible fixes:
- Test-side: reach the list through the portal's own navigation instead of a second
page.goto('purposes'), so there's no extra full sign-in. Small, but it only hides the glitch.
- Portal-side: retry a silent sign-in once on
error=invalid_request before showing
loadFailed. This also protects real users.
2. Deep-linked goto() sometimes lands on /dashboard (undiagnosed)
Recorded under "Known flakiness → Still open" in dpdp-integration-test-suite/TEST-SCENARIOS.md:
a deep-linked goto() occasionally ends up on /dashboard, and the test times out on a page that
never rendered. It hasn't recurred in the 44 runs above.
It goes through the same step as case 1: a full page load followed by the portal's silent
sign-in. It may be the same root cause, or a sibling; worth diagnosing both together.
A lead from the code, not verified: the portal saves the requested route
(rememberReturnPath in utils/authClient.ts, kept in sessionStorage) only on the path where no
sign-in code is already waiting. If a code left over from an earlier, interrupted sign-in is still
parked when a deep link loads, the route is never saved, and the app opens its home page. Confirming
it needs a new occurrence with its trace.
3. Event publish returns 500 EN-5001 on MySQL (tracked in #333)
09.07.05 and 09.08.06 failed their first attempt while seeding an event: POST /events returned
500 {"code":"EN-5001","message":"Event publish failed"} after 0.1s, and the retry passed. 3 times in
22 MySQL runs, never on PostgreSQL or h2. It's a server-side bug, likely a MySQL lock conflict in the
publish transaction. See #333 for the analysis and job links.
4. 08.09.01 — one test timeout on MySQL
08.09.01 hit its 60s test timeout once, in the fork's nightly
36308707757
(MySQL), and passed on retry. Seen once, so not diagnosed yet.
Suggested approach
- Measure. Run the full suite several times with
--retries=0 --workers=2 (the configured
local worker count), on MySQL and on PostgreSQL. Use the JSON reporter or HTML report to list
first-attempt failures, and keep each failed test's trace (trace: 'retain-on-failure' is
already set).
- Diagnose each case from its trace, as done for case 1: the action timeline, the page's
network requests, the last screencast frame, and the IS log at that timestamp.
- Fix at the source. Record each outcome in
TEST-SCENARIOS.md "Known flakiness", as the
04.05.06 entry does.
- Gate it. Once a few consecutive runs show 0 flaky, enable
failOnFlakyTests for CI. Keep
retries so that a run still produces a trace on failure, but no longer passes because of a
retry.
Problem
The E2E suite retries failed tests:
retries: process.env.CI ? 1 : 2indpdp-integration-test-suite/playwright.config.ts. A test that fails and then passes on retry isreported as flaky, and the run still goes green. Retries hide intermittent failures instead of
surfacing them, and some of those failures are real product behaviour that users can hit too.
Goal
Every test passes on its first attempt:
the server config.
failOnFlakyTestsin config and--fail-on-flaky-testson the CLI, so CI could fail a run thatonly passed because of a retry.
Known cases
Measured across 44 recent CI database runs (2026-09-22 to 2026-09-28): 22 MySQL, 11 PostgreSQL and
11 h2, from the upstream PR E2E runs and the fork's nightlies and release dry runs.
03.03.01silent sign-in fails/dashboardEN-5001on event publish (09.07.05,09.08.06)08.09.01test timeout1.
03.03.01— purpose search by partial name (diagnosed, rare)Observed: once, in a
super-tenantrun against a fresh IS 7.3.0 + PostgreSQL install(172 passed, 1 flaky, 3 skipped). It passed on retry. It hasn't recurred in the 44 runs above.
Cause, from the Playwright trace and the IS logs:
/purposes, twice, about asecond apart. Each full page load makes the portal sign in again silently with
GET /oauth2/authorize, reusing the IS login session.error=invalid_request&error_description=Invalid authorization request. At that moment ISlogged:
authorization.loadFailed, with a Tryagain button). The purpose list never rendered.
Search by purpose namefield until the 30-second test timeout. Nosearch request was ever sent; every API call that did go out returned 200 in 16–115 ms.
Not database-related: the warning appeared exactly once, at the failure. It is absent from a
full 356-test PostgreSQL run and from a MySQL-backed server's log.
Two things to check:
is not verified; it needs a look at IS's authorize flow.
App.tsxgoes straight to theloadFailedscreen, anduseCurrentUserQueryisretry: false. A real user who reloads twicequickly could hit the same screen.
Possible fixes:
page.goto('purposes'), so there's no extra full sign-in. Small, but it only hides the glitch.error=invalid_requestbefore showingloadFailed. This also protects real users.2. Deep-linked
goto()sometimes lands on/dashboard(undiagnosed)Recorded under "Known flakiness → Still open" in
dpdp-integration-test-suite/TEST-SCENARIOS.md:a deep-linked
goto()occasionally ends up on/dashboard, and the test times out on a page thatnever rendered. It hasn't recurred in the 44 runs above.
It goes through the same step as case 1: a full page load followed by the portal's silent
sign-in. It may be the same root cause, or a sibling; worth diagnosing both together.
A lead from the code, not verified: the portal saves the requested route
(
rememberReturnPathinutils/authClient.ts, kept insessionStorage) only on the path where nosign-in code is already waiting. If a code left over from an earlier, interrupted sign-in is still
parked when a deep link loads, the route is never saved, and the app opens its home page. Confirming
it needs a new occurrence with its trace.
3. Event publish returns
500 EN-5001on MySQL (tracked in #333)09.07.05and09.08.06failed their first attempt while seeding an event:POST /eventsreturned500 {"code":"EN-5001","message":"Event publish failed"}after 0.1s, and the retry passed. 3 times in22 MySQL runs, never on PostgreSQL or h2. It's a server-side bug, likely a MySQL lock conflict in the
publish transaction. See #333 for the analysis and job links.
4.
08.09.01— one test timeout on MySQL08.09.01hit its 60s test timeout once, in the fork's nightly36308707757(MySQL), and passed on retry. Seen once, so not diagnosed yet.
Suggested approach
--retries=0 --workers=2(the configuredlocal worker count), on MySQL and on PostgreSQL. Use the JSON reporter or HTML report to list
first-attempt failures, and keep each failed test's trace (
trace: 'retain-on-failure'isalready set).
network requests, the last screencast frame, and the IS log at that timestamp.
TEST-SCENARIOS.md"Known flakiness", as the04.05.06entry does.failOnFlakyTestsfor CI. Keepretriesso that a run still produces a trace on failure, but no longer passes because of aretry.