Problem
On MySQL, POST /events sometimes fails with a 500:
{"code":"EN-5001","message":"Event publish failed","description":"Failed to publish event."}
The same request succeeds when repeated immediately. A real publisher would get the same 500, so
this is a product bug, not only a flaky test.
Evidence from CI
Across 44 recent CI database runs (2026-09-22 to 2026-09-28: upstream PR E2E runs plus the fork's
nightlies and release dry runs), it happened 3 times in 22 MySQL runs, and 0 times in 22
PostgreSQL and h2 runs:
The failure is in each job's Playwright report (the playwright-report-mysql artifact), and in the job log's
annotations. Each time the test was seeding an event through publishMarkedEventViaApi
(dpdp-integration-test-suite/utils/eventNotificationSetup.ts:95) and got 500 instead of 201. The
retry passed. It hits different tests, so it is not tied to one test's data.
Where the error comes from
EventPublishServiceImpl.persistEvent wraps every RuntimeException from the publish transaction
in this EN-5001, so any database error during a publish looks like this. The real exception is
logged with LOG.error("Failed to publish event [...]") in wso2carbon.log. None of these jobs
failed, so e2e.yml uploaded no server logs, and the underlying exception hasn't been seen yet.
Likely cause (not yet verified)
Failing within 0.1s points to an immediate database error, most likely a deadlock, not a
lock-wait timeout (MySQL's default innodb_lock_wait_timeout is 50s).
The publish transaction takes two SELECT … FOR UPDATE locks
(EventNotificationCommonDBQueries.java):
- The topic:
getActiveTopicByOrgAndNameForUpdateQuery,
… FROM TOPIC WHERE ORG_ID = ? AND LOWER(NAME) = LOWER(?) AND LOWER(STATUS) = 'active' FOR UPDATE.
LOWER(NAME) can't use an index, so MySQL scans the org's range in UQ_TOPIC_ORG_ID and, in
InnoDB's default REPEATABLE READ, locks every topic row in the org it scans, plus the gaps
between them. The index that would find the one row, UQ_TOPIC_ORG_ACTIVE_NAME (ORG_ID, ACTIVE_NAME), isn't used because the query filters on LOWER(NAME), not ACTIVE_NAME.
- The subscriptions it fans out to:
getActiveSubscriptionsForFanOutQuery,
… FROM SUBSCRIPTION s WHERE ORG_ID = ? AND EXISTS (… SUBSCRIPTION_TOPIC …) AND STATUS = 'active' ORDER BY SUBSCRIPTION_ID FOR UPDATE.
Registering a subscription goes the other way. It writes the SUBSCRIPTION row, then
inserts SUBSCRIPTION_TOPIC rows, whose foreign key FK_ST_TOPIC makes InnoDB take shared locks on
the referenced TOPIC rows. A publish holding every topic in the org, waiting for a subscription
row, and a subscription change holding that row, waiting for a topic, is a deadlock. The E2E suite
runs event tests in parallel in one org, so this is exactly the load that would trigger it.
This would also explain why only MySQL fails: PostgreSQL and H2 lock only the rows that match, and
don't take gap locks.
How to confirm
- Capture the real exception: reproduce on MySQL, or have CI upload
carbon-logs-<db> when a test
was flaky, not only when the job fails. Look for Failed to publish event in wso2carbon.log.
- Right after a failure, run
SHOW ENGINE INNODB STATUS and read LATEST DETECTED DEADLOCK. It
names both transactions and the locks each held and wanted.
- A local reproduction: on a MySQL-backed install, publish events on one topic in a loop while
another loop registers subscriptions on other topics in the same org.
Possible fixes (once confirmed)
- Lock only the one topic row: look the topic up by
(ORG_ID, ACTIVE_NAME) so
UQ_TOPIC_ORG_ACTIVE_NAME is used, or read it without FOR UPDATE and lock it by TOPIC_ID.
- Take locks in one consistent order in every transaction that touches both topics and
subscriptions.
- Retry the publish transaction once on a deadlock (MySQL SQLState
40001, error 1213). That's
the standard handling for InnoDB deadlocks, and safe here because the whole transaction rolls
back.
Related
Found while checking #301 (flaky E2E tests). Once this is fixed, 09.07.05 and 09.08.06 should
stop being flaky on MySQL.
Problem
On MySQL,
POST /eventssometimes fails with a 500:The same request succeeds when repeated immediately. A real publisher would get the same 500, so
this is a product bug, not only a flaky test.
Evidence from CI
Across 44 recent CI database runs (2026-09-22 to 2026-09-28: upstream PR E2E runs plus the fork's
nightlies and release dry runs), it happened 3 times in 22 MySQL runs, and 0 times in 22
PostgreSQL and h2 runs:
3637980346009.07.05(multi-tenant)36409658934, E2E (mysql)09.07.05(super-tenant)3641720591009.08.06(multi-tenant)The failure is in each job's Playwright report (the
playwright-report-mysqlartifact), and in the job log'sannotations. Each time the test was seeding an event through
publishMarkedEventViaApi(
dpdp-integration-test-suite/utils/eventNotificationSetup.ts:95) and got 500 instead of 201. Theretry passed. It hits different tests, so it is not tied to one test's data.
Where the error comes from
EventPublishServiceImpl.persistEventwraps everyRuntimeExceptionfrom the publish transactionin this
EN-5001, so any database error during a publish looks like this. The real exception islogged with
LOG.error("Failed to publish event [...]")inwso2carbon.log. None of these jobsfailed, so
e2e.ymluploaded no server logs, and the underlying exception hasn't been seen yet.Likely cause (not yet verified)
Failing within 0.1s points to an immediate database error, most likely a deadlock, not a
lock-wait timeout (MySQL's default
innodb_lock_wait_timeoutis 50s).The publish transaction takes two
SELECT … FOR UPDATElocks(
EventNotificationCommonDBQueries.java):getActiveTopicByOrgAndNameForUpdateQuery,… FROM TOPIC WHERE ORG_ID = ? AND LOWER(NAME) = LOWER(?) AND LOWER(STATUS) = 'active' FOR UPDATE.LOWER(NAME)can't use an index, so MySQL scans the org's range inUQ_TOPIC_ORG_IDand, inInnoDB's default REPEATABLE READ, locks every topic row in the org it scans, plus the gaps
between them. The index that would find the one row,
UQ_TOPIC_ORG_ACTIVE_NAME (ORG_ID, ACTIVE_NAME), isn't used because the query filters onLOWER(NAME), notACTIVE_NAME.getActiveSubscriptionsForFanOutQuery,… FROM SUBSCRIPTION s WHERE ORG_ID = ? AND EXISTS (… SUBSCRIPTION_TOPIC …) AND STATUS = 'active' ORDER BY SUBSCRIPTION_ID FOR UPDATE.Registering a subscription goes the other way. It writes the
SUBSCRIPTIONrow, theninserts
SUBSCRIPTION_TOPICrows, whose foreign keyFK_ST_TOPICmakes InnoDB take shared locks onthe referenced
TOPICrows. A publish holding every topic in the org, waiting for a subscriptionrow, and a subscription change holding that row, waiting for a topic, is a deadlock. The E2E suite
runs event tests in parallel in one org, so this is exactly the load that would trigger it.
This would also explain why only MySQL fails: PostgreSQL and H2 lock only the rows that match, and
don't take gap locks.
How to confirm
carbon-logs-<db>when a testwas flaky, not only when the job fails. Look for
Failed to publish eventinwso2carbon.log.SHOW ENGINE INNODB STATUSand readLATEST DETECTED DEADLOCK. Itnames both transactions and the locks each held and wanted.
another loop registers subscriptions on other topics in the same org.
Possible fixes (once confirmed)
(ORG_ID, ACTIVE_NAME)soUQ_TOPIC_ORG_ACTIVE_NAMEis used, or read it withoutFOR UPDATEand lock it byTOPIC_ID.subscriptions.
40001, error1213). That'sthe standard handling for InnoDB deadlocks, and safe here because the whole transaction rolls
back.
Related
Found while checking #301 (flaky E2E tests). Once this is fixed,
09.07.05and09.08.06shouldstop being flaky on MySQL.