Repository navigation
Events with equal timestamps reload out of append order (SqliteSessionService, DatabaseSessionService), breaking Workflow replay #7286
Description
Activity
Picking this one up now — opening a PR shortly. Flagging it here so nobody duplicates the work; if someone is already on it, say so and I will drop mine.
- addedservices[Component] This issue is related to runtime services, e.g. sessions, memory, artifacts, etc[Component] This issue is related to runtime services, e.g. sessions, memory, artifacts, etc
on Sep 28, 2026 Reproduced on main for both services @rishwanth-th. For SqliteSessionService it's actually a regression since v2.7.0: v2.6.3 still kept append order and the id tie-break added in 4ca975e is what changed it.
Until there's a fix for both a workaround that worked: a small append_event override that bumps a tied timestamp to last_event + 1µs before calling super(). Order survives a reload from a separate service instance too so it should also cover the cross-process case. Could you try it on your workflow loop?
@chelsealong the rowid change in #7287 looks right for SQLite. Could you switch "Fixes" to "Part of #7286" so this stays open for DatabaseSessionService and add a note about the rowid/VACUUM assumption? Please re-run the sessions tests after that before asking for review. FYI test_append_event_calls_rollback_on_commit_failure is flaky on Windows on main too so it's not from your change.
Thanks @surajksharma07, tried it on the workflow loop. It fixes the loop on both services, with two changes to hold on
DatabaseSessionServiceand across instances: compare with<=, and move in whole microseconds.Setup. The question loop from the report (a resumable
Workflow: wait → classify → answer → wait), ADK 2.9.2, Python 3.12.14, Windows 11. ADK's clock is forced to a 15.625 ms tick throughset_time_providerso events tie, and every resume gets a fresh session service and Runner, as after a restart or on another instance. 30 loops per run. "Two instances" alternates resumes between two clocks, one of them behind (0.5 s; 5 s on Postgres, where a resume takes longer than 0.5 s and nothing would move).SqliteSessionServiceDatabaseSessionService(aiosqlite)DatabaseSessionService(Postgres 17)no override failed 5 of 7 runs failed 2 of 3 passed 3 of 3 override with ==(only an equal timestamp moves),+ 1e-6failed 2 of 3 failed 1 of 3 passed 3 of 3 override with <=,+ 1e-6passed 3 of 3 passed 3 of 3 passed 3 of 3 the same, two instances passed 3 of 3 failed 3 of 3 (loop 6) failed 3 of 3 (loops 6–12) override with <=, next whole microsecond, with and without two instancespassed 6 of 6 passed 6 of 6 passed 9 of 9 Failures are all
Replay divergence detected.Why the two changes:
<=: once a tied event has been moved tolast + 1e-6, the next event in the same tick is behind the last one, not tied, so==leaves it in place. A clock that runs behind (another instance) does the same.- Whole microseconds: at today's epoch a float steps 2.4e-7 s, so
last + 1e-6adds 0.954 µs.DatabaseSessionServiceorders by itstimestampcolumn, which holds whole microseconds (the exact float inevent_datais only used after ordering), so a run of moves drifts until two events share a stored microsecond and reload by id.SqliteSessionServicestores the float, so it isn't affected.
async def append_event(self, session, event): if not event.partial and session.events: last_us = round(session.events[-1].timestamp * 1e6) if round(event.timestamp * 1e6) <= last_us: event.timestamp = (last_us + 1) / 1e6 return await super().append_event(session, event)
With that version our service passes its end-to-end suite: 24 scenarios twice, and three scenarios with two instances on one Postgres 17.
Two notes for the fix and its tests:
- A deterministic test for the
DatabaseSessionServicecase: append 40 events with the same timestamp and ids that sort in reverse, then reload. With+ 1e-6two of them swap at index 6; with whole microseconds they reload in append order. The loop test only fails when a swapped pair matters to replay, so it passes some runs. - For
DatabaseSessionServiceacross instances, ordering by a value assigned at insert (like the rowid in fix(sessions): preserve append order for tied timestamps in SqliteSessionService #7287 for SQLite) would take clocks out of it entirely.
Good catch on both points @rishwanth-th, and you're right that the 1µs version didn't really cover the cross-instance case. Checked your 40-event test on 2.10.0 with DatabaseSessionService (aiosqlite): last + 1e-6 swaps a pair every run (index 5 to 19 depending on the clock) while <= with whole microseconds reloads in append order every time. SqliteSessionService is fine with both since it stores the float.
So your version is the workaround to use until there's a fix. Both services still order by (timestamp, id) in 2.10.0 so no release has it yet.
@chelsealong could you switch #7287 to "Part of #7286" and add the rowid/VACUUM note when you get a chance? It's still at 4367fd9. For the DatabaseSessionService follow-up an insert-assigned sequence column (as you noted, a schema migration) is the only option that also survives clock skew between instances and the 40-event reverse-id test above would make a good deterministic regression test for it.
- added a commit that references this issue
on Oct 2, 2026 I reproduced the remaining DatabaseSessionService ordering problem on main
539a0d071, after #7287 landed. I appended 40 events with identical timestamps and reverse-sorted IDs, then read them through a separate service instance:DatabaseSessionService (sqlite+aiosqlite) append first: event_39; reload first: event_00 last five appended: event_04, event_03, event_02, event_01, event_00 num_recent_events=5: event_35, event_36, event_37, event_38, event_39The event timestamps themselves round-trip correctly. SqliteSessionService now preserves the tied-timestamp append order. With decreasing timestamps, both services still reorder the history, so using an insertion index only as a timestamp tie-break would leave clock skew unresolved. These checks used local SQLite; I have not tested PostgreSQL.
I also reproduced the impact with a model-free resumable Workflow (
first -> second -> wait_for_answer), rebuilding the Runner, service and graph after the pause. The six-scenario reproducer uses the public clock/ID providers and the unmodified 15-second replay barrier. DatabaseSessionService fails on tied timestamps/reverse IDs; its forward-ID and increasing-clock controls complete, as does SqliteSessionService with tied timestamps. Decreasing clocks cause replay divergence in both services. I ran it in a fresh Python 3.12 environment installed from the pinned commit, with no dev/test extras.I would like to implement the DatabaseSessionService follow-up. I found no open PR covering its insertion-order column. My proposed scope is:
- A new schema version with a per-session append sequence, allocated atomically in the same transaction as the event and state update. Event loading and stale-session checks would use that sequence, including the storage revision token: I also confirmed on SQLite that a stale snapshot is accepted when a later append has the same timestamp as the already-stored event, because the current timestamp marker does not advance.
- Preserve Event.timestamp and retain after_timestamp as a time filter; num_recent_events would select the latest appended events.
- Keep v0/v1 readable and provide an explicit source-to-destination migration through the existing migration runner, rather than changing old tables during startup. For historical events, backfill in the existing
(timestamp, id)order: the original append order is not recoverable from those schemas. - Regression tests for tied timestamps, a separate reader/service instance, recent-event limits, timestamp filters, stale writers, rollback, and migration. I would also validate SQLite and PostgreSQL separately.
Before I change the schema, is this direction welcome, and should the sequence define the primary event order or only break timestamp ties? I prefer primary append order for replay correctness, but that changes existing behavior for deliberately out-of-order timestamps. Please let me know if this follow-up is already being developed internally.
Reacted by JabrailConfirmed on current main @jabrailkhalil. Your 40-event reverse-id test still reloads out of order through DatabaseSessionService (aiosqlite) from a fresh instance while SqliteSessionService now keeps append order. Note that f33343f landed just after the v2.11.0 tag so no release has the SQLite fix yet. @rishwanth-th's append_event override is still the workaround on 2.11.0 for both services.
The stale-writer point holds too. The revision marker is session update_time which is set from event.timestamp so two writers loaded from the same snapshot both get through if they append at the same timestamp (a later timestamp is rejected as expected). That fits in your scope and deserves its own regression test.
The direction sounds right but the schema change and primary order vs tie-break only are the maintainers' call. @DeanChensj could you weigh in? Primary order gives the replay correctness you're after but it changes how deliberately out-of-order timestamps load and tie-break only leaves the clock-skew case open.
Thanks for confirming. I’ll include a stale-writer regression with identical timestamps. I’m keeping the schema change pending the maintainers’ decision on primary append order versus a tie-break.
Nothing has moved on main since (7d56ef8) @jabrailkhalil. DatabaseSessionService still breaks ties on (timestamp, id) and the 40-event reverse-id test still reloads out of order from a fresh instance while SqliteSessionService keeps append order. No release after v2.11.0 yet so @rishwanth-th's override is still the workaround.
While the schema question is open a draft PR with only the regression tests (tied timestamps across a fresh reader, num_recent_events and the same-timestamp stale writer) marked xfail for DatabaseSessionService would give the maintainers something concrete to decide on and those tests hold whichever order they pick.
Thanks @surajksharma07 — the tests-only draft is now open as #7492. It covers the fresh-reader 40-event history, recent-event limits and the same-timestamp stale writer, with strict assertion-only xfails for the five known failures. The Linux Python 3.11–3.14 validation matrix is green; production code and schemas are unchanged. I’m leaving it draft while the primary-order versus tie-break decision is open.
🔴 Required Information
Describe the Bug:
Session services reload a session's events ordered by
(timestamp, id). Event idsare random (
Event.new_id()), so when two events carry the same timestamp theycome back in a random-but-stable order, not in the order they were appended. This
happens in
SqliteSessionService(ORDER BY timestamp DESC, id DESC, with acomment saying the id tie-break makes the order stable across reads) and in
DatabaseSessionService(order_by(timestamp.desc(), id.desc())).For a resumable graph
Workflowthis breaks resume.ReplayManagerbuilds thereplay sequence barrier from the reloaded events, so a swapped pair gives the
barrier a completion order the workflow can never follow, and the resume fails
after the 15 s barrier timeout:
Equal timestamps are common: Python 3.12 on Windows ticks every 15.6 ms by default,
and fast function nodes emit several events per tick on any OS. It is a different
cause from #7027 (re-emitted outputs reordering the sequence in memory, fixed in
#7028): this one is in storage, and the repro below involves no workflow at all.
Steps to Reproduce:
pip install "google-adk[db]==2.9.2"(plusasyncpgfor Postgres)repro.pypython repro.py sqlite, orpython repro.py "sqlite+aiosqlite:///./r.db", orpython repro.py "postgresql+asyncpg://user:pass@localhost:5432/db"Expected Behavior:
get_sessionreturns events in append order, including events with equaltimestamps.
Observed Behavior:
The same for
DatabaseSessionServiceon SQLite and on Postgres 17.In a resumable
Workflow(a question loop: wait → classify → answer → wait,SqliteSessionService), forcing ties with a 15.625 ms clock throughgoogle.adk.platform.time.set_time_providermade resume fail with the divergenceabove at loop 7; the same workflow with a strictly increasing clock passed every
loop.
Environment Details:
Model Information:
🟡 Optional Information
Minimal Reproduction Code:
How often has this issue occurred?:
Additional Context:
Possible fixes: order by insertion rather than by a random id on ties (SQLite
rowid, or an autoincrement / sequence column inDatabaseSessionService), or makeevent ids monotonic.
User-side workaround: a strictly increasing clock through
google.adk.platform.time.set_time_provider(no two events in a process share atimestamp). It does not help across processes.