What happened
After a shard reload triggered by a persistence timeout, transfer tasks for in-flight workflows are never delivered to the matching service.
The transfer queue reader's exclusiveReaderHighWatermark advances past these tasks' taskIds, so the queue processor considers them already processed. Affected workflows become permanently stuck — they do not self-heal even after 16+ hours.
Environment
- Temporal Server: v1.27.2 (self-hosted)
- Persistence: PostgreSQL + pgpool
- Affected service: history
Steps to reproduce
Not reliably reproducible on demand. Occurs under the following sequence:
- A shard (e.g. shard 321) fails to load due to a persistence timeout.
- The shard is re-acquired with an incremented range ID (e.g. 5 → 6).
- New or in-progress workflows on that shard generate transfer tasks (WorkflowTask / ActivityTask).
- Those transfer tasks are never delivered to matching.
Observed behavior
- Mutable state
workflowTaskRequestId: "emptyUuid"
- Workflow task state:
PENDING_WORKFLOW_TASK_STATE_SCHEDULED, never transitions to STARTED
- Workflow task attempt: Stuck at 1 (occasionally retries up to 4, still stuck)
- Matching task queue
approximateBacklogCount: 0 — matching never received the task
- Transfer queue
exclusiveReaderHighWatermark.taskId: Greater than the stuck task's transfer taskId (e.g. watermark = 6291553, stuck task ≈ 6291544)
- Workflow status: Permanently
RUNNING, or FAILED due to ScheduleToStart timeout
- Workers: Online, actively polling, receiving no tasks
Expected behavior
After a shard reload, the transfer queue processor should deliver all pending transfer tasks to matching,
including those created around the time of the reload. The exclusiveReaderHighWatermark should not advance past tasks that have not been successfully dispatched.
Relevant history service logs (in order)
Failed to load shard GetOrCreateShard: failed to get ShardID 321: context deadline exceeded Range updated for shardID 321 Task key range updated
Root cause analysis
The transfer queue reader's exclusiveReaderHighWatermark is set to a value that skips over transfer tasks that exist in persistence but were never delivered to matching. Once the watermark advances past a task's taskId, the queue processor will never attempt to deliver it again.
This creates a permanent task loss — not a transient delay.
This is distinct from the known RPC-level gap described in #9118 (crash between task creation and matching dispatch).
Here, the tasks are persisted but the queue reader state itself is incorrect after shard re-acquisition.
Possibly related issues
What happened
After a shard reload triggered by a persistence timeout, transfer tasks for in-flight workflows are never delivered to the matching service.
The transfer queue reader's
exclusiveReaderHighWatermarkadvances past these tasks'taskIds, so the queue processor considers them already processed. Affected workflows become permanently stuck — they do not self-heal even after 16+ hours.Environment
Steps to reproduce
Not reliably reproducible on demand. Occurs under the following sequence:
Observed behavior
workflowTaskRequestId:"emptyUuid"PENDING_WORKFLOW_TASK_STATE_SCHEDULED, never transitions toSTARTEDapproximateBacklogCount: 0 — matching never received the taskexclusiveReaderHighWatermark.taskId: Greater than the stuck task's transfertaskId(e.g. watermark = 6291553, stuck task ≈ 6291544)RUNNING, orFAILEDdue toScheduleToStarttimeoutExpected behavior
After a shard reload, the transfer queue processor should deliver all pending transfer tasks to matching,
including those created around the time of the reload. The
exclusiveReaderHighWatermarkshould not advance past tasks that have not been successfully dispatched.Relevant history service logs (in order)
Failed to load shard GetOrCreateShard: failed to get ShardID 321: context deadline exceeded Range updated for shardID 321 Task key range updated
Root cause analysis
The transfer queue reader's
exclusiveReaderHighWatermarkis set to a value that skips over transfer tasks that exist in persistence but were never delivered to matching. Once the watermark advances past a task'staskId, the queue processor will never attempt to deliver it again.This creates a permanent task loss — not a transient delay.
This is distinct from the known RPC-level gap described in #9118 (crash between task creation and matching dispatch).
Here, the tasks are persisted but the queue reader state itself is incorrect after shard re-acquisition.
Possibly related issues
Unavailableblip resets History queue backoff, causing a sustained retry storm #11547 — Queue processor saturation after persistence errors (related: ack-level starvation, but opposite direction — ack level fails to advance rather than advancing too far)