Skip to content

Transfer tasks permanently lost after shard reload — exclusiveReaderHighWatermark advances past undelivered tasks #12300

Description

@zx1986

What happened

After a shard reload triggered by a persistence timeout, transfer tasks for in-flight workflows are never delivered to the matching service.
The transfer queue reader's exclusiveReaderHighWatermark advances past these tasks' taskIds, so the queue processor considers them already processed. Affected workflows become permanently stuck — they do not self-heal even after 16+ hours.

Environment

  • Temporal Server: v1.27.2 (self-hosted)
  • Persistence: PostgreSQL + pgpool
  • Affected service: history

Steps to reproduce

Not reliably reproducible on demand. Occurs under the following sequence:

  1. A shard (e.g. shard 321) fails to load due to a persistence timeout.
  2. The shard is re-acquired with an incremented range ID (e.g. 5 → 6).
  3. New or in-progress workflows on that shard generate transfer tasks (WorkflowTask / ActivityTask).
  4. Those transfer tasks are never delivered to matching.

Observed behavior

  • Mutable state workflowTaskRequestId: "emptyUuid"
  • Workflow task state: PENDING_WORKFLOW_TASK_STATE_SCHEDULED, never transitions to STARTED
  • Workflow task attempt: Stuck at 1 (occasionally retries up to 4, still stuck)
  • Matching task queue approximateBacklogCount: 0 — matching never received the task
  • Transfer queue exclusiveReaderHighWatermark.taskId: Greater than the stuck task's transfer taskId (e.g. watermark = 6291553, stuck task ≈ 6291544)
  • Workflow status: Permanently RUNNING, or FAILED due to ScheduleToStart timeout
  • Workers: Online, actively polling, receiving no tasks

Expected behavior

After a shard reload, the transfer queue processor should deliver all pending transfer tasks to matching,
including those created around the time of the reload. The exclusiveReaderHighWatermark should not advance past tasks that have not been successfully dispatched.

Relevant history service logs (in order)

Failed to load shard GetOrCreateShard: failed to get ShardID 321: context deadline exceeded Range updated for shardID 321 Task key range updated  

Root cause analysis

The transfer queue reader's exclusiveReaderHighWatermark is set to a value that skips over transfer tasks that exist in persistence but were never delivered to matching. Once the watermark advances past a task's taskId, the queue processor will never attempt to deliver it again.
This creates a permanent task loss — not a transient delay.
This is distinct from the known RPC-level gap described in #9118 (crash between task creation and matching dispatch).
Here, the tasks are persisted but the queue reader state itself is incorrect after shard re-acquisition.

Possibly related issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions