Queue reliability

Worker shutdown: stop taking jobs before draining

Graceful worker shutdown starts by stopping new job acquisition. Then drain owned work within a bounded deadline and close dependencies. In this simulation, closing admission avoids two interruptions that a longer polling window does not prevent.

Model
Two slots, six jobs, integer ticks
Shutdown signal
Tick 5
Stop polling; drain to 16
Four complete, none interrupted, two unclaimed
Scope
No process, broker or acknowledgement test
A closed amber gate stops new parcels while blue parcels already inside continue toward completion.
Conceptual illustration of stopping admission while already-owned jobs finish.

Put an admission boundary before the drain

A queue worker cannot drain a fixed set of jobs while it continues claiming new ones. Separate shutdown into an admission decision, completion of already-owned work and dependency cleanup. Give each phase an observable boundary so a deployment can prove which work was accepted before termination.

The local model in this article has two worker slots and six queued jobs. At tick 5, shutdown begins. Stopping admission and allowing completion through tick 16 finishes four jobs, interrupts none and leaves two unclaimed. Continuing to poll until the same deadline also finishes four jobs, but interrupts two newly claimed jobs. Waiting longer did not compensate for keeping admission open.

This is a discrete-time simulation. It runs no operating-system processes, Kubernetes Pods or message broker. Its purpose is to expose an ordering problem before a platform-specific implementation makes it difficult to see.

Draw the three boundaries your deployment must cross

First, stop acquiring new responsibility. For a queue consumer, that means stopping polling or subscription delivery according to the broker contract. An HTTP readiness response does not necessarily stop a background consumer from claiming messages. A process serving both traffic types has more than one admission path.

Second, identify the jobs the worker already owns. Let them complete within the remaining shutdown budget or follow the documented handoff policy. Finally, close connections and flush the operational records needed to explain the result. Closing the database pool before its jobs finish converts an orderly signal into application failures.

STOP ADMISSION / Stop polling / delivery / Record the cutoff / Track already-owned jobs; DRAIN OWNED WORK / Finish within deadline / Record interrupted work / Preserve recovery contract; CLOSE DEPENDENCIES / Flush required records / Close connections last / Stay inside total grace
Figure 1. Proposed architecture. Proposed deployment sequence; the fixture simulates acquisition and completion. View full-size figure.

BullMQ's graceful-shutdown documentation describes closing a worker so it stops taking new jobs and waits for current jobs to finish. It also warns that this close operation does not time out by itself. The process hosting a worker therefore still needs a bounded termination design. BullMQ: graceful shutdown.

Kubernetes documents a termination grace period, lifecycle hooks and forced termination after the available grace expires. A pre-stop hook consumes time within that lifecycle. Do not budget the hook, the application drain and cleanup as though each receives a separate full grace interval. Kubernetes: Pod termination.

Read the schedule, not only the completed counter

Jobs A through F have durations of 2, 4, 7, 11, 17 and 23 ticks. A and B start at tick 0. C starts when A finishes at tick 2, and D starts when B finishes at tick 4. When shutdown arrives at tick 5, C and D are still active.

Six jobs across four executed termination policies
Shutdown policyCompletedInterruptedStill unclaimed
Exit at tick 5222
Keep polling until tick 16420
Stop polling at 5, drain to 16402
Stop polling at 5, drain only to 12312

With admission closed, C finishes at tick 9 and D at tick 15. The tick-16 boundary leaves one tick after the last completion for proposed cleanup. Cleanup itself is not simulated. With polling still open, E starts at tick 9 and F at tick 15; both remain active at forced exit.

EXIT AT 5 / 2 complete / 2 interrupted / 2 unclaimed; POLL TO 16 / 4 complete / 2 interrupted / 0 unclaimed; STOP AT 5, DRAIN / To 16: 4 / 0 / 2 / To 12: 3 / 1 / 2
Figure 2. Executed local model. Counts follow complete / interrupted / unclaimed order. View full-size figure.

The short-window case is equally useful: correct admission handling cannot make D finish by tick 12. A termination budget must fit the allowed remaining work, or the design must support a safe interruption and resumption path. Setting a larger timeout without bounding job duration simply moves the uncertainty.

Download the scheduler model and four-case output. Run python3 experiment.py in a writable folder. Completion events at the deadline are processed before the model stops; the report states integer ticks, not measured seconds.

Preserve the message contract across interruption

An interrupted job is not automatically a lost job. A broker may redeliver it after ownership expires. Conversely, a completed business action is not automatically a completed message if the worker dies before acknowledging it. The model deliberately does not simulate acknowledgements or external effects, so its completed counter cannot establish exactly-once delivery.

Place acknowledgement after the required durable result according to the queue's delivery semantics. Make repeated execution safe where redelivery is possible. For a job that changes external state, test a crash after the effect but before acknowledgement. Review the existing background-job fencing-token example when a previous owner can return after a replacement starts.

Do not release ownership early while the original operation continues without protection. Two workers can then act on the same job. If cancellation is cooperative, identify the points where the task actually observes it and whether it can leave a partial result. A function receiving a cancellation request is different from the function having stopped.

Size the budget around the remaining work

Use the maximum supported remaining job time, not only average runtime, to decide whether full completion is a reasonable shutdown promise. A report-generation job that can run for hours needs checkpoints or a different lifecycle from a short notification task.

Reserve time for cleanup and for the platform's signal-delivery path. Measure those components in the real environment. The model's one-tick cleanup allowance illustrates a boundary; it is not a recommended grace period for any deployment.

If a worker multiplexes many job types, consider separate pools with different shutdown contracts. That prevents a long-running export from defining the termination behavior of every short task. The database connection-pool saturation analysis is relevant when draining workers and replacement workers temporarily compete for the same database capacity.

Prove that no new ownership appears after the cutoff

Record the shutdown signal, the admission cutoff and every job acquisition using a consistent clock or ordered event log. The first release criterion is that no acquisition begins after the cutoff. The next is that every previously acquired job has a recorded completion or a deliberate recovery disposition.

  • Start jobs with known remaining durations, trigger a normal deployment and compare the acquisition log with the cutoff.
  • Force termination before the longest job completes and observe redelivery and duplicate-effect handling.
  • Make dependency cleanup slow and verify that it stays inside the process's total grace budget.
  • Repeat while replacement workers start, so shared database and queue limits are part of the test.

Publish the supported job-duration and recovery contract in the worker runbook. A deployment can then fail for a specific reason, such as a post-cutoff acquisition or an unaccounted job, instead of treating every process exit with code zero as successful draining.

Sources

Documentation checked .

  1. Kubernetes: pod lifecycle and termination
  2. BullMQ: graceful shutdown

Continue the conversation

Comments (0)

    Leave a comment

    Your name and comment stay in this page and are cleared after the spam check.

    10–2,000 characters. Keep the discussion relevant to this article.

    Spam protection verification
    Spam protection loads when you begin the form.

    JavaScript is required to use this form and its spam protection.