Put an admission boundary before the drain
A queue worker cannot drain a fixed set of jobs while it continues claiming new ones. Separate shutdown into an admission decision, completion of already-owned work and dependency cleanup. Give each phase an observable boundary so a deployment can prove which work was accepted before termination.
The local model in this article has two worker slots and six queued jobs. At tick 5, shutdown begins. Stopping admission and allowing completion through tick 16 finishes four jobs, interrupts none and leaves two unclaimed. Continuing to poll until the same deadline also finishes four jobs, but interrupts two newly claimed jobs. Waiting longer did not compensate for keeping admission open.
This is a discrete-time simulation. It runs no operating-system processes, Kubernetes Pods or message broker. Its purpose is to expose an ordering problem before a platform-specific implementation makes it difficult to see.
Draw the three boundaries your deployment must cross
First, stop acquiring new responsibility. For a queue consumer, that means stopping polling or subscription delivery according to the broker contract. An HTTP readiness response does not necessarily stop a background consumer from claiming messages. A process serving both traffic types has more than one admission path.
Second, identify the jobs the worker already owns. Let them complete within the remaining shutdown budget or follow the documented handoff policy. Finally, close connections and flush the operational records needed to explain the result. Closing the database pool before its jobs finish converts an orderly signal into application failures.
BullMQ's graceful-shutdown documentation describes closing a worker so it stops taking new jobs and waits for current jobs to finish. It also warns that this close operation does not time out by itself. The process hosting a worker therefore still needs a bounded termination design. BullMQ: graceful shutdown.
Kubernetes documents a termination grace period, lifecycle hooks and forced termination after the available grace expires. A pre-stop hook consumes time within that lifecycle. Do not budget the hook, the application drain and cleanup as though each receives a separate full grace interval. Kubernetes: Pod termination.
Read the schedule, not only the completed counter
Jobs A through F have durations of 2, 4, 7, 11, 17 and 23 ticks. A and B start at tick 0. C starts when A finishes at tick 2, and D starts when B finishes at tick 4. When shutdown arrives at tick 5, C and D are still active.
| Shutdown policy | Completed | Interrupted | Still unclaimed |
|---|---|---|---|
| Exit at tick 5 | 2 | 2 | 2 |
| Keep polling until tick 16 | 4 | 2 | 0 |
| Stop polling at 5, drain to 16 | 4 | 0 | 2 |
| Stop polling at 5, drain only to 12 | 3 | 1 | 2 |
With admission closed, C finishes at tick 9 and D at tick 15. The tick-16 boundary leaves one tick after the last completion for proposed cleanup. Cleanup itself is not simulated. With polling still open, E starts at tick 9 and F at tick 15; both remain active at forced exit.
The short-window case is equally useful: correct admission handling cannot make D finish by tick 12. A termination budget must fit the allowed remaining work, or the design must support a safe interruption and resumption path. Setting a larger timeout without bounding job duration simply moves the uncertainty.
Download the scheduler model and four-case output. Run python3 experiment.py in a writable folder. Completion events at the deadline are processed before the model stops; the report states integer ticks, not measured seconds.
Preserve the message contract across interruption
An interrupted job is not automatically a lost job. A broker may redeliver it after ownership expires. Conversely, a completed business action is not automatically a completed message if the worker dies before acknowledging it. The model deliberately does not simulate acknowledgements or external effects, so its completed counter cannot establish exactly-once delivery.
Place acknowledgement after the required durable result according to the queue's delivery semantics. Make repeated execution safe where redelivery is possible. For a job that changes external state, test a crash after the effect but before acknowledgement. Review the existing background-job fencing-token example when a previous owner can return after a replacement starts.
Do not release ownership early while the original operation continues without protection. Two workers can then act on the same job. If cancellation is cooperative, identify the points where the task actually observes it and whether it can leave a partial result. A function receiving a cancellation request is different from the function having stopped.
Size the budget around the remaining work
Use the maximum supported remaining job time, not only average runtime, to decide whether full completion is a reasonable shutdown promise. A report-generation job that can run for hours needs checkpoints or a different lifecycle from a short notification task.
Reserve time for cleanup and for the platform's signal-delivery path. Measure those components in the real environment. The model's one-tick cleanup allowance illustrates a boundary; it is not a recommended grace period for any deployment.
If a worker multiplexes many job types, consider separate pools with different shutdown contracts. That prevents a long-running export from defining the termination behavior of every short task. The database connection-pool saturation analysis is relevant when draining workers and replacement workers temporarily compete for the same database capacity.
Prove that no new ownership appears after the cutoff
Record the shutdown signal, the admission cutoff and every job acquisition using a consistent clock or ordered event log. The first release criterion is that no acquisition begins after the cutoff. The next is that every previously acquired job has a recorded completion or a deliberate recovery disposition.
- Start jobs with known remaining durations, trigger a normal deployment and compare the acquisition log with the cutoff.
- Force termination before the longest job completes and observe redelivery and duplicate-effect handling.
- Make dependency cleanup slow and verify that it stays inside the process's total grace budget.
- Repeat while replacement workers start, so shared database and queue limits are part of the test.
Publish the supported job-duration and recovery contract in the worker runbook. A deployment can then fail for a specific reason, such as a post-cutoff acquisition or an unaccounted job, instead of treating every process exit with code zero as successful draining.
Sources
Documentation checked .

Continue the conversation
Comments (0)