BullMQ stalled jobs: what they are and how to find them
Checked against BullMQ 6.3.4 on 4 October 2026 · All guides
A queue shows active: 8. Three of those eight have no worker behind them anymore, and nothing in BullMQ's states tells you which. That is a stalled job: the most common reason a BullMQ job runs twice, and the one that is hardest to see.
Short answer. A stalled job is an active job whose worker stopped renewing its lock. BullMQ notices on its next stalled-job check, moves the job back to wait and another worker runs it again. Stalled is not a state, so you find it through the stalled event, job.stalledCounter, or the <prefix>:<queue>:stalled SET in Redis.
How a job stalls
When a worker takes a job, it writes a lock key next to it (bull:payments:42:lock) that expires after lockDuration, 30 seconds by default. While the job runs, the worker renews that lock every lockRenewTime, half of lockDuration by default.
The job stalls when the lock expires before it is renewed. In practice there are three causes:
- The worker process died: a deploy that kills pods without draining, an out-of-memory kill, a crash.
- The event loop was blocked for longer than
lockDuration. The renewal is a timer on the same thread as your processor, so 40 seconds of synchronous work (parsing a huge file, a CPU-heavy loop) starves it. - The worker lost Redis long enough for the lock to run out.
What BullMQ does about it
Every worker runs a stalled-job check every stalledInterval (30 seconds by default). It is one Lua script, moveStalledJobsToWait, and it works in two passes:
- It puts every id currently in
activeinto a SET,<prefix>:<queue>:stalled. - On the next round, every id still in that SET whose
:lockkey is gone is really stalled. BullMQ removes it fromactive, increments the job'sstcfield, pushes it back towait, emits astalledevent and clears the SET.
Then the job runs again. If it already sent half the emails or charged the card before the worker died, those side effects happen twice. BullMQ is at-least-once: a processor that can stall must be idempotent.
A job that stalls more than maxStalledCount times (1 by default) is not retried forever. In BullMQ 6 the script marks it to fail and still moves it to wait; when a worker picks it up, it fails with job stalled more than allowable limit. Jobs produced by a job scheduler are the exception: they are never failed for stalling.
Why you cannot see it in the job states
stalled is not one of BullMQ's job states. job.getState() on a stalled job returns active, and getJobCounts() counts it as active. A dashboard that only shows states shows you a healthy-looking active: 8 while three of them are dead.
How to find stalled jobs
In code, listen for the event. On a worker it fires in the process that recovered the job; QueueEvents sees it from anywhere:
import { QueueEvents } from "bullmq";
const events = new QueueEvents("payments", { connection });
events.on("stalled", ({ jobId }) => {
logger.warn({ jobId }, "job stalled: its worker lost the lock, it will run again");
});
On a job, stalledCounter is the durable trace. Anything above 0 is the answer to "why did this run twice?":
const job = await queue.getJob("42");
if (job.stalledCounter > 0) {
console.log(`job 42 stalled ${job.stalledCounter} time(s)`);
}
In Redis, the SET holds the active ids being checked between two rounds. It is a transient reading, but it is O(1) to count:
SCARD bull:payments:stalled
On BullMQ 6's Postgres backend the same mechanism lives in columns of the job table instead of keys. The lock is locked_until_ms, the stalled check marks active jobs with stalled_marked on one pass and reclaims the ones still marked with an expired lock on the next, and stalled_count is the job's stalledCounter. An expired lock is one query on the partial index BullMQ keeps for active jobs:
SELECT id, stalled_count
FROM bullmq.job
WHERE queue = 'payments'
AND state = 'active'
AND locked_until_ms < extract(epoch FROM now()) * 1000;
How to make it stop
- Drain workers on shutdown. Call
await worker.close()onSIGTERMso active jobs finish before the process exits, and give the pod a grace period longer than your longest job. - Do not block the event loop. Move CPU-heavy work to a sandboxed processor or a worker thread, or break it into steps that yield.
- Raise
lockDurationonly for jobs that really need long synchronous stretches. A longer lock also means a dead worker's job waits longer before it is picked up again. - Make processors idempotent, because some stalls will always happen: an idempotency key per side effect, or a check of what was already done before doing it.
Seeing it in Bullpane
Bullpane does not invent a "stalled" tab, because that would contradict BullMQ's model. On a queue's active tab it warns how many of those jobs are stalled right now ("3 of these are stalled (worker lost the lock)"), read in the same round trip as the counts, on Redis and on BullMQ 6's Postgres backend alike. In the job list, every job that ever stalled carries a badge with its stalledCounter, so the job that ran twice is visible without digging through logs.
Run Bullpane on your queues Open the live demo
Sources
- BullMQ 6.3.4 source:
commands/moveStalledJobsToWait-9.lua(the two-pass check,stc, the stall limit and the scheduler exception) andclasses/worker.js(defaults:lockDuration30000,stalledInterval30000,maxStalledCount1). - BullMQ docs: Stalled jobs
Something here wrong for your BullMQ version? Write to hello@bullpane.com and it gets fixed.