Alerting on a BullMQ backlog before it hurts your throughput

Checked against BullMQ 6.3.4, on Redis and on Postgres, on 4 October 2026 · All guides

The first backlog alert everybody writes is "waiting above 1,000". It fires every morning on the queue that bursts 50,000 emails and drains them in ten minutes, and it never fires on the billing queue that has 200 jobs stuck behind a worker doing two a minute. The number of waiting jobs is not the problem. How long they wait is.

Short answer. Alert on three things instead of a count: the age of the oldest waiting job (what your users feel), the time to drain the backlog at the current throughput, and no progress: jobs waiting while nothing completes. Skip paused queues, require two bad readings in a row, and turn on metrics in your workers or there is no throughput to measure.

Why a count threshold fails

Ten thousand waiting jobs on a queue that completes 5,000 a minute is two minutes of delay. Two hundred on a queue that completes two a minute is a hundred minutes. Any single threshold is too low for the first queue and too high for the second, and setting one per queue just moves the guessing.

The three signals

Count prioritized jobs in the backlog: they live in their own set and are taken before wait. A flood of high-priority jobs is exactly what makes the oldest ordinary job wait, and the oldest-job signal catches that starvation. Leave delayed out: those are not late, they are scheduled.

Measuring throughput correctly

Throughput comes from BullMQ's metrics, which are only collected when workers are created with metrics. Do not sum getMetrics().data: BullMQ only writes a minute into that list once a job finishes in a later minute, so the current minute is never there and a quiet queue reads zero. Sample the running total, meta.count, and divide by the time between samples.

import { Worker, MetricsTime } from "bullmq";

new Worker("emails", processor, { connection, metrics: { maxDataPoints: MetricsTime.ONE_WEEK } });

If any deployment processes the queue without metrics, the jobs it completes are never counted, and the no-progress check below fires while everything is fine. Turn it on everywhere before you alert on it.

The checker

One function measures a queue. It makes three small reads (the counts, one metrics point, one job), cheap enough to run every minute on every queue. It works the same on Redis and on BullMQ's Postgres backend:

const previous = new Map();   // queue name -> { at, completed }

async function measure(queue) {
  const [counts, metrics] = await Promise.all([
    queue.getJobCounts("wait", "prioritized"),
    queue.getMetrics("completed", 0, 0),       // meta.count is the running total of completed jobs
  ]);
  // The oldest waiting job is the last one in the default, newest-first order.
  const [oldest] = counts.wait > 0 ? await queue.getJobs(["wait"], counts.wait - 1, counts.wait - 1) : [];
  const backlog = counts.wait + counts.prioritized;
  const now = Date.now();
  const last = previous.get(queue.name);
  previous.set(queue.name, { at: now, completed: metrics.meta.count });
  if (!last) return null;                       // the first reading only sets the baseline

  const perMinute = ((metrics.meta.count - last.completed) / (now - last.at)) * 60_000;
  return {
    backlog,
    perMinute,
    drainMinutes: backlog === 0 ? 0 : perMinute > 0 ? backlog / perMinute : Infinity,
    oldestWaitMinutes: oldest ? (now - oldest.timestamp - (oldest.delay ?? 0)) / 60_000 : 0,
  };
}

Another one decides. It skips paused queues, waits for two bad readings in a row so a burst does not page anyone, sends one message when a problem starts and one when it ends:

const LIMITS = { oldestWaitMinutes: 5, drainMinutes: 15 };
const BREACHES_BEFORE_ALERT = 2;
const breaches = new Map();
const firing = new Set();

function problem(m) {
  if (m.backlog > 0 && m.perMinute === 0)
    return `${m.backlog} jobs waiting and nothing completed since the last check: are the workers up?`;
  if (m.oldestWaitMinutes > LIMITS.oldestWaitMinutes)
    return `oldest job has waited ${m.oldestWaitMinutes.toFixed(1)} min (limit ${LIMITS.oldestWaitMinutes})`;
  if (m.drainMinutes > LIMITS.drainMinutes)
    return `${m.backlog} waiting at ${Math.round(m.perMinute)}/min: ${Math.round(m.drainMinutes)} min to drain (limit ${LIMITS.drainMinutes})`;
  return null;
}

async function notify(text) {
  await fetch(process.env.SLACK_WEBHOOK_URL, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ text }),
  });
}

async function check(queue) {
  if (await queue.isPaused()) return;          // a paused queue grows on purpose
  const m = await measure(queue);
  if (!m) return;
  const issue = problem(m);
  const count = issue ? (breaches.get(queue.name) ?? 0) + 1 : 0;
  breaches.set(queue.name, count);
  if (count === BREACHES_BEFORE_ALERT && !firing.has(queue.name)) {
    firing.add(queue.name);
    await notify(`:rotating_light: ${queue.name}: ${issue}`);
  } else if (!issue && firing.has(queue.name)) {
    firing.delete(queue.name);
    await notify(`:white_check_mark: ${queue.name} is keeping up again`);
  }
}

setInterval(() => Promise.all(queues.map(check)).catch(console.error), 60_000);

Why not getJobs(["wait"], 0, 0, true) for the oldest job? On Redis it works, but on BullMQ 6.3.4's Postgres backend getJobs ignores asc and returns the newest job either way. Reading the last index of the default order is right on both.

Run against a real Redis and a real Postgres with no workers, 50 waiting jobs, then a worker coming back, it posts exactly two messages on each:

:rotating_light: emails: 50 jobs waiting and nothing completed since the last check: are the workers up?
:white_check_mark: emails is keeping up again

Run it in one place, not in every worker: one checker per queue, or the same alert arrives once per replica. The state lives in memory, so a restart only costs one reading of baseline.

Choosing the limits

Doing this in Bullpane

Bullpane Pro has alert rules that send to Slack or a webhook, with cooldowns and a resolved message, scoped to a queue, a folder, a connection or every queue at once, for queues on Redis and on BullMQ 6's Postgres backend. For backlogs today it has a waiting above rule, which counts waiting and prioritized jobs and leaves paused queues out, next to failure count, failure rate and p95 processing time. The oldest-job and time-to-drain signals in this guide are not rules in Bullpane yet; until they are, the checker above runs fine next to it.

Install Bullpane See Pro pricing

Sources

Something here wrong for your BullMQ version? Write to hello@bullpane.com and it gets fixed.