How to build a dashboard to monitor BullMQ

Checked against BullMQ 6.3.4 on 4 October 2026 · All guides

Most BullMQ dashboards show job counts per state and stop there. Counts tell you how many jobs exist. They do not tell you whether the queue keeps up, whether a worker died holding jobs, or whether Redis is about to delete your data. This guide lists what a dashboard has to show, the BullMQ call behind each number and what it costs Redis.

Short answer. Per queue: counts by state, throughput, failure rate, the age of the oldest waiting job, stalled jobs, paused, workers. Per Redis: memory, eviction policy, evicted keys. Read them with bounded calls, never KEYS, and turn on metrics in your workers, or there is no throughput to show.

What to show, and how to read it

SignalWhy it mattersHow to read itCost
Jobs by stateThe baseline: waiting, active, delayed, failed, completed, prioritized, paused.queue.getJobCounts()One Lua script, an LLEN or ZCARD per state.
ThroughputWhether the queue keeps up. Counts cannot answer this.getMetrics("completed"), see the trap belowCheap with a bounded range.
Failure rateFailures over finished jobs in a window, not the size of the failed set.getMetrics("failed") against getMetrics("completed")Same.
Oldest waiting jobThe delay your users feel. The best single latency number.queue.getJobs(["wait"], n - 1, n - 1), n = waiting count (why not asc)One job.
Stalled jobsActive jobs whose worker is gone. Invisible in the states. GuideSCARD <prefix>:<queue>:stalled, job.stalledCounterO(1).
PausedA paused queue has a growing backlog on purpose. Do not alert on it.queue.isPaused()O(1).
Workers connectedZero workers and a backlog is an outage.queue.getWorkersCount()CLIENT LIST; some managed Redis services block it.
Redis memory and evictionAny policy other than noeviction lets Redis delete job keys under memory pressure, which corrupts queues.INFO memory and INFO stats: used_memory, maxmemory, maxmemory_policy, evicted_keysO(1).

Running BullMQ 6 on Postgres instead of Redis? The signals are the same, the reads are not: see Monitoring BullMQ on Postgres.

BullMQ itself checks the eviction policy when it connects and prints IMPORTANT! Eviction policy is allkeys-lru. It should be "noeviction". That warning scrolls by once in a log. A dashboard should keep it on screen.

The metrics trap: data is empty

Throughput and failure rate come from BullMQ's metrics, and two things make them read zero:

meta.count, the running total, is always current. Sample it twice and divide:

import { Worker, MetricsTime } from "bullmq";

new Worker("payments", processor, {
  connection,
  metrics: { maxDataPoints: MetricsTime.ONE_WEEK },   // one point per minute, kept for a week
});

// elsewhere, on a timer:
const { meta } = await queue.getMetrics("completed", 0, 0);   // 0, 0: read one point, not the whole list
const perMinute = ((meta.count - last.count) / (Date.now() - last.at)) * 60_000;
last = { count: meta.count, at: Date.now() };

Rules that keep the dashboard from hurting production

Option 1: Prometheus and Grafana

If you already run Prometheus, queue.exportPrometheusMetrics() gives the exposition text for one queue: bullmq_job_count per state, plus completed and failed totals. One trap: concatenating it for several queues repeats the HELP and TYPE lines, and Prometheus rejects the whole scrape (second HELP line for metric name "bullmq_job_count"). Group the lines by metric first:

function mergePrometheus(texts) {
  const families = new Map();   // metric name -> { meta, samples }
  for (const line of texts.join("\n").split("\n")) {
    if (!line) continue;
    const name = line.startsWith("#") ? line.split(" ")[2] : line.slice(0, line.search(/[{ ]/));
    const family = families.get(name) ?? { meta: new Set(), samples: [] };
    families.set(name, family);
    if (line.startsWith("#")) family.meta.add(line);
    else family.samples.push(line);
  }
  return [...families.values()].flatMap((f) => [...f.meta, ...f.samples]).join("\n") + "\n";
}

http.createServer(async (req, res) => {
  if (req.url !== "/metrics") return res.writeHead(404).end();
  const texts = await Promise.all(queues.map((q) => q.exportPrometheusMetrics({ env: "production" })));
  res.writeHead(200, { "content-type": "text/plain; version=0.0.4" }).end(mergePrometheus(texts));
}).listen(9464);

Two things it does not cover: the oldest waiting job and stalled jobs are not in the export, so add your own gauges for them. And exportPrometheusMetrics reads the whole per-minute metrics list on every scrape to use one counter from it. With a week of retention that is up to 10,080 entries per queue, per scrape: keep the scrape interval sensible.

Grafana then gives you graphs and alerting. It does not give you the jobs: when the failure rate spikes, you still need somewhere to see which jobs failed and why.

Option 2: embed bull-board

bull-board mounts inside your Express, Fastify, NestJS or Hono app, behind the auth you already have. It is the quickest way to see jobs, retry them and read a stack trace. It is a viewer: no throughput, no failure rate over time, no alerts, no search inside job data. Full comparison.

Option 3: a dashboard that already does this

Bullpane covers most of this guide out of the box, on Redis and on BullMQ 6's Postgres backend: every queue with counts, success rate and throughput charts, stalled jobs on the active tab, server health (Redis memory, eviction policy and evicted keys; Postgres connections, table sizes and BullMQ's untrimmed event table), and the jobs themselves with search inside job data. Reads follow the rules above (no KEYS, one round trip per read, payloads truncated on the server), so it is safe to point at production. It runs next to your stack, no database to set up:

docker run -d -p 3000:3000 -v bullpane-data:/data bullpane/bullpane
# or, without Docker:
npx bullpane --redis redis://localhost:6379
npx bullpane --postgres postgres://user:pass@localhost:5432/app

Install the free edition Open the live demo

Sources

Something here wrong for your BullMQ version? Write to hello@bullpane.com and it gets fixed.