How to build a dashboard to monitor BullMQ
Checked against BullMQ 6.3.4 on 4 October 2026 · All guides
Most BullMQ dashboards show job counts per state and stop there. Counts tell you how many jobs exist. They do not tell you whether the queue keeps up, whether a worker died holding jobs, or whether Redis is about to delete your data. This guide lists what a dashboard has to show, the BullMQ call behind each number and what it costs Redis.
Short answer. Per queue: counts by state, throughput, failure rate, the age of the oldest waiting job, stalled jobs, paused, workers. Per Redis: memory, eviction policy, evicted keys. Read them with bounded calls, never KEYS, and turn on metrics in your workers, or there is no throughput to show.
What to show, and how to read it
| Signal | Why it matters | How to read it | Cost |
|---|---|---|---|
| Jobs by state | The baseline: waiting, active, delayed, failed, completed, prioritized, paused. | queue.getJobCounts() | One Lua script, an LLEN or ZCARD per state. |
| Throughput | Whether the queue keeps up. Counts cannot answer this. | getMetrics("completed"), see the trap below | Cheap with a bounded range. |
| Failure rate | Failures over finished jobs in a window, not the size of the failed set. | getMetrics("failed") against getMetrics("completed") | Same. |
| Oldest waiting job | The delay your users feel. The best single latency number. | queue.getJobs(["wait"], n - 1, n - 1), n = waiting count (why not asc) | One job. |
| Stalled jobs | Active jobs whose worker is gone. Invisible in the states. Guide | SCARD <prefix>:<queue>:stalled, job.stalledCounter | O(1). |
| Paused | A paused queue has a growing backlog on purpose. Do not alert on it. | queue.isPaused() | O(1). |
| Workers connected | Zero workers and a backlog is an outage. | queue.getWorkersCount() | CLIENT LIST; some managed Redis services block it. |
| Redis memory and eviction | Any policy other than noeviction lets Redis delete job keys under memory pressure, which corrupts queues. | INFO memory and INFO stats: used_memory, maxmemory, maxmemory_policy, evicted_keys | O(1). |
Running BullMQ 6 on Postgres instead of Redis? The signals are the same, the reads are not: see Monitoring BullMQ on Postgres.
BullMQ itself checks the eviction policy when it connects and prints IMPORTANT! Eviction policy is allkeys-lru. It should be "noeviction". That warning scrolls by once in a log. A dashboard should keep it on screen.
The metrics trap: data is empty
Throughput and failure rate come from BullMQ's metrics, and two things make them read zero:
- Metrics are opt-in, per worker. Nothing is collected unless the worker that finishes the job was created with
metrics. If two deployments process the same queue and only one has it, the other one's jobs are never counted. - The current minute is never in
data.getMetrics()returns per-minute counts indata, newest first, but BullMQ only writes a minute into that list once a job finishes in a later minute. Read during a busy minute, or on a quiet queue, anddatais empty.
meta.count, the running total, is always current. Sample it twice and divide:
import { Worker, MetricsTime } from "bullmq";
new Worker("payments", processor, {
connection,
metrics: { maxDataPoints: MetricsTime.ONE_WEEK }, // one point per minute, kept for a week
});
// elsewhere, on a timer:
const { meta } = await queue.getMetrics("completed", 0, 0); // 0, 0: read one point, not the whole list
const perMinute = ((meta.count - last.count) / (Date.now() - last.at)) * 60_000;
last = { count: meta.count, at: Date.now() };
Rules that keep the dashboard from hurting production
- Never
KEYS. It blocks Redis for the whole keyspace. Every queue has a<prefix>:<queue>:metahash; find them withSCAN:
Cache the result for a while. Queues are created rarely; the scan does not need to run on every refresh.async function discoverQueues(redis, prefix = "bull") { const names = new Set(); let cursor = "0"; do { const [next, keys] = await redis.scan(cursor, "MATCH", `${prefix}:*:meta`, "COUNT", 1000); for (const key of keys) names.add(key.slice(prefix.length + 1, -":meta".length)); cursor = next; } while (cursor !== "0"); return [...names]; } - On BullMQ 6's Postgres backend the cost model changes. Discovery is cheap, with one catch:
bullmq.metahas a row for every queue aQueueorWorkerwas created for, but a queue fed only by aFlowProducerhas none until a worker shows up, so also look for queue names inbullmq.job. Counts are not cheap: there is noLLEN, so a count reads every row of the state. Keep every query on a single state, which is what BullMQ's partial indexes are built for, cache counts for a few seconds, and do not let a dashboard query run in parallel on several cores of the database your workers share. - Always pass a range.
getJobs,getMetricsand the job lists all default to ranges that can mean "everything". A failed set with 200,000 jobs is not a list to render. - Do not load payloads in lists. Job data can be megabytes. Show the payload when someone opens the job, not on every row.
- Read-only by default. A dashboard that can retry, drain and obliterate needs a login and an audit trail, or it should not have those buttons.
Option 1: Prometheus and Grafana
If you already run Prometheus, queue.exportPrometheusMetrics() gives the exposition text for one queue: bullmq_job_count per state, plus completed and failed totals. One trap: concatenating it for several queues repeats the HELP and TYPE lines, and Prometheus rejects the whole scrape (second HELP line for metric name "bullmq_job_count"). Group the lines by metric first:
function mergePrometheus(texts) {
const families = new Map(); // metric name -> { meta, samples }
for (const line of texts.join("\n").split("\n")) {
if (!line) continue;
const name = line.startsWith("#") ? line.split(" ")[2] : line.slice(0, line.search(/[{ ]/));
const family = families.get(name) ?? { meta: new Set(), samples: [] };
families.set(name, family);
if (line.startsWith("#")) family.meta.add(line);
else family.samples.push(line);
}
return [...families.values()].flatMap((f) => [...f.meta, ...f.samples]).join("\n") + "\n";
}
http.createServer(async (req, res) => {
if (req.url !== "/metrics") return res.writeHead(404).end();
const texts = await Promise.all(queues.map((q) => q.exportPrometheusMetrics({ env: "production" })));
res.writeHead(200, { "content-type": "text/plain; version=0.0.4" }).end(mergePrometheus(texts));
}).listen(9464);
Two things it does not cover: the oldest waiting job and stalled jobs are not in the export, so add your own gauges for them. And exportPrometheusMetrics reads the whole per-minute metrics list on every scrape to use one counter from it. With a week of retention that is up to 10,080 entries per queue, per scrape: keep the scrape interval sensible.
Grafana then gives you graphs and alerting. It does not give you the jobs: when the failure rate spikes, you still need somewhere to see which jobs failed and why.
Option 2: embed bull-board
bull-board mounts inside your Express, Fastify, NestJS or Hono app, behind the auth you already have. It is the quickest way to see jobs, retry them and read a stack trace. It is a viewer: no throughput, no failure rate over time, no alerts, no search inside job data. Full comparison.
Option 3: a dashboard that already does this
Bullpane covers most of this guide out of the box, on Redis and on BullMQ 6's Postgres backend: every queue with counts, success rate and throughput charts, stalled jobs on the active tab, server health (Redis memory, eviction policy and evicted keys; Postgres connections, table sizes and BullMQ's untrimmed event table), and the jobs themselves with search inside job data. Reads follow the rules above (no KEYS, one round trip per read, payloads truncated on the server), so it is safe to point at production. It runs next to your stack, no database to set up:
docker run -d -p 3000:3000 -v bullpane-data:/data bullpane/bullpane
# or, without Docker:
npx bullpane --redis redis://localhost:6379
npx bullpane --postgres postgres://user:pass@localhost:5432/app
Install the free edition Open the live demo
Sources
- BullMQ 6.3.4 source:
commands/getCounts-1.lua,classes/queue-getters.js(getMetrics,getWorkers,exportPrometheusMetrics),classes/redis-connection.js(the eviction policy warning). - BullMQ docs: Metrics · Going to production
- The Prometheus merge was checked with
promtool check metrics.
Something here wrong for your BullMQ version? Write to hello@bullpane.com and it gets fixed.