Monitoring BullMQ on Postgres
Checked against BullMQ 6.3.4 on PostgreSQL 16, 4 October 2026 · All guides · On Redis: How to build a dashboard to monitor BullMQ
BullMQ 6 can keep its queues in PostgreSQL instead of Redis. The queues behave the same; monitoring them does not. There is no LLEN: a count is a real count. There is a table that never shrinks. And the dashboard reads from the same database your workers write to, so a careless query costs them throughput.
Short answer. Everything is in one schema (bullmq by default): jobs in job with a state column and a partial index per state, events in event, totals in metrics, the paused flag in meta. Read with a read-only role, a statement timeout and no parallel query; refresh big counts slowly; and delete old rows from event on a schedule, because BullMQ never does.
What BullMQ creates
runMigrations(), or a worker started with migrate: true, creates these tables in the schema you pass to createPostgresBackend:
| Table | What is in it |
|---|---|
job | One row per job: queue, id, state, data, timestamps (added_at_ms, process_at_ms, processed_at_ms, finished_at_ms), the lock (locked_until_ms) and stalled_count. A partial index per state. |
event | The event stream (QueueEvents). Never trimmed: see below. |
metrics | Per queue and kind (completed, failed): a running count and per-minute deltas. |
meta | Queue flags, paused among them. |
scheduler, job_log, job_dependency, dedup, rate_limit | Job schedulers, job logs, flows, deduplication, rate limiting. |
The state values are waiting, active, delayed, completed, failed and waiting-children. Prioritized jobs are waiting rows with a priority above 0, and pausing a queue does not move rows: it sets a flag.
The signals, in SQL
The BullMQ API (getJobCounts, getMetrics, getWorkersCount) works on Postgres too, and the backlog checker runs unchanged on it. SQL is for a Grafana panel, a psql session during an incident, or a dashboard that should not load BullMQ at all.
Jobs per state. Each state is counted on its own partial index, an index-only scan:
SELECT state, count(*) AS jobs
FROM bullmq.job
WHERE queue = 'emails'
GROUP BY state;
Oldest waiting job, the delay your users feel. A delayed job is ready from process_at_ms, everything else from added_at_ms:
SELECT id,
(extract(epoch FROM now()) * 1000)::bigint
- greatest(added_at_ms, coalesce(process_at_ms, 0)) AS waited_ms
FROM bullmq.job
WHERE queue = 'emails' AND state = 'waiting'
ORDER BY greatest(added_at_ms, coalesce(process_at_ms, 0))
LIMIT 1;
Through the API, do not use getJobs(["wait"], 0, 0, true) for this: in 6.3.4 the Postgres backend ignores asc and returns the newest job. The last index of the default order is the oldest one on both backends:
const { wait } = await queue.getJobCounts("wait");
const [oldest] = wait > 0 ? await queue.getJobs(["wait"], wait - 1, wait - 1) : [];
Stalled jobs: active jobs whose lock ran out. On Redis this is a SET between two rounds of a check; here it is a column (what stalled means):
SELECT count(*) AS stalled
FROM bullmq.job
WHERE queue = 'emails' AND state = 'active'
AND locked_until_ms < (extract(epoch FROM now()) * 1000);
Paused. Paused jobs stay waiting in the table, so the SQL counts above do not know the queue is paused. Check the flag before alerting on a backlog:
SELECT value = '1' AS paused FROM bullmq.meta WHERE queue = 'emails' AND field = 'paused';
Throughput. metrics.count is the running total of finished jobs, collected only by workers created with metrics. Sample it twice and divide by the time between samples:
SELECT kind, count FROM bullmq.metrics WHERE queue = 'emails';
Connected workers. Each worker names its LISTEN connection after the queue, emails or emails:w:<name>:
SELECT count(*) AS workers
FROM pg_stat_activity
WHERE datname = current_database()
AND (application_name = 'emails' OR application_name LIKE 'emails:w:%');
Two blind spots: workers that connect through PgBouncer or another pooler show up as the pooler's connections, so the count can read 0 while they work; and the name has no schema in it, so two schemas with a queue of the same name share the count.
The table that never shrinks
BullMQ 6's trimEvents() is not implemented on Postgres, and nothing else deletes from event. In our test, 100 jobs added and completed left 401 rows: about four per job. A queue doing a million jobs a day adds around four million rows a day, forever. Watch its size and delete old rows on a schedule:
SELECT pg_size_pretty(pg_total_relation_size('bullmq.event'));
DELETE FROM bullmq.event
WHERE created_at_ms < (extract(epoch FROM now() - interval '7 days') * 1000);
Those rows are the event stream QueueEvents listens to. Keep as many days as you would want to look back on, and run the delete off-peak: on a big table it is a lot of dead rows for autovacuum.
Reading without slowing the workers
On Redis a dashboard competes for a single thread. On Postgres it competes for the same CPU, I/O and connections as your workers. What keeps it harmless:
- A read-only role. Enough for every read above, and a mistake in a dashboard query cannot touch a job:
CREATE ROLE dashboard_ro LOGIN PASSWORD '…'; GRANT USAGE ON SCHEMA bullmq TO dashboard_ro; GRANT SELECT ON ALL TABLES IN SCHEMA bullmq TO dashboard_ro; - A statement timeout and no parallel query on the dashboard's sessions. Without the second, one count over a large state can take several cores at once:
SET statement_timeout = '10s'; SET max_parallel_workers_per_gather = 0; - Its own
application_name, so it is easy to find inpg_stat_activityand is never mistaken for a worker. - Slow refresh for big states. Waiting and active counts are small and worth reading often. Counting ten million completed jobs costs about a second of one core; once a minute is plenty.
- Vacuum. Counts are index-only scans only while autovacuum keeps the visibility map current. Right after a bulk load they visit the table and get much slower. A queue table with heavy churn benefits from a lower
autovacuum_vacuum_scale_factoronjob. - Behind a transaction pooler (PgBouncer, Supavisor, RDS Proxy), session settings do not stick. Run each read in its own transaction with
SET LOCAL.
Doing it with Bullpane
Bullpane 0.6 reads and operates BullMQ queues on Postgres with the same pages, actions and alerts as on Redis, following the rules above: a read-only role is enough to browse, reads run with a timeout and no parallel query, counts on big states are refreshed less often, it works behind PgBouncer and over TLS, and a health card shows connections, database and table sizes and warns when event passes 1 GiB. Measured on 1M jobs across 20 queues, the stats for every queue take 60 ms; with ten browser tabs open on a steady workload, worker throughput stayed within noise.
npx bullpane --postgres postgres://user:pass@localhost:5432/app
# or
docker run -d -p 3000:3000 -v bullpane-data:/data bullpane/bullpane
Install the free edition Read the Postgres notes
Sources
- BullMQ 6.3.4 source:
postgres/migrations/0001_schema.sql(tables, thejob_stateenum and the partial indexes),postgres/postgres-queue-backend.js(trimEventsnot implemented, workers named throughapplication_name). - Every query and snippet here was run against BullMQ 6.3.4 on PostgreSQL 16, with real workers, including the read-only role.
- Bullpane: BullMQ on Postgres, with the load and lock measurements.
Something here wrong for your BullMQ version? Write to hello@bullpane.com and it gets fixed.