Monitoring BullMQ on Postgres

Checked against BullMQ 6.3.4 on PostgreSQL 16, 4 October 2026 · All guides · On Redis: How to build a dashboard to monitor BullMQ

BullMQ 6 can keep its queues in PostgreSQL instead of Redis. The queues behave the same; monitoring them does not. There is no LLEN: a count is a real count. There is a table that never shrinks. And the dashboard reads from the same database your workers write to, so a careless query costs them throughput.

Short answer. Everything is in one schema (bullmq by default): jobs in job with a state column and a partial index per state, events in event, totals in metrics, the paused flag in meta. Read with a read-only role, a statement timeout and no parallel query; refresh big counts slowly; and delete old rows from event on a schedule, because BullMQ never does.

What BullMQ creates

runMigrations(), or a worker started with migrate: true, creates these tables in the schema you pass to createPostgresBackend:

TableWhat is in it
jobOne row per job: queue, id, state, data, timestamps (added_at_ms, process_at_ms, processed_at_ms, finished_at_ms), the lock (locked_until_ms) and stalled_count. A partial index per state.
eventThe event stream (QueueEvents). Never trimmed: see below.
metricsPer queue and kind (completed, failed): a running count and per-minute deltas.
metaQueue flags, paused among them.
scheduler, job_log, job_dependency, dedup, rate_limitJob schedulers, job logs, flows, deduplication, rate limiting.

The state values are waiting, active, delayed, completed, failed and waiting-children. Prioritized jobs are waiting rows with a priority above 0, and pausing a queue does not move rows: it sets a flag.

The signals, in SQL

The BullMQ API (getJobCounts, getMetrics, getWorkersCount) works on Postgres too, and the backlog checker runs unchanged on it. SQL is for a Grafana panel, a psql session during an incident, or a dashboard that should not load BullMQ at all.

Jobs per state. Each state is counted on its own partial index, an index-only scan:

SELECT state, count(*) AS jobs
FROM bullmq.job
WHERE queue = 'emails'
GROUP BY state;

Oldest waiting job, the delay your users feel. A delayed job is ready from process_at_ms, everything else from added_at_ms:

SELECT id,
       (extract(epoch FROM now()) * 1000)::bigint
         - greatest(added_at_ms, coalesce(process_at_ms, 0)) AS waited_ms
FROM bullmq.job
WHERE queue = 'emails' AND state = 'waiting'
ORDER BY greatest(added_at_ms, coalesce(process_at_ms, 0))
LIMIT 1;

Through the API, do not use getJobs(["wait"], 0, 0, true) for this: in 6.3.4 the Postgres backend ignores asc and returns the newest job. The last index of the default order is the oldest one on both backends:

const { wait } = await queue.getJobCounts("wait");
const [oldest] = wait > 0 ? await queue.getJobs(["wait"], wait - 1, wait - 1) : [];

Stalled jobs: active jobs whose lock ran out. On Redis this is a SET between two rounds of a check; here it is a column (what stalled means):

SELECT count(*) AS stalled
FROM bullmq.job
WHERE queue = 'emails' AND state = 'active'
  AND locked_until_ms < (extract(epoch FROM now()) * 1000);

Paused. Paused jobs stay waiting in the table, so the SQL counts above do not know the queue is paused. Check the flag before alerting on a backlog:

SELECT value = '1' AS paused FROM bullmq.meta WHERE queue = 'emails' AND field = 'paused';

Throughput. metrics.count is the running total of finished jobs, collected only by workers created with metrics. Sample it twice and divide by the time between samples:

SELECT kind, count FROM bullmq.metrics WHERE queue = 'emails';

Connected workers. Each worker names its LISTEN connection after the queue, emails or emails:w:<name>:

SELECT count(*) AS workers
FROM pg_stat_activity
WHERE datname = current_database()
  AND (application_name = 'emails' OR application_name LIKE 'emails:w:%');

Two blind spots: workers that connect through PgBouncer or another pooler show up as the pooler's connections, so the count can read 0 while they work; and the name has no schema in it, so two schemas with a queue of the same name share the count.

The table that never shrinks

BullMQ 6's trimEvents() is not implemented on Postgres, and nothing else deletes from event. In our test, 100 jobs added and completed left 401 rows: about four per job. A queue doing a million jobs a day adds around four million rows a day, forever. Watch its size and delete old rows on a schedule:

SELECT pg_size_pretty(pg_total_relation_size('bullmq.event'));

DELETE FROM bullmq.event
WHERE created_at_ms < (extract(epoch FROM now() - interval '7 days') * 1000);

Those rows are the event stream QueueEvents listens to. Keep as many days as you would want to look back on, and run the delete off-peak: on a big table it is a lot of dead rows for autovacuum.

Reading without slowing the workers

On Redis a dashboard competes for a single thread. On Postgres it competes for the same CPU, I/O and connections as your workers. What keeps it harmless:

Doing it with Bullpane

Bullpane 0.6 reads and operates BullMQ queues on Postgres with the same pages, actions and alerts as on Redis, following the rules above: a read-only role is enough to browse, reads run with a timeout and no parallel query, counts on big states are refreshed less often, it works behind PgBouncer and over TLS, and a health card shows connections, database and table sizes and warns when event passes 1 GiB. Measured on 1M jobs across 20 queues, the stats for every queue take 60 ms; with ten browser tabs open on a steady workload, worker throughput stayed within noise.

npx bullpane --postgres postgres://user:pass@localhost:5432/app
# or
docker run -d -p 3000:3000 -v bullpane-data:/data bullpane/bullpane

Install the free edition Read the Postgres notes

Sources

Something here wrong for your BullMQ version? Write to hello@bullpane.com and it gets fixed.