SYS—01Job-Scheduler
Distributed job scheduler
A task queue built from scratch — the machinery BullMQ, Sidekiq and Celery run under the hood.
The problem
Background jobs fail, workers die mid-task, and two workers must never run the same job at the same time. The goal: at-least-once delivery with retries, priorities and delayed jobs — without a job ever being silently lost.
Decisions & trade-offs
- Postgres is the truth. Redis is only an index.
Instead of
Using Redis as the queue itselfEvery job lives durably in Postgres. Redis holds one sorted set per priority tier, scored by run-at time, so “what’s due next?” is an in-memory lookup — not a table scan.
Cost: Every path that makes a job eligible has to write to both stores. Forgetting one is the easiest bug to introduce in this codebase.
- Claim with a guarded UPDATE, not row locks.
Instead of
SELECT … FOR UPDATE SKIP LOCKEDOnce Redis has picked a single candidate there is nothing left to lock against. UPDATE … WHERE status = 'pending' is the guard; zero rows back means a stale entry, so drop it and try again.
Cost: The claim path now has a hard dependency on Redis — by design, there is no fallback to scanning the table.
- Shut down gracefully, don’t hope.
Instead of
Letting the process die on SIGTERMThe worker awaits its in-flight job before exiting, so scaling down never strands work. The API closes every WebSocket before server.close() — an open socket blocks it forever, which testing caught.
Cost: Scaling down now takes as long as the slowest in-flight job — shutdown waits for it.
The crux, in code
async function claimNextJob() {
const redis = await getRedisClient();
for (let tier = MIN_TIER; tier <= MAX_TIER; tier++) {
while (true) {
const jobId = await redisQueue.peekDueJob(redis, tier);
if (!jobId) break;
const claimed = await pool.query(
`UPDATE jobs
SET status = 'processing', attempts = attempts + 1
WHERE id = $1 AND status = 'pending'
RETURNING *`,
[jobId]
);
await redisQueue.removeJob(redis, tier, jobId);
if (claimed.rows.length > 0) return claimed.rows[0];
// stale Redis entry — already claimed elsewhere, try again
}
}
return null;
}
What breaks — honestly
- No lease or visibility timeout yet — a SIGKILL’d worker strands its job in “processing”.
- Exponential backoff (5 s · 2ⁿ) has no jitter and no cap, so mass failures retry in lockstep.
- No priority aging: a steady flood of priority-0 jobs can starve lower tiers.
- At-least-once, not exactly-once — job handlers have to be idempotent.
Build log
- Postgres-backed queue + submission API
- Separate worker process + handler registry
- Retries, exponential backoff, dead-letter queue
- Priorities + delayed jobs
- Redis sorted-set hot claim path
- Live dashboard over WebSockets + Pub/Sub
- Docker, N workers, graceful shutdown
