Backpressure: Bounding Queues Before They Bound You
Introduction
Backpressure is the practice of letting a slow stage in a system push back on the stage feeding it, so that work cannot pile up faster than it can be finished. The most common form is a bounded queue: the buffer in front of a worker is given a fixed maximum size, and once it is full the system refuses new work instead of accepting it. That refusal is the backpressure, and it travels upstream, telling the producer to slow down, retry later, or give up.
The alternative is an unbounded queue, which is what you get by default from an in-memory array, an unbuffered channel left to grow, or a message consumer that accepts everything offered. An unbounded queue looks safe because it never rejects anything, and it is the single most common way a system that was merely slow becomes a system that is down. In the simulation below, a service running 30% over capacity with an unbounded queue reaches a six-minute request latency and keeps climbing; the same service with a bounded queue holds latency at 168 milliseconds by refusing the 23% of requests it was never going to serve anyway.
This post shows why unbounded queues fail, derives the one equation that predicts it, builds a bounded work queue in TypeScript, and covers how to choose the limits and shed load well. It continues the thread of the last four posts: the LRU cache decides what to keep, rate limiting decides what callers may send, Bloom filters decide what to skip, consistent hashing decides which node holds a key, and backpressure decides what a node does when work arrives faster than it can finish it.
The Unbounded Queue Is a Latency Bomb
Picture one worker that handles requests taking 10 milliseconds each, so it can finish 100 per second. Offer it fewer than 100 per second and it keeps up. The interesting question is what happens as the arrival rate approaches that ceiling.
The ratio of arrival rate to service rate is the utilization, written as the Greek letter ρ (rho). At ρ = 0.5 the worker is busy half the time. The instinct is that latency climbs gently as ρ rises toward 1, and the instinct is wrong. Simulating a million requests at each utilization, with random arrival timing, gives the mean time each request spends in the system, W, and the number of requests in the system at any moment, L:
One worker, 10 ms service (μ = 100/s), unbounded queue:
| 0.50 | 19.8 | 0.99 | 91 |
|---|---|---|---|
| 0.80 | 49.9 | 3.99 | 234 |
| 0.90 | 101.4 | 9.14 | 491 |
| 0.95 | 193.8 | 18.38 | 965 |
| 0.99 | 819.0 | 80.99 | 3,322 |
Latency does not rise linearly, it rises hyperbolically. Going from 90% to 99% utilization is a 10% increase in load and roughly an eightfold increase in latency. This is not a property of any particular system; it is the shape of the function W = (1/μ) / (1 − ρ), which has a vertical asymptote at ρ = 1. The last few percent of a server's theoretical capacity are unusable in practice, which is why running at 70-80% utilization is a deliberate choice rather than a failure to optimize.
Everything so far assumes ρ stays below 1. The moment arrivals exceed service capacity, even briefly, the queue stops having a steady state at all.
Little's Law
One equation governs every queue, and it is simple enough to keep in your head:
The average number of items in a system (L) equals the average arrival rate (λ) times the average time each item spends there (W). It holds for any stable system regardless of arrival distribution, service distribution, or scheduling order. The simulation confirms it exactly: three workers at 20ms service under 120 requests per second measured an average of 5.069 requests in the system, against a predicted 120.1/s × 0.0422 s = 5.069.
Little's Law is what makes an unbounded queue's failure inevitable rather than unlucky. If λ exceeds the rate at which items leave, then items accumulate, L grows without bound, and because L = λ × W with λ roughly fixed, W grows without bound alongside it. A queue that never rejects is a queue that converts excess load directly into latency, and unbounded latency is indistinguishable from failure: clients time out, retry, and add still more load.
Read the other direction, the same equation is a design tool. Fix the maximum L you are willing to hold and you have fixed the maximum W a request can suffer. Bounding the queue is bounding the latency.
Bounding the Queue
Give the queue a maximum size and the behavior under overload changes completely. Running the same worker at 130 requests per second against a capacity of 100, so 30% more load than it can serve, with an unbounded queue and then with bounded ones:
λ = 130/s, μ = 100/s (30% overload), one worker:
| unbounded | ~6 minutes | ~11 minutes | 0.0% | 100/s |
|---|---|---|---|---|
| bounded, cap 50 | 465.9 ms | 654 ms | 23.3% | 100/s |
| bounded, cap 20 | 168.3 ms | 300 ms | 23.5% | 100/s |
| bounded, cap 10 | 74.5 ms | 171 ms | 24.5% | 98/s |
The unbounded row is not a typo. With arrivals permanently exceeding service, the queue grows for as long as the overload lasts; in a finite simulation it reached a six-minute average latency and was still climbing when the run ended. Every one of those requests is being served long after the client gave up, so the throughput column reads a useless 100 per second of work nobody is waiting for anymore.
The bounded rows tell the opposite story. Throughput is identical, because the worker was always going to complete 100 requests per second and no queue changes that. What the bound changes is the fate of the 30 extra requests per second: instead of joining a queue that guarantees them a timeout, they are rejected immediately, and the requests that are accepted get a fast, predictable answer. A smaller cap trades a slightly higher rejection rate for a lower latency ceiling, and that is the only real knob. The rejections were not caused by bounding the queue; they were caused by the overload. Bounding the queue only decides whether the caller finds out in 168 milliseconds or 6 minutes.
Every request reaching a bounded queue takes one of two routes, and the split happens at the capacity check:
What happens to 76.5% of requests at 30% overload with a cap of 20: there is room in the queue, so the request waits briefly and is served. Bounding the queue is what keeps that wait short.
A request arrives. 130 per second are coming in; the worker can finish 100.
All 5 steps
Use the zoom buttons, or focus the diagram and press plus or minus to zoom, arrow keys to pan, and 0 to reset.
A Bounded Work Queue in TypeScript
The same idea inside a single process is a concurrency limiter: at most N tasks running at once, at most M more waiting, and anything beyond that rejected. This is the shape you want in front of a database pool, an outbound API client, or any resource that degrades when too many callers hit it at once.
export class QueueFullError extends Error { constructor(limit: number) { super(`work queue is full (limit ${limit})`); this.name = "QueueFullError"; }} export class BoundedQueue { private active = 0; private readonly queue: Array<() => void> = []; private readonly maxConcurrency: number; private readonly maxQueueDepth: number; constructor(maxConcurrency: number, maxQueueDepth: number) { if (maxConcurrency < 1) { throw new RangeError("maxConcurrency must be at least 1"); } if (maxQueueDepth < 0) { throw new RangeError("maxQueueDepth must be at least 0"); } this.maxConcurrency = maxConcurrency; this.maxQueueDepth = maxQueueDepth; } get depth(): number { return this.queue.length; } run<T>(task: () => Promise<T> | T): Promise<T> { // Shed load: both the running slots and the wait queue are full. if ( this.active >= this.maxConcurrency && this.queue.length >= this.maxQueueDepth ) { return Promise.reject( new QueueFullError(this.maxConcurrency + this.maxQueueDepth) ); } return new Promise<T>((resolve, reject) => { const attempt = () => { this.active++; Promise.resolve() .then(task) .then(resolve, reject) .finally(() => { this.active--; this.pump(); }); }; if (this.active < this.maxConcurrency) attempt(); else this.queue.push(attempt); }); } private pump(): void { if (this.active < this.maxConcurrency && this.queue.length > 0) { const next = this.queue.shift()!; next(); } }}Three details carry the correctness. The rejection is synchronous and happens before any promise is created, so an overloaded system spends almost nothing deciding to say no. The .finally runs whether the task resolved or threw, so a failing task always frees its slot; without it, one thrown error would permanently shrink the pool until nothing ran at all. And pump starts exactly one waiting task per completion, which keeps the running count at or below the limit and preserves first-in-first-out order for everything waiting.
Choosing the Limits
The two numbers are not guesses; Little's Law sizes both.
Concurrency is the number of tasks running at once, and it should match the concurrency the downstream resource actually wants. A Postgres pool of 20 connections wants a maxConcurrency of 20, because the twenty-first concurrent query only queues inside the database instead of inside your process, where you can see it. Rearranging L = λ × W, the throughput you can sustain is maxConcurrency / serviceTime: twenty slots at 25ms each is 800 requests per second, and no setting of the queue depth raises that ceiling.
Queue depth controls how much latency you will tolerate before shedding. A waiting task's added delay is roughly queuePosition × serviceTime / maxConcurrency, so a queue depth of 20 in front of twenty 25ms slots adds at most about 25ms of waiting before a task even starts. Set the depth from the latency budget you are willing to spend, then let everything past it be rejected. A queue depth in the low tens is common; a queue depth of ten thousand is an unbounded queue wearing a disguise, since no client waits for the ten-thousandth position.
The habit worth forming: pick concurrency from the downstream limit, pick queue depth from the latency budget, and never pick either by watching how large the queue grows in production and raising the cap to match. That last move is how unbounded queues get reintroduced one incident at a time.
Shedding Load Well
Rejecting work is a feature, but a rejection is a message to a caller, and the message matters.
- Return the right signal. Over HTTP, a shed request is a 503 with a
Retry-Afterheader, distinct from the 429 a rate limiter returns. 429 means "you personally have sent too much"; 503 means "the service as a whole is saturated right now." Clients back off differently for each, and conflating them hides which problem you have. - Shed the cheapest thing that helps. Not all work is equal. Dropping a prefetch, a recommendation panel, or an analytics write protects the checkout path that shares the same worker pool. A tiered limiter that sheds low-priority work first keeps the important requests flowing while the system is over capacity.
- Beware the retry storm. A client that retries immediately on a 503 doubles the load precisely when the load is the problem. Backpressure only works if callers honor it, which means exponential backoff with jitter on their side, and a circuit breaker that stops calling a saturated dependency entirely for a while.
- Keep a timeout as the backstop. Even a bounded queue can hold a task longer than its answer is worth. A deadline on each task, checked before it starts running, discards work whose caller has already timed out, so the worker never spends its scarce capacity computing an answer no one will read.
Backpressure Across a System
A single bounded queue protects one stage. Real systems are pipelines, and backpressure has to propagate along the whole chain or it just relocates the unbounded queue to the next stage that lacks one.
The canonical example is already under your code: TCP flow control. A receiver advertises a window of how many bytes it can accept, the sender may not exceed it, and a slow reader shrinks the window until the fast writer blocks. The unboundedness has nowhere to accumulate because every stage refuses to over-accept. Reactive-streams libraries generalize the same contract to application data, letting a slow consumer signal demand upstream so a fast producer never overruns it.
Message queues are where teams most often lose this property. Putting a broker between a fast producer and a slow consumer feels like backpressure, but a broker with unbounded retention is just the unbounded in-memory queue moved to disk, now measured in hours of lag instead of milliseconds of latency. The broker needs its own bound, a maximum depth or retention, and a defined policy for what happens when it fills, or the pipeline has no backpressure at all, only a very large buffer waiting to overflow.
When You Do Not Need It
Bounding a queue costs a rejection path and a decision about limits, and not every queue is a liability.
- Utilization stays low. A system that never approaches its capacity has short queues by definition, and the bound never triggers. It still belongs there as a safety limit, but it will not change day-to-day behavior.
- The producer is already paced. A cron job processing yesterday's records generates work at a rate you control, so there is no faster producer to push back on.
- The queue is durable and the work must not be lost. Some workloads need every item processed eventually, and there rejection is wrong. The answer is a bounded durable queue plus autoscaling the consumers, not an unbounded one; the bound is what tells the autoscaler it is falling behind.
- A single in-order consumer. If work is already produced and consumed by the same synchronous loop, it cannot outrun itself and there is nothing to bound.
Conclusion
An unbounded queue is a promise to accept work you have no way to finish, and Little's Law turns that promise into unbounded latency the instant load exceeds capacity. Bounding the queue does not create the rejections that overload requires; it only decides whether the caller learns the answer is no in milliseconds or in minutes, and it keeps the accepted requests fast while it does. The two limits come straight from the same equation: concurrency from the resource downstream, queue depth from the latency you can spend. A system that bounds its queues fails by refusing a knowable fraction of its load, which is the only kind of overload failure a client can actually plan around.
Key Takeaways
- Backpressure lets a slow stage refuse new work so queues cannot outgrow the rate they drain, and a bounded queue is its most common form.
- Latency rises hyperbolically with utilization: the jump from 90% to 99% load roughly octupled measured latency, because W = (1/μ)/(1 − ρ) has an asymptote at full load.
- Little's Law, L = λ × W, guarantees that a queue accepting more than it drains grows without bound in both length and delay.
- A 30% overload sent an unbounded queue to a six-minute latency; a bounded queue held it to 168ms by shedding the 23% it could never have served.
- Size concurrency from the downstream limit and queue depth from the latency budget; never raise a cap just because the queue keeps filling.
- Shed with a 503 and Retry-After, drop the cheapest work first, defend against retry storms, and keep a per-task deadline as the backstop.

Steven Brown
Software Engineer
I am a Software Engineer based in the United States, passionate about writing code and developing applications. My journey into tech followed a unique path, beginning with a 9-year enlistment as a Russian Cryptologic Linguist in the US Army. This experience has fueled my unwavering commitment to excel in all aspects of software engineering.
Thanks for reading! If you found this helpful, check out more articles below or head back to the blog.
Back to Blog
