· 9 min read
How to Handle Millions of Requests on a Server
Our server was getting thousands of requests in the same minute and the database ran out of connections. Here are the 4 problems we faced and the solution we made for each one.

What happens when your server gets a million requests? Does it handle all of them, or does it go down?
I work on a server that receives webhooks. A webhook is basically a message that another platform sends to your server when something happens. For example, Meta sends one every time someone writes a comment on a Facebook page.
Most of the time these messages come slowly. But sometimes thousands of them come in the same minute. When that happened, our server had problems.
In this post I will show you every problem we faced and the solution we made for it. The main solution is only 40 lines of code. It is called a concurrency limiter. You can copy it and use it today.
First, what is the real problem?
Most people think the problem is the number of requests. It is not.
The real problem is that all the requests start their work at the same time. Look at this code:
const results = await Promise.all(
jobs.map((job) => process(job))
);
It looks fine. But map starts the work for every job immediately. Promise.all does not limit anything. It only waits for them to finish.
So if you have 5,000 jobs, 5,000 of them are running at the same moment. Every job needs memory. Every job needs a database connection. Your server does not have 5,000 of those.
Keep this in mind, because every problem below comes from it.
Problem 1: The sender does not wait for you
When a platform sends you a webhook, it waits for your answer. But it only waits for a few seconds. Meta waits around 5 seconds.
What happens if your server is slow and answers late? The platform thinks the message failed. So it sends the same message again. Now your busy server has even more work, and you can also process the same message twice.
The solution: answer first, work later
We answer immediately and do the real work in the background.
function receiveWebhook(payload: Payload, res: Res) {
// 1. Answer immediately, so the sender does not send it again
res.status(200).send("EVENT_RECEIVED");
// 2. Do the real work in the background
void processInBackground(payload);
}
Now the sender is happy. It got its answer in a few milliseconds and it will not send the message again.
But this solution creates the next problem.
Problem 2: The database runs out of connections
We answer fast now. So the platform can send us messages very fast. And every message starts a background job immediately.
A database has a limited number of connections. This is called the connection pool. Let's say the pool has 20 connections and 5,000 jobs start together. All 5,000 jobs ask for a connection. 20 of them get one. The others wait, then they time out, then they fail.
So the server did not crash because of the traffic. It crashed because we let all the work start at once.
The solution: a concurrency limiter
We put a limit on how many jobs can run at the same time. We chose 10.
const processingLimit = createConcurrencyLimit(10);
function receiveWebhook(payload: Payload, res: Res) {
res.status(200).send("EVENT_RECEIVED");
// The work waits for its turn. Only 10 run at the same time.
void processingLimit(() => processInBackground(payload));
}
Only one line changed. Now 10 jobs run and the others wait for their turn. We already answered the sender, so nobody is waiting on us. The work just stands in a line.
This one change is what stopped the database from running out of connections.
How a concurrency limiter works
Here is the whole idea in three sentences:
- Every job lines up in a queue.
- A gate lets only a few of them run at the same time.
- When one job finishes, the next one in the line starts.
And here is the full code. It is 48 lines and it has zero dependencies.
export interface ConcurrencyLimit {
<T>(task: () => Promise<T>): Promise<T>;
readonly activeCount: number;
readonly pendingCount: number;
}
export function createConcurrencyLimit(maxConcurrency: number): ConcurrencyLimit {
if (!Number.isInteger(maxConcurrency) || maxConcurrency < 1) {
throw new RangeError(
`maxConcurrency must be a positive integer, got ${maxConcurrency}`
);
}
const queue: Array<() => void> = [];
let activeCount = 0;
const startNext = (): void => {
if (activeCount >= maxConcurrency) return;
const next = queue.shift();
if (!next) return;
activeCount += 1;
next();
};
const limit = <T>(task: () => Promise<T>): Promise<T> => {
return new Promise<T>((resolve, reject) => {
const run = (): void => {
Promise.resolve()
.then(task)
.then(resolve, reject)
.finally(() => {
activeCount -= 1;
startNext();
});
};
queue.push(run);
startNext();
});
};
Object.defineProperties(limit, {
activeCount: { get: () => activeCount },
pendingCount: { get: () => queue.length },
});
return limit as ConcurrencyLimit;
}If you like to watch, this video explains the same code step by step.
Now let's understand it part by part.
1. You give it a function, not a promise
export interface ConcurrencyLimit {
<T>(task: () => Promise<T>): Promise<T>;
readonly activeCount: number;
readonly pendingCount: number;
}This is the most important thing to understand. A promise has already started. If you write limit(fetch(url)), the request is already running before the limiter sees it. So the limiter cannot control anything.
You have to write limit(() => fetch(url)). Now you are giving it a function, and the limiter decides when to call it.
2. It checks the limit you pass
if (!Number.isInteger(maxConcurrency) || maxConcurrency < 1) {
throw new RangeError(
`maxConcurrency must be a positive integer, got ${maxConcurrency}`
);
}Why do we need this check? Because a wrong number does not give an error. It fails silently, and that is worse.
- If you pass
0, nothing ever starts. Every job waits forever. - If you pass
NaN, the limit stops working and everything runs at once. - If you pass
2.5, it behaves like 3.
So we accept only a positive whole number, and we throw an error for anything else.
3. It keeps only two things in memory
const queue: Array<() => void> = [];
let activeCount = 0;A queue and a counter. That is all the state. The queue does not hold the jobs. It holds small functions that start a job. The counter is how many jobs are running right now.
4. startNext is the gate
const startNext = (): void => {
if (activeCount >= maxConcurrency) return;
const next = queue.shift();
if (!next) return;
activeCount += 1;
next();
};- If all the slots are busy, do nothing. The job stays in the queue.
- If there is a free slot, take the first job in the line.
- If the line is empty, we are done.
- Increase the counter first, and then run the job.
We increase the counter before we run the job. So the gate can never let in more than the limit.
5. limit puts every job in the line
const limit = <T>(task: () => Promise<T>): Promise<T> => {
return new Promise<T>((resolve, reject) => {
const run = (): void => {
Promise.resolve()
.then(task)
.then(resolve, reject)
.finally(() => {
activeCount -= 1;
startNext();
});
};
queue.push(run);
startNext();
});
};limit gives you a promise immediately. But that promise finishes only when the job gets its turn and completes.
Look at Promise.resolve().then(task). Why not just call task()? Because if the job throws an error before it returns a promise, the slot would never be free again. You would lose one slot forever. Starting inside a promise chain turns that error into a normal failure.
And look at finally. It runs when the job succeeds and also when it fails. It frees the slot and starts the next job. So every job that finishes pulls the next one. This is the engine of the whole thing.
6. It shows you what is happening
Object.defineProperties(limit, {
activeCount: { get: () => activeCount },
pendingCount: { get: () => queue.length },
});
return limit as ConcurrencyLimit;activeCount is how many jobs are running. pendingCount is how many are waiting. They are getters, so they always give you the live number. Remember these two, because they solve Problem 4.
Problem 3: One slow job blocks everything
In the start we used the same limiter for different kinds of work. Some work was fast, like reading details from an API. Some work was slow, like asking an AI service to analyse text.
One day the AI service became slow. Every call was taking around 25 seconds. Let's do the maths:
- We had 20 slots, and each call took 25 seconds.
- So we could finish around 2,900 calls in one hour.
- But around 5,000 new items were coming every hour.
- So the line was growing by around 2,000 every hour.
After a few hours, more than 10,000 items were waiting. New work was stuck behind all of them. Things that should take 3 minutes were taking 5 hours.
And the fast work was stuck too, because it was standing in the same line as the slow work.
The solution: a separate line for each kind of work
// One line for each kind of work
const webhookLimit = createConcurrencyLimit(10);
const detailsLimit = createConcurrencyLimit(25);
const aiLimit = createConcurrencyLimit(20);
Now if the AI service is slow, only the AI line is slow. Webhooks still get processed. Details still get loaded.
We also learned a second thing. Most of our items did not need the AI at all. But they were waiting for it anyway. So now we check the cheap things first, and we send an item to the slow line only if it really needs it.
Problem 4: The line is growing and nobody knows
The limiter protects your server. But it cannot make the work faster. If work comes faster than it finishes, the line grows and grows, silently.
In our case 10,000 items were waiting and nobody knew. We found out only when users told us that things were late.
The solution: look at the two numbers
This is why the limiter has activeCount and pendingCount. We show them on a health page:
function getQueueStats() {
return {
active: detailsLimit.activeCount,
pending: detailsLimit.pendingCount,
limit: 25,
};
}
And we write a warning in the logs when too many jobs are waiting. We do it maximum one time per minute, so the logs do not fill up.
let lastWarningAt = 0;
function warnIfTooManyWaiting() {
const pending = detailsLimit.pendingCount;
if (pending < 50) return;
const now = Date.now();
if (now - lastWarningAt < 60_000) return;
lastWarningAt = now;
console.warn(`Backlog: ${pending} jobs are waiting`);
}
So how do you read these numbers? If active is equal to the limit and pending keeps growing, you have a problem. Now you can see it before your users do.
What this does not solve
I want to be honest with you. The limiter is a strong tool, but it is not everything.
- The queue is in memory. If the server restarts, the waiting jobs are lost. For work that we cannot lose, we save it in a database table first, and a worker takes it in small batches.
- The queue has no maximum size. If the work never slows down, the memory will finish one day. You can add a maximum and reject the extra work.
- There is no timeout. A job that never finishes keeps its slot forever. Give every job a time limit.
- It works on one server only. If you have 4 servers and each one has a limit of 10, the real limit is 40.
For real millions of requests you also need more servers, a load balancer and caching. But all of those still need this idea. Do not let all the work start at the same time.
If you do not want to write it yourself, you can use p-limit. It is one of the most used packages on npm and it works on the same idea. The difference is that now you know what it is doing.
All the problems and solutions
- The sender does not wait. Answer first, do the work in the background.
- The database runs out of connections. Put a limit on how many jobs run together.
- One slow job blocks everything. Use a separate line for each kind of work.
- The line grows and nobody knows. Watch the active and pending numbers.
What is next
A concurrency limiter controls how many jobs run at the same time. But it does not control how many requests one user can send in one second. That is a different tool. It is called a rate limiter, and I will explain it in the next post.
I post every topic as a short video first. You can follow @adil_thewebdev on Instagram for the next one. And if you are building something that has to scale, you can contact me here.