1. What it is
Throughput is how much work a system actually gets done in a given amount of time, such as the number of requests it handles per second.
2. Everyday analogy
Imagine a neighborhood car wash. (This is a made-up picture to build intuition, not a real business.) Each car takes ten minutes to go through. That ten minutes is like latency: how long one car waits from start to finish.
Now ask a different question: how many cars come out clean every hour? That number is throughput. And it can change without making any single wash faster. Open a second wash bay, and up to twice as many cars can be cleaned per hour, even though each car still takes ten minutes. Or let one slow step, like the single person drying cars by hand, hold everyone up, and the whole line slows to that person's pace.
In computing, latency asks "how long does one request take?" Throughput asks "how much gets handled per second?" They are related, but not the same.
3. How it actually works
Cloudflare's guide separates three ideas that often get mixed up. Bandwidth is the maximum amount of data that could pass through a network at any given time. Think of it as the size of the pipe. Throughput is the average amount of data that actually passes through over a period of time. Latency is a measure of time, not of amount. Throughput and bandwidth are not necessarily the same, Cloudflare notes, because throughput is affected by latency and other factors.
Throughput is usually counted per second. For a network, that might be megabytes per second. For a website, Google's book on Site Reliability Engineering says system throughput is typically measured in requests per second. For a messaging system, it might be messages per second. The unit depends on what kind of work you are counting.
A system's throughput can be held back by a single part. Engineers call that part the bottleneck: a component or resource that limits the performance of the whole system. Picture the narrow neck of a bottle, which limits how fast you can pour. Microsoft's architecture guide gives a clear example: if the database is the bottleneck, adding more web servers will not help. You have to find and fix the narrow point first.
Two ways to raise throughput come up in this lesson. The first is to do more work side by side, for example by adding more machines, like opening more wash bays. The second is batching: grouping many small pieces of work into fewer, larger chunks, so the fixed overhead of handling each piece is shared. Throughput and latency can also pull against each other. If a system waits for extra confirmation before replying, each request takes longer, and that extra waiting can reduce how much gets done per second.
4. Real-world example
In 2014, Jay Kreps published a benchmark on LinkedIn's engineering blog. He tested Apache Kafka, a system for passing messages between programs. The post describes Kafka as originally built at LinkedIn and, by then, part of the Apache Software Foundation.
The test used small, 100-byte messages on a Kafka cluster of three machines, with three more machines used to generate the test load and run supporting software. He mostly used default settings, on purpose, to show "off the shelf" performance. Some of the reported results:
- One sender, no extra copies of the data: 821,557 messages per second.
- Three senders on three different machines, with three copies of each message kept: 2,024,032 messages per second. That is the "2 million writes per second" in the post's title.
These are 2014 results from one test setup. They are a snapshot from that benchmark, not a measure of how Kafka performs today.
The post credits two design choices for these numbers: writing to disk in long, continuous runs instead of jumping around, and batching small messages into larger chunks. It also shows the trade-off from section 3. With one sender and three copies of each message, throughput was 786,980 messages per second when the system replied as soon as its own copy was saved, but 421,823 when it waited for all copies to confirm first. Kreps noted that this extra waiting "does seem to affect our throughput."
The same post measured latency separately: a median of 2 milliseconds for a message to travel from sender to receiver. That is a good reminder that throughput and latency are two different measurements of the same system.
Source: LinkedIn Engineering, "Benchmarking Apache Kafka: 2 Million Writes Per Second (On Three Cheap Machines)" (April 27, 2014): https://engineering.linkedin.com/kafka/benchmarking-apache-kafka-2-million-writes-second-three-cheap-machines
5. Diagram
Requests in Done per second
===> [ step A ] ==> [ step B ] ==> [ step C ] ===> (throughput)
fast SLOW fast
(bottleneck)
The whole line can only move as fast as step B.
Add a second B side by side, and more work gets through:
/==> [ step B ] ==\
==> [ A ] ==< >==> [ C ] ==>
\==> [ step B ] ==/6. When you'd care
Throughput matters when lots of things happen at once. Think of a ticket sale opening at noon, a busy shopping day, or a system that processes many payments or messages. In those moments, the question is not "how fast is one request?" but "can we keep up with all of them?"
When someone says "we need to handle more traffic," the first step is to find the bottleneck.
A common mistake: treating throughput and latency as the same thing. Improving one does not automatically improve the other, and sometimes they trade off, as the Kafka example shows. Always ask which one you actually need to improve.
7. Check yourself
Q1. An imagined shop actually serves 120 customers per hour. Which idea does "120 per hour" describe?
- A) Latency
- B) Throughput
- C) Bandwidth
- D) DNS
Show answer to question 1
Answer: B. Throughput is the amount of work actually completed per unit of time.
Q2. A website's database is the slowest part of the system. What happens if you add more web servers in front of it?
- A) Throughput doubles automatically
- B) Throughput likely stays about the same, because the database is still the bottleneck
- C) The database gets faster
- D) Latency disappears
Show answer to question 2
Answer: B. Microsoft's guide notes that if the database is the bottleneck, adding more web servers won't help.
Q3. In the 2014 Kafka benchmark, what happened to throughput when the system waited for all copies to confirm before replying?
- A) It went up
- B) It stayed exactly the same
- C) It went down
- D) It was not measured
Show answer to question 3
Answer: C. It fell to 421,823 messages per second for one sender, versus 786,980 when the system did not wait.
8. Sources
- LinkedIn Engineering, "Benchmarking Apache Kafka: 2 Million Writes Per Second (On Three Cheap Machines)" (2014): https://engineering.linkedin.com/kafka/benchmarking-apache-kafka-2-million-writes-second-three-cheap-machines
- Cloudflare Learning Center, "What is latency?" (section on latency, throughput, and bandwidth): https://www.cloudflare.com/learning/performance/glossary/what-is-latency/
- Microsoft Azure Architecture Center, "Design to scale out": https://learn.microsoft.com/en-us/azure/architecture/guide/design-principles/scale-out
- Google, Site Reliability Engineering, chapter "Service Level Objectives": https://sre.google/sre-book/service-level-objectives/
- Oracle Database Quality of Service Management User's Guide, Glossary ("bottleneck"): https://docs.oracle.com/cd/E11882_01/server.112/e24611/glossary.htm