1. What it is
Scalability is how well a system can handle more work when you give it more resources.
2. Everyday analogy
Picture a small neighborhood bakery. (This is an imagined example, not a real shop.) One baker, one oven. On a normal morning, it keeps up. Then a local news story mentions the bakery, and the line runs out the door.
The owner has two choices. She can buy a bigger, faster oven. That helps, but ovens only come so big, and while the old one is swapped out, nothing gets baked. Or she can add a second oven and a second baker, then a third, and keep adding as the line grows. If one oven breaks, the others keep going.
Computer systems face the same two choices, and they have names. Making one machine bigger is called vertical scaling. Adding more machines is called horizontal scaling. The bakery is only a picture, though. In real systems, adding machines is not enough on its own. The work also has to be designed so that many machines can share it, which the next section explains.
3. How it actually works
Start with the two kinds of scaling, as Microsoft's architecture guide defines them. Vertical scaling, also called "scaling up," means increasing the capacity of one resource, such as moving an application to a larger machine. It often requires taking the system offline for a while during the switch, and Microsoft notes there may be limits to how much you can scale up. Horizontal scaling, also called "scaling out," means adding more copies, called instances, of a resource. The application keeps running while new instances are added, and extra instances can be removed when demand drops.
Microsoft's guide also gives a useful way to measure scalability: compare how much more work you get done with how many more resources you added. In an ideal system, doubling the resources doubles the throughput. In practice, scalability is typically limited by bottlenecks, parts that limit the performance of the whole system.
Horizontal scaling works best when any instance can handle any request. If one server keeps important information only in its own memory, requests have to keep going back to that same server, which limits how far you can spread the work. The guide's advice is plain: "Make sure that any instance can handle any request." Some parts are harder to spread out than others. Databases are one example. Microsoft notes that scaling a database horizontally usually involves partitioning, which means dividing the data into smaller parts that are stored separately, and that this is generally not automated.
Cloud providers let companies add capacity as demand rises and remove it when it is no longer needed. Microsoft's guide calls this elastic scaling and describes it as a primary advantage of the cloud.
4. Real-world example
Netflix has written publicly about why it moved its systems into the cloud. In a 2016 post, Netflix leaders explained that the journey began in August 2008, when a major database corruption left the company unable to ship DVDs to its members for three days.
In their words, that is when they "realized that we had to move away from vertically scaled single points of failure, like relational databases in our datacenter, towards highly reliable, horizontally scalable, distributed systems in the cloud." A single point of failure is a part that the rest of the system depends on, so its failure can take down the whole thing. In 2008, that database's failure meant Netflix could not ship DVDs for three days. Netflix chose Amazon Web Services (AWS) as its cloud provider.
The post reports that, at the time of writing, Netflix had eight times as many streaming members as in 2008, and that overall viewing had grown by three orders of magnitude, about a thousandfold, in eight years. Netflix wrote that supporting this growth from its own data centers would have been extremely difficult: "we simply could not have racked the servers fast enough." In the cloud, it said, it could add thousands of virtual servers and petabytes of storage within minutes.
The move was not quick. Netflix said it took seven years, finishing in early January 2016. Instead of moving its old systems over unchanged, it rebuilt almost all of its technology, breaking one large application into hundreds of smaller services.
Source: Netflix, "Completing the Netflix Cloud Migration" (February 12, 2016): https://about.netflix.com/en/news/completing-the-netflix-cloud-migration
5. Diagram
VERTICAL (scale up) HORIZONTAL (scale out)
[ server ] [ server ] [ server ] [ server ]
| \ | /
v \ | /
[ BIGGER SERVER ] one system, shared work
+ suits apps that are hard to split + add or remove as demand changes
- has an upper size limit + one failure needn't stop all
- often needs downtime - work must be designed to share6. When you'd care
Scalability matters when growth is likely or uneven. Some services see sudden surges in traffic, while others have a smaller, more predictable load. Scalability also matters for reliability, because spreading work across several machines means one failure does not have to stop everything.
If you work with a technical team, a useful question is not just "does it scale?" but "what happens if demand doubles?" Google's book on Site Reliability Engineering suggests asking whether a service can properly handle double its traffic.
A common mistake: assuming more servers will always fix slowness. Microsoft's guide warns that scaling out "isn't a magic fix for every performance issue." If the database is the bottleneck, adding web servers won't help. Find the bottleneck first.
7. Check yourself
Q1. A company replaces its server with a single, much larger one. What kind of scaling is this?
- A) Horizontal scaling
- B) Vertical scaling
- C) Load balancing
- D) Caching
Show answer to question 1
Answer: B. Vertical scaling means increasing the capacity of one resource, such as moving to a larger machine.
Q2. Microsoft's guide advises: "Make sure that any instance can handle any request." Why does that matter for horizontal scaling?
- A) So that work can be spread freely across all the machines
- B) So that each machine can be smaller
- C) So that the website looks the same on every phone
- D) So that the servers use less electricity
Show answer to question 2
Answer: A. If requests must keep returning to one specific server, the work can't be spread out freely.
Q3. What event did Netflix say started its move to the cloud?
- A) A new product launch in 2016
- B) A major database corruption in 2008 that stopped DVD shipments for three days
- C) A power outage at a streaming studio
- D) A request from a government agency
Show answer to question 3
Answer: B. Netflix wrote that its cloud journey began after that August 2008 database corruption.
8. Sources
- Netflix, "Completing the Netflix Cloud Migration" (2016): https://about.netflix.com/en/news/completing-the-netflix-cloud-migration
- Microsoft Azure Architecture Center, "Autoscaling guidance": https://learn.microsoft.com/en-us/azure/architecture/best-practices/auto-scaling
- Microsoft Azure Architecture Center, "Design to scale out": https://learn.microsoft.com/en-us/azure/architecture/guide/design-principles/scale-out
- Microsoft Azure Well-Architected Framework, "Architecture strategies for optimizing scaling and partitioning": https://learn.microsoft.com/en-us/azure/well-architected/performance-efficiency/scale-partition
- Microsoft Azure Well-Architected Framework, "Architecture strategies for optimizing scaling costs": https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/optimize-scaling-costs
- Google Cloud Architecture Center, "Design reliable infrastructure for your workloads": https://docs.cloud.google.com/architecture/infra-reliability-guide/design
- Google, Site Reliability Engineering, chapter "Monitoring Distributed Systems": https://sre.google/sre-book/monitoring-distributed-systems/
- Oracle Database Quality of Service Management User's Guide, Glossary ("bottleneck"): https://docs.oracle.com/cd/E11882_01/server.112/e24611/glossary.htm