buildwith.rextech & cs, animated Follow
All posts

Vertical vs horizontal scaling: when to use which

Bigger server or more servers? Both are right in the right place. The interview answer, real AWS prices, and the two mistakes where money quietly burns.

Cloud & systems · 7 Oct 2026 · 4 min read

It's results day and the college portal freezes. One person says "get a bigger server". Another says "add four more". In an interview the same question arrives more formally — how would you scale this system? — and both answers are correct. They're just correct in different places.

Two ways to add capacity

Vertical scaling (scaling up) means running the same application on a bigger machine: more CPU, more memory. The code doesn't change. The cost is a short window of downtime while you move to the larger machine, and a hard ceiling — at some point there is no bigger machine to buy.

Horizontal scaling (scaling out) means running more copies of the application on more machines, with a load balancer in front that sends each request to whichever server is free. It only works if any server can answer any request.

A photocopy shop outside a college in exam week is a useful picture. When the queue grows, you can replace the machine with a faster one (scale up), or run three machines and have the person at the counter send each job to a free machine (scale out). That person at the counter is the load balancer.

Where the money actually goes

A common belief is that big machines are expensive. On the major clouds, doubling a machine's size usually doubles its hourly price, so the price per CPU stays roughly the same. Here are AWS's general-purpose m7i instances in the Mumbai region:

InstancevCPU / RAMUSD per hour
m7i.large2 / 8 GB0.106
m7i.xlarge4 / 16 GB0.212
m7i.2xlarge8 / 32 GB0.424
m7i.4xlarge16 / 64 GB0.848

Source: AWS EC2 on-demand pricing feed, Linux, ap-south-1, checked 7 October 2026. About $0.053 per vCPU-hour on every row.

So the waste isn't in the size of the machine. It's in capacity you pay for and don't use. Two mistakes cause most of it:

  1. A big machine idling. A 16-vCPU machine running around the clock costs about $619 a month. If it spends the night at 10–20% utilisation, most of that money pays for idle cores. Autoscaling — adding and removing smaller machines as traffic changes — means you pay for peak capacity only while the peak lasts.
  2. A fleet for a small app. One m7i.large costs about $77 a month. Three of them plus a load balancer cost more than $250. If one machine was enough, that's roughly three times the bill, plus more moving parts to maintain. (Load-balancer pricing here uses AWS's published US-East rate; the Mumbai rate may differ slightly.)

The limits of each approach

Vertical scaling has a ceiling. The largest instance in the m7i family, the 48xlarge, has 192 vCPUs and 768 GB of memory. And however large a single machine is, it is still one machine: if it fails, everything stops. That's a single point of failure.

Horizontal scaling has a different problem: state. If a user's login session is stored in one server's memory, their next request may land on a different server that has never heard of them, and they appear logged out. In the photocopy shop, your PDF was saved on machine 1's desktop, and next time you're sent to machine 2. The fix is to move state out of the servers — sessions into Redis or a database, files into object storage — so every server is stateless and interchangeable.

Why the database is different

Application servers are easy to scale out because they can be made stateless. A database's whole job is to hold state, which makes splitting it hard. So the usual path is to scale the database up for as long as practical, add read replicas (read-only copies) when reads dominate — writes still go to a single primary — and only then consider sharding, which splits data across separate databases and makes joins and rebalancing harder.

A real example: in 2016 Stack Overflow described running nine web servers alongside SQL Server machines with 384 to 768 GB of memory — the web tier scaled out, the database scaled up. (Nick Craver, “Stack Overflow: The Architecture – 2016 Edition”.)

The answer to give in an interview

  1. Keep application servers stateless and scale them out behind a load balancer.
  2. Scale the database up for as long as you reasonably can.
  3. Add read replicas if reads dominate.
  4. Shard last, when writes or data outgrow the largest machine you can afford.
  5. Size everything from measured usage, not guesses.

Want to feel it rather than read it? The scaling simulator on the home page lets you add traffic, servers and machine size, and shows the bill as you go.