buildwith.rextech & cs, animated Follow
All posts

What is system design? A beginner's map

Ten users and everything works. A million at once and the server falls over. A plain-language map of the five pieces every large system is built from.

Cloud & systems · 7 Oct 2026 · 2 min read

You build a website. Ten people use it and everything is fine. Then it gets shared, a million people arrive at once, and the server stops responding. System design is the work of deciding, ahead of time, how a system should be shaped so that moment doesn't take it down.

Most large systems are combinations of the same handful of parts. Learn what each one is for and you can read almost any architecture diagram.

1. The server

A server is the computer that runs your application and answers requests. One request at a time is easy; thousands per second is where design starts to matter. When a server has more work than it can handle, requests queue up, slow down and eventually fail — often with an HTTP 503 Service Unavailable.

2. Scaling

When load grows, you add capacity. You can make the server bigger (vertical scaling) or add more servers (horizontal scaling). Bigger is simpler but has a ceiling and leaves you with a single point of failure; more servers survive a failure but require the application to be stateless. We cover the trade-off in depth in Vertical vs horizontal scaling.

3. The load balancer

With several servers, something has to decide which one handles each request. That's the load balancer. The simplest strategy is round robin — take turns — while smarter ones send work to the least busy server or route by region. The load balancer also checks server health and stops sending traffic to a server that has failed.

4. The cache

Many requests ask for the same thing: the home page, a popular product, a profile. A cache keeps recent answers in fast memory so the system doesn't recompute them every time. It's often the single biggest speed-up available. The catch is freshness: cached data can go stale, so entries are given an expiry time (a TTL) or are cleared when the underlying data changes.

5. The database

The database is where data actually lives. Keeping it separate from the application servers means a server can crash without losing anything. Databases are the hardest part to scale because they hold state, which is why systems usually scale the database up first, add read-only replicas for heavy read traffic, and split data across machines (sharding) only when nothing else is enough.

Putting it together

A request from a user reaches the load balancer, which forwards it to one of several stateless servers. The server checks the cache first; on a miss it asks the database, stores the answer in the cache, and responds. Each piece removes a different bottleneck:

  • More servers handle more requests at once.
  • The load balancer spreads requests and routes around failures.
  • The cache avoids repeating expensive work.
  • The database keeps data safe and consistent.

How to practise

Pick an app you use every day and sketch it with these five boxes. Where does the traffic come in? What would you cache? What happens when one server dies? That habit — naming the bottleneck and the part that removes it — is most of what an interviewer is looking for.