Servers and the fleet
Five roles, the processes they run, the thresholds that colour them, and how the rack grows.
Last updated:
You manage named machines — web-01, db-02, lb-01. Each one carries CPU,
RAM, disk and temperature, a status derived from those, and a list of running
processes that top will show you.
The five roles
| Role | Joins from | Baseline CPU / RAM | Bandwidth | Processes you will see |
|---|---|---|---|---|
web | night 1 | 38% / 46% | 1500 req/s each | nginx, api-worker |
db | night 1 | 32% / 61% | — | postgres, checkout_tx |
lb | night 3 | 44% / 30% | 3000 req/s each | haproxy, keepalived |
cache | night 4 | 22% / 55% | — | redis-server |
worker | night 6 | 51% / 44% | — | celery, backup-agent |
Only web and lb carry bandwidth capacity. That is worth internalising: a
database is never the reason your network bar is red, and adding cache machines
does nothing for a traffic problem.
Process names are stable, so they become part of the vocabulary. A line naming
backup-agent at three in the morning means something quite different from a
line naming nginx.
Status thresholds
Status is derived from whichever metric is worst:
| Metric | WARN at | CRIT at |
|---|---|---|
| CPU | 80% | 95% |
| RAM | 85% | 95% |
| Disk | 90% | 98% |
| Temperature | 85 °C | 95 °C |
DOWN is different — it is not derived from a threshold and it is sticky.
A machine that goes down stays down until something puts it back.
Temperature is clamped to a 40–120 °C band, so a cooling problem has a real ceiling rather than running away to nonsense values.
How the rack grows
The fleet is rebuilt each night from the night's plan. It is not random and it is not a formula you need to derive — it is a table:
| Night | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Machines | 3 | 4 | 5 | 7 | 7 | 9 | 11 | 14 | 14 | 16 |
Night one is exactly web-01, web-02, db-01. By night ten you are running
sixteen machines across all five roles, with two databases, two load balancers
and three workers.