Module 5 · Cloud & DevOps Engineer Capstone: Ready for the Sale
Capacity for this year's peak
Turn last year's peak and this year's forecast into a capacity target, read a load test to find what limits each configuration, and use Little's law to see why a connection pooler fixes last year's problem.
About 25 minutes
The problem
Marketing expects this year's sale to bring 1.6 times last year's peak traffic. The team ran a load test in staging (loadtest.csv): for each configuration, the highest request rate it handled while keeping p95 latency under 800 ms and errors under 1%. How much capacity does Kasuwa need, and which configuration delivers it?
The concept
A capacity target
Target = last year's peak × expected growth × headroom. Headroom (here 30%) covers forecast error, uneven traffic within a minute, and losing an instance or two at the worst moment.
Find the limiting resource
Each configuration hits a different wall: app instances, database connections or database CPU. Adding more of something that isn't the limit does nothing, as last year showed.
Little's law
Average number of things in a system = arrival rate × time each spends there. For a database: connections busy at once = queries per second × seconds per query. It tells you how many connections you actually need, and it's usually far fewer than 20 per instance.
Example
The target:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/platform/"
metrics = pd.read_csv(base + "sale_metrics.csv")
load = pd.read_csv(base + "loadtest.csv")
peak_2025 = metrics["requests_per_s"].max()
target = peak_2025 * 1.6 * 1.3
print(f"2025 peak {peak_2025:.1f} req/s × 1.6 growth × 1.3 headroom = target {target:.0f} req/s")2025 peak 160.3 req/s × 1.6 growth × 1.3 headroom = target 333 req/sWhat each configuration can handle:
table = load.pivot_table(index="instances", columns=["pooler", "db_class"], values="max_rps_within_slo")
print(table, "\n")
load.groupby(["pooler", "db_class"])["limited_by"].agg(lambda s: ", ".join(dict.fromkeys(s)))pooler no yes
db_class db.r6g.xlarge db.r6g.2xlarge db.r6g.xlarge
instances
6 74.0 73.0 69.0
9 104.0 106.0 106.0
12 118.0 140.0 145.0
15 119.0 180.0 178.0
18 119.0 211.0 217.0
21 120.0 251.0 249.0
24 119.0 281.0 267.0
27 120.0 339.0 262.0
30 119.0 373.0 253.0
pooler db_class
no db.r6g.xlarge app instances, database connections
yes db.r6g.2xlarge app instances
db.r6g.xlarge app instances, database CPU
Name: limited_by, dtype: objectWithout a pooler, capacity stops at about 120 requests a second, however many instances you add: the database connections run out, exactly as on sale day. With the pooler, capacity grows with instances until the current database's CPU becomes the limit. Only the larger database lets capacity keep growing. The smallest configuration that meets the target:
enough = load[load["max_rps_within_slo"] >= target].sort_values("instances")
enough.head(3)pooler db_class instances max_rps_within_slo limited_by
25 yes db.r6g.2xlarge 27 339 app instances
26 yes db.r6g.2xlarge 30 373 app instancesSo the plan is: the pooler, the larger database, and autoscaling allowed up to 30 instances, since the smallest passing size leaves almost no room beyond the headroom. That matches the max_size in PR 214.
Why does the pooler work? With Little's law, at the target rate, and assuming each request spends about 60 ms in the database:
db_seconds_per_request = 0.060
busy_connections = target * db_seconds_per_request
print(f"Connections busy at once at {target:.0f} req/s: about {busy_connections:.0f}")
print(f"Connections last year's setup would open at 30 instances: {30 * 20}")Connections busy at once at 333 req/s: about 20
Connections last year's setup would open at 30 instances: 600The work needs about twenty connections at a time. Last year's design opened twenty per instance, whether busy or idle. A pooler lets many app instances share a small set of real database connections.
Walkthrough
- Run the cells.
- Recompute the target with growth of 1.4 and 2.0. Which configuration does each need?
- Plot capacity against instances for the three configurations, with a horizontal line at the target.
- List what the load test didn't cover: for example, the payment gateway's own limits.
- Write the capacity plan (the task below).
Practice
Practice
What is the capacity target in requests per second? Whole number.
Practice
What is the smallest number of instances in the load test that meets the target?
Task
7 minWrite the capacity plan (60 to 140 words): the target and how you got it, the configuration you need and why the others fail, the role of the pooler, and one limit of the load test.
Your work is checked for
- Gives the target
- Explains growth and headroom
- Names the configuration (pooler, database size, instances)
- Explains why others fail (connections, CPU)
- Covers the pooler
- A limit of the test
- Between 60 and 140 words
Check your understanding
Answer every question to check.