The change looks finished#
The pull request changes one line. The API service on Amazon ECS goes from four tasks to twelve, because traffic grew and latency followed it.
resource "aws_ecs_service" "api" {
name = "api"
desired_count = 12 # was 4
}
terraform plan shows one attribute changing in place. Nothing in the repository is wrong, and a careful reviewer approves it. The change is still dangerous, and the reason is not in the diff.
The arithmetic lives somewhere else#
Each task opens a database connection pool when it starts. Most pools have a fixed ceiling: HikariCP defaults to 10 connections, node-postgres to 10, and SQLAlchemy to 5 plus an overflow of 10. Teams raise these under load, and 20 per task is common.
The database has its own ceiling. PostgreSQL refuses new clients past max_connections, which on Amazon RDS comes from the instance's parameter group, and it holds a few of those back for superusers. Suppose the limit is 200 and three are reserved:
| Connections | |
|---|---|
| Four tasks at 20 each | 80 |
| Twelve tasks at 20 each | 240 |
| Available to the application | 197 |
The steady state does not fit. The rollout is worse. ECS replaces tasks with a rolling deployment, and its default lets the service run up to twice its desired count while old and new tasks overlap. The migration job, the queue workers and the admin console draw on the same limit.
How it fails#
New tasks start, open their pools, and PostgreSQL answers with FATAL: sorry, too many clients already. The tasks fail their health checks, so ECS stops them and starts replacements, which fail the same way. The deployment stalls. Depending on how connections are retried, tasks that were healthy before the change can lose their connections too.
The incident review will find the cause quickly.
Why review misses it#
The three numbers that decide the outcome sit in three places:
- the task count, in the Terraform diff;
- the pool size, in the application's configuration, often in a different repository;
- the connection limit, in a live parameter group that may have drifted from any file that describes it.
A person reviewing the pull request sees the first. A code reviewer that reads the repository may find the second. Neither can see the third, which is a setting of the environment the change deploys into.
What a review needs to catch it#
Catching this class of failure takes two pinned inputs: the exact commit under review, and a recorded snapshot of the target environment that includes the database's limit and the other clients that share it. With both, the failure is arithmetic. Without the snapshot, the honest answer is that the review is incomplete, not that the change looks fine.
This is the review Proofline posts on the pull request for this change, with illustrative values:
Scale API to 12 tasks #42
main from scale-api@@ -17,4 +17,4 @@ resource "aws_ecs_service" "api" { name = "api"− desired_count = 4+ desired_count = 12View evidence
- Code
infra/ecs/api.tf:19sets 12 tasks. The application opens its 20-connection pool before reporting healthy.- Environment
- The pinned snapshot of
orders-prodrecordsmax_connections = 200and 3 reserved connections.
The extra tasks can fail their health checks and stall the deployment. Reduce the task count or pool size, or add a connection pooler.
}The finding names the trigger, the mechanism and the evidence behind each number, so the author can check it instead of trusting it.
Ways to fix it#
- Shrink the pool. Twelve tasks at 15 connections fit under 197, if nothing else connects.
- Add a connection pooler. Amazon RDS Proxy or PgBouncer multiplexes many client connections onto fewer database connections.
- Raise the limit deliberately. A higher
max_connectionscosts memory on the database instance, so size it with the instance class. - Bound the rollout. A lower maximum percent on the ECS deployment keeps old and new tasks from doubling the count during a deploy.
Database connections run out before all tasks become healthy
Each task opens 20 connections before passing its health check. With 12 tasks, the service needs 240 connections. The production database allows 197.