All posts

A one-line scale-up can exhaust your database

Raising a service's task count is a correct change in the diff. Whether production can absorb it depends on numbers that live in the environment, outside the repository.

The change looks finished#

The pull request changes one line. The API service on Amazon ECS goes from four tasks to twelve, because traffic grew and latency followed it.

resource "aws_ecs_service" "api" {
  name          = "api"
  desired_count = 12 # was 4
}

terraform plan shows one attribute changing in place. Nothing in the repository is wrong, and a careful reviewer approves it. The change is still dangerous, and the reason is not in the diff.

The arithmetic lives somewhere else#

Each task opens a database connection pool when it starts. Most pools have a fixed ceiling: HikariCP defaults to 10 connections, node-postgres to 10, and SQLAlchemy to 5 plus an overflow of 10. Teams raise these under load, and 20 per task is common.

The database has its own ceiling. PostgreSQL refuses new clients past max_connections, which on Amazon RDS comes from the instance's parameter group, and it holds a few of those back for superusers. Suppose the limit is 200 and three are reserved:

Connections
Four tasks at 20 each80
Twelve tasks at 20 each240
Available to the application197

The steady state does not fit. The rollout is worse. ECS replaces tasks with a rolling deployment, and its default lets the service run up to twice its desired count while old and new tasks overlap. The migration job, the queue workers and the admin console draw on the same limit.

How it fails#

New tasks start, open their pools, and PostgreSQL answers with FATAL: sorry, too many clients already. The tasks fail their health checks, so ECS stops them and starts replacements, which fail the same way. The deployment stalls. Depending on how connections are retried, tasks that were healthy before the change can lose their connections too.

The incident review will find the cause quickly.

Why review misses it#

The three numbers that decide the outcome sit in three places:

  • the task count, in the Terraform diff;
  • the pool size, in the application's configuration, often in a different repository;
  • the connection limit, in a live parameter group that may have drifted from any file that describes it.

A person reviewing the pull request sees the first. A code reviewer that reads the repository may find the second. Neither can see the third, which is a setting of the environment the change deploys into.

What a review needs to catch it#

Catching this class of failure takes two pinned inputs: the exact commit under review, and a recorded snapshot of the target environment that includes the database's limit and the other clients that share it. With both, the failure is arithmetic. Without the snapshot, the honest answer is that the review is incomplete, not that the change looks fine.

This is the review Proofline posts on the pull request for this change, with illustrative values:

your-org / your-repo

Scale API to 12 tasks #42

Openyou want to merge 1 commit into main from scale-api
infra/ecs/api.tf+1 −1
@@ -17,4 +17,4 @@
1717 resource "aws_ecs_service" "api" {
1818 name = "api"
19− desired_count = 4
19+ desired_count = 12
prooflinebotcommented

Database connections run out before all tasks become healthy

Each task opens 20 connections before passing its health check. With 12 tasks, the service needs 240 connections. The production database allows 197.

View evidence
Code
infra/ecs/api.tf:19 sets 12 tasks. The application opens its 20-connection pool before reporting healthy.
Environment
The pinned snapshot of orders-prod records max_connections = 200 and 3 reserved connections.

The extra tasks can fail their health checks and stall the deployment. Reduce the task count or pool size, or add a connection pooler.

2020 }
Illustrative review in an example repository.

The finding names the trigger, the mechanism and the evidence behind each number, so the author can check it instead of trusting it.

Ways to fix it#

  • Shrink the pool. Twelve tasks at 15 connections fit under 197, if nothing else connects.
  • Add a connection pooler. Amazon RDS Proxy or PgBouncer multiplexes many client connections onto fewer database connections.
  • Raise the limit deliberately. A higher max_connections costs memory on the database instance, so size it with the instance class.
  • Bound the rollout. A lower maximum percent on the ECS deployment keeps old and new tasks from doubling the count during a deploy.

Review your next pull request against its infrastructure.

Connect one repository and the environment it deploys to. Proofline reviews each change against recorded evidence from that environment.

Get Started with GitHub

More from the blog