All engineering cases
CASE 02 / Scalability · Live commerce

Handling a 12k-user spike during a live commerce event.

A live shopping event drew more than 12,000 concurrent users, about twice the usual peak and a load the platform had never been tested at. After keeping the event running, we reproduced the spike with load tests and tuned Cloud Run autoscaling until tests sustained about 20,000 users within the same infrastructure budget.

  • Node.js
  • Google Cloud Run
  • GCP
  • Taurus
  • Load testing
  • Autoscaling

Context

A live commerce platform used during live shopping events for a large Brazilian fashion brand. Events usually peaked around 6,000 concurrent users; one passed 12,000. The backend ran on Google Cloud Run after a recent migration from AWS to GCP, and the platform had never been tested at that load.

What failed

  • The database saturated first, mostly on CPU and memory, and the pressure cascaded through the rest of the infrastructure.
  • Capacity could not grow fast enough. Thousands of people joined at once, and between cold starts and the autoscaling configuration, new instances arrived after the demand was already there.
  • The migration was recent, but the problem only showed up at this level of traffic.

Immediate response

The event was live, so stability came before diagnosis. Resources were raised as an emergency measure to keep the platform up through the event. That bought time, not an explanation.

Reproducing the failure

  • After the event, we reproduced the load under controlled conditions with Taurus.
  • Developers simulated thousands of concurrent users, raising the load step by step.
  • Metrics and logs at each step showed which limits were reached first.

Changes

With the infrastructure team, we adjusted infrastructure and Cloud Run autoscaling, retested, and repeated. Each round looked for a balance between enough capacity to absorb thousands of users arriving at once and the cost of keeping that capacity available.

My contribution

As a Senior Developer, I worked with the infrastructure team on:

  • Running the Taurus load tests and reading the results against metrics and logs.
  • Analyzing where the limits were and taking part in each round of autoscaling and capacity changes.
  • Retesting after every change.

Result

After the adjustments, load tests sustained about 20,000 concurrent users without exceeding the infrastructure budget defined for that scenario. This is a load-test result, not a later production event.

Engineering takeaway

A spike is a different problem from growth. Autoscaling reacts to demand that already exists, so when thousands of people arrive in the same minute, cold starts decide whether capacity shows up in time. A load test that ramps up gently will pass on a system that fails at a live event; it has to reproduce simultaneous entry. From there, performance is a trade-off between capacity, how fast it reacts, and what it costs to keep ready.