When the Platform Had to Handle 20× More Traffic

Impact

~20× more user capacity

When demand spiked far beyond what our monolithic backend could handle, we didn't rewrite the infrastructure — we distributed traffic across instances and multiplied capacity by 20×.

Software Developer

During the pandemic, demand for our platform exploded almost overnight. Our infrastructure had been enough for normal traffic, but growth combined with promotional spikes started pushing it far past what it was built for. The backend was essentially a monolith — each instance had limited processing capacity, and vertical scaling alone would eventually hit a physical and economic limit. We had no Kubernetes-based autoscaling ready to react to those spikes, and we needed a fix fast, without turning a performance emergency into a full infrastructure rewrite.

Scaling horizontally with what we had

The first decision was changing how we distributed traffic. Instead of trying to make a single instance handle the whole load, we spun up multiple API instances, both on AWS and on servers we already had internally, and used Apache as a load balancer to distribute requests between them. The idea was simple — if one instance could handle a given amount of traffic, several instances running in parallel gave us considerably more total capacity without having to modify the application right away.

The frontend had its own quirk — our web infrastructure was deployed on AWS, so we also needed to distribute those requests across different servers. We implemented a DNS-level distribution strategy to route traffic to different destinations, so it stopped depending on a single entry point. It was not a fully elastic infrastructure platform, but we did not need to build one to solve the immediate problem — we needed to buy capacity quickly and in a controlled way.

One of the most important decisions was not trying to solve every architectural problem we would eventually face right then. We did not migrate everything to Kubernetes, did not rewrite the backend, did not turn the platform into a fully distributed architecture. We built a distribution layer around the infrastructure we already had, which turned a single-instance capacity problem into a load-distribution problem across several instances, and we managed to do it in a very short period.

Result

The strategy worked — a series of relatively small configuration changes let us increase the platform's user capacity by roughly 20× compared to the original setup. This mattered most during promotional campaigns, where we could go from normal traffic to huge spikes almost instantly, and the infrastructure distributed that load across the available servers. The platform kept good performance during one of the periods of highest growth and demand we had experienced.

What I learned

This was one of those problems where the most technically sophisticated solution was not necessarily the right one. In hindsight, we probably could have designed a much more complex, fully automated infrastructure, but the problem in front of us was not building the perfect infrastructure — it was keeping the platform from collapsing while demand grew exponentially. The solution was pragmatic, relatively cheap, and fast to implement, and it gave us something we needed at that moment — time. Time for the product to keep growing, and to later think about a deeper evolution of the infrastructure. To me, this case still represents an idea that matters when working with production systems — good engineering is not always about building the most advanced architecture. Sometimes it is about finding the smallest change capable of solving a huge problem.