<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>codetheorem.dev</title><description>a site for my genai/agentic systems worklogs and nonsensical thoughts</description><link>https://codetheorem.github.io/</link><language>en-us</language><item><title>Why We Need Distributed Systems</title><link>https://codetheorem.github.io/blog/why-we-need-distributed-systems/</link><guid isPermaLink="true">https://codetheorem.github.io/blog/why-we-need-distributed-systems/</guid><description>Where one server stops being enough, what you actually trade away when you go distributed, and the three properties every design argument eventually reduces to.</description><pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2 id=&quot;the-one-server-setup-nobody-wants-to-admit-they-miss&quot;&gt;The one server setup nobody wants to admit they miss&lt;/h2&gt;
&lt;p&gt;You have a web app on a single server. Nginx, your application code, and Postgres, all on the same box. A few dozen people use it during the day. Response times are good. When something breaks you SSH in, tail the logs, and usually know what went wrong inside a minute.&lt;/p&gt;
&lt;p&gt;This is a nice place to be. Most teams leave it earlier than they should.&lt;/p&gt;
&lt;p&gt;The reason it’s nice isn’t nostalgia. A single machine gives you things that are genuinely expensive to buy back later. Your app talks to your database through a function call measured in microseconds, not a network round trip measured in milliseconds. There’s exactly one copy of your data, so “which version is correct” isn’t a question anyone has to ask. There’s one set of logs. Deploys are push, restart, done.&lt;/p&gt;
&lt;p&gt;None of that is nothing. Every one of those properties have a price tag once you go distributed, and you pay it forever.&lt;/p&gt;
&lt;p&gt;So the honest framing isn’t “when do I get to build a distributed system.” It’s “what forces me off this machine.”&lt;/p&gt;
&lt;h2 id=&quot;what-actually-forces-you-off&quot;&gt;What actually forces you off&lt;/h2&gt;
&lt;p&gt;Four things, and they arrive in a fairly predictable order.&lt;/p&gt;
&lt;p&gt;CPU goes first for most applications. Every server has a fixed number of cores, and once your business logic, JSON serialization, and template rendering saturate them, requests start queuing. The symptom isn’t a crash. It’s response times that creep from 80ms to 400ms over a few weeks while everyone argues about whether it’s the database.&lt;/p&gt;
&lt;p&gt;Memory goes next, and it goes badly. Your app and your database are competing for the same RAM. When you run out, the OS starts swapping, and memory access that took nanoseconds now takes milliseconds. Performance doesn’t degrade gracefully here. It falls off a cliff, and it usually does so at 2am on a Saturday.&lt;/p&gt;
&lt;p&gt;Disk I/O is the quiet one. SSDs pushed this ceiling much further out than it used to be, but a single drive still has a throughput limit, and a Postgres instance doing a few thousand transactions per second will find it.&lt;/p&gt;
&lt;p&gt;Network bandwidth matters if you’re streaming, serving large payloads, or holding thousands of open WebSocket connections. One machine, one NIC.&lt;/p&gt;
&lt;p&gt;And then there’s the fifth thing, which isn’t a resource limit at all. That box is a single point of failure. Hardware fault, kernel panic, a bad deploy, someone filling the disk with a runaway log file. Doesn’t matter which. Everything is down. Every user, every request.&lt;/p&gt;
&lt;p&gt;You can throw money at the first four. You can’t throw money at the fifth.&lt;/p&gt;
&lt;h2 id=&quot;buying-a-bigger-box-works-until-it-doesnt&quot;&gt;Buying a bigger box works until it doesn’t&lt;/h2&gt;
&lt;p&gt;The instinctive fix is vertical scaling. More cores, more RAM, faster disks. And this works, genuinely works, for longer than distributed systems evangelists like to admit.&lt;/p&gt;
&lt;p&gt;What kills it is the cost curve. Eight cores to sixteen might cost you double. Sixteen to thirty two, maybe four times. Past that you’re shopping for specialised hardware at ten to twenty times commodity pricing, and you’re doing it to buy a machine that is still, at the end of all that spending, one machine that can die.&lt;/p&gt;
&lt;p&gt;Vertical scaling is a strategy with a known expiry date. The only question is whether you hit it before your growth stalls.&lt;/p&gt;
&lt;h2 id=&quot;going-out-instead-of-up&quot;&gt;Going out instead of up&lt;/h2&gt;
&lt;p&gt;The alternative is to split the work across several smaller machines.&lt;/p&gt;
&lt;p&gt;The canonical first step is unglamorous. A load balancer in front of two or three app servers, with the database moved to its own dedicated box and a read replica behind it. Same traffic, but each server now sits at 35% CPU instead of 95%. There’s headroom. When you need more, you add another commodity server and the load balancer picks it up.&lt;/p&gt;
&lt;p&gt;That’s the whole idea. It’s not complicated. What’s complicated is everything that follows from it.&lt;/p&gt;
&lt;h2 id=&quot;what-you-get&quot;&gt;What you get&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Capacity that scales with your budget, not with hardware limits.&lt;/strong&gt; The ceiling becomes how many machines you’re willing to run, not how fast the fastest available CPU is.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failures that don’t take everything down.&lt;/strong&gt; One server dies, the others keep serving. Users might see a blip during failover. They don’t see a maintenance page.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The ability to scale parts independently.&lt;/strong&gt; Your database and your app servers almost never hit their limits at the same time. Distributed architecture lets you add read replicas without touching the app tier, or add app servers without touching the database.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Proximity to users.&lt;/strong&gt; A request from Tokyo served out of Tokyo is 20ms. Served out of Virginia it’s 200ms. For an interactive product, that difference is the entire feel of the thing.&lt;/p&gt;
&lt;h2 id=&quot;what-you-give-up-which-is-more-than-people-expect&quot;&gt;What you give up, which is more than people expect&lt;/h2&gt;
&lt;p&gt;This is the part that gets undersold in every architecture diagram.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Every function call becomes a network request.&lt;/strong&gt; That request can fail. It can time out. It can arrive out of order. Worst of all, it can succeed on the receiving end while the response gets lost on the way back, so your caller believes the operation failed when it actually happened. Every interaction between machines now has failure modes that simply did not exist on one box. Idempotency stops being something nice to have and becomes a requirement in places you won’t anticipate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failure stops being binary.&lt;/strong&gt; On one server, you’re up or you’re down. In a distributed system, Server A is healthy, Server B is thrashing, the database is mid backup and slow, and roughly 8% of requests are failing for reasons nobody can reproduce. “Is the site up?” becomes a question without a clean answer, which means your monitoring has to get much better before your architecture gets much bigger.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The network is not reliable and will remind you.&lt;/strong&gt; Packets drop. Latency spikes. Partitions happen, and they don’t happen cleanly, you rarely get a total split. You get the app server that can reach the database but not the cache, or the two app servers that can’t see each other but can both see the database and are now both convinced they’re the leader.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State gets genuinely hard.&lt;/strong&gt; Two copies of the data on two servers. A user updates their profile against Server A. The next request lands on Server B. What does it see? This sounds like a trivial question and it is not. It’s one of the hardest problems in the field, it has formal machinery built around it called consistency models, and it deserves it’s own post rather than a paragraph here.&lt;/p&gt;
&lt;p&gt;If your workload fits on one machine, none of the above is a price worth paying. Say that out loud in the design review.&lt;/p&gt;
&lt;h2 id=&quot;three-properties-permanently-in-tension&quot;&gt;Three properties, permanently in tension&lt;/h2&gt;
&lt;p&gt;Nearly every architecture argument you’ll have reduces to one of three things.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scalability&lt;/strong&gt; is whether you can absorb more load by adding resources, without rewriting anything. That last clause is the whole point. Plenty of systems can handle 10x traffic if you’re willing to spend three months restructuring them first. That isn’t scalability, that’s a rebuild.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Resiliency&lt;/strong&gt; is whether the system keeps working when things break, and things break constantly. Servers crash, disks fill, deploys ship bugs. Resiliency isn’t about preventing failure, it’s about containing how far a failure spreads and recovering without a human in the loop. Degraded but functioning beats down.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Maintainability&lt;/strong&gt; is whether people can understand, operate, and change the system. This one gets treated as a soft concern and it absolutely isn’t. A system nobody can debug in production, nobody can deploy without dread, and nobody can modify without breaking something adjacent is a liability regardless of how well it scales.&lt;/p&gt;
&lt;p&gt;The tension is the interesting part. Redundancy and failover make a system more resilient and less maintainable, less moving parts is not something you get for free. Caching and partitioning make it more scalable and introduce entirely new failure modes. There’s no configuration where all three are maximised. You’re picking a point on a surface based on what your product actually needs, and being honest about that beats copying whatever architecture was in the last conference talk you watched.&lt;/p&gt;
&lt;h2 id=&quot;the-pieces-youll-end-up-with&quot;&gt;The pieces you’ll end up with&lt;/h2&gt;
&lt;p&gt;Most modern distributed systems converge on roughly the same components.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Clients.&lt;/strong&gt; Browsers, mobile apps, other services. They’re outside your control, they’re on worse networks than you assume, and they will retry aggressively at the worst moment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CDN.&lt;/strong&gt; A globally distributed cache for static assets. Your JS, CSS, and images get served from a node near the user. This absorbs the easy traffic so your servers can spend their cycles on requests that actually need computing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Load balancer.&lt;/strong&gt; Distributes incoming requests across your app servers. Works at different layers depending on how much you need it to know, whether that’s routing based on DNS, Layer 4 routing on IP and port, or Layer 7 reading HTTP headers and making decisions based on path or host.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Application services.&lt;/strong&gt; Your business logic, often split into several independently deployable pieces. Auth, orders, payments. They talk to each other synchronously over HTTP or gRPC, or asynchronously through a queue. Splitting these too early is one of the most common and most expensive mistakes in this space.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Database.&lt;/strong&gt; Usually Postgres or MySQL for anything transactional. A primary handling writes, one or more replicas handling reads. You get fault tolerance, promote a replica, and read scaling, spread the queries, from the same arrangement.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cache.&lt;/strong&gt; Redis or Memcached sitting between the app and the database. Microsecond responses instead of millisecond ones. The tax is that you now maintain two copies of the truth, and reconciling them is your problem forever.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Message queue.&lt;/strong&gt; Kafka, SQS, RabbitMQ. Lets Service A hand off work and move on instead of blocking on Service B. This decouples services in time, which means they no longer have to be healthy simultaneously. That property is worth more than the throughput gains people usually cite.&lt;/p&gt;
&lt;p&gt;None of these are mandatory. Add each one when something is actually hurting, not because the diagram looks incomplete without it.&lt;/p&gt;</content:encoded></item></channel></rss>