Opinion / Cloud

The Cloud Is Not Dead

You just need to know what you are paying for.

E
E
11 min read

In 2023, David Heinemeier Hansson wrote that 37signals had left the cloud.

And I loved the article.

Which might sound strange coming from someone about to disagree with much of its conclusion.

I sell enterprise infrastructure. Part of my job is literally making the case that customers should buy hardware instead of continuing to rent compute from a hyperscaler. I have spent much of my career around servers, storage, networking, virtualization and data centers.

So when someone says that owning hardware can be dramatically cheaper than renting the equivalent capacity from the cloud, you won’t hear much disagreement from me.

DHH’s numbers are compelling.

37signals spent roughly $500,000 on servers providing around 4,000 vCPUs, 7.68TB of RAM and 384TB of NVMe capacity. The company expected the move to save at least $1.5 million every year compared with its previous cloud spending, and it planned to amortize that hardware over five years.

That’s a great infrastructure story.

But it isn’t the death of cloud computing.

If anything, it demonstrates something infrastructure architects have understood for years:

The right infrastructure depends on the workload and the organization running it.

And 37signals happens to be a very good candidate for owning hardware.

I Ran the Numbers Too

After reading DHH’s article, I became curious.

What would something resembling those resources cost today if I tried to build them using AWS, Azure or Google Cloud?

I ran my own rough cost analysis.

The cloud lost.

Badly.

Depending on architecture, commitments, storage, traffic and services, reproducing substantial dedicated compute capacity in a hyperscale cloud can easily push a five-year bill well north of $2 million.

So let’s get one thing out of the way:

Cloud can be expensive.

There is no magic accounting trick that changes that.

AWS, Microsoft and Google are operating enormous global infrastructures, and customers pay for access to them.

But comparing a cloud invoice directly against a server purchase order misses part of what the cloud invoice is buying.

A server is not infrastructure.

It is a component of infrastructure.

Congratulations. You Bought a Server.

Let’s say you leave the cloud.

Your AWS bill disappears.

Great.

Now someone has to run the environment replacing it.

You need servers.

You need storage.

You need networking.

You need racks.

You need power.

You need cooling.

You need firewalls.

You need monitoring.

You need backups.

You need redundancy.

You need hardware support.

You need firmware management.

You need security.

You need someone who understands why the storage array started screaming at 2:13 in the morning.

And unless you are colocating the equipment, you need the physical facility surrounding all of it.

Most importantly, you need people who know how these systems work.

That last line gets left out of a surprising number of cloud-versus-on-premises calculations.

Employees are expensive.

A small startup can have a developer or cloud architect provision databases, compute, object storage, networking and application services without ever touching a physical server.

Move that environment completely on premises and suddenly you may need expertise in virtualization, networking, storage, systems administration, backup, security and hardware lifecycle management.

Maybe those aren’t six separate employees.

But those responsibilities don’t disappear simply because you bought the servers.

The cloud provider was performing some of them for you.

AWS explicitly describes this division through its shared-responsibility model: AWS manages the underlying facilities, hardware, networking and virtualization infrastructure while customers retain responsibility for different layers depending on the services they consume. Microsoft describes essentially the same progression: in an on-premises environment, the customer owns the entire stack, while Azure assumes responsibility for physical hosts, physical networking and the data center as workloads move into its cloud.

You aren’t merely renting CPUs.

You’re transferring operational responsibility.

And operational responsibility has a price.

37signals Is Actually a Great Example

This is where DHH’s story gets particularly interesting.

37signals didn’t buy $500,000 worth of servers and wheel them into the office.

The company bought hardware and shipped it directly to two data centers, where remote-hands personnel rack the equipment. 37signals then manages the environment remotely using KVM, Docker and Kamal.

That’s smart.

It also demonstrates that infrastructure isn’t binary.

Cloud or on-premises isn’t the only decision.

You can own hardware and colocate it.

You can lease hardware.

You can consume managed infrastructure.

You can build private cloud.

You can run hybrid cloud.

You can keep steady-state workloads on owned equipment and burst into public cloud when necessary.

You can use SaaS for one part of the business, PaaS for another, public-cloud IaaS somewhere else and physical servers for the workloads where the economics make sense.

The infrastructure world has spent years trying to turn this into a religious argument.

It shouldn’t be one.

But Don’t Forget the Cluster

There is another question I had when looking at those beautiful hardware numbers.

How much of that capacity is actually usable?

If I buy a server with 1TB of memory, I don’t necessarily have 1TB available to applications after designing the production environment around it.

Production infrastructure needs resilience.

Depending on the architecture, clustering, replication, failover capacity, storage protection and reserved headroom can consume a meaningful portion of the raw resources you’ve purchased.

A design capable of surviving the loss of a node cannot normally operate every node at 100% utilization and still honestly call itself highly available.

That doesn’t mean you automatically lose 50% of your infrastructure. The amount varies enormously by architecture.

But raw CPU, memory and storage numbers are not the same thing as usable resilient capacity.

The cloud has this issue too—you still have to architect your applications correctly for availability.

But you aren’t buying the spare physical server sitting in the next rack.

The provider owns that infrastructure.

You’re consuming capacity from it.

That difference matters when comparing costs.

The Most Expensive Feature of Cloud Is Also One of Its Best

DHH makes a point I completely agree with: being able to provision enormous amounts of infrastructure within minutes is incredible.

He argues that 37signals simply doesn’t have enough unpredictable demand to justify paying the premium for that capability.

Exactly.

37signals doesn’t need it.

Someone else might.

Imagine two developers starting a company.

They don’t know whether they’ll have 500 customers next year or 5 million.

Should they buy enough infrastructure for 5 million?

Absolutely not.

Should they buy enough for 500 and hope they can order, ship, rack, configure and integrate additional infrastructure quickly enough if their product suddenly takes off?

Probably not.

They can open a cloud portal.

Deploy the application.

Start small.

If the idea fails, destroy the resources.

If it succeeds, scale them.

That flexibility has tremendous value when uncertainty is high.

You aren’t necessarily paying for the cheapest CPU cycle.

You’re paying for optionality.

This Is Why Startups Love Cloud

Think about how absurdly powerful modern infrastructure provisioning has become.

I can have an idea in the morning and have infrastructure supporting it running before lunch.

I don’t need a purchase order.

I don’t need to call a hardware vendor.

I don’t need to wait for freight.

I don’t need to rack anything.

I don’t need to cable anything.

I don’t need to download an ISO.

I don’t need to configure a RAID controller.

I don’t need to wonder whether somebody remembered to update the firmware.

I provision what I need.

If the project works, great.

If it doesn’t, I delete it.

The infrastructure becomes disposable.

That’s an extraordinary capability for experimentation.

And experimentation is what startups do.

Maybe I’ve Just Become Lazy

I have a home lab.

If I want a clean environment there, I can log into my virtualization platform, carve out a new VM, assign compute and memory, find the appropriate installation media, install the operating system, patch it, configure networking and then start working on whatever idea caused me to create the VM in the first place.

None of this is particularly difficult.

I’ve been doing it for years.

But now I’m spoiled.

I can log into a cloud portal, select what I want, click a few buttons and start working.

Or better yet, describe the infrastructure as code and deploy the whole environment repeatedly.

Maybe I’ve become lazy.

I prefer to think I’ve become impatient.

Either way, my time has value too.

Hardware Doesn’t Live Forever

This is another part of the equation that deserves more attention.

37signals said it intended to amortize its new servers over five years.

Five years is perfectly reasonable.

But I’ve worked with organizations that refresh infrastructure considerably faster.

I’ve also worked with organizations at the opposite extreme.

Some equipment seems determined to achieve archaeological significance.

I have encountered customers still running IBM AS/400-era systems old enough that replacement parts become their own adventure.

Eventually someone is asking whether a drive they found on eBay might work.

If that machine is running something critical to your business, don’t be that customer.

Owning hardware means owning the hardware lifecycle.

At some point you have to refresh it.

That means another capital purchase.

Migration planning.

Support contracts.

Firmware compatibility.

Potential application compatibility.

Implementation labor.

Risk.

And eventually disposal.

Cloud customers have plenty of problems of their own, but they generally don’t wake up wondering whether a failed physical disk is still under warranty.

The provider handles the physical infrastructure underneath them.

That has economic value even though it doesn’t appear as a line item labeled “Things I No Longer Have to Worry About.”

The Refresh Cycle Changes the Math

This is why I would be careful with a five-year comparison that simply looks like this:

Servers: $500,000

versus

Cloud: millions

Suppose your organization follows a shorter hardware lifecycle.

Now that $500,000 isn’t necessarily a one-time five-year expense.

Perhaps you refresh part or all of the environment during that period.

Add support contracts.

Add colocation.

Add networking.

Add power.

Add backup infrastructure.

Add implementation services.

Add staff.

Add spare capacity required for failure scenarios.

Add the engineering time spent maintaining the environment.

Suddenly the gap starts narrowing.

Maybe cloud still loses.

For many stable workloads, it absolutely will.

But now we’re finally comparing architectures instead of invoices.

Cloud Isn’t Automatically Highly Available Either

There is one place where I would modify my original instinct.

It is tempting to say that cloud “already comes resilient.”

That’s not quite true.

Cloud provides the infrastructure and services required to build resilient systems.

Customers still have to architect them properly.

Put a poorly designed application on one virtual machine in one availability zone and the cloud will not magically make it highly available.

AWS explicitly keeps responsibility for guest operating systems, applications and many configuration decisions with the customer.

What cloud changes is the availability of the building blocks.

Need another availability zone?

It’s there.

Need object storage designed around enormous distributed infrastructure?

It’s there.

Need another geographic region?

It’s there.

Need another hundred servers?

Give it a few minutes.

Try asking your hardware vendor for another hundred servers by this afternoon.

Tell me how that goes.

The Cloud Tax Is Real

None of this excuses bad cloud economics.

Cloud waste is real.

Idle instances are real.

Forgotten resources are real.

Oversized virtual machines are real.

Data-egress charges are real.

Organizations absolutely spend fortunes because nobody has seriously examined what they’re consuming.

And once an application becomes mature, predictable and large enough, owning the underlying hardware can become extremely attractive.

That’s where I agree with DHH.

If your company knows its baseline demand, has experienced infrastructure engineers, can amortize equipment across several years and doesn’t need enormous elasticity, you should absolutely run the numbers.

You may discover you’re paying a very expensive cloud tax for flexibility you rarely use.

But that’s different from saying the cloud itself was a mistake.

Do the TCO, Not the Headline Math

The debate shouldn’t be:

Cloud versus on-premises.

It should be:

What is the total cost of operating this workload at the service level the business requires?

Calculate the cloud bill.

Then calculate the hardware.

And the storage.

And networking.

And facilities.

And power.

And support.

And software licensing.

And redundancy.

And backup.

And disaster recovery.

And security.

And refresh cycles.

And the people required to operate all of it.

Then calculate something much harder to put into a spreadsheet:

What is flexibility worth to your organization?

For 37signals, apparently not $3.2 million per year.

That’s entirely reasonable. Its workload is mature, its demand is predictable, its technical team is sophisticated, and it has enough scale to make hardware ownership economically attractive. DHH explicitly acknowledges that cloud can still make sense for young companies and highly variable workloads.

For a six-person startup trying to figure out whether its product will even exist two years from now?

The answer can look very different.

The Cloud Was Never Supposed to Win Every Workload

Maybe this is where the industry went wrong.

We spent years talking about “cloud transformation” as though cloud were the destination.

It isn’t.

Cloud is an infrastructure consumption model.

Owned hardware is another.

Colocation is another.

Managed hosting is another.

Hybrid architecture is another.

None deserves ideological loyalty.

I sell hardware.

I would love to sell you servers.

But if you’re a three-person startup asking me whether you should spend your limited capital building an enterprise infrastructure stack before you’ve even proven your product?

I’ll probably tell you to keep your credit card and open a cloud account.

Come see me when your cloud bill starts hurting.

Then we’ll run the numbers.

Because David Heinemeier Hansson is right about something extremely important:

You should do your own math.

I would just add one line.

Make sure you’re counting everything.