← Field Notes
Homelab Homelab Operations
Published 8 min read

I Had Enough Hardware for a Cluster. I Chose Not to Build One.

Why having spare machines, spare disks, and a more elaborate architecture available was not enough reason to replace a homelab that already worked.

In this article
  1. Available hardware is not a requirement
  2. The proposed architecture did have advantages
  3. I had already paid for complexity once
  4. Storage efficiency was the most convincing temptation
  5. The homelab changed when people started depending on it
  6. More nodes would not automatically make the system more resilient
  7. I used a migration test instead
  8. Choosing not to migrate is still an engineering decision
  9. Spare hardware can stay spare
  10. Complexity needs to earn its place
  11. Publication and privacy note

I had enough hardware sitting around to make the idea tempting.

There were multiple computers, several disks with more capacity than their operating systems actually needed, and enough spare parts to start imagining a cleaner, more centralized environment. I could consolidate things. I could turn the machines into nodes. I could make the unused portions of those disks feel less wasted.

On paper, rebuilding the homelab around a cluster sounded like progress.

I decided not to do it.

That decision ended up being more useful than the cluster would have been.

Available hardware is not a requirement

This is an easy trap in a homelab.

You acquire another machine, so you start looking for a job for it. You find unused disk capacity, so you start designing a storage layer that can consume it. You learn about another virtualization platform, orchestration system, or distributed filesystem, and suddenly the existing environment starts looking primitive.

Nothing actually broke.

The architecture just stopped being interesting enough.

I had to separate two questions:

What could I build with the hardware I have?

and:

What problem does my current environment need me to solve?

Those are not the same question.

The first one is great for learning. The second one is how I should make operational decisions.

I already wrote about defining a service’s job before adding it in Before I Add a Service, I Write Down the Job. I realized the same rule needed to apply to architecture.

If I cannot describe the problem the migration solves, the migration itself is probably the project.

The proposed architecture did have advantages

That matters, because I did not reject the idea because it was bad.

A clustered virtualization environment could have given me a cleaner way to allocate compute resources. Centralized storage could have made spare capacity easier to use. Moving workloads between systems could have become more structured. A unified management layer would have been interesting to learn.

Those are legitimate benefits.

But benefits do not exist in isolation.

Every new abstraction also becomes another thing I need to understand when something fails.

That cost matters more in my homelab now than it did when everything was purely experimental.

I had already paid for complexity once

The strongest argument against rebuilding came from experience.

I had previously worked through situations where applications depended on storage somewhere else, containers needed to agree on how data was presented, and one apparently simple service problem turned out to be a dependency problem several layers below it.

None of that made distributed systems bad.

It taught me that every boundary creates an operational contract.

If application workloads live on one machine while their data lives on another, the application host needs storage to be available and mounted correctly before the workload is truly healthy.

If several systems participate in a service, each system becomes part of the troubleshooting path.

If a management layer creates additional virtual networks, storage abstractions, or node relationships, I need to understand those too.

That is why The Storage Host Became a Contract, Not Just a Server and The Application Host Is Replaceable on Purpose became important lessons for me.

Abstraction is useful when it buys me something.

It is expensive when I add it only because I can.

Storage efficiency was the most convincing temptation

The disks bothered me more than the compute.

A boot disk with hundreds of gigabytes free can feel wasteful when another machine needs storage. Looking at several systems like that makes pooling the capacity seem obvious.

But raw utilization is not the only measure that matters.

I also care about:

  • how obvious it is which machine owns the data;
  • what happens when one machine is offline;
  • how many systems need to be healthy before an application can read its files;
  • whether I can recover a failed host without first rebuilding an entire storage layer;
  • how easily I can explain the environment six months later.

A more efficient pool of capacity can still be a worse operational design for my needs.

I would rather leave some disk space unused than create a dependency I do not need.

Unused capacity looks inefficient on a diagram.

Unnecessary complexity feels inefficient every time something breaks.

The homelab changed when people started depending on it

This is the part that changed my decision the most.

There was a time when rebuilding everything over a weekend was part of the fun. If a service disappeared while I experimented, that was mostly my problem.

That stopped being entirely true.

Once a homelab provides things other people in the house actually use, architectural experiments acquire a maintenance window whether I call it that or not.

I wrote about that transition in When the Homelab Became Household Infrastructure.

The threshold for a rebuild has to get higher when the environment has users.

“That would be cool” is still a valid reason to build something in a lab.

It is no longer automatically a valid reason to replace something that is working.

More nodes would not automatically make the system more resilient

This was another assumption I had to challenge.

A diagram with several nodes looks more resilient than a diagram with two ordinary servers.

It might be.

But node count is not resilience.

If every node depends on the same storage layer, that storage layer can still be the critical dependency. If recovery requires knowledge of a cluster manager, a distributed storage system, virtual networking, and the applications themselves, I may have improved failover while making recovery harder.

That can be a good trade.

I just need to make it deliberately.

For my environment, I care a lot about replaceability. I want to be able to lose an application host, rebuild it, reconnect the persistent pieces, and understand what happened.

That simplicity has value even when it does not look sophisticated.

I used a migration test instead

Rather than asking whether the new architecture was better in theory, I started asking what I would need to prove before replacing the existing one.

The questions looked something like this:

QuestionWhat I needed to know
What problem disappears?A concrete limitation, not aesthetic dissatisfaction
What new dependencies appear?Storage, networking, quorum, management, or node relationships
What happens during one-host failure?Which services continue and which stop
What happens during storage failure?Whether compute redundancy still matters
How do I restore it?A recovery path I can actually perform
What gets easier day to day?A recurring operational improvement
What is the rollback plan?How I return to the known-good design

I could answer some of those questions in favor of the rebuild.

I could not answer the first one strongly enough.

That was enough to stop.

It is the same principle behind The Test Matrix I Use Before I Trust a Change: define success before the change, not after you have already invested enough effort to want the change to succeed.

Choosing not to migrate is still an engineering decision

This is the part I think homelabs make easy to forget.

Building is visible.

Not building is invisible.

A new dashboard, cluster, rack layout, storage pool, or migration produces screenshots and a story. Keeping the current architecture can feel like doing nothing.

But I did not do nothing.

I evaluated a different design against the environment I actually operate. I compared its benefits with its dependencies. I considered recovery, troubleshooting, storage ownership, and the people who use the services.

Then I rejected the migration because the operational case was not strong enough.

That is a decision.

More importantly, it is a reversible decision.

If the current architecture develops a real limitation later, I can revisit the cluster idea with a concrete requirement instead of a vague desire to use every piece of hardware efficiently.

Spare hardware can stay spare

I still have machines I can experiment with.

That is actually better.

A spare computer does not need to become a permanent production node to justify owning it. It can be a test environment, a temporary migration target, a recovery machine, a place to learn something destructive, or simply hardware available when another system fails.

That fits the model I described in The Three Jobs My Homelab Actually Serves.

Not every machine needs a permanent workload.

Not every disk needs to be 100 percent allocated.

Not every capability needs to become infrastructure.

Complexity needs to earn its place

I still like the cluster idea.

I may build one eventually.

The difference is that I no longer consider the ability to build it a reason to deploy it.

If a future workload needs easier migration between hosts, stronger isolation, different availability characteristics, or a storage model my current setup cannot provide cleanly, then the calculation changes.

Until then, the existing environment is understandable, recoverable, and doing its job.

That is enough.

One of the hardest habits I have had to learn in the homelab is knowing when technical curiosity should create a separate experiment instead of a production migration.

Sometimes the most mature upgrade is the one I decide not to make.

That is also the line between deliberate maintenance and endless rebuilding that I explored in The Difference Between Maintenance and Endless Tinkering.

Publication and privacy note

This Field Note is based on a real architecture decision. The published version intentionally generalizes machine identities, hardware quantities, storage capacities, network design, mount points, management interfaces, application inventory, account information, and other details that would turn the lesson into a map of the environment.

The architecture decision is the useful part. The private topology is not.

JO

Written by

Jessie Owens

I run Eldritch IT and write about the systems, repairs, infrastructure decisions, and business lessons behind the work.