Terabytes describe how much data a system can hold. They do not describe whether the system can be trusted.
Early in my homelab, storage mostly meant capacity.
How many drives could I add? How large would the pool become? How quickly would the media library consume the available space?
Those questions matter, but they are not the questions that determine whether storage is dependable.
A large pool is not useful when applications cannot mount it, permissions drift between systems, nobody knows which copy is authoritative, or the recovery plan amounts to hoping every drive survives.
Shared storage forced me to stop thinking in totals and start thinking in contracts.
Capacity is only the visible number
Drive capacity is easy to compare because it produces one clean number.
Reliability depends on decisions that are less visible:
- Which system owns the authoritative data?
- How are drives grouped, and which failures can the layout tolerate?
- How is corruption detected?
- Which datasets or directories serve which purposes?
- How do clients mount the storage?
- Which identities receive read or write access?
- What should dependent services do when storage is unavailable?
- Which data exists in an independent backup?
- What evidence proves that restoration works?
Adding another drive answers only one of those questions.
That is the first lesson: capacity is a property of the hardware, while dependability is a property of the design.
Shared storage creates a chain of dependencies
A common homelab pattern separates storage from application compute. One system owns the pool and exposes selected data to another system, where applications consume it.
That division can be useful. Storage and compute have clearer responsibilities, application maintenance does not necessarily require moving the data, and multiple services can use one controlled source.
It also creates a dependency chain:
| Boundary | What must be true |
|---|---|
| Drives to pool | Devices are available and the pool is healthy enough to serve data |
| Pool to dataset | The intended dataset exists with the correct properties |
| Dataset to export | Only the approved data is shared to approved clients |
| Export to client | The client mounts the expected remote filesystem |
| Client to container | The service receives the correct subpath and identity |
| Container to application | The application can perform only the reads or writes its role requires |
A failure at any point can appear in the application as the same vague symptom: the files are missing.
That is why troubleshooting only from the media server or container dashboard is not enough. The application is the final consumer of a storage path assembled across several layers.
The mount is part of the application
A remote mount can fail without removing the local mount-point directory.
That distinction matters.
Suppose an application expects its library beneath a directory used as a network mount. If the remote filesystem is unavailable, the directory itself may still exist as an ordinary empty local path. A container can start successfully against that path because, from the container runtime’s perspective, the directory is present.
The application then sees an empty library rather than an obvious storage failure.
Worse, a service with write access may begin creating new data in the unmounted local directory. When the remote filesystem returns, that locally written data becomes hidden beneath the mount, creating two different locations that appeared to have the same path at different times.
The safer design treats storage readiness as an application prerequisite:
- Verify that the expected filesystem is actually mounted.
- Confirm that it is the intended source, not merely an existing directory.
- Prevent dependent services from starting when the mount is absent.
- Alert on mount failure rather than waiting for an application symptom.
- Revalidate the complete path after reboots, network interruptions, and maintenance.
Startup ordering alone is not enough if it only means “try this unit first.” The design needs a meaningful readiness check and a deliberate failure state.
Permissions cross every boundary
Shared storage also turns identity into architecture.
Access may be evaluated at the dataset, export, client mount, host filesystem, container mapping, and application process. A file can be visible at one layer while remaining unreadable or unwritable at the next.
The tempting fix is broad permission changes. That may make the immediate error disappear, but it also removes the boundaries that explain who is supposed to modify the data.
I need consistent answers to questions such as:
- Which identity owns newly created files?
- Which group represents the applications that share access?
- Are numeric user and group identities consistent where they need to be?
- Which clients are permitted to mount the export?
- Which services need read-only access?
- Which service is allowed to rename, import, or delete files?
- What should detect and correct permission drift?
A successful test performed as an administrator proves very little about the account used by the actual service.
When troubleshooting, I test from the same system, container, path, and effective identity used by the application. That separates a real authorization problem from a path or mount problem.
ZFS improves integrity, but it does not remove responsibility
ZFS changed how I think about storage because it treats integrity as a first-class concern.
OpenZFS documents that data is protected by end-to-end checksums. Reads verify those checksums, and a redundant pool can repair damaged data when a valid second copy is available. A scrub walks stored data to find latent corruption and repairs it where redundancy permits.
That is valuable, but each capability has a boundary.
Checksums can detect corruption without being able to repair it when no good copy exists. Redundancy can keep data available through defined device failures without protecting against every other loss. A scrub can validate the pool without proving that an application database was captured consistently.
Snapshots also have a specific job. OpenZFS defines a snapshot as a read-only, point-in-time view of a dataset. Snapshots can make accidental changes easier to reverse, but they normally remain attached to the same storage system and depend on the same pool.
They do not automatically protect against:
- destruction of the storage system;
- theft, fire, or another site-level event;
- compromised administrative access;
- a mistake that deletes both live data and local snapshots;
- missing encryption or recovery information;
- an application state that was inconsistent when captured.
The phrase is familiar because it remains important:
Redundancy is not backup, and backup is not proven until restoration works.
The authoritative copy must be explicit
Shared storage becomes dangerous when several directories look like plausible sources.
A migration, temporary copy, old mount, synchronization job, and backup restore can leave behind multiple versions of the same data. Without a documented authority, a later repair may preserve the wrong copy or synchronize stale data over the correct one.
For every important dataset, I want to know:
| Question | Why it matters |
|---|---|
| Where is the authoritative copy? | Prevents two systems from accepting conflicting writes |
| Which applications may change it? | Defines ownership and limits damage |
| Which copies are replicas or backups? | Prevents a secondary copy from silently becoming production |
| How is freshness verified? | Avoids restoring or syncing stale data |
| What is the recovery order? | Ensures dependencies return in a usable sequence |
| What proves recovery succeeded? | Replaces “the files are visible” with an actual validation result |
This record does not need to expose private names publicly. Role-based labels such as storage host, application host, and backup target are enough to explain the architecture while the exact implementation remains in private documentation.
Performance belongs to the entire path
A fast pool cannot overcome every network, application, or client limit.
For media playback, the experienced result can depend on:
- drive and pool activity;
- network-filesystem behavior;
- the link between storage and compute;
- the application host’s CPU, memory, and local cache;
- container path configuration;
- the media format;
- whether the client can direct-play it;
- whether transcoding introduces another workload;
- the final client network connection.
A slow copy or buffering stream therefore needs evidence before storage receives the blame.
I work through the path in layers:
- Is the pool healthy?
- Is the storage system under abnormal load?
- Is the network path stable?
- Is the remote mount present and responsive?
- Is the application reading the expected file from the expected path?
- Is another operation, such as transcoding, actually causing the delay?
- Can a controlled test reproduce the problem outside the application?
Storage performance is not experienced at the drive. It is experienced at the end of the dependency chain.
Organization is part of recovery
Directory and dataset structure should reflect different operational needs.
Application configuration, databases, original files, replaceable media, downloads, backups, and temporary data do not deserve identical retention, permissions, snapshot schedules, or recovery priority.
Combining everything because space is available makes future migrations and restores harder. It also encourages backup policies that are either wasteful or incomplete.
A useful structure makes it easier to answer:
- What must be backed up independently?
- What can be regenerated?
- What needs an application-aware backup procedure?
- What should never be writable by a particular service?
- What can be restored later without blocking critical service recovery?
Organization does not replace documentation, but it gives the documentation a stable model to describe.
The storage contract I would define first
Before connecting the first application, I would now document a storage contract containing:
| Area | Decision |
|---|---|
| Authority | The system and dataset that own the live data |
| Purpose | What belongs there and what explicitly does not |
| Access | Clients, service identities, and read/write requirements |
| Mount behavior | How the client mounts it and what happens on failure |
| Service dependency | Which applications must wait for storage readiness |
| Integrity | Pool-health checks, scrubs, and alerting |
| Change history | Snapshot purpose and retention expectations |
| Backup | Independent destination and protected recovery information |
| Recovery | Restore order and validation procedure |
| Review trigger | Events that require the document to be updated |
Capacity planning would still matter, but it would follow the decisions that determine whether the data remains understandable, accessible, and recoverable.
Storage became much more useful once I stopped treating it as a number on a dashboard.
Related articles
- The Backup Readiness Review That Changed My Build Order
- Why Container Paths Have to Match Across the Media Stack
- The Definitive Guide to Building a Homelab from Scratch
- The Definitive Guide to Jellyfin on Ubuntu with Docker
Sources reviewed
Technical behavior was checked on July 28, 2026 against current primary documentation:
- Checksums and Their Use in ZFS — OpenZFS
- Scrub and Resilver — OpenZFS
- Snapshots, Clones and Bookmarks — OpenZFS
- Send and Receive — OpenZFS
Security note
This article describes the architecture using generic roles and example relationships. Exact hostnames, addresses, pool and dataset names, exported paths, mount points, user and group identifiers, permissions, backup destinations, retention values, and recovery locations remain in private documentation.
AI transparency
AI assisted with structure, copy editing, and checking current OpenZFS documentation. The storage relationships, mount-failure lessons, permission issues, troubleshooting approach, and conclusions come from operating and documenting my own homelab.