System administrators know that backup integrity is paramount. In a hard-linked file-level backup scheme, an older directory (yesterday’s backup) is expected to hold an “approved” content V1, even when the current directory (today’s backup) is allowed to change to a new content V2. If after an rsync update the file in “yesterday” has changed to V2, we must decide if the older generation still truly contains V1. The key is to trace the inode identity: if the two directory paths share the same inode, they are effectively the same file, and an update can propagate through all names. This playbook shows how to test which rsync mode preserves or breaks such sharing, and how to confirm whether each generation matches its intended manifest. We use a minimal local test: simple regular files, no symlinks, a single filesystem, and tools like GNU rsync, cp/stat, sha256sum on Linux (with rsync 3.4.1 observed). All commands are run by one user in a throwaway directory to avoid external interference. The goal is a clear decision on whether yesterday’s backup must be ACCEPTED as unchanged, REPAIRED by fixing the process, HELD pending outside sources, or RECOVERED from a known good copy.
Declare what an older backup generation must retain
A backup generation (say G1) is a directory of files that should each preserve their approved bytes (V1). A later/current directory (say G2) is allowed to update some files to new content V2. Crucially, G1 and G2 are two sets of hard-linked files on one filesystem (not a kernel snapshot). By Unix semantics, if a file in G1 and its counterpart in G2 share the same inode (same device and inode number), they are literally the same file. In that case, changing it in G2 also changes it in G1. Thus the contract is: each file in G1 must remain linked only to its own data V1. If G1 and G2 share inodes, G1 does not isolate the old bytes. As a result, evidence of yesterday’s state is only valid if G1’s names point to inodes that do not get overwritten.
Backup integrity is a routine admin responsibility. We denote “approved bytes” V1 (e.g. the configuration text at snapshot time) and the updated bytes V2. The older backup directory (e.g. previous/ or “yesterday/”) must continue to present exactly V1 for each file in its manifest. By contrast, the current directory (“today”) is allowed to change content to V2. However, just naming a directory “yesterday” does not guarantee its contents are safe; we must verify inode identity. The Linux inode manual reminds us that each file has a unique (device,inode) pair per filesystem, and a link count nlink telling how many names point to it. Two paths with identical device:inode and nlink>1 are merely aliases. We must treat an older backup directory as a potential alias unless proven otherwise. In short: a generation must retain its original data or it fails to meet the backup contract.
Create a disposable same-filesystem laboratory
To test this without risk, make a fresh temp directory and run everything there. For example:
ROOT=$(mktemp -d)
cd "$ROOT" # work in a new scratch dir
echo "Kernel: $(uname -sr)" > lab-env.txt
echo "Filesystem: $(df -T .)" >> lab-env.txt
rsync --version | head -1 >> lab-env.txt
cp --version | head -1 >> lab-env.txtThese commands log the kernel and filesystem type (all files must be on the same FS) and tool versions. Ensure the FS is not a network or pseudo-FS (hard-link scope is one filesystem). Next, verify hard-link support:
echo "test" > testfile
ln testfile linkfile || { echo "Hard links unsupported!" >&2; exit 1; }
stat -c "%d:%i %h %n" testfile linkfile > lab-env.txtThis writes the dev:inode and link count (%h) for the two names, confirming they refer to the same inode with nlink=2. If any check fails (no hard links, unexpected FS, missing tools), abort. Every step (source files, current and previous directories) will stay under this root. This separation ensures we “identify the actual filesystem path in use” and avoid contaminating production data.
Prove two paths name the same file
We must explicitly show that two filenames refer to the same inode. After linking, run stat with both device and inode. For example, if we’ve created current/config.txt and previous/config.txt as hard links, do:
stat -c "%d:%i %h %n" current/config.txt previous/config.txt
sha256sum current/config.txt previous/config.txtThe stat output will show the same device:inode for both and an nlink=2 (after linking), which proves they’re the same file. The SHA-256 hashes verify their bytes match, but note: matching hashes alone do NOT prove identity (two separate copies of V1 would have the same hash). Only the combination of device ID, inode number, and a link count >1 proves a shared inode. We use SHA-256 as a content check (“bytes equal”) and device:inode as identity. Keep this distinction clear throughout.
Build V1 and an independent expected manifest
Now prepare our initial content (V1) and expected digests. For example:
mkdir source
printf "ASCII mode=V1\n" > source/config.txt
echo "stable" > source/unchanged.txt
rsync -a source/ current/
cp -al current/ previous/At this point, current/ and previous/ are hard-linked copies of the source/ tree. We record the original V1 bytes outside both:
printf "ASCII mode=V1\n" > original_config.txt
sha256sum original_config.txt > expected_manifest.txt
sha256sum source/unchanged.txt >> expected_manifest.txtThe file expected_manifest.txt now holds the SHA-256 digest of the approved V1 contents. We deliberately keep original_config.txt outside the current/ and previous/ directories (“oracle” data) so it remains untouched during tests. This allows us later to distinguish real file recovery from mere hash checking. The unchanged.txt file (“sentinel”) will help detect if the entire link-tree got carried over correctly.
Update the current generation with --inplace
We first test rsync’s in-place update mode. Start from a fresh copy of the above state:
rm -rf current previous
rsync -a source/ current/
cp -al current/ previous/Now change only the source config to V2 (longer content forces detection):
printf "ASCII mode=V2\n" > source/config.txt
rsync -a --itemize-changes --inplace source/ current/ 2>&1 | tee inplace-update.log
echo "Exit code: $?"This runs rsync with the --inplace flag. According to the rsync documentation, this causes the destination file to be opened and updated directly. The observed behavior (exit code 0) should be:
current/config.txt is updated in place to V2.
previous/config.txt also changes, because it shares the same inode.
unchanged.txt in both dirs remains V1 (the sentinel is unaffected).
For example, after the update we might see in inplace-update.log an itemized change for config.txt. Check the new state:
stat -c "%d:%i %h %n" current/config.txt previous/config.txt
sha256sum current/config.txt previous/config.txtBoth lines will show the same device:inode (nlink stayed 2) and both SHA-256 digests now equal the V2 digest. This matches the documented caveat: if a file has extra hard links, updating it will break that linkage unless carefully managed. In short, --inplace caused the “yesterday” generation to unexpectedly hold V2, violating our contract. (The unchanged.txt lines would show identical digests V1 in both dirs, nlink=2, as expected.)
Repeat with ordinary replacement writes
Next, rerun the scenario without --inplace so rsync uses its default behavior (new-temp-then-rename):
rm -rf current2 previous2
rsync -a source/ current2/
cp -al current2/ previous2/
echo "ASCII mode=V2\n" > source/config.txt
rsync -a --itemize-changes source/ current2/ 2>&1 | tee normal-update.log
echo "Exit code: $?"Now rsync -a (no --inplace) replaces files by renaming temporary files. The observed outcome is:
current2/config.txt is replaced with V2, but because the original current2/config.txt was removed first, its old inode is freed.
previous2/config.txt (the linked copy) is untouched and still holds V1.
The two names no longer share an inode: current2/config.txt got a new inode, previous2/config.txt kept the original inode (now nlink=1 for each).
unchanged.txt remains V1 in both with nlink=2.
Verify:
stat -c "%d:%i %h %n" current2/config.txt previous2/config.txt
sha256sum current2/config.txt previous2/config.txtcurrent2/config.txt will have a different inode (and nlink=1), and its digest equals V2. previous2/config.txt shows the original inode (nlink=1) and digest V1. This is the desired behavior: the old generation kept V1 independent of the update. The link between them was “split” by the replacement write. In other words, a non-inplace update correctly preserved the older bytes.
Separate source-link preservation from historical isolation
It’s important to understand what rsync -a (archive) does and does not do. The -a option implicitly includes -H, --hard-links, meaning that if multiple files in the source tree are hard-linked, the destination will preserve those links. However, this is strictly about linking files within the transfer, not about guarding previous backups. The rsync docs emphasize that -H makes source-linked files link at the destination, but it does not go back and unlink or protect any preexisting files outside that set. In effect, -H is about preserving content relationships in one sync run, not enforcing immutability of older snapshots.
Likewise, --inplace is meant to avoid using a temp file, not to preserve history. In fact, the rsync manual warns that if a file has “extra hard-link connections to files outside the transfer” (like our previous generation), updating it will break that linkage. The manual explicitly cautions: using --inplace will update all names of a shared inode, so it’s dangerous if you have multiple aliases.
For building an isolated new snapshot, the proper rsync tool is --link-dest. This option tells rsync: “Copy all files to a new directory, but if a file is unchanged, hard-link it from a given reference directory instead of sending it.” Unlike --inplace, --link-dest leaves the reference (old snapshot) untouched. From the manual:
“--link-dest=DIR ... unchanged files are hard linked from DIR to the destination directory.”
That is how many backup scripts create incremental generations: each rsync --link-dest=prev backup makes a new tree of hard links pointing back to the previous tree, and never mutates the old one. In summary: -H only preserves links for the current transfer, --inplace updates in place (affecting all aliases), and --link-dest creates a new tree of links. We will not attempt to repair an old generation by re-running rsync with --link-dest on it; instead, we create a completely new generation if needed.
Reconcile every generation against the declared bytes
We summarize the evidence in this table. For each case (inplace update vs ordinary update) and each file in each generation, we list the device:inode, link count, and SHA-256 before/after the update, plus the expected content and final verdict. Exit codes 0 indicate the rsync ran “successfully” even if it broke our backup contract.
Case | Generation | Path | Device:Inode | Nlinks | Digest (before → after) | Expected | Exit | Verdict |
inplace | current | current/config.txt | 2050:100 | 2 | abcdef1… → 1234567… | V2 bytes | 0 | PASS |
inplace | previous | previous/config.txt | 2050:100 | 2 | abcdef1… → 1234567… | V1 bytes | FAIL | |
inplace | current | current/unchanged.txt | 2050:101 | 2 | 7654321… → 7654321… | same | PASS | |
inplace | previous | previous/unchanged.txt | 2050:101 | 2 | 7654321… → 7654321… | same | PASS | |
ordinary | current | current/config.txt | 2050:200 | 1 | abcdef1… → 1234567… | V2 bytes | 0 | PASS |
ordinary | previous | previous/config.txt | 2050:100 | 1 | abcdef1… → abcdef1… | V1 bytes | PASS | |
ordinary | current | current/unchanged.txt | 2050:101 | 2 | 7654321… → 7654321… | same | PASS | |
ordinary | previous | previous/unchanged.txt | 2050:101 | 2 | 7654321… → 7654321… | same | PASS |
(Table: “inplace” = rsync with --inplace; “ordinary” = rsync without it. Digests are truncated examples. “Expected” shows which content we intended in each generation.)
This table exposes the critical difference. In the inplace case, the previous/config.txt ended up with V2 (digest 1234567…) instead of V1, so it failed the manifest check. The unchanged.txt files stayed as expected. In the ordinary case, previous/config.txt kept V1 (abcdef1…), so it matches expectation. Note that in the inplace case, current and previous share an inode (2050:100) even after the update (nlink=2), whereas in the ordinary case they split (different inodes, nlink=1). The exit code 0 in both cases indicates rsync itself saw no I/O errors, but that alone does not ensure backup integrity.
Make assertions fail on changed historical content
We can automate these checks. For example, after the inplace update we could run:
# Compare previous generation file to the original V1
expected=$(sha256sum original_config.txt | cut -d' ' -f1)
prev_hash=$(sha256sum previous/config.txt | cut -d' ' -f1)
if [ "$prev_hash" != "$expected" ]; then
echo "ERROR: previous generation has unexpected content!" >&2
fiThis assertion fails because previous/config.txt holds V2, not the expected V1. In contrast, doing the same after the ordinary update would pass (prev_hash matches expected). We should always compare against our independent “oracle” (original_config.txt) or the expected manifest, not against current/config.txt, since that could be corrupted too. Additional assertions include checking the link count and device:inode:
# Ensure current and previous are not erroneously linked after update
id_current=$(stat -c "%d:%i" current/config.txt)
id_previous=$(stat -c "%d:%i" previous/config.txt)
if [ "$id_current" = "$id_previous" ]; then
echo "ALERT: current and previous are still linked!" >&2
fiIf this triggers, we know a shared inode remains. Similarly, confirm all files are present:
# Check that the unchanged sentinel is still in both trees
[ -e current/unchanged.txt ] && [ -e previous/unchanged.txt ] || \
echo "ERROR: sentinel file missing!"Each check that fails should raise an alarm; that is exactly our goal: detecting when history diverged.
Construct a new generation without rewriting the old one
If the update process is deemed unsafe (as with --inplace), we need a repair plan: build a fresh backup generation from scratch. The safest approach is to rsync the approved inputs into a new empty directory, using the reviewed (non-in-place) method. For example:
rsync -a source/ new_current/This “new_current/” now contains V2 of config.txt (and all other approved files), and we never touched previous/. We can optionally use --link-dest on an unchanged file to save space, but again never run any option that modifies previous/. Next, validate the manifest of new_current/ against expected values:
new_hash=$(sha256sum new_current/config.txt | cut -d' ' -f1)
[ "$new_hash" = "1234567…" ] || echo "ERROR: new backup content incorrect"Assuming it checks out, we then atomically publish new_current/ as the updated backup (for example, swapping symlinks or renaming directories outside this lab). All the work of atomic publication (locking, moving, cleaning) lies outside this playbook. The key point: we built a proper new snapshot and validated it.
Do not call a replacement write a complete backup transaction
It’s important not to oversell this solution. Rewriting files at the file level is not a full backup system. We have not addressed multi-file consistency, writer quiescence, or crash-safety across the directory. For instance, if source/ were a running database, a single rsync -a would not guarantee a consistent snapshot of all tables. We also haven’t covered preserving ownership, ACLs, xattrs, or any backup retention policy. In practice, a robust backup involves these concerns too. In our context, we explicitly limit scope: we assume small ordinary files with no concurrent changes. Beyond that, the replacement write we performed should be viewed as a fix for the script’s behavior, not a silver-bullet “atomic backup”.
Check metadata and other writers separately
Our tests focused on file content. Metadata and permissions could also leak changes across hard links. For example, if we ran:
chmod 444 previous/config.txt
ls -l current/config.txt previous/config.txtwe would see that both names now appear read-only. This is because mode (and owner, mtime, etc.) are stored in the inode and shared by all links. Making one alias read-only does not freeze the underlying data against writes by root or by using the other alias. Indeed, any process with write permission (e.g. root) could still open and write the file via current/config.txt. We emphasize: metadata changes propagate through links just like data changes. There is no safe “chmod trick” to protect an old backup: one must rely on separate inodes instead.
Contain damage without overwriting the remaining evidence
Once a problematic update is discovered, stop the offending process immediately. Preserve all logs and evidence in place. For instance, do not overwrite or delete the “previous” directory even if it now contains wrong data. Do not attempt an immediate in-place restore on those same files. Instead, rename the old directories, copy logs, and flag that generation as suspect.
Disable or roll back the script/cron job that ran rsync.
Save its command line and stderr/stdout logs. (For example, use a pipeline that collects "$@" from cron or redirect output to files.) Make sure to preserve the exit code and output.
Do not delete any generation or overwrite any file yet. Inventory what was touched.
Mark the timeline: note which generation(s) are in doubt and which files changed unexpectedly.
At this point, our goal is to preserve the state for forensic examination. Do not run a new restore on the same directory structure or use a hard link path to restore bytes, as that could overwrite previous/. Also, do not “clean up” the backup tree by deleting mismatches or placing dummy files; that could cover up the problem. We are still in damage control mode.
Recognize when old bytes are no longer recoverable here
If inspection shows that all hard-linked names now contain V2 and no standalone copy of V1 exists (other than hashes), we cannot reconstruct the lost bytes from within this filesystem. For example, if rsync --inplace was used when the link count was 2, both links now show V2. In that case, we must HOLD. No amount of rsync flags on this broken tree will bring back V1. We must alert stakeholders that the older bytes are gone and rely on a trusted outside copy (if any exists). Remember: a hash alone cannot recreate lost content. The decision at this point is beyond a local script: involve data owners, look for other backups, and document the loss.
Recover into an independent destination and retest
If we have preserved an independent copy of V1 (our original_config.txt or any other source of truth), we can demonstrate actual recovery. Create a fresh directory (completely outside the current/previous trees) and populate it using that independent copy:
mkdir recover
cat original_config.txt > recover/config.txtNow recover/config.txt is a new file with V1, unlinking it from any old inode. Check:
stat -c "%d:%i %h %n" recover/config.txt
sha256sum recover/config.txtThis should show nlink=1 and the digest matching our expected V1 (e.g. abcdef1…). If so, we have proved that V1 can be restored given an independent source. The recovered config.txt is isolated from any other file, fulfilling the contract for a restored generation. We also test a couple of edge cases:
If some files had remained unchanged across generations, they would appear in both current and previous links. In recovery, you could optionally hard-link them in the new tree to save space, but this is purely an optimization.
If we intentionally use a mismatched “manifest” (say we edited original_config.txt differently and it doesn’t match expected_manifest.txt), our assertion above would catch it. In that scenario, recovery reveals the inconsistency and triggers a HOLD, because the recovered content doesn’t match the documented history.
In summary, recovering to a new location and re-running the validation proves we can restore the old bytes when we have them. It also underscores that this is separate from the forward-update workflow: fixing the backups (recovering V1) does not itself fix the broken update script.
Choose accept, repair, hold or recover
Based on the tests above, we reach one of four outcomes:
ACCEPT: If every file in the previous generation matches its V1 manifest and the current generation was updated in an approved (tested) way, we can accept. That means the historical bytes are intact under a verified workflow. Future updates can proceed (possibly after minor fixes to logging or command flags).
REPAIR: If the historical backup is out of spec (as with --inplace) but we have a reliable way to recover V1 (e.g. from source or an external archive), then we repair the process. This means fixing the rsync options (e.g. remove --inplace or use --link-dest), rerunning the updated backup into a new directory, and then continuing. The bad generation is still suspect, but we can replace it.
HOLD: If the integrity is breached and we cannot authenticate or recover the lost data, we must hold and escalate. No changes are applied. For example, if we find previous/config.txt has the wrong bytes and we have no independent V1, we must treat the old generation as compromised. The fix is to stop, investigate, and possibly restore from a different backup system.
RECOVER: If the old bytes are lost locally but we do possess a trusted independent copy of V1, we perform a recovery. We restore the affected files into a new directory (as above) and substitute that directory as the accepted “previous” generation. The new generation is then validated.
A key point: we should never rely on a dry-run or summary output as proof of preservation. A dry-run (or --itemize-changes) is merely diagnostic. Actual validation requires checking content and identity after the update. Only when files truly survived intact or were successfully reconstructed should we tick the “pass” box.
Separate repair permission from restore permission
Finally, note that correcting the backup script is not the same as authorizing data changes. The person who owns the backup script (or runs rsync) should not unilaterally decide to delete or overwrite data without oversight. Ideally, define roles: one person or team owns the backup infrastructure, another owns the data. The backup admin can fix the rsync flags, but the data owner must approve discarding or restoring any historical data. In practice, this means even if we decide to “repair” or “recover,” the step to remove the suspect generation or write new bytes should only be taken with explicit agreement.
Assign ownership and regression triggers
After an incident, document the status of each backup generation. Keep logs outside the backup directories (for example, in a separate /var/log/backup-validation/). Include: versions of rsync and cp used, filesystem type, and the exact command lines that were run (record argv). For each generation, note its verdict (e.g. “Generation 2026-09-30: MATCH” or “SUSPECT”). List affected paths and whether an independent recovery source is available. Finally, specify the next authorized action (Accept, Repair, Hold, or Recover) and who may execute it.
Plan to retire old evidence after this incident to avoid confusion. Set triggers for future testing: any change in the rsync version, update flags, backup script, filesystem type, or even a modification to the source data writing process should prompt a full re-validation. A good practice is to automate this validation as part of any change management (e.g. after a system update or a config change). Maintain regression-test fixtures outside the live backup path (our original_config.txt is an example).
In short, hand off a concise summary: “Backup dated 2026-09-30: old file changed unexpectedly during update. Hard link update via --inplace caused the issue. Independent copy of V1 is available in original_config.txt. Next: rebuild backup with rsync -a (no --inplace), then deprecate the broken version. Script has been corrected.”
Build the administration foundations behind recovery evidence
This exercise depends on low-level filesystem knowledge and command-line proficiency, exactly the kind of skills covered in Refonte Learning’s System Administration Program. That certificate course (6 months, ~10–12 hr/week) teaches Windows and Linux system fundamentals, including command-line work, backup/recovery practices, virtualization, troubleshooting, and hands-on projects. It assumes learners have basic computer networking knowledge and are working toward a bachelor’s degree or higher. Success grants a Training Certificate and an optional internship (outcomes depend on meeting the program’s criteria).
Understanding inodes, hard links, and tools like stat and sha256sum is part of that curriculum. The rigorous validation steps here illustrate the kind of evidence-based administration the program advocates. We encourage interested readers to consult the current System Administration Program details for curriculum and prerequisites.
In conclusion, accept only the generation proven to hold its approved bytes. If our tests show the older backup is still intact under a reviewed workflow, we accept it. If not, we repair our process or hold the data, recovering from an independent source if available. This disciplined approach, rooted in filesystem mechanics and actual observations, is the surest way to avoid silent data loss in hard-linked backup schemes.
