Platform engineer reviewing Kubernetes StatefulSet PVC retention settings and storage bindings on a computer screen.

Before Scaling a StatefulSet Down, Prove Which PVCs Survive

Thu, Oct 8, 2026

Consider a scenario where a Kubernetes StatefulSet is deliberately scaled down from 2 replicas to 1 and then back up to 2. After the upscale, the pod at ordinal-1 returns, but the data from the original volume is missing. The key question is which PersistentVolumeClaim (PVC) and underlying PersistentVolume (PV) that replacement pod is actually using. Did we reattach the original storage (with its UID and data), or did the system provision a fresh volume under the same name?

The procedure compares two StatefulSet configurations. In both, the StatefulSet starts with 2 replicas and uses the same dynamic storage class; the only difference is the whenScaled retention policy. One arm (the keep arm) uses whenScaled: Retain, the other (the drop arm) uses whenScaled: Delete, while both use whenDeleted: Retain. We keep each actual PV’s persistentVolumeReclaimPolicy set to Retain so that deleting a PVC does not automatically delete the backend volume. We prepare unique marker bytes on each pod’s volume before any scale-down.

From these observations, there are three possible verdicts after the test:

  •     ACCEPT_ORIGINAL_REATTACHMENT: The original PVC/PV (with the same UID) and marker bytes reappear on upscale. Original storage returned.

  •     CLASSIFY_NEW_ALLOCATION: The scaled-up pod gets a new PVC with the same name but a new UID, bound to a new PV, and the original marker is absent from the new volume. Only new allocation was returned.

  •     HOLD/INVESTIGATE: Evidence is incomplete or contradictory (e.g. missing reads or mismatched identities). The behavior cannot be confirmed.

We do not rely on success of the kubectl scale command alone. Instead, we examine the evidence: controller configuration, API UIDs, the PVC→PV binding (via spec.volumeName and claimRef UIDs), CSI volume identifiers if present, and the exact marker file contents. This synthetic test uses no real application data, only a disposable namespace and approved dynamic storage provisioner. The outcomes below are documented expectations and illustrative states, not recorded cluster results. An operator must record the actual cluster version, storage driver, and observations when running the procedure. No live cluster test was performed for this article.

Separate the replica count from the storage claim

When reviewing a proposed StatefulSet scale-down, treat the pod count change separately from the storage retention outcome. An ordinary StatefulSet Pod replacement, while the replica count remains unchanged, should reuse its existing PVC and storage. However, a replica count reduction may delete PVCs (if policy says so) rather than just detaching them. We must answer three questions independently: (1) Which Pod appears for ordinal-1 after the change? (2) Which historical PVC and PV are bound (by name and UID)? (3) Which data bytes are present on that volume?

Maintaining an unchanged ordinal-0 pod provides a stable control: ordinal-0’s PVC and data should remain intact in both arms. New pods replacing existing ones (due to failure or manual deletion) should reattach the same PVC/PV by default, but scaling down reduces replica count and may delete PVCs depending on policy. Don’t confuse this storage question with availability decisions like PodDisruptionBudgets. For example, node-drain acceptance focuses on whether service uptime met the contract, not whether data was retained. Here we focus strictly on which storage is attached to the pod after scale changes, which is a different control objective.

Define the three identities that names cannot prove

Kubernetes object names alone are not unique identifiers over time. We identify objects by name plus their UID (and namespace, for namespaced resources). For example, a PVC named data-keep-1 in namespace pvc-retention-xyz is not the same as a later PVC with the same name; each has a distinct .metadata.uid. Thus we record PVC name + namespace + UID, and for PVs (which are cluster-scoped) we use PV name + UID.

To verify the binding, we use the PVC’s .spec.volumeName (which names the PV) and the PV’s .spec.claimRef (which includes the bound claim’s name, namespace, and UID). Together these confirm which PVC is bound to which PV at each stage. If the CSI driver provides a persistent volumeHandle or similar locator, we log it as additional evidence of the physical volume. Absence of a stable driver identifier (e.g. in some in-tree plugins) means we cannot peer into the storage backend, so we stick to Kubernetes IDs and file contents only.

Freeze the two retention decisions before testing

We explicitly set both retention policies for clarity, separating them from the PV reclaim policy. In apps/v1, the persistentVolumeClaimRetentionPolicy field has two subfields: whenDeleted (applies when the whole StatefulSet is deleted) and whenScaled (applies when its replica count is reduced). Each can be Retain or Delete. By default, both fields default to Retain. We configure whenDeleted: Retain for both arms so that deleting the StatefulSet itself will not remove PVCs (this keeps deletion-of-workload logic out of our comparison). The only difference between the two StatefulSets is whenScaled: the “keep” arm uses Retain, the “drop” arm uses Delete. In all cases we use a StorageClass with reclaimPolicy: Retain so that deleting a PVC moves its PV to Released rather than deleting the volume. (Dynamically provisioned PVs inherit their StorageClass’s reclaim policy, which defaults to Delete, so we confirm the class explicitly has Retain.)

Read the effective policy instead of trusting a filename

Before proceeding, we verify the cluster’s effective configuration. After applying our manifest, we kubectl get statefulset on each to inspect .spec.persistentVolumeClaimRetentionPolicy. We also fetch the StorageClass JSON and each actual PV JSON. We record the server version and CSI provisioner version. If any policy was not applied as intended (for example, if an unsupported or disabled retention feature prevents the fields from taking effect), we abort the test. Similarly, if any PV’s persistentVolumeReclaimPolicy is not Retain, we cannot run the downscale test since a deleted PVC might auto-delete the volume. (Our test environment is a dedicated disposable namespace with no other pods or prebound claims, and we assume sufficient quota and topology. We do not use nodeName in pods, avoiding binding interference.) A mismatch between intended and actual policy is a test-preparation failure.

Build two isolated StatefulSets with one changed field

We create two identical StatefulSets (and matching headless Services for DNS stability) named keep and drop in a new namespace. The only difference is whenScaled. Below is the setup script. (All variables like KCTX (kubectl context), SC (approved Retain StorageClass name), and FIXTURE_IMAGE (digest-pinned pod image) must be provided by the operator.) The script generates the namespace and two StatefulSets with 2 replicas each, each with a single volumeClaimTemplate named data. Both sets have whenDeleted: Retain and differ in whenScaled: Retain vs Delete. We then apply the combined JSON with kubectl apply.

set -euo pipefail
: "${KCTX:?set authorized context}" "${SC:?set approved Retain class}"
: "${FIXTURE_IMAGE:?set approved Python workload digest}"
export RUN="$(python3 -c 'import uuid; print(uuid.uuid4().hex[:12])')"
export NS="pvc-retention-${RUN}"
export KCTX SC FIXTURE_IMAGE
mkdir "evidence-${RUN}"
cd "evidence-${RUN}"
k() { kubectl --context "$KCTX" --namespace "$NS" "$@"; }
existing="$(kubectl --context "$KCTX" get namespace "$NS" --ignore-not-found -o name)"
test -z "$existing" || { printf '%s\n' 'namespace already exists' >&2; exit 1; }
kubectl --context "$KCTX" version -o json > versions.json
kubectl --context "$KCTX" get storageclass "$SC" -o json > storageclass.json
import json, os
ns, run, sc, workload = (os.environ[k] for k in
                        ("NS", "RUN", "SC", "FIXTURE_IMAGE"))
if "@sha256:" not in workload:
    raise SystemExit("An approved digest-pinned workload is required")
items = [{"apiVersion": "v1", "kind": "Namespace",
          "metadata": {"name": ns, "labels": {"retention-run": run}}}]
for arm, policy in (("keep", "Retain"), ("drop", "Delete")):
    labels = {"retention-run": run, "retention-arm": arm}
    meta = {"name": arm, "namespace": ns, "labels": labels}
    # Headless Service for stable DNS identity
    items.append({"apiVersion": "v1", "kind": "Service", "metadata": meta,
                  "spec": {"clusterIP": "None", "selector": labels,
                           "ports": [{"port": 80, "name": "unused"}]}})
    # StatefulSet
    items.append({"apiVersion": "apps/v1", "kind": "StatefulSet",
      "metadata": meta, "spec": {
        "serviceName": arm, "replicas": 2,
        "persistentVolumeClaimRetentionPolicy": {
            "whenDeleted": "Retain", "whenScaled": policy},
        "selector": {"matchLabels": labels},
        "template": {"metadata": {"labels": labels}, "spec": {
          "automountServiceAccountToken": False,
          "terminationGracePeriodSeconds": 10,
          "securityContext": {"runAsNonRoot": True, "runAsUser": 1000,
                              "runAsGroup": 1000, "fsGroup": 1000,
                              "seccompProfile": {"type": "RuntimeDefault"}},
          "containers": [{"name": "holder", "image": workload,
            "command": ["python3", "-c",
              "import signal, sys; "
              "signal.signal(signal.SIGTERM, "
              "lambda signum, frame: sys.exit(0)); signal.pause()"],
            "securityContext": {"allowPrivilegeEscalation": False,
                                "capabilities": {"drop": ["ALL"]}},
            "resources": {"requests": {"cpu": "10m", "memory": "32Mi"},
                          "limits": {"memory": "64Mi"}},
            "volumeMounts": [{"name": "data", "mountPath": "/data"}]}]}},
        "volumeClaimTemplates": [{"metadata": {"name": "data", "labels": labels},
          "spec": {"accessModes": ["ReadWriteOnce"], "volumeMode": "Filesystem",
                   "storageClassName": sc,
                   "resources": {"requests": {"storage": "1Gi"}}}}]}})
with open("fixture.json", "x") as f:
    json.dump({"apiVersion": "v1", "kind": "List", "items": items}, f, indent=2)
Save the generator to manifest.py, run python3 manifest.py, and then kubectl --context "$KCTX" apply -f fixture.json. The JSON List contains our two StatefulSets (keep and drop) clearly showing whenScaled: Retain vs Delete. We have also created headless Services, but they are not used by an application; they simply anchor the StatefulSet’s .spec.serviceName. No application containers produce data yet; they will wait (due to signal.pause() in the command) until we seed /data/marker.bin by exec.
Establish a binding and marker baseline
Once applied, Kubernetes will create PVCs and PVs. For this example, verify that the approved StorageClass uses volumeBindingMode: WaitForFirstConsumer. This delayed volume binding postpones provisioning and binding until a Pod using the claim is scheduled. We wait (with a timeout) for each StatefulSet to have 2 Ready pods (keep-0, keep-1; drop-0, drop-1) and for each of the 4 PVCs to reach Bound. We then verify each bound PV’s .spec.persistentVolumeReclaimPolicy is indeed Retain (as inherited from the class). This prevents PVC deletion from automatically requesting backend deletion through the PV reclaim policy. We also confirm the attached volume of each pod is writable: no nodeName or other binding overrides are set, and there is no contention (ReadWriteOnce limits read/write mounting to one node and can permit multiple Pods on that node; this fixture assigns each claim to one Pod).
Save the expected bytes before any storage mutation
Before any downscale, we write unique marker data to each pod’s volume. This establishes a ground-truth “original” file content to compare later. We define the content for each (arm, ordinal) as <RUN>|<arm>|<ordinal>|original\n in UTF-8. First, we compute locally and save their base64 and SHA-256 hashes in baseline-expected.json:
import base64, hashlib, json, os, subprocess
expected = {}
for arm in ("keep", "drop"):
    for ordinal in (0, 1):
        pod = f"{arm}-{ordinal}"
        raw = f"{os.environ['RUN']}|{arm}|{ordinal}|original\n".encode()
        expected[pod] = {"base64": base64.b64encode(raw).decode(),
                         "sha256": hashlib.sha256(raw).hexdigest()}
with open("baseline-expected.json", "x") as f:
    json.dump(expected, f, indent=2)
Then we execute a Python snippet inside each container to write the marker file and verify it. We use kubectl exec from the host to avoid shell quoting issues:
seed = """import base64, os, sys
raw = base64.b64decode(sys.argv[1], validate=True)
with open('/data/marker.bin', 'xb') as f:
    f.write(raw); f.flush(); os.fsync(f.fileno())
with open('/data/marker.bin', 'rb') as f:
    actual = f.read()
if actual != raw:
    raise SystemExit('marker readback mismatch')
"""
for pod, value in expected.items():
    subprocess.run(["kubectl", "--context", os.environ["KCTX"],
        "--namespace", os.environ["NS"], "exec", pod, "-c", "holder", "--",
        "python3", "-c", seed, value["base64"]], check=True)

This writes marker.bin in each pod’s /data directory. The exclusive-open 'xb' ensures we do not overwrite if rerunning (if something went wrong). After writing, the code reads back the file and errors out if it does not match. In effect, we now have four marker files with known checksums: keep-0, keep-1, drop-0, drop-1. We record these baseline hashes for later comparison.

By the end of this stage we have 4 pods (2 in each arm) running, 4 PVCs Bound to 4 Retain-PV volumes, and each has a unique marker file. Ready status alone is not sufficient evidence of storage correctness: we have demonstrated storage read/write validation and recorded the content. This baseline will allow us to spot any changes after scaling.

Record a complete identity ledger

We now capture all object identities and states before changing replica count. For each PVC (names are data-keep-0, data-keep-1, data-drop-0, data-drop-1), we note: its UID, .status.phase, and .spec.volumeName (the PV bound). Then for each PV, we note: name, UID, phase, claimRef (including bound PVC UID if any), persistentVolumeReclaimPolicy, and if possible the CSI volume handle or other driver identifier. We also note each StatefulSet’s .metadata.uid, .status.observedGeneration, and each pod’s UID. We log ownerReferences and finalizers for PVC/PV, as well as any deletionTimestamp.

Practically, this means running kubectl get -o json on each namespace object (StatefulSets, Pods, PVCs) and cluster objects (the referenced PVs, the StorageClass) and saving them with timestamps. We keep stderr and nonzero exits. For example:

k get statefulset/keep -o json > keep-statefulset.json
k get statefulset/drop -o json > drop-statefulset.json
k get pod -l "retention-run=$RUN" -o json > pods.json
k get pvc -l "retention-run=$RUN" -o json > pvcs.json
python3 - <<'PY'
import json
with open("pvcs.json") as f:
    claims = json.load(f)["items"]
names = sorted({p["spec"]["volumeName"] for p in claims})
if len(claims) != 4 or len(names) != 4 or not all(names):
    raise SystemExit("Expected four bound claims and distinct PVs")
with open("baseline-pv-names.txt", "x") as f:
    f.write("".join(name + "\n" for name in names))
PY
mkdir -p pvs
while IFS= read -r pv; do
    kubectl --context "$KCTX" get pv "$pv" -o json > "pvs/${pv}.json"
done < baseline-pv-names.txt

Filter the namespaced fixture objects by the unique retention-run label, following any pagination to ensure completeness. Fetch each PV by the name in its PVC’s .spec.volumeName, because PVC labels need not be copied to PVs. The StorageClass JSON was captured earlier. Preserve all raw JSON, and use separate snapshot directories for later stages so the baseline is not overwritten. Continue querying the original PV names after a claim disappears, and also capture newly bound PVs after scale-up.

Our ledger records columns like: arm, ordinal, StatefulSet UID, Pod UID, PVC name/UID, PV name/UID, persistentVolumeReclaimPolicy, claimRef UID, marker SHA256, etc. These values let us later join policy and content data. Any missing piece (e.g. we cannot query the filesystem except via logs) is noted. We do not infer anything beyond what the API shows: for example, labels on a PVC do not necessarily appear on its PV.

Note: this step is about collecting evidence, not changing anything. We ensure we have proof of the original state of IDs and bytes, akin to freezing a baseline. This follows the same discipline as verifying a known-good state in Kubernetes user-namespace storage validation. A pod’s readiness doesn’t prove it can read our marker file; by actually writing and hashing, we have definitive evidence of the mounted storage content.

Replace a Pod without changing the replica count

As a negative control, we perform an ordinary pod replacement (without altering the replica count) in each arm. This simulates a failure-and-restart situation, which should not trigger PVC deletion under either policy.

For each arm (keep and drop), do:

for arm in keep drop; do
    k delete pod "${arm}-1" --wait=true
done

This deletes the -1 pod. We expect the StatefulSet controller to create a new pod -1 with a different Pod UID. We wait until the old pod’s UID disappears and a ready pod with the same name appears.

We then check the PVCs and marker files. Expected outcome: In both arms, the original PVC UID and PV UID for ordinal-1 should remain unchanged, and the marker.bin content should still match the original SHA256. This is because scaling was still 2, so these pods were merely replaced, not removed by downscaling. If instead we saw a new PVC UID or altered data, that indicates an unexpected change and we would stop the experiment to inspect.

According to Kubernetes documentation, deleting a pod that’s part of a StatefulSet will not delete its PVC (and thus will reuse the storage). Indeed, our negative control should confirm that behavior: each replacement pod re-mounts the same data volume and reads the old marker. If any pod replacement causes a change to its PVC/PV or data, that suggests something is misconfigured (in which case we cannot proceed to interpret downscaling). This step ensures that only a replica count change (not a pod restart) can trigger whenScaled behavior.

Scale down and observe what actually disappears

With the baseline and negative control satisfied, we now scale down each StatefulSet from 2 to 1 replica:

k scale statefulset keep --replicas=1
k scale statefulset drop --replicas=1

We record the responses (usually a success message) but more importantly we watch the cluster state. We expect the controller to observe the generation change and terminate one pod in each set (keep-1 and drop-1). We ensure ordinal-0 pods stay Ready.

For the keep arm (whenScaled: Retain): the original ordinal-1 pod (keep-1) is terminated, but its PVC and PV should remain intact and bound. The PVC (data-keep-1) should still exist with the same UID and volumeName, and the PV remains Bound. This means the “keep” arm has effectively done nothing to its storage during scale-down.

For the drop arm (whenScaled: Delete): the original ordinal-1 pod (drop-1) is terminated, and its PVC is deleted through the controller’s owner references and garbage collection. We verify that drop-1 pod is gone. The PVC object data-drop-1 should no longer appear, and the bound PV should transition to Released status (since persistentVolumeReclaimPolicy=Retain). The spec.claimRef on that PV should still refer to the old claim (by name and UID). We do not delete or reuse the PV; we simply note that it is Released with its data still on disk, requiring manual cleanup by the storage owner.

We poll the API (for up to, say, 300 seconds) to catch the final state, checking Pod, PVC, and PV status with timestamps. We do not forcibly remove any resources or remove finalizers. If the drop PVC goes to Terminating and sticks, the finalizer could still be blocking. We wait until it either vanishes or times out the check. If after our local timeout it still exists, we declare a HOLD and preserve the state. We deliberately do not push a deadline from theory; finalizers might delay deletion, and a deletion request does not establish when finalizer processing will finish.

We tie this observation to StatefulSet retention and garbage collection. Finished-Job cleanup uses the separate ttlSecondsAfterFinished field on Jobs, including Jobs created by CronJobs; it does not govern this StatefulSet comparison. After scale-down, we should see in the ledger data that keep has both original PVC/PV still Bound, whereas drop has the PVC gone and the old PV in Released. This confirms that under whenScaled: Delete, the system dropped the drop-1 claim as documented.

Classify deletion and release without guessing

At this point, the evidence is in the API: the data-drop-1 PVC should be absent. We explicitly check it with k get pvc data-drop-1 --ignore-not-found -o name. A successful run with no output (exit code 0) indicates a true NotFound (PVC deleted). Any other error (Forbidden, etc.) is treated as UNKNOWN and escalates to HOLD. Note: if the PVC has a deletionTimestamp but is Terminating, it is not yet gone; we treat that as still present (so it should appear with a timestamp in k get pvc data-drop-1 -o json).

For PVs, a key fact is that retaining the backend is not an automatic “cleanup success.” Even if a PV is in Released state, the data still sits on disk. We keep the old PV object as evidence: it retains its UID and claimRef (pointing to the old PVC UID). We do not clear claimRef: doing so can make a retained PV available for rebinding and requires a separate recovery decision. We also do not delete the PV, because removing its Kubernetes object does not prove that the retained backend data was erased. The presence of a Released PV with old data means we have work to do if we need that storage back. Neither finalizer processing nor Released status proves that backend data has been erased. Released means the volume is no longer bound to a live claim, and manual steps would be required to clean or rebind it.

Thus, we record: Keep-arm still has 2 PVCs/PVs, Drop-arm now has 1 PVC/PV bound plus 1 PV in Released. We have raw API dumps for these observations. We explicitly avoid assumptions: no PV is automatically Available, and Release did not erase data. We treat any failure to observe expected deletions (for example, data-drop-1 still listed as Terminating) as HOLD, preserving logs. We do not delete any finalizers to force completion; we accept the cluster’s natural behavior.

Scale up and test original storage versus new allocation

Now we restore the replica counts to 2 for both StatefulSets:

k scale statefulset keep --replicas=2
k scale statefulset drop --replicas=2

We again wait for new pods and updated bindings. Each set will create a new pod at ordinal-1. We check Pod UIDs to ensure they are fresh. Then we analyze storage:

·    Keep arm: With the original PVC still present, the new keep-1 pod should reattach the same PVC and PV. We expect the exact same PVC UID to reappear bound, and the same PV (UID, volumeName) with the old marker bytes. We kubectl exec into keep-1 and read /data/marker.bin. It should match the original hash from baseline. If for some reason it does not (e.g. we see different bytes), something went wrong; that would go to INVESTIGATE. Assuming all is well, we record the original PV UID, PVC UID, and content hash as in our baseline. This corresponds to “original storage reattached,” a PASS for this arm.

  •     Drop arm: The original drop-1 PVC was deleted. Thus the StatefulSet controller recreates the PVC with the same claim name and a new UID. With the original PV left Released and no other matching Available PV, dynamic provisioning should supply a fresh PV for the new drop-1 pod. We expect a new PVC UID and a different PV name+UID. We confirm data-drop-1 now exists again but with a new UID. Its .spec.volumeName refers to a new PV object (not the Released one). We check that this new PV’s claimRef points to the new PVC UID. Importantly, reading /data/marker.bin in drop-1 should now yield file not found. We test this inside the container (for example, using Python with exception handling) and record that missing-marker as intended. For robustness, we also try writing a different new file (e.g. /data/probe-<timestamp>) to confirm the volume is writable (if that fails, it’s an ERROR). We do not recreate or fix marker.bin; its absence is the expected outcome for a new volume.

We then compare PV UIDs: the new PV is separate from the old Released one. The old PV (from before) is still in the cluster (but not attached to any Pod) and we keep its record. Any surprising result (for example, if the new PVC somehow bound to the original Released PV, or if the drop-arm reads old bytes) would trigger INVESTIGATE. Otherwise, the drop arm is a PASS meaning “new allocation occurred”. We explicitly retain the old Released PV in our evidence; the cluster did not do anything to its claimRef. A complete drop-arm result yields a new PVC/PV and missing marker, confirming that whenScaled: Delete did not recover the old data.

Prove marker absence without confusing it with an access error

When checking drop-1’s filesystem, we must distinguish “file not found” from other failures. We do something like:

import hashlib

try:
    with open("/data/marker.bin", "rb") as f:
        data = f.read()
except FileNotFoundError:
    print("MARKER_MISSING")
except Exception as e:
    print("ERROR", e)
    raise
else:
    print("UNEXPECTED_CONTENT", hashlib.sha256(data).hexdigest())
    raise SystemExit(1)

We ensure we catch only FileNotFoundError as the expected outcome. If, say, permissions were wrong (PermissionError) or disk I/O fails, that is an error; we label that HOLD, because it is not the expected result of a retention check. We also verify we can write a different file on /data to prove the volume is truly mounted and writable. All new identifiers (PVC UID, PV UID, and file hashes) are logged. Again, we do not attempt to put marker back or auto-pass; the presence of a new volume and absence of the old file is our evidence of “new allocation” under drop.

Reconcile both arms against an explicit oracle

The following decision table states the expected outcomes to compare with the operator’s recorded observations. Each row checks a stage from baseline through final observation. For brevity, we show the essential parts:

Stage

keep (arm)

drop (arm)

Interpretation

Eligible baseline

Original PVC/PV (Bound, UIDs) + marker

Original PVC/PV (Bound, UIDs) + marker

All objects created and data seeded

Ordinary Pod replacement

New Pod UID, original PVC/PV/bytes

New Pod UID, original PVC/PV/bytes

Replacement restart, not downscale

Completed downscale

Original PVC/PV remain Bound

Original PVC gone, old PV Released

whenScaled differs between arms

Completed upscale

Original PVC/PV/bytes returned

New PVC UID, new PV, marker missing, old PV retained

Name reuse != original data

Incomplete/contradictory

HOLD

HOLD

Observations missing or inconsistent

  •     In the baseline, both arms must show the expected original storage (PVC/PV UIDs and marker hashes).

  •     The ordinary replacement row confirms that deleting a pod (without changing replicas) did not affect storage in either arm.

  •     After downscale, the keep arm should still have the original PVC/PV, while the drop arm will have lost its PVC and moved the PV to Released.

  •     After upscale, the keep arm should have exactly the original PVC/PV and marker bytes again, whereas the drop arm should show a new PVC/PV and no original data.

  •     Any row that fails those expectations, or if any step was unobserved (e.g. a PVC stuck terminating), leads to HOLD.

We only consider the experiment passed for each arm if all expected invariants hold. Note that a successful kubectl command does not guarantee those conditions; we require checking the table of observed values. We do not grant automatic success just from seeing the table; we explicitly verify each field’s value against what the spec and baseline require.

Decide whether to proceed, pause or request separate recovery

Based on the above, we define our conclusions. A complete and correct keep-arm result leads to ACCEPT_ORIGINAL_REATTACHMENT (original storage returned). A complete drop-arm result leads to CLASSIFY_NEW_ALLOCATION (only new storage was returned). Any incomplete steps (timeout waiting, error, mismatch) lead to HOLD for further investigation. Any conflicting evidence (e.g. seeing data when not expected) is INVESTIGATE.

These correspond to decisions by different stakeholders. If the workload owner requires data persistence (they declared “keep my ordinal-1 data”), then ACCEPT_ORIGINAL_REATTACHMENT means the controller met that requirement in the keep-arm scenario. If the policy was “we can tolerate new storage,” then CLASSIFY_NEW_ALLOCATION would be acceptable. If original data is required but our evidence showed new allocation, then we must escalate: the storage owner must know the old PV's identity to attempt recovery from backup or manual reclamation. The ledger of UIDs is critical here. In general, restoring a replica count cannot resurrect a deleted PVC’s UID. Once a PVC is deleted, a later PVC of the same name is a distinct object. Thus a rollback or policy change cannot bring back the old UID. Changing whenScaled after the fact does not recover the lost PVC.

The operational decisions are:

  •     If keep succeeded, we accept that original volume was reattached.

  •     If drop succeeded, we acknowledge that a new allocation was used (meeting the retention policy but not recovering old data).

  •     Either way, if evidence is missing, we hold.

  •     If something unexpected happened, we investigate.

No matter which outcome, we now have the exact identity mapping needed to either continue operation with the chosen path or to manually recover using the old PV asset (if that was needed and we had retention).

Preserve evidence and clean up the synthetic assets

Before tearing down the namespace, we archive all evidence: JSON dumps, hashes, timestamps, logs. Then we clean up explicitly and carefully. First, delete the StatefulSets (k delete statefulset keep drop), which will terminate pods. Wait for Pod termination and verify the corresponding unmount or detach state before continuing. Now delete only the PVCs we recorded as part of the fixture (for example, by scripting k get pvc -l retention-run=$RUN). Because our policy was whenDeleted: Retain, deleting the StatefulSet did not remove any PVCs, so we must delete them manually if desired. We do not delete the entire namespace at once until all data is saved, and we do not bulk-delete all PVs or storage-class resources.

Each original PV has persistentVolumeReclaimPolicy: Retain, so even after deleting PVC objects, the PV objects remain in the cluster with Released status. Those volumes (the storage assets) are now the responsibility of the storage owner to clean up. We do not delete PVs from the cluster to prove asset deletion; doing so would only remove Kubernetes metadata, not the external disk. Proper cleanup (e.g. wiping disks) must follow the organization’s process. Until then, we report CLEANUP_PENDING for those assets.

Throughout cleanup, we operate with least privileges as per production hardening guidance: only delete the resources we created, and leave system objects untouched. We ensure that RBAC grants the test user the required permissions for these namespaced objects, preserve the applicable admission controls, and avoid wildcard deletions. We do not remove all PVCs or PVs in the class, because that could inadvertently affect unrelated workloads. We preserve namespace and PV objects while saving their metadata, then delete the StatefulSets and their PVCs by name. The namespace itself can be deleted at the end once all evidence is confirmed saved.

Turn the result into an owned scaling change

Finally, we summarize and document the entire exercise as a formal Change Review checklist. Items include:

  •     Requirement: the retention policy under test (whenScaled=Retain vs Delete).

  •     Effective policies: the actual values of whenScaled and whenDeleted from the StatefulSet objects, reclaimPolicy from storageclass.json, and persistentVolumeReclaimPolicy from every actual PV JSON snapshot.

  •     Cluster/provisioner versions: Kubernetes client and server versions in versions.json, with provisioner and CSI-driver versions recorded separately.

  •     Baseline passed: that 2 Pods per arm were running and marker files written.

  •     Negative control: that deleting a pod left PVC/PV untouched in both arms.

  •     Identity mapping: the original vs new PVC UIDs and PV UIDs for ordinal-1 in both arms.

  •     Marker results: the SHA-256 of the returned file in keep-1, and the absence in drop-1.

  •     Control check: ordinal-0 PVC/PV/bytes were unchanged.

  •     Any errors or missing evidence: e.g. if a PVC stayed Terminating, mark HOLD.

  •     Which scenario is accepted: e.g. “Accept original storage returned” or “Original storage not preserved; new allocation was used.”

  •     Cleanup owner: who will reclaim the Released PVs (typically the storage admin).

We record which artifacts (JSON files, logs) in our evidence directory correspond to each bullet.

We explicitly note the scope: a completed test validates one specific storage workflow on the given cluster and class. If anything changes (e.g. a different CSI driver, storage class, or server version), one should rerun the comparison. Reusing the same namespace/run ID or re-seeding with old data would invalidate the test; a fresh run must be done for a repeat.

This synthetic trial does not guarantee durability under arbitrary failures, nor does it test multi-node or cloud-specific cases. A completed run shows how PVC retention policies behaved in that specific environment. We link each observed change to UIDs and bytes to make the decision reproducible.

Build repeatable storage practice with Refonte Learning

This procedure illustrates a disciplined DevOps approach: record effective configuration, reproduce changes, and verify outcomes. Practical labs and capstone projects help build the Kubernetes foundations needed for this work. Refonte Learning’s DevOps Engineering Program runs for 3 months with an expected commitment of 12–14 hours per week. Its published curriculum covers CI/CD, cloud skills, Docker and Kubernetes, and provides foundations for designing controlled infrastructure experiments. Applicants must be engaged in bachelor’s or postgraduate studies.

For further Kubernetes practice, explore the DevOps Engineering Program overview and its practical learning activities. The program lists Oskar Eriksson as a lead instructor and mentor; the published information does not confirm that this exact storage-retention lab is part of the current curriculum. Teams with an authorized Kubernetes environment and dynamic storage can also adapt this procedure: record policies and UIDs with kubectl, preserve the original marker bytes, and verify every observed transition.

Before trusting a scaled-up StatefulSet to recover data, we must prove it with evidence. A returned pod name alone is not enough proof that the “right” disk came back; only matching volume UID and matching file contents can do that. By following the detailed steps above, a platform team can confidently accept the change (or pause and recover) based on concrete identity and content checks, not guesswork.