Stacked S3 delete markers can hide data in a version-enabled bucket even after some markers are removed. For example, consider an object with two distinct data versions (V1 and V2) and two successive simple deletes. The first DeleteObject call inserts a delete marker D1 (making the key invisible) and then a second simple delete when D1 is current will insert another delete marker D2. Removing only the latest marker (D2) exposes D1, still blocking visibility. In other words, even though a DeleteObject call returned success (204), the object remains hidden behind an older delete marker.
Our goal is to prove that the approved version V2 is ultimately restored and readable, and to record exactly which markers were removed by the API. A correct recovery “contract” means confirming: the requested marker transitions occurred; exactly V2 (the intended version) is now current and its bytes are returned by a GET; V1 remains as a noncurrent version; and our unrelated control object is unchanged. A successful HTTP 204 alone does not guarantee this outcome. Instead, we must inventory versions, inspect permissions, and match content digests to declare success or a HOLD state. Unlike bulk-delete reconciliation in unversioned buckets, here we care not just about visibility but also about which specific version was undeleted.
This playbook (Python 3.12, Boto3 1.43.108, Botocore with total_max_attempts=1) walks through a complete example: defining expectations, setting up a test bucket, writing two payloads and a control file, deleting twice, selectively removing markers, and validating at each step. We gather evidence from the AWS API, including version IDs and checksums, and feed it to a deterministic verifier that issues an ACCEPT or HOLD decision. We illustrate why an orphaned marker (D1) prevents recovery until it is removed, and provide a decision matrix and cleanup guidance. This ensures cloud developers and storage operators can safely “restore the right version” in S3 with auditable checks and no confusion between delete markers and actual data versions.
Define the recovery contract before deleting anything
First, clarify the actors. In S3 terms, an object key is the name (path) of the object. Each immutable data version of that key has a unique VersionId. A delete marker (D) is also an object version (with its own ID) but contains no data. When a simple DELETE (no versionId) is called on a versioned bucket, S3 creates a delete marker as the current version. From a user’s perspective, this hides the object: a plain GET returns 404 Not Found and x-amz-delete-marker: true in the header. Importantly, all older versions (data or markers) are still present and continue to incur storage charges. Only a DeleteObject with a specific versionId permanently removes that version (data or marker).
In our scenario, we declare V2 (the “approved” content) as the target version to restore. We will create V1 (draft) and V2, then issue two simple deletes, yielding markers D1 and D2. The “recovery contract” is to remove exactly D1 and D2 (by their VersionIds) so that the key becomes visible again. Successful recovery is defined as: exactly D1 and D2 were removed, V2 is now current, and a GET returns the exact V2 bytes; V1 remains noncurrent and readable; and a control key (same prefix) is untouched.
It is crucial to separate visibility restoration (making an object visible again) from version disposal (permanently deleting a data version). In legal/compliance scenarios, one might want to permanently delete sensitive versions, but here we explicitly avoid permanently removing data versions. We only remove markers to restore visibility. This differs from an unversioned bucket use case (like our bulk-delete reconciliation article), where deletes remove content outright. Here, version IDs are tracked and we never discard a data version we might want.
Before we start, we link this scenario to related problems: Reconcile unversioned deletes covers per-key checks after a bulk delete; here we instead focus on stacked markers and version recovery. Also, see the S3 Object Lock release controls discussion for retention-specific workflows, which are excluded here. Object Lock protects data versions, not delete markers, so it is separate from the marker removal covered here. In summary, our baseline is: bucket is versioning-enabled, two data versions V1/V2 exist, and no delete markers yet. We define V2 as the target to be current. All acceptance criteria (current version ID and content hash, plus intact old version and control file) are written down before any mutation.
Pin a version-enabled, disposable environment
We use a dedicated nonproduction bucket with versioning enabled. In a terminal:
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install "boto3==1.43.108"
python -m pip freeze > requirements.lock
This pins Boto3 1.43.108 (released Oct 2, 2026, requiring Python >=3.10) and its Botocore dependency (resolved via pip). Record the exact resolved Botocore, JMESPath, and other dependency versions in requirements.lock. The AWS CLI is not used here. We will run code under a profile with appropriate IAM permissions. The script expects these environment variables: S3LAB_BUCKET, S3LAB_OWNER (expected AWS account ID), S3LAB_REGION, and a unique prefix string for this run. It writes evidence into a local directory (e.g. evidence/).
We also specify server-side encryption SSE-S3 for the payloads (small text); this is to avoid any issues and ensure contents are encrypted at rest, even though we did not embed KMS keys. We record all actual versions (Python patch level e.g. 3.12.0), Botocore version (from import botocore after install), OS (like Ubuntu 22.04 or Amazon Linux 2023), AWS account ID and IAM principal (e.g. ARN from sts:GetCallerIdentity), Region, and Boto3 retry configuration (we set retries={"mode":"standard","total_max_attempts":1} to disable retries).
Verify the bucket state and the reader’s authority
Before writing, ensure the bucket’s Versioning=Enabled. A quick boto3.client("s3").get_bucket_versioning(Bucket=bucket) should show 'Status': 'Enabled'. (AWS recommends waiting ~15 minutes after turning on versioning before new writes for full propagation.) The bucket must not be an Object Lock or MFA-Delete protected bucket for this playbook; we assume no retention/encryption beyond SSE-S3. Likewise, ensure no lifecycle or cross-region replication rules will interfere.
The IAM credentials must allow:
s3:PutObject, s3:DeleteObject on the target bucket (with ExpectedBucketOwner restriction),
s3:GetObject, s3:GetObjectVersion, s3:ListBucket, s3:ListBucketVersions for reading, plus s3:GetBucketVersioning for the preflight check,
s3:DeleteObjectVersion for deleting specific versions. We do not need PutBucketVersioning or any global actions. We do not use Object Lock or legal holds, but if they were, our script would need PutObjectLegalHold etc., which we omit. (For retention, see the S3 Event Holds release workflow article.) We also do not involve multi-factor (MFA) delete. In code, before mutating, we separately check versioning status (and skip or abort if not enabled) and confirm the caller’s identity via STS. Then all actions are recorded with ExpectedBucketOwner to avoid cross-account mistakes.
Declare two data versions and a similar-prefix control
We define three payloads in code:
payloads = {
"V1": b"revision=1;status=draft\n",
"V2": b"revision=2;status=approved\n",
"C1": b"control=must-remain\n",
}
target = prefix + "report.txt"
control = prefix + "report.txt.control"These are static bytes; we record each content’s length and SHA-256 digest locally (e.g. manifest JSON) so we can verify GET responses. The target is a key like prefix+"report.txt", and the control object uses the same prefix but name "report.txt.control". We PUT these with s3.put_object, capturing their VersionIds and ensuring none are empty or duplicate. If the keys already exist, we abort to avoid interfering with pre-existing data. The control object (C1) is just to verify our list operations didn’t corrupt anything unintended.
Notice we set the key prefix the same (except suffix) so that Prefix=prefix in listing will gather both the target and its control. The target and control keys have no common ancestor “folder” (since S3 is flat), but listing with Prefix as above groups them. We avoid any multipart uploads or incomplete objects; these payloads are small, single-part, complete objects. (This distinguishes our scenario from multipart cleanup issues in S3: see the multipart upload cleanup article, which deals with stray parts, whereas here we treat completed object versions.)
At this point, we declare “baseline” expected state: key target has versions V1 and V2, with V2 as current. No delete markers exist yet. Our manifest notes: content lengths and SHA-256 for V1 (24 bytes), V2 (27 bytes), and C1 (20 bytes). We will later use these values to check that GET returns exactly the right bytes, not just a readable object.
Record requests and read every version page
We implement a small harness s3_marker_lab.py to log all API calls. Key parts:
import hashlib, json
from datetime import datetime, timezone
import boto3
from botocore.config import Config
from botocore.exceptions import BotoCoreError, ClientError
s3 = boto3.client(
"s3", region_name=region,
config=Config(
retries={"mode": "standard", "total_max_attempts": 1},
connect_timeout=10, read_timeout=30,
),
)
def call(operation, parameters):
request = {"Bucket": bucket, "ExpectedBucketOwner": owner, parameters}
# Remove raw Body for logging, but include its SHA256
safe = {k: v for k, v in request.items() if k != "Body"}
if "Body" in request:
safe["BodySHA256"] = hashlib.sha256(request["Body"]).hexdigest()
record = {
"at": datetime.now(timezone.utc).isoformat(),
"operation": operation, "request": safe,
}
try:
result = getattr(s3, operation)(**request)
# If there's a streaming Body, read it fully for digest
stream = result.pop("Body", None)
if stream is not None:
try:
body = stream.read()
finally:
stream.close()
result["ObservedSHA256"] = hashlib.sha256(body).hexdigest()
result["ObservedLength"] = len(body)
except ClientError as error:
result = error.response
except BotoCoreError as error:
result = {"TransportError": type(error).__name__}
record["response"] = result
with log.open("a", encoding="utf-8") as out:
out.write(json.dumps(record, default=str) + "\n")
return result
def require_status(response, status):
actual = response.get("ResponseMetadata", {}).get("HTTPStatusCode")
if actual != status or "Error" in response or "TransportError" in response:
raise RuntimeError("HOLD: unexpected or unknown request outcome")
return responseThis call function logs each S3 request and its JSON response (including headers like x-amz-version-id) to a file. We deliberately record the Body’s SHA-256 for GET operations, but never log raw Body bytes. We set retries to a total of 1 attempt to avoid a hidden retry creating an unwanted second marker. A network transport error is recorded as such, and require_status will abort with a HOLD if the HTTP status isn’t exactly as expected (e.g. not 200 or 204). We explicitly do not suppress errors: on any unexpected response or exception, we bail out.
Preserve both data-version and marker records
After any change, we need to inventory all versions of our keys. We use ListObjectVersions with pagination. Pseudo-code:
pages = []
kwargs = {"Bucket": bucket, "ExpectedBucketOwner": owner, "Prefix": prefix, "MaxKeys": 1}
while True:
resp = call("list_object_versions", **kwargs)
require_status(resp, 200)
pages.append(resp)
if not resp.get("IsTruncated"):
break
kwargs["KeyMarker"] = resp.get("NextKeyMarker")
kwargs["VersionIdMarker"] = resp.get("NextVersionIdMarker")We use MaxKeys=1 per page deliberately so each response contains at most one Version or DeleteMarker, forcing multiple pages and exercising continuation tokens. We save all pages in the log. This way, even if some pages contain only delete markers or only versions, none are skipped. We then combine all <Version> entries and <DeleteMarker> entries from every page (using whichever of the Versions and DeleteMarkers collections are present, with continuation markers when the response is truncated). We ensure there are no missing or duplicate entries across pages, and that the last page has IsTruncated=False. (Any missing NextKeyMarker or mismatch is an error causing HOLD.)
From the combined list, we filter records whose Key equals our target or control. For target, we expect both Versions (kind "data") and DeleteMarkers (kind "marker"). Each record has VersionId and IsLatest. Data-version records also include Size; marker identity comes from membership in the DeleteMarkers collection, not a DeleteMarker: true field in each record. We capture them for evaluation.
Notably, a plain LIST Objects (ListObjectsV2) would not show keys that are currently hidden by a delete marker. But ListObjectVersions reveals everything (data or marker). This mirrors AWS CLI’s list-objects-v2 vs list-object-versions behavior. (See the AWS CLI pagination and filtering guidance for a broader discussion of paginated listings.) Using our call logger ensures we know the exact VersionId and IsLatest flag of every item at each stage.
Establish the current and historical read baseline
Now we execute the baseline writes in order: PUT V1, then PUT V2, then PUT C1 (control). Each put_object(Bucket, Key, Body, ServerSideEncryption='AES256') returns an HTTP 200 and a VersionId in the response headers. We call require_status to ensure 200. We record those VersionIds (e.g. VID1, VID2, VIDc). At this point, VID2 is the current version of report.txt. We confirm by doing:
Ordinary GET report.txt (no versionId): require_status(200), should retrieve VID2 content (27 bytes, matching our sha256).
Explicit GET report.txt?versionId=VID1: should require_status(200) and return V1 content (24 bytes).
Explicit GET report.txt?versionId=VID2: should require_status(200) and return V2 content.
GET report.txt.control (no version): returns control body (20 bytes).
We check that Result['ObservedSHA256'] and ObservedLength match our manifest for V1, V2, C1. We do not rely on ETag (which can be an MD5 digest for single-part SSE-S3 uploads, but is not a universal content digest across encryption and multipart-upload modes). After these, our inventory (via ListObjectVersions) should list two versions for target: V1 and V2, both with IsLatest=False for V1 and True for V2; and no delete markers. The control key should have one version. We print or log these to evidence:
# Expected state
target versions = {V1 (IsLatest: False), V2 (IsLatest: True)}, markers = {}
Any deviation here (missing version, wrong content, etc.) would be a HOLD, since we haven’t yet done any deletes.
Issue a simple delete and prove what remains
We perform the first mutation:
del_resp = call("delete_object", Key=target)
require_status(del_resp, 204)
With no VersionId, in a versioned bucket, AWS creates a delete marker. The 204 response includes "DeleteMarker": true and a new VersionId (say D1). We capture D1 = del_resp["VersionId"]. Now we inventory versions again. We expect to see:
Data versions: V1, V2 (same IDs as before)
Delete marker: D1 (with IsLatest: True). So the list of entries for key=target should be V1 (non-current), V2 (non-current), D1 (current marker). The ordinary GET on report.txt should now return 404 Not Found and a header x-amz-delete-marker: true, since D1 is now the latest version. We verify this by calling GetObject without versionId: it should not be 200. (For a missing object, credentials with s3:ListBucket receive 404; without that permission, S3 can return 403. A 403 response is inconclusive and requires HOLD; it does not prove that a marker is current.) We also do explicit GETs: GetObject(VersionId=VID1) returns V1 body; GetObject(VersionId=VID2) returns V2; GetObject(VersionId=D1) returns error 405 (MethodNotAllowed) with header x-amz-delete-marker: true. We should see an HTTPStatusCode of 405 for that call.
We record:
Before delete: {V1, V2}, current=V2.
After delete once: {V1, V2, D1}, current=D1, plain GET returns 404 with marker. This illustrates that the object is hidden but not removed. A key detail: a normal LIST Objects would not show “report.txt” at all (since its current version is a marker), but our ListObjectVersions still lists V1 and V2 as "old versions". We do not rely on LIST; we check via versions.
Repeat the delete to create a second marker
We immediately do another simple delete on the same key:
del_resp2 = call("delete_object", Key=target)
require_status(del_resp2, 204)
Because D1 is currently the latest version and we did not specify VersionId, AWS will not remove D1. Instead, it creates another delete marker with a different VersionId (say D2). We verify D2 != D1. Now inventory should show V1, V2 (non-current) and two delete markers D1, D2, where D2 is current. The ordinary GET is still 404. If we try a GET with VersionId=D2, we expect a 405 with marker header. We explicitly check that: yes, it should not return a zero-byte object; it must show delete-marker: true. (If it returned zero-length 200, that would be incorrect behavior.) At this point, evidence shows two markers stacked, current=D2, and both original versions still stored.
We do not attempt to delete again without versionId, because on error we should not blindly retry. If the second delete had returned something unexpected (like 304 or error), we would abort. For a successful second simple delete, expect HTTP 204 and a new marker when one already exists.
Remove D2 and observe why recovery still fails
Now we remove only the latest marker D2 by specifying it:
rm2_resp = call("delete_object", Key=target, VersionId=D2)
require_status(rm2_resp, 204)
This is a version-specific deletion: it will permanently delete the version D2 (which is a delete-marker version). The response has x-amz-delete-marker: true and the VersionId of the deleted marker (D2). We now inventory again. We expect D2 gone, but D1 remains and becomes current (IsLatest=True). Data versions V1, V2 are still there, and V2 is not current. In other words, state is {V1, V2, D1}, current=D1.
Crucially, ordinary GET is still 404 because D1 is still a marker. The object is not visible, even though we just successfully removed a marker. One must remove D1 as well to restore. We must emphasize: do not chalk this up to “eventual consistency”; our list confirms D1 is current now. The visibility barrier is D1.
Make the misleading success fail an acceptance test
An inexperienced automation might think “we did a delete-object (204), so the object is restored”. To show why that’s wrong, consider a naive check:
if rm2_resp.get("ResponseMetadata", {}).get("HTTPStatusCode")==204:
print("ACCEPT: visibility restored!")
else:
print("HOLD")
This script would output ACCEPT. But if we then do a GET, we still get 404. A correct evaluator sees: currentVersion=D1≠V2, and the ordinary GET still does not return V2 content; control file still okay. Thus the verdict must be HOLD. In other words, the action succeeded (D2 was removed), but our goal was not met because we still have a delete marker. We will illustrate in our classifier logic that “HTTP 204” without further evidence is not enough.
Remove D1 and verify the restored V2 bytes
Now we proceed to remove the remaining marker:
rm1_resp = call("delete_object", Key=target, VersionId=D1)
require_status(rm1_resp, 204)
We expect D1 to be gone. Re-listing versions should now show only {V1, V2} (no markers) and V2 should again have IsLatest=True. The ordinary GET report.txt should now return 200 and yield the V2 payload exactly. We verify the SHA-256 and length against our manifest. An explicit GET for V1 still works, yielding V1 content. The control key is unchanged at every stage (GET control always returns C1 bytes).
We now have an expected table for the target key at each stage, with current version and GET outcome:
Stage | Target data versions | Target markers | Latest target | Ordinary target GET |
Baseline | V1, V2 | None | V2 | 200, V2 bytes |
Delete once | V1, V2 | D1 | D1 | 404, marker header |
Delete twice | V1, V2 | D1, D2 | D2 | 404, marker header |
Remove D2 | V1, V2 | D1 | D1 | 404, marker header |
Remove D1 | V1, V2 | None | V2 | 200, V2 bytes |
This is our expected state table (not from AWS, but our design). After deletion, only at the final recovery stage do we get a 200 with V2 content.
Require identity, bytes, and unaffected controls together
Finally, we assert that recovery is complete only if all checks align: current version is V2, and that retrieving the object indeed returned the expected V2 bytes (via independent SHA256). One might see a readable object and mistakenly assume it is correct. For example, if somehow a third version V3 (unknown) had appeared, or if the content were corrupted, reading content alone would be misleading. We compare the observed (Key, VersionId, data/marker flag, IsLatest, digest) against exactly the expected sets built from baseline IDs. We also check the control record to ensure it hasn’t changed. Any discrepancy (wrong versionId, bytes mismatch, or missing noncurrent V1) triggers a HOLD state. We explicitly do not say “the object is restored” until we have both identity and content proof for V2, plus intact other records. If anything is off, we record HOLD and do not proceed to a cleanup.
Reject invalid state records with a local classifier
We encapsulate validation in a pure function evaluate(stage, pages, probes, manifest) that returns (decision, reasons). It checks:
Collection shape: Expect exactly one record per version or marker per key with no duplicates. If list pages are missing or duplicate continuation tokens appear, reject.
Version identities: From the log pages, gather sets of (Key, VersionId, kind) where kind is “Version” or “DeleteMarker”. There should be no duplicate (Key,VersionId) entries. If duplicates exist, HOLD.
Expected keys/versions: Compare against our known expected set for that stage (using captured V1,V2,C1,D1,D2). If an unexpected Key or VersionId appears or an expected one is missing, HOLD.
IsLatest flags: Exactly one record per key has IsLatest=True. If more or less, or wrong element flagged, HOLD.
Probes: These are GET/GET Version responses. We expect either status 200 with correct body, or documented error (404 or 405 with marker) for each probe (e.g. GET with no version, GET V1, GET V2, GET C1). Any missing probe, or 403/timeout (TransportError) is treated as “inconclusive evidence” → HOLD.
We do not infer stage from the responses (e.g. we don’t assume if V2 is current, we were at “Remove D1” stage; rather we have the stage name and expected IDs given from our own state machine to compare against). In code we compare sets of (Key,VersionId,Kind) so ordering doesn’t hide duplicates. For example, after second delete, we expect {(target, V1, Version), (target, V2, Version), (target, D1, Marker), (target, D2, Marker), (control, C1, Version)}. If one marker is missing, or a wrong version shows up, evaluation fails.
Test the classifier without claiming an AWS experiment
To validate our evaluate logic itself, we write synthetic JSON fixtures (mocked lists and probes) for all positive cases (the five table stages) and negative cases:
Omit D1: Invent pages missing D1 after “Delete once”. The classifier should return HOLD with reason “expected marker D1 not found”.
Duplicate D2: Simulate pages listing D2 twice (maybe from a bug). HOLD with “duplicate versionId”.
Add V3: Extra version not in manifest. HOLD for “unexpected version”.
Reverse IsLatest: Tag V2 as current instead of D2 in stage 3. HOLD due to logical mismatch.
Wrong control bytes: Change control’s probe digest. HOLD “control changed”.
403 on GET: E.g. GET target returns 403 due to missing ListBucket. This is “inconclusive missing-object error” and should result in HOLD (we don’t assume object is gone or visible).
TransportError: If probe had a read timeout or s3:error, that’s HOLD “transport uncertainty”.
Truncate listing early: If pages stop early (IsTruncated false prematurely). HOLD “truncated listing”.
Run these through evaluate, and ensure it flags them as HOLD without silently fixing them. These checks test whether the classifier can detect anomalies without live AWS calls. (These are offline Python unit tests using crafted dictionaries; they do not call AWS.) Passing these synthetic tests provides evidence about the classifier, not evidence that the AWS recovery sequence was executed.
Distinguish version removal from visibility restoration
It might be tempting to remove V1 (the draft version) when trying to “restore” the object, thinking “no one needs V1”. But deleting V1 would be irreversible content loss. The correct recovery path is just removing the markers. If one wanted to truly purge V1 now that V2 is in place, that would be a separate, deliberate step, e.g. delete_object(VersionId=VID1). That action permanently deletes V1’s bytes and is irreversible through normal S3 version recovery. Since our playbook is not about secure deletion, we stop after visibility restore. We do note: if one does delete V1 at the end, you must explicitly target the specific VersionId=VID1 and then verify only V2 and control remain. But this final purging step is out of scope for our acceptance logic; we only record that it would be permanent. We avoid any “purge script” or claims of hardware erasure.
Apply a decision matrix that names the remaining work
Based on our evidence, we decide among four outcomes:
Decision | Required evidence | Next action |
ACCEPT VISIBILITY REMOVAL | Key is hidden (ordinary GET 404), and version list matches baseline (with markers) | Document retained versions (who is current) |
RESTORE EXACT MARKERS | Desired version ID (V2) and marker IDs (e.g. D1,D2) are known and unchanged, and key is hidden only by those markers | Remove the next blocking marker via versioned delete, then re-evaluate |
ACCEPT RESTORED | V2 is current (IsLatest), GET returns V2 bytes (digest verified), V1 and control intact, full inventory as expected | Close the recovery record as successful |
HOLD | Any missing or denial response, mismatch in expected versions/markers, control changed, or concurrent writes detected | Troubleshoot (reconcile inventory) before further changes |
Important: an isolated 204 response (like after delete calls) is never enough for ACCEPT RESTORED. For instance, after deleting D1 we get 204, but we must still verify V2’s content. Similarly, we decouple the business choice (“should we restore the object?”) from the technical proof (“did our targeted change actually happen?”). The first two decision lines show that we might intentionally stop at hiding the key (if legal policy prefers it), but that choice is made by policy, not by the code. Our code simply verifies the chosen transition. The matrix provides explicit labels for each path with “Next action” guidance, ensuring nothing is assumed or skipped.
Own reruns, interrupted recovery, and bounded cleanup
Assign clear roles: the application owner (or version owner) states that V2 is the desired version to keep visible, the storage operator oversees the delete-marker removal under the given authority, and the reviewer (another engineer or auditor) checks the evidence logs against the decision. This is similar to disaster recovery planning, which emphasizes ownership and repeatability. One should rebuild evidence logs if AWS CLI or SDK versions change.
If a run is interrupted (script failure, lost log file, crash) we do not assume partial state is safe. We preserve the last recorded state (pages/probes), and re-run ListObjectVersions to reconcile. Unknown or stale markers must be identified (e.g. by listing). We do not assume we can just replay a deletion if we lost the response; that risks deleting the wrong record. Likewise, deleting the current latest version by mistake is irreversible. Thus we log every step before and after each delete, so we always know what was last executed.
Reruns must start fresh: use a new prefix and directory, or confirm the old run is cleaned up. The cleanup (to return the bucket to empty for these keys) must remove only the exact VersionIds we created (V1, V2, C1, D1, D2). It should never do a “delete all versions under prefix” with a simple DeleteObject call, because that could add new markers and disturb other data. A bounded cleanup function would explicitly use delete_object(Key=target, VersionId=<each ID>) for V1 and V2, and delete_object(Key=control, VersionId=C1) for the control, if disposal is authorized after removing the markers. These version-specific deletes permanently remove the fixture data. We leave that separate and small, outside of the evidence tape. We do not attempt automatic bucket purges as part of this recipe.
Reconcile an unknown outcome before another mutation
If a single API call’s outcome is unclear (network error, missing log), we stop and reconcile. For example, if we sent the second delete but didn’t capture D2’s ID, we list versions to find it. If some response said D2 was removed but that didn’t appear in logs, we list to verify. At no point do we proceed with another destructive call (like deleting D1) without confirming the current inventory. Missing or wrong version IDs must be made explicit in the evidence (our CSV or JSON log) and an analyst must update them.
Important: “restoring visibility” (removing delete markers) is not a roll-back of an irreversible delete of a version. Once V2 has been permanently deleted by DeleteObject with VersionId=VID2, removing a marker cannot bring it back. That data-version deletion is outside the recovery sequence. V1 is retained as a readable earlier version, although it cannot reproduce the approved V2 bytes. The recovery sequence removes only markers and preserves data versions; the separate optional cleanup permanently deletes the selected fixture versions.
Build the cloud foundations behind reliable recovery
Precise resource identity and scoped authority are key principles of cloud development. Here, we treat each version and marker as an exact resource (via VersionId). We measure outcomes (status codes, digests) rather than guess. We minimize permissions to just what’s needed, and we log everything. This aligns with the Cloud Development Program’s curriculum, which emphasizes careful infrastructure and code (see Cloud Development Program details). By practicing these disciplined recovery steps in a sandbox, we hone skills that underpin resilient cloud applications. For example, knowing AWS S3’s consistency guarantees and versioning rules lets us build more reliable services.
The Cloud Development Program covers cloud app deployment, CI/CD, containers, and security; our S3 version-recovery playbook is one concrete exercise in cloud data operations. It complements foundational coursework: rigorous testing of failure scenarios, audit log creation, and validation. While the program doesn’t explicitly teach this exact S3 workflow, its coverage of performance monitoring and optimization provides related foundations for gathering logs and metrics and making evidence-based decisions. Enrollment in that program requires commitment (~12–15 hours/week for 3 months) and provides mentorship (e.g. from MSc Charlotte Smith), but if you want to strengthen your recovery engineering, building hands-on proficiency like this playbook is a great step toward that.
