When a developer issues an Amazon S3 DeleteObjects request, an HTTP 200 (OK) response may conceal a mix of successes and failures. The service returns a “DeleteResult” containing per-key <Deleted> entries for objects it removed and <Error> entries for keys it did not delete (for example due to permission issues). A bulk-delete envelope status alone cannot guarantee each key’s fate. Before considering the batch operation “closed,” an authorized developer must reconcile every requested key: verify the reported result, check for actual absence, and handle retries only where safe.
This article defines that reconciliation process and the four possible outcomes (ACCEPT, REPAIR, RETRY, or HOLD) for unversioned S3 buckets.
We assume an owned test bucket (no versioning, locks, or retention) with exactly three inert keys in scope: one allowed for deletion, one explicitly denied, and one already absent. All SDK logic is local (using Boto3 with a Botocore Stubber) to assert the client’s parsing rules, separate from any live AWS calls.
Our scope is strictly the service contract for DeleteObjects on unversioned data, including both verbose and quiet responses. We will not cover versioned buckets, retention holds, or lifecycle jobs. We use documented API behaviors and S3’s data consistency model as the baseline.
The key challenges are: (a) confirming that every request key is reported exactly once (either in Deleted or Errors), (b) distinguishing reported delete successes from proof of prior existence, (c) handling the quiet-mode contract correctly, and (d) verifying actual absence under correct read permissions.
A completed DeleteObjects call with no errors is necessary but not sufficient to close the case. We’ll show how to parse and validate the raw response, test for inconsistent reports (like duplicate or missing entries), and then separately use a head-object check to prove that each supposed deletion took effect (or recognize if a denied key remains).
Finally, we tie these checks to a structured decision: we can ACCEPT the operation only when all keys are accounted for and absent (within a controlled writer context); otherwise we either REPAIR the client’s reporting logic, RETRY a smaller subset (denied keys only, once permissions are fixed and writer side is held), or HOLD if outcomes are unknown or contradictory. The result is a clear procedural playbook, complete with code to stub and inspect every scenario, that lets a cloud developer close out a bulk delete job with confidence.
Define what closing a deletion actually means
When an S3 bulk delete API returns, it has provided a report of the attempted operation, not a binding contract that all those objects are now gone. The HTTP envelope (200 OK) means “we processed your request,” but the detailed contents (the per-key results) carry the real status. In a mixed result, some keys appear in <Deleted> (meaning S3 considers them deleted) and some in <Error> (meaning S3 failed to delete them, e.g. AccessDenied). For example, a DeleteObjects request for keys A, B, C might yield:
<Deleted><Key>A</Key></Deleted>
<Error><Key>B</Key><Code>AccessDenied</Code></Error>
<Deleted><Key>C</Key></Deleted>This mixed DeleteObjects result means A and C were removed (or were already gone) and B was not removed (S3 returned a 403 code for B). A single Boolean summary (“Did the request fail?”) is inadequate. Instead, closing the operation means verifying four facets:
Envelope status vs per-key result: We check the HTTP 200 status (or caught exception). A 200 with a valid body means the request was accepted; an HTTP 403/400 means the whole batch was denied or malformed (no per-key info).
Complete key reconciliation: Every requested key must appear exactly once in either a Deleted list or an Error list. No unrequested key should appear. We enforce that Deleted ∪ Error = RequestedKeys and the two sets are disjoint.
Prior-existence vs reported deletion: Seeing a key in Deleted does not prove it existed before. S3’s contract is that if an object is already absent, it will still list it in <Deleted>. Thus we must not infer from “Deleted” that data was removed, only that “it isn’t there now”.
Actual absence verification: With writers held, we must check each key with a HeadObject call. A 404 (Not Found) after a delete generally means the object is gone; but a 403 (Forbidden) from HeadObject under insufficient ListBucket permission is not confirmation of absence. Strong consistency for deletes means a 404 should be immediate if the key was removed, but we also account for the possibility that another writer could have overwritten the key since deletion, which would show up as present.
Only when all keys are accounted and absent (given writer control) can the developer ACCEPT the outcome. If a key was reported deleted but is still present on head-check (or vice versa), that indicates a bug or race. If any response conditions are malformed (duplicates, missing data, same key in both lists, etc.), the report itself requires REPAIR (the local parsing logic or SDK is wrong).
If some keys failed (Errors) and all others are clean, the developer should consider RETRY for the failed subset only, after explicitly correcting IAM or bucket policy for those keys, ensuring only those are retried.
If any outcome is ambiguous (e.g. head-check inconclusive due to 403) or contradictory, then the operation is HOLD: do not mark as done until the ambiguity is resolved by higher-level authority. A stale "client stub says OK" alone cannot close a real deletion job; we need fresh, authoritative evidence for each key.
Note: We treat DeleteObjects as a parallel batch API, not a queue or job. It is neither guaranteed idempotent (because an unversioned key deleted, then overwritten, might cause unintended deletes on retry) nor intrinsically atomic. This procedure is inspired by cloud data-pipeline architecture principles: always reconcile each item’s status and document results before advancing.
Gate the bucket, prefix and deletion authority
Preconditions: Use an owned general-purpose S3 bucket. It must never have had versioning enabled and have no object locks or retention policies on the keys we will use. Choose a specific bucket prefix (folder) with a controlled set of test objects and no other content. In production terms, your IAM or bucket policy should explicitly allow s3:DeleteObject on the intended objects, deny it on test-denied keys, and ensure your head-check identity has s3:GetObject (and ideally ListBucket) on the bucket.
These are fundamental AWS identity and security foundations: separate the permissions for write-delete and for verification. Do not proceed if you lack clear authority over the bucket or prefix. For a lab, explicitly create or prove the prefix and keys beforehand, then assign a separate verifier role that can HEAD but not necessarily delete.
By contrast, do not mix in any retention or versioning workflows to simulate failures. For example, do not attempt to “fail” the delete by setting an object retention period. That changes the conditions. We assume: deleting an unversioned current object is a different contract from deleting a version or circumventing a lock. If the bucket had Object Lock or retention enabled, that would require MFA or governance bypass, which is out of scope. If versioning were on, a DeleteObjects would add a delete marker rather than truly removing data, which is also out of scope.
We are focusing only on deleting current live objects in an unversioned bucket. (In other words, we won’t use a lifecycle rule or version delete marker to manufacture an error; we want an actual per-key error from permissions only.)
Keep retention and versioned-object workflows outside the fixture
Do not repurpose retention holds, legal holds, or version-delete failures to simulate errors. An unversioned delete simply removes or confirms absence of the object. A versioned-bucket scenario risks conflating “delete marker” semantics with bulk delete logic. Those belong in separate guides. Our test objects exist as ordinary unprotected objects or plain placeholders, and we delete them with normal s3:DeleteObject permission.
Create an independent three-key request manifest
Before calling the API, explicitly construct and record the manifest of objects you plan to delete. In our design this is exactly three keys under the allowed prefix:
allowed-key: an inert file that exists and we have permission to delete.
denied-key: an inert file that exists but our caller’s policy explicitly denies delete (so it should fail).
absent-key: a key name that does not exist prior to the call. (You should verify it is absent with a pre-check or ensure you never created it.)
For each key in the manifest, also record metadata (for example, a hash of its content or known ETag) so we can prove later that “allowed-key” and “denied-key” were indeed present before the delete. This manifest is your ground truth and should be cryptographically stable (e.g. record exact object digests and creation timestamps).
The manifest must be treated as authoritative and immutable during the test. We will never derive the list of keys from the API response; instead, all parsing and verification must reference the original manifest data.
Validate the manifest strictly: all three keys must be unique (no duplicates by name) and within the approved prefix. Assign a case ID or similar to correlate this test. For example, you might write a JSON or YAML file:
case: delete-test-123
Bucket: my-test-bucket
Manifest:
- Key: "stage1/allowed-key.txt"
ETag: "abc123"
- Key: "stage1/denied-key.txt"
ETag: "def456"
- Key: "stage1/absent-key.txt"
Present: falseHere we explicitly note that absent-key.txt should not exist (Present: false) and record ETags for the others. By keeping this data separate, we ensure that when the response says “allowed-key.txt is deleted,” we do not mistakenly treat that as evidence that it existed before; the manifest already has that proof.
Likewise, if the response shows an absent key in <Deleted>, we must remember that we knew it was absent.
After manifest validation (no accidental duplicates, correct case sensitivity, etc.), then issue the DeleteObjects call. Never assume that a key not in the manifest can appear, or that names are case-insensitive (S3 is case-sensitive). If you see an unexpected key in the response, treat it as an error.
Key point: The requested set and the manifest must match exactly. A proper reconciliation requires that DeletedKeys ∪ ErrorKeys = ManifestKeys. Any deviation from this (extra or missing names) must trigger a REPAIR path.
Build a network-free SDK contract harness
We implement the above logic in code to verify the DeleteObjects contract before hitting real S3. Here’s how to set up a Botocore Stubber harness to simulate DeleteObjects calls without any actual network. We create a Boto3 S3 client with fake credentials and a dummy endpoint, and wrap it in a Stubber. This ensures no real AWS call is made; the stub provides pre-canned responses. (The endpoint URL can be a localhost or example; the Stubber intercepts calls regardless.)
We then programmatically define the expected request parameters and stubbed outputs.
For example, create a Python script delete_verify.py:
# delete_verify.py
import boto3
from botocore.stub import Stubber
from botocore.config import Config# Use fake credentials and a dummy endpoint
session = boto3.Session(
aws_access_key_id="FAKEACCESSKEY",
aws_secret_access_key="FAKESECRETKEY",
region_name="us-west-2"
)
s3 = session.client(
's3',
endpoint_url='http://localhost:4566',
config=Config(signature_version='s3v4')
)
stubber = Stubber(s3)bucket_name = "my-test-bucket"
request_keys = ["allowed-key.txt", "denied-key.txt", "absent-key.txt"]
expected_params = {
'Bucket': bucket_name,
'Delete': {
'Objects': [{'Key': k} for k in request_keys],
'Quiet': False
}
}# Define a stubbed verbose response (HTTP 200) with mixed results
response = {
'Deleted': [
{'Key': 'allowed-key.txt'},
{'Key': 'absent-key.txt'} # S3 returns absent as 'Deleted'
],
'Errors': [
{
'Key': 'denied-key.txt',
'Code': 'AccessDenied',
'Message': 'Access Denied'
}
]
}
stubber.add_response('delete_objects', response, expected_params)
stubber.activate()try:
result = s3.delete_objects(
Bucket=bucket_name,
Delete={
'Objects': [{'Key': k} for k in request_keys],
'Quiet': False
}
)
print("Stubbed verbose result:", result)
finally:
stubber.deactivate()Run this script (for example, with $ python3 delete_verify.py). The stubber ensures the client will see exactly response as above. In practice, result will be a Python dict identical to response. (We assume Boto3 version around 1.43.x to match the stub shapes.)
Key points in this setup:
We used explicit fake keys and signature just to form a client; no environment creds are picked up.
We specified expected_params to force the stub to check that the call exactly matches our manifest.
We activated the stubber to intercept the call.
We printed out the returned structure.
We included assert_no_pending_responses() implicitly by deactivating the stubber; if our call didn’t match the stubbed one exactly, the stubber would complain.
This is the heart of our test fixture, a miniature simulation of the client side. We can similarly stub the quiet mode case by changing 'Quiet': True and defining an expected response with only errors. For brevity, that could be another stub call in the same script (or a second invocation):
# Stub for Quiet=True
quiet_params = expected_params.copy()
quiet_params['Delete']['Quiet'] = True
quiet_response = {
'Errors': [
{
'Key': 'denied-key.txt',
'Code': 'AccessDenied',
'Message': 'Access Denied'
}
]
}
stubber.add_response('delete_objects', quiet_response, quiet_params)
stubber.activate()
try:
result_quiet = s3.delete_objects(
Bucket=bucket_name,
Delete={
'Objects': [{'Key': k} for k in request_keys],
'Quiet': True
}
)
print("Stubbed quiet result:", result_quiet)
finally:
stubber.deactivate()Again, we would print(result_quiet) to see that for quiet mode only the errors list shows up. If no exception is thrown, the stubbed result_quiet dict will have {'Errors': [...]} and no Deleted key in the body.
This approach, using a local stub, exemplifies robust cloud operating practices: we run the exact same client-side logic (Boto3 parsing, business rules) offline.
We also record the Python, boto3 and botocore versions explicitly (e.g. via print(boto3.__version__, botocore.__version__)) to document the environment. Because Stubber enforces the expected parameters, we guard against falling back to a real AWS call: if our code accidentally tried a real call, it would fail the stub expectations.
Inspect a mixed verbose response instead of its status alone
With the stubbed result from above, we can now simulate how our client logic should interpret it. For our example response:
{
"Deleted": [
{"Key": "allowed-key.txt"},
{"Key": "absent-key.txt"}
],
"Errors": [
{"Key": "denied-key.txt", "Code": "AccessDenied", "Message": "Access Denied"}
]
}A naive report might say: “Deleted: 2, Errors: 1,” but we need to break it down per key. A robust reconciliation ledger lists every requested key, what the service claimed for it, and what we verify post-delete. For readability, consider a table like:
Key | Reported Result (DeleteObjects) | Error Code | Remarks |
allowed-key.txt | Deleted | (none) | OK as expected |
denied-key.txt | (no Deleted entry) Error | AccessDenied (403) | Confirmed not deleted |
absent-key.txt | Deleted | (none) | Key was absent pre-delete |
From the raw stub above, we’d parse result['Deleted'] to extract {"allowed-key.txt", "absent-key.txt"} and result['Errors'] for {"denied-key.txt"}. One temptation is to treat anything in Deleted as a successful delete. But note the third row: absent-key.txt appears in “Deleted” even though it wasn’t present. This is per the S3 contract: if you ask to delete an already-nonexistent key, S3 confirms deletion by listing it as Deleted.
Hence, the Deleted list is not a proof of prior existence. In our ledger, we should flag that with a remark: we knew it was absent before the call, so “Deleted” here means S3 did nothing but confirm the object is (and remains) gone.
Contrast a flawed summary versus the corrected ledger. Simplistic code might just print “Allowed and Absent keys were deleted, Denied key was denied.” But a better output, capturing key identity, might look like:
normalized_ledger = {
"allowed-key.txt": {
"Deleted": True, "Status": "deleted", "Error": None
},
"denied-key.txt": {
"Deleted": False, "Status": "error", "Error": "AccessDenied"
},
"absent-key.txt": {
"Deleted": True, "Status": "deleted", "Error": None
}
}Here we clearly see each key, whether it was in the Deleted list or returned an Error, and what that error code was. This preserves the raw client-reported result keyed by object name. (Internally, code might do something like deleted_keys = {d['Key'] for d in result.get('Deleted', [])}, error_map = {e['Key']: e['Code'] for e in result.get('Errors', [])}, then iterate over the original request list to build a consistent mapping.)
By keeping “allowed-key.txt” separate from “absent-key.txt,” even though both are marked Deleted, we avoid the false inference that absent-key.txt was overwritten or present. We record in our evidence that absent-key.txt had no prior data.
The vital insight here is: the DeleteObjects response is purely service-reported, not an independent audit log. As one AWS example response explains: it includes a <Deleted> for each item successfully deleted and an <Error> for each item not deleted.
The phrase successfully deleted includes the no-op case. Thus our acceptance criteria will later require us to separately prove that those “Deleted” keys are indeed now absent (if we had permission to read them).
Separate reported success from proof of prior existence
The absent-key example highlights a subtle but critical point: S3’s Deleted list is not a table of previously-existing objects. It’s a table of objects now gone or confirmed gone. Even a green light in Deleted for a key only says “we don’t see it”.
We must treat “reported delete success” and “prior existence” as orthogonal. In practice, this means our evidence checklist has separate columns: “Was object there before?” and “Did service claim it deleted it?”.
For the absent-key, we might note:
WasPresentBeforeDelete: No (per manifest)
ReportedDeleted: Yes (per response)
VerifiedNowAbsent: Yes (will confirm with HEAD=404)
That shows a mismatch: Deleted but WasPresent is false. For the allowed-key, both columns are yes (it existed, then was deleted).
We should never collapse these columns into one; mixing them could hide an error (e.g. if somehow an absent key was not reported as Deleted, that would be a protocol violation).
In summary, when inspecting a verbose DeleteObjects response, do not assume that “Deleted count == objects that were there.” Always cross-reference with your known pre-state. This separation of concerns is part of building trustworthy operational checks: we process the raw service output and then apply an independent verification step (next section) under a known identity.
S3’s DeleteObjects API explicitly says if the object isn’t found, it returns the result as deleted. We use that contract rather than adding our own rule. It ensures that a 200 OK batch response alone cannot serve as evidence that an object ever existed, only that it is now non-present (at least under current permissions).
Interpret quiet mode without inventing missing evidence
Now consider the same scenario with Quiet=true. In quiet mode, by definition, the response only includes errors. For a complete successful deletion, S3 will return an empty <Errors> list and no <Deleted> entries. In our stub harness, we simulate:
# Expected parameters (same keys, quiet)
quiet_params = {
'Bucket': bucket_name,
'Delete': {
'Objects': [{'Key': k} for k in request_keys],
'Quiet': True
}
}quiet_response = {
# Since denied-key fails, Errors list has it; others produce no output
'Errors': [{
'Key': 'denied-key.txt',
'Code': 'AccessDenied',
'Message': 'Access Denied'
}]
}stubber.add_response('delete_objects', quiet_response, quiet_params)
stubber.activate()quiet_result = s3.delete_objects(
Bucket=bucket_name,
Delete={
'Objects': [{'Key': k} for k in request_keys],
'Quiet': True
}
)
print("Quiet mode output:", quiet_result)
stubber.deactivate()The reported ledger in quiet mode then comes from taking all requested keys minus the ones in the Errors list. In our case, the service only tells us about denied-key.txt. We interpret that as “denied-key.txt had an error; all other requested keys were successfully deleted (or absent already)”. Formally:
error_keys = {e['Key'] for e in quiet_result.get('Errors', [])}
# {'denied-key.txt'}
reported_deleted = set(request_keys) - error_keys
# {'allowed-key.txt', 'absent-key.txt'}We must be explicit that this logic only applies after verifying that the 200 envelope is present and the response structure is valid. In other words, for quiet mode we rely on:
The HTTP status code is 200 (or equivalently, no exception was raised).
The body is parseable as the documented structure (even if it has no Deleted key).
There are no unexpected fields.
Only then do we say: “Quiet mode didn’t list X or Y, so they succeeded.” We should not reverse engineer quiet mode from a missing output in a failing envelope. For example, if our stub returned an empty dict or threw an exception (which it might if we misuse Stubber), that is not a valid quiet success. We do not treat absence of data as automatic success unless it conforms to the documented format.
A safe implementation checks if 'Errors' is present (perhaps empty or not) and that the API call did not raise. If errors are only partially returned or the result is not as expected, we abort and investigate (REPAIR the client or report).
To illustrate correct vs incorrect interpretation, imagine a broken parser that simply counted only Deleted entries. In quiet mode, a broken parser would see zero Deleted entries and (wrongly) assume “everything was quiet-success”. That would hide the fact that some keys (denied-key) actually failed. We must avoid such mistakes. A valid quiet-mode reconciliation looks like:
Envelope code = 200 (quiet_result did not raise).
The response body has an 'Errors' list (possibly empty) and no 'Deleted' list.
All keys not in Errors are marked as “Reported Deleted/Absent”.
In our data above: quiet mode led us to mark allowed-key.txt and absent-key.txt as deleted, and denied-key.txt as error. We also carry forward the same notion that “absent-key.txt” was previously absent. In output form:
# Quiet mode
normalized_ledger = {
"allowed-key.txt": {"Status": "deleted", "Error": None},
"denied-key.txt": {"Status": "error", "Error": "AccessDenied"},
"absent-key.txt": {"Status": "deleted", "Error": None}
}with reported success set = {"allowed-key.txt", "absent-key.txt"}.Finally, note: if a quiet-mode call returns without an HTTP exception but the response is missing an expected 'Errors' key altogether, we should treat that as incomplete. The documented API says it will at least provide an <Errors> element for any failures.
So our logic should flag any deviation from the documented schema as needing REPAIR of our parsing assumptions. For instance, if Boto3 quietly returned an empty dict instead of {'Errors': []}, we should not blithely assume success. We would log that as a malformed API response (likely an SDK or stub bug).
Distinguish request exceptions from embedded key errors
Up to now, we assumed the request got an HTTP 200 with mixed content. But in reality, DeleteObjects can also fail at the request level, for example if the caller lacks s3:DeleteObject permission across the board, or if the bucket name was wrong, or if MFA was required and missing. In these cases, Boto3 will not return a normal response dict; it will raise a botocore.exceptions.ClientError. We must separate that scenario from the per-key error scenario.
Case A (200 OK + Error entries): As above, the operation succeeded for some keys (maybe including absent ones) and failed for others. No exception is raised. We handle this by examining result['Errors'] vs result['Deleted'].
Case B (ClientError exception): This is a total failure of the request. Common examples:
A bucket policy explicitly denied s3:DeleteObject on all objects. S3 might return HTTP 403 AccessDenied for the whole request. In our stub harness we can simulate:
stubber.add_client_error(
'delete_objects',
service_error_code='AccessDenied',
service_message='Access Denied',
expected_params=expected_params
)
stubber.activate()
try:
s3.delete_objects(
Bucket=bucket_name,
Delete={
'Objects': [{'Key': k} for k in request_keys],
'Quiet': False
}
)
except Exception as e:
print("Caught ClientError:", e)
finally:
stubber.deactivate()In this scenario, no per-key parsing is possible, because the service never returned per-key data at all. All keys in this case are unresolved: we do not know if any were deleted, because the request was aborted. This is fundamentally different from having a 200 with Errors. We must treat the whole operation as failed and consider next steps. Possibly the developer would fix the permissions (since AccessDenied means “you don’t have authority”), but we do not assume any subset succeeded.
The entire manifest remains to be addressed. The final outcome here is not ACCEPT; it’s either REPAIR (if the request should have been allowed but our client logic made a bad call), or HOLD (if we literally hit a permission wall and need owner intervention).
Case C (Unexpected/malformed response): Suppose neither an exception nor a valid response dict is returned (e.g. due to a network timeout, or stub returning None or an empty string). This is also a case where no reliable per-key result exists. Our logic should preserve this as a separate state. For now, we note it as a system error. We would log it and escalate. In staged negative tests, record “ClientError or no data” and do not run further parsing, because the operation was incomplete. Do not invent missing data: if the stub yields [] or crashes, flag it.
We must emphasize: No exception thrown does not mean “no errors.” The absence of a Python exception just means “service answered with 200.” Even then, there might be error codes in the returned structure. Conversely, catching a ClientError means no keys are resolved. Our client code should look like:
from botocore.exceptions import ClientError
try:
resp = s3.delete_objects(
Bucket=bucket,
Delete={'Objects': objs, 'Quiet': False}
)
# Process resp['Deleted'] and resp['Errors']
except ClientError as err:
code = err.response['Error']['Code']
if code == 'AccessDenied':
print("Request-level AccessDenied")
else:
print("Other request error:", code)
# Handle as overall failure (no per-key results)We rely on the Boto3 error-handling model that all service-side errors come through as ClientError exceptions. Thus, if we catch one, we know the call never returned a normal payload.
Preserve ambiguous completion as an explicit state
If the call timed out or we got a non-200 without an Error list, that is an unknown outcome. We should log it as “DeleteObjects request did not complete cleanly.”
Do not assume the worst (all keys failed) or best (all succeeded); escalate as HOLD. For example, a network glitch causing s3.delete_objects() to hang or a stub returning {} should result in a HOLD flag in our records.
Reject incomplete and contradictory reconciliation ledgers
Before we trust the parsed result, enforce strict ledger validation. Our ledger must have unique, disjoint, complete coverage of keys. Implement checks such as:
No duplicate keys in the request manifest (precheck fail). E.g., if "allowed-key.txt" appears twice in Objects, error out immediately.
No duplicate keys across Deleted or Errors lists. If a key appears twice, reject the response as corrupt.
No overlap: a key cannot be in both Deleted and Errors. If stubbed data had {"Deleted": [{"Key":"X"}], "Errors":[{"Key":"X",...}]}, that is an inconsistent API response; we must fail and REPAIR (or HOLD if discovered at runtime).
Coverage: After parsing, Deleted_keys ∪ Error_keys must equal the set of requested keys. If something is missing or extra (for example, stub erroneously included a key not requested, or left out a key entirely), the logic should detect that and reject the reconciliation. This might happen if a developer passed a wrong bucket name in expected_params, causing stub validation to skip or add unknown keys.
Here's a snippet in our harness code to enforce these rules (conceptually):
requested = set(request_keys)
deleted_keys = {item['Key'] for item in result.get('Deleted', [])}
error_keys = {item['Key'] for item in result.get('Errors', [])}# Unique/disjoint check
if deleted_keys & error_keys:
raise RuntimeError(
"Key in both Deleted and Errors: {}".format(
deleted_keys & error_keys
)
)
if deleted_keys | error_keys != requested:
raise RuntimeError(
"Mismatch: reported keys {} vs requested {}".format(
deleted_keys | error_keys, requested
)
)If any check fails, we do not proceed to calling HEAD. Instead, we flag REPAIR (the code that parsed or called the API must be fixed).
Let’s outline some negative test cases we would precheck:
Missing fields: e.g. stub response {'Errors': [...], 'Deleted': None} or empty dict. We should catch this and treat as invalid structure. (The spec requires Deleted to exist as a list, possibly empty. A None or absent key should be flagged.)
Unrequested keys: e.g. stub with 'Deleted': [{'Key': 'extra.txt'}] not in manifest. Fails coverage rule.
Duplicate input keys: if the original Objects had repeated "allowed-key.txt", we should reject before calling at all.
Same key in Deleted and Errors: contradictory output; fail out.
For these first four inconsistencies, we mark them as locally prechecked errors that must be fixed in the deletion logic or manifest (REPAIR). For completeness, we note other possible anomalies, such as a requested key missing from both lists, as still requiring handling. The coverage rule should have caught that case.
We do not silently “repair” by dropping rows. Instead, treat the ledger as unreliable. For example, if stub returns the overlapping key case, drop nothing: log and halt. If any count mismatch occurs, do not proceed to HEAD.
Human intervention is needed to understand why the S3 response did not follow the API contract.
In summary, only a perfectly self-consistent response passes to the next stage. This is similar to how Stubber response validation or external API error handling would reject malformed payloads. Our checks ensure that before we trust the API, it has honored the rules: each key was accounted for once.
Verify current absence with known read authority
Assuming we have a valid ledger now, the final step is to check what’s actually in S3 after the delete. This requires a separate call under a known identity, e.g. an IAM user or role that is allowed to list or get objects in the bucket. For each key in our manifest (the same three we requested), we do something like:
from botocore.exceptions import ClientError
for key in request_keys:
try:
s3.head_object(Bucket=bucket_name, Key=key)
current_state = 'exists'
except ClientError as err:
code = err.response['Error']['Code']
current_state = (
'absent' if code == '404'
else 'forbidden' if code == '403'
else 'error'
)
print(key, "is", current_state)According to the HeadObject API, if the object truly does not exist, and our principal has ListBucket permission, S3 returns 404 Not Found; if we lack list permission, we get 403 Forbidden. We interpret those as:
404 Not Found: the object is absent. Good evidence that a delete took effect.
403 Forbidden: we cannot tell if it’s absent or just not visible. Not treated as success. (Remember: a 403 here isn’t a deletion proof; it usually means “you don’t have permission to see it”, which could be either because it’s there and locked or just hidden. We must escalate or treat this as HOLD.)
200 OK: implies the object still exists. This is a conflict with a previous “Deleted” claim, and means the delete likely didn’t happen or was overwritten. We must treat this as a failure.
We keep track of each outcome. For example, our verification ledger might look like:
Key | Post-Head Result | Decision |
allowed-key.txt | Absent (delete confirmed) | |
denied-key.txt | 200 OK (still present) | Not deleted (expected, file left for deny) |
absent-key.txt | 404 Not Found | Absent (was already absent) |
Here, “allowed-key.txt” yields 404, as desired, so we accept that deletion. “denied-key.txt” yields 200; the file is still there. This matches our expectation that S3 did not delete it, so we’ll plan a RETRY for that key (with permissions fixed) later. “absent-key.txt” yields 404 (or perhaps also 404 by definition); that’s fine; it confirms nothing was there, consistent with the report.
We must explicitly assume strong consistency for this step: after a DeleteObjects (assuming no other writes happened in between), a head should immediately reflect the deletion. AWS documentation states that “after a successful delete of an existing object, any subsequent read request immediately receives the latest version”.
However, this guarantee only holds if no other writer interfered. If another writer (outside our control) could have put an object with the same key between our delete and our head, then the head might show the new object. In that case, we would be seeing the writer’s new content, not the original.
This is why we assume a “writer boundary”: we hold off other writes during this check. If we cannot guarantee that, the result is ambiguous and we must HOLD.
Read-after-delete is not a lock on the key name
Imagine a writer separately does PutObject(Bucket, Key=allowed-key.txt) after we deleted it. The service-level strong consistency only says “head will see the latest object,” which might now be the new one. If that happens, our head_object will return 200 with the new content, which could be misread as “delete failed.”
We do not roll back or restore the old content (unversioned S3 has no backup). Instead, we document that “The key was re-written by another actor; original delete is moot.” This scenario forces a HOLD because we can’t safely repeat a delete (we would be deleting someone else’s new object).
We must notify the data owner about the replacement, since recovering the original unversioned bytes is impossible (unlike a versioned bucket or a queue message retry). In our controlled test lab, we would not simulate this unless explicitly testing the logic path; but the article’s point is to distinguish new writes from the old target.
Finally, consider the denied key. If head_object(Bucket, 'denied-key.txt') returns 200, that’s exactly what we want (it remains). If it returned 404, that would be surprising: it would mean S3 might have let it go even though it was denied? Unlikely in normal policy semantics, but it would indicate a potential bug or an eventual-consistency delay.
We should be ready: if the denied key somehow disappears (404), we do not assume “successful delete” because we never asked for it. That would instead be an out-of-band deletion. We should treat that as an anomaly (REPAIR or HOLD).
In short, the post-delete read step provides independent evidence:
If a key was reported deleted (or absent) and we see 404, that is consistent.
If a key was reported deleted but we see 200, that’s a logical contradiction (HOLD).
If a key was reported errored (not deleted) and we still see 200, that’s consistent with “delete was denied” (safe).
If an errored key yields 404, this means someone else deleted it; that’s a business-policy violation (someone deleted data when they shouldn’t). The normal safe action is HOLD and escalate.
At the end of this section, we either have all 404s for keys that should be gone, or we have identified a problem. Only if each “Deleted” key is now absent (404) and each “Error” key is present (200 or 403) in a way that matches the issue can we be comfortable moving on.
Approve a failed-subset retry rather than a whole-batch replay
Assuming we found exactly one type of failure, the “denied-key” case, we now have a partial success: the allowed-key is gone, denied-key is still present, absent-key remains absent. The bulk delete call partially succeeded (2/3) and partially failed (1/3).
The question: what next? We do not re-send the entire manifest to delete again. That risks re-deleting allowed-key if it was somehow re-created. We only want to retry the denied key, and only after its situation is remedied.
Concretely: to retry denied-key.txt, the data owner must explicitly fix the permission issue (e.g. update the IAM policy or ACL to allow deletion of that key) AND must confirm “no new object has been written to that key since we last checked.” Only then should we issue a new DeleteObjects for just that key. We do it as a separate operation, not part of an unbounded loop.
E.g., once the policy is corrected, one might run:
s3.delete_objects(
Bucket=bucket_name,
Delete={
'Objects': [{'Key': 'denied-key.txt'}],
'Quiet': False
}
)Quiet mode could also be used for a one-key delete, though that is unnecessary. A single-key DeleteObject would suffice too; this example uses DeleteObjects to keep the bulk-delete workflow consistent.
Importantly, because we did not use versioning, the new call is not idempotent. If another write happened to denied-key.txt between now and the fix, then deleting it would remove the wrong object. Therefore we must have “writer control” again.
If we cannot ensure that (for example, the denied key owner might have written something else), then we should HOLD the situation and not retry blindly. The only fully safe context to retry is when we know exactly the object we’re deleting is the one that failed last time.
After retry, we again verify with HEAD. If it returns 404, ACCEPT that key’s delete now too, and include it in the final evidence record.
If it still fails (odd, but maybe the retry hit a new denial or condition), we would escalate. But ideally, after fixing the denials, the retry either succeeds or raises an exception (e.g. another AccessDenied), in which case we mark HOLD or REPAIR as needed.
Notice we explicitly avoided re-deleting the allowed-key. It’s already gone, so reissuing it would create an “object-not-found” success, which would artificially inflate our success count. That’s misleading because a new DeleteObjects would get all keys minus the allowed one, potentially returning “absent-key” (which we know is absent) as Deleted and ignoring the possibility that something new appeared. Keeping a minimal retry subset prevents that error.
So the rule is: Retry only those keys that previously errored, and only under controlled conditions. This is a narrower operation. It’s akin to best practice in external API handling.
We are not implementing an exponential backoff or multiple automatic attempts; we do exactly one authorized retry with one key, under supervision. The rest (allowed-key) remain marked done.
If there were multiple failed keys, we would similarly submit a DeleteObjects with that sub-list. But here it’s one by design.
Handle replacements and mistakes without a rollback promise
In some cases, you may find that you cannot safely proceed. If any key’s deletion semantics are not clear, escalate. For example, if the post-check revealed that a key reported as deleted is still present (200) or returns 403, you have conflicting evidence. Do not try to “fix” it by re-adding something. Instead, HOLD the operation and bring in the data owner or ops team.
The record you hand off should clearly show: “Key X expected deleted but is present/forbidden.” That indicates a violation of assumptions (maybe someone else inserted new data, maybe the policy logic changed concurrently, etc.).
In our lab scope, if we encounter such a contradiction (which we didn’t simulate with the stub), we simply record it as HOLD. We would not attempt any rollback script (because for unversioned, rollback is impossible unless you had snapshots or backups outside S3).
After the exercise is done (or if we reset the test), we can re-create our synthetic objects by a PUT if needed. But it’s crucial to state: a fresh PUT to recover objects is not an “undo” of a failed deletion. It’s creating new data. If this were a real failure, a data recovery process might look different (e.g. restore from backup or versioned bucket). Our lab PUT only resets the test state for the next run; it is not part of normal retry logic.
A stable key string is not an immutable generation identifier
If a user deletes an object, then later re-uploads a file with the same key, the name is the same but the content and identity are different. S3 does not prevent reusing key names in unversioned buckets. Unlike versioned buckets or Azure Blob leases, there’s no built-in “protect this name until further notice.”
In tests, you could simulate a replacement by having the allowed-key come back (in a controlled way). But we do not treat that as an automatic rollback. We either identify it as a new object (if we can, via metadata differences) or treat it as a sign to HOLD. (Versioning would handle this differently by giving new version IDs; but that’s a separate feature set.)
The key takeaway is that once you lose the old object, only external backup or logs could restore it; the S3 API itself doesn’t.
Make the ACCEPT, REPAIR, RETRY or HOLD decision
Now we synthesize all evidence. The final state machine has four outcomes:
ACCEPT: Achieved when every key’s outcome is understood and consistent:
All requested keys appeared in the final report (unique/disjoint),
Any key listed as Deleted has been verified absent now,
Any key with Error is still present (or at least not deleted),
No unknown or missing responses.
The writer boundary is respected (we held or isolated writers). In short, we can confidently close the batch. In ACCEPT, the operation is done; one can safely say “these objects have been removed as intended.” The code and report were correct; evidence matches reality.
REPAIR: Chosen if we detect a flaw in our own processing or in the API report format. Examples: duplicate or missing keys, same key in both lists, unexpected schema, or even a bug in the local validation (like misparsing quiet mode).
In practice, REPAIR means “fix our deletion-handling code or the input before proceeding.” It does not mean doing another delete against AWS immediately. For instance, if Deleted_keys ∪ Error_keys didn’t match the manifest, that’s our client logic bug, so we debug it. Once repaired, the developer can re-run the test harness. The underlying S3 call was likely correct (or simulated), so no further AWS action is taken until this is fixed.
RETRY: Applicable when some keys failed but others succeeded, with no contradictions. In our scenario, that means the only failing keys are those in Errors and those errors are clearly re-runnable (e.g. AccessDenied due to fixable IAM).
Here we do one more delete-only-those-keys, as described above. We keep the previously deleted keys out of that request. We also keep the original successes and evidence on record (so the final report will say e.g. “2 of 3 removed; retried 1 more”). After retry, we re-verify absence. The final ACCEPT can only happen if the retry succeeds (i.e. the errored key is now also gone). Until then, we treat the original delete as partially done.
HOLD: Used when something unexpected or unresolved is present. Examples:
The absence check returned 403 or some status where we cannot conclude absence/presence.
Contradictory evidence (e.g. head says present for a Deleted key).
A request-level error (ClientError) that cannot be immediately fixed (maybe an intermittent network issue).
Any scenario outside the narrow success path. HOLD means do not dismiss the operation. Document the evidence in detail, notify stakeholders, and leave it for someone with the appropriate authority to make a call or investigate further. It might later transition to RETRY or REPAIR after human action.
Here’s a simplified decision matrix tying conditions to actions:
Condition | Action |
No errors; all Deleted keys now absent | ACCEPT |
Mixed success/errors; all reported deletions absent; failed keys eligible for retry (permissions fixable) | RETRY (subset) |
Missing/overlapping keys in report | REPAIR |
Presence of reported-deleted key (HEAD 200) or HEAD 403 on key without list permission | HOLD (requires owner review) |
Total request denied (ClientError AccessDenied) | HOLD/REPAIR (fix policy or caller) |
The HeadObject documentation explains how to interpret 403 vs 404 on head, which directly influences HOLD vs ACCEPT for absence. The Boto3 error-handling guide reminds us that a thrown 403 (ClientError) is a service-level error (classify as HOLD/REPAIR, not a per-key error).
The gist: Don’t close the book on this delete until the evidence is consistent. If the API stub says everything’s OK but the ledger has holes, do not accept it (call for REPAIR). If only one key failed and we have control, we allow a retry on that small piece (and keep the others accepted).
We never automatically retry an entire batch because that changes semantics. If there’s any doubt left (unknown head status, unfixable error, key reuse uncertainty), we lock it down and escalate (HOLD).
Finally, note that an accepted client logic run (in Stubber) does not itself imply an accepted real delete. The stubber validated our local parser branches, but the real-world state is still separate.
In any reporting or logging, include a flag that says “Verified under Stubbed simulation.” A real delete done by AWS would need its own evidence.
Hand off a deletion record that another operator can review
When the operation is closed (whether ACCEPT, RETRY, or HOLD), we must produce a complete record for someone else to audit. This is akin to handing off a completed pipeline job in a cloud-native workflow. The record should include:
Request Manifest Digest: A hash (or copy) of the original key list and prefixes, bucket name, and manifest ID. This ties to the exact keys approved for deletion.
SDK/Client Details: The full AWS SDK version (boto3, botocore), the region and endpoint used, and any relevant flags (e.g. quiet mode on/off). This helps reproduce the run if needed.
Raw Response: The JSON/XML body as returned from AWS (or stub). If quiet mode, note the mode. If a ClientError, log its full response['Error'] code/message and metadata (request ID).
Normalized Outcome Table: The reconciled per-key ledger we built (as above, showing for each key: “Requested -> Reported Deleted? -> Error? -> Verified Absent?”). This is the core of the evidence.
Post-State Checks: For accepted deletes, record each head_object status (404, 200, etc.) along with timestamps. Include the identity under which the head was performed.
Retry Approval Info: If RETRY was done or will be done, note what permission change happened and who authorized the retry. Also note the subset of keys retried.
Uncertainty/Next Steps: If HOLD, clearly list what is unresolved (e.g. “denied-key returned 403 on head; cannot confirm deletion; needs owner resolution”). If ACCEPT, state that all checks passed.
Request IDs & Logs: Optionally, include AWS request IDs (from the response metadata) if this were real. (Our stub logs are placeholders.) In practice, these IDs help AWS support trace issues.
Time stamps: Mark when the delete request was issued and when the verification checks ran, to anchor against writes.
Authorization Identities: Note which AWS IAM identities were used for delete vs head. For instance, “Delete executed by IAM user dev-batch-deleter; HEAD done by IAM user dev-data-verifier.” If they differ, highlight that separation as part of the security boundary.
This “deletion record” should then be passed to whichever system or person needs it. It is not a simple success/fail signal; it’s documentation. One might include it in a ticket, or attach it to audit logs.
In an automated cloud data pipeline, you could imagine this being a JSON artifact that downstream quality-check tasks ingest. (For further reading on pipeline handoffs, see Refonte Learning’s Cloud Native Data Engineering guide. Preserve output and context when one step finishes.)
Revalidate when the SDK, policy or bucket contract changes
A cautionary note: this entire process depends on our environment (SDK behavior, IAM policies, bucket type) staying the same. If tomorrow we upgrade boto3, or change the bucket to have versioning, or alter the IAM role, we should rerun the checks.
For example, if a policy change now denies Quiet mode or expects an extra header, our parser might break. Also, if bucket versioning is accidentally turned on, DeleteObjects will add markers. All of this can invalidate the assumptions in our ledger.
In those cases, we would mark the old record invalid and start a fresh reconciliation under the new conditions. In other words, we treat the test as contract-bound: any change in the contract triggers a new agreement step.
Develop cloud skills around verifiable operation results
This entire exercise, reconciling a multi-object delete down to each object and state, exemplifies the kind of thorough cloud development and DevOps mindset you would learn in a structured training program. It requires understanding AWS service contracts, careful coding of idempotent or at-least-once operations, error-handling for external APIs, and clear documentation.
If you are interested in systematically building these skills, consider the Refonte Learning Cloud Development Program. It is a 3-month, 12–15 hour/week course that covers cloud architecture design, containerization (Docker, Kubernetes), infrastructure as code, security fundamentals, monitoring, and DevOps practices.
Such a program can help solidify best practices like the ones demonstrated here, ensuring that you not only write cloud code, but also verify its outcomes with evidence.
In summary, before accepting a bulk delete operation as done, we must cross-check every key. The acceptance criteria include correct per-key reporting, verified absence or presence, and controlled retry logic. With this playbook, a team can turn a single S3 API call into a fully auditable deletion record, meeting both operational and compliance needs without guessing at any object’s fate.
