Cloud engineer reviewing S3 multipart upload cleanup and checking for remaining parts at a workstation.

S3 Multipart Upload Cleanup: Verify the Remaining Parts

Thu, Oct 1, 2026

In a cloud data pipeline, an operator might issue an abort for a stalled or abandoned multipart upload, only to realize afterward that it’s not guaranteed all parts disappeared. Amazon S3 does not automatically complete or abort multipart uploads; each upload’s parts are charged until explicitly freed. Moreover, an abort request itself only signals S3 to free storage, but in-flight part uploads might still succeed and leave fragments behind. Before we can “close the case” on a given upload (the exact tuple of bucket, object key and UploadId), we must verify several things: no new parts arrived, all in-progress parts have finished, and each remaining part is gone. This article defines four possible outcomes of a cleanup decision (close, retry, hold, escalate) and how to test for them. It emphasizes that the unit of work is a single UploadId, not a prefix or file name. For illustration, we use a disposable S3 bucket in us-east-1 with known settings and fixed credentials (region, CLI version, retry policy, and a narrow IAM role). The fixture creates one completed “sentinel” object at a key, two concurrent multipart uploads (target and control) at that same key, and one upload at a different key. By tracking all issued requests in a controller ledger, stopping producers, and fully enumerating S3 listings, we produce auditable evidence before deciding if the upload is truly cleared.

The key question is "Did the abort actually clear all parts for that specific upload ID?" A mere HTTP 204 from AbortMultipartUpload is not enough. Instead, we need to collect end-to-end evidence: stop any upload workers, walk through all pages of ListMultipartUploads and ListParts for that exact (bucket, key, UploadId), and compare with the producer’s log of parts. We treat a successful abort as just one step; only an empty parts list and absence of the UploadId from listings (with controls intact) can let us confidently close. If we find stray parts, a hidden page, or contradictory signals, we must either retry with bounds or hold for human review. Throughout, we avoid assumptions about physical deletion or billing; we only rely on the S3 API’s view. Finally, we document a machine-readable manifest of the fixture, worker chronology, raw responses and a decision matrix, so another engineer can independently confirm the cleanup verdict. This approach turns an ambiguous “success” signal into a precise validation of S3’s multipart-abort contract.

Define what an upload cleanup decision must prove

An aborted multipart-upload cleanup routine must conclusively prove that exactly the targeted upload (bucket, key, UploadId) has no remaining parts and is no longer in progress. This means first identifying the target upload fully: not just by key name (since multiple UploadIds can share a key) but by the unique UploadId string returned at initiation. We also must ensure no new parts are being submitted to this upload. Practically, this requires a controller ledger: before any requests are sent, record each CreateMultipartUpload (giving UploadId), each UploadPart (part number, size, ETag) and their timestamps. That way, after stopping the upload, we can compare API listings against this authoritative record. Equally important, we must stop any upload workers or clients from sending further parts to any upload (target or controls) at the same key. Since AWS allows parts to trickle in after a stop request, we explicitly wait (join threads or processes) until all in-flight uploads finish. Only then do we call abort. This separates the “producer side” work from the “cleanup verification” phase.

After calling abort, the controller must inspect the S3 state. We query all multipart uploads in the bucket to confirm the target UploadId is absent (and note any remaining uploads for controls). We then call ListParts on the target ID. A truly cleaned upload should yield either an empty parts list or a NoSuchUpload error (indicating no parts and no record of the upload). We also verify that the pre-existing sentinel object (the completed object we placed at the key) is untouched: its bytes or checksum should match what we recorded. This distinguishes the cleanup of parts from any object lifetime issues. (Lifecycle rules or billing are separate topics: we are not measuring physical erasure or final billing here.)

In summary, the cleanup check must prove: the intended upload is gone, no parts of it remain, producers have stopped, and all other objects/uploads (controls) are intact. These prove an operational close-out, not just a storage cost or compliance event. The independent controller ledger (the recorded requests and results) is crucial: it lets us catch any listing or race anomalies. If the ledger shows a part was supposed to upload but still appears in ListParts, cleanup has failed. If the target UploadId still appears in ListMultipartUploads, it hasn’t been fully aborted. By requiring concrete API evidence against our ledger, we treat “cleanup success” as a conditional acceptance criterion, not as an inherent S3 promise. Only when every check passes (and controls are unaffected) do we declare the upload cleared. This disciplined approach fits into any robust cloud data-pipeline architecture, where each stage’s outcome must be independently verifiable.

Pin the bucket, client and permission baseline

First, we establish a fully specified environment for repeatability. For example, we might choose AWS Region us-east-1 and a unique bucket named example-bucket-refonte-mpu-lab-12345 (randomized to avoid collisions). We disable bucket versioning and set no server-side encryption, so that objects overwrite normally and parts behave predictably. The AWS CLI is pinned to a specific version (say, aws-cli/2.9.1 Python/3.11.2 on Linux) and boto3 to version 1.26.100 (botocore 1.29.100). These versions should be noted at the top of logs or code. In the AWS CLI config (~/.aws/config), keep default retry settings (5 retries, 60s read timeout) or specify them if the script changes them. For clarity, we record the exact AWS principal in use (e.g. IAM role arn:aws:iam::123456789012:role/MPUCleanupRole) which has a scoped policy allowing only the S3 actions needed on this bucket: s3:CreateMultipartUpload, s3:UploadPart, s3:ListMultipartUploads, s3:ListParts, s3:AbortMultipartUpload, s3:GetObject, and lifecycle checks. No wildcard or administrative permissions are used. The lab prefix (e.g. lab-run-20260930-01/) embeds a UTC timestamp for uniqueness. Every output timestamp in the test is recorded in ISO8601 UTC format (e.g. 2026-09-30T12:00:00Z).

Separate the upload identity from the object key

It is important to note that S3 keys and upload IDs are distinct identifiers. Two multipart uploads can share the same bucket and key, but have different UploadId values and different states. For example, our fixture will have a target upload at key lab-run-.../data.bin with UploadId TARGETID, and a same-key control upload also at data.bin but with UploadId CONTROL1ID. Even though they share the key, aborting the target’s UploadId does not affect the control upload (and vice versa). AWS’s ListMultipartUploads sorts by key and then initiation time when listing multiple uploads under the same key. Likewise, S3’s CompleteMultipartUpload will, if versioning is disabled, simply replace the object at that key if called. In contrast, our completed sentinel object at that key is an actual S3 object (as if a previous multipart had completed). This object stands apart from any in-progress uploads. When verifying cleanup, we must treat the object presence separately: we must preserve the sentinel object’s data and not confuse it with upload parts. In other words, the cleanup routine should not delete or modify the key, only abort or list the specific UploadId. A completed object at the key (sentinel) is outside the lifecycle of multipart parts and won’t be targeted by an AbortMultipartUpload (which affects only the identified upload). In summary, our acceptance logic focuses on the specific UploadId (not the key alone); we verify that the UploadId is cleared while all other keys/versions remain as before.

Create a small fixture with controls worth preserving

We now build our test fixture using AWS CLI and synthetic parts. For brevity we show key steps (actual scripts would be more verbose). First, create a small sentinel object:

# Create a completed "sentinel" object to preserve:
printf 'Sentinel data' > sentinel.txt
aws s3api put-object --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin --body sentinel.txt

This writes a known object at lab-run-20260930-01/data.bin. We record its size (e.g. 13 bytes) and ETag (say "e3b0c44298fc1c14..."). Next, create the multipart uploads:

# Create target multipart upload on same key "data.bin"
target_upload_id=$(aws s3api create-multipart-upload \
    --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin \
    --query UploadId --output text)

# Create same-key control upload at same key
same_key_upload_id=$(aws s3api create-multipart-upload \
    --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin \
    --query UploadId --output text)

# Create other-key control upload at a different key "data2.bin"
other_key_upload_id=$(aws s3api create-multipart-upload \
    --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data2.bin \
    --query UploadId --output text)
These calls return three UploadIds (e.g. TARGETID, CONTROL1ID, CONTROL2ID) which we save in our controller ledger. We now upload parts to each:
# Prepare two small parts (6 MiB and 5 MiB) for the target
dd if=/dev/zero bs=1M count=6 of=part1.bin
dd if=/dev/zero bs=1M count=5 of=part2.bin
# Parts for the same-key control (6 MiB)
dd if=/dev/zero bs=1M count=6 of=control1_part1.bin
# Part for the other-key control (6 MiB)
dd if=/dev/zero bs=1M count=6 of=other_part1.bin

# Upload parts to target upload (UploadId TARGETID)
aws s3api upload-part --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin --part-number 1 \
    --body part1.bin --upload-id $target_upload_id
aws s3api upload-part --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin --part-number 2 \
    --body part2.bin --upload-id $target_upload_id

# Upload part to same-key control (UploadId CONTROL1ID)
aws s3api upload-part --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin --part-number 1 \
    --body control1_part1.bin --upload-id $same_key_upload_id

# Upload part to other-key control (UploadId CONTROL2ID)
aws s3api upload-part --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data2.bin --part-number 1 \
    --body other_part1.bin --upload-id $other_key_upload_id

Each upload-part returns an ETag in the response. We record each part’s number, size (in bytes) and ETag in our ledger. For instance, the target had Part 1 = 6291456 bytes (ETag "etag1"), Part 2 = 5242880 bytes (ETag "etag2"). The controls similarly have known part uploads. At this point the precondition holds: the bucket contains a completed object data.bin with the sentinel content, and three in-progress uploads (two at data.bin, one at data2.bin) with the above parts. Bucket versioning is off, so these parts are isolated from the sentinel object. We’ve kept the fixture minimal and self-contained.

Demonstrate the one-abort false green

A naive cleanup script might simply issue AbortMultipartUpload on the target UploadId and assume “all done” if it returns success. For example:

aws s3api abort-multipart-upload --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin --upload-id $target_upload_id
# If this returns 204 No Content, the script might stop here.
echo "Abort returned $?"  # 0 means HTTP 200-level in AWS CLI terms

This alone is insufficient. AWS itself warns that aborting an upload does not guarantee all in-flight parts are gone; parts that were still uploading might finish after the abort request. In fact, the AbortMultipartUpload API will happily return a success (HTTP 204) even if some parts slip through, because they were already en route. Our ledger would then show two parts issued, but if we rely solely on abort status, we haven’t confirmed S3 removed them.

To illustrate: suppose the abort call succeeded in turning off that upload, but one part was still uploading in the background. The abort operation frees storage only after all parts associated with the upload are finalized. If we ignore listing, we might get a false positive “clean” while there’s a leftover part or a subsequent new multipart. Therefore, a valid acceptance must query S3 after abort to confirm emptiness, as AWS recommends.

Keep the race experiment observational

In testing, we should run producers (the upload-part calls) and the abort in parallel to reflect real concurrency. For example, we might start the part uploads in background threads or processes, then issue the abort in the main thread without waiting a fixed time. Crucially, we log the actual chronology: when each part upload started and finished relative to the abort invocation. This makes the experiment observational rather than contrived. We might use threading or background processes and record events like:

[2026-09-30T12:00:01Z] ProducerThread1: Uploaded part 1 for target
[2026-09-30T12:00:02Z] ProducerThread2: Uploaded part 1 for control1
[2026-09-30T12:00:03Z] Controller: Issued AbortMultipartUpload for UploadId TARGETID
[2026-09-30T12:00:04Z] ProducerThread1: Upload part 2 for target succeeded

This log (part of the controller ledger) shows that one of the target’s parts completed after the abort call, exactly the race AWS warns about. We do not artificially wait for all parts to finish before aborting; that would hide the race effect. Instead, we let the abort race against any remaining uploads and carefully note what happens. If a post-abort part arrives, our later reconciliation (ListParts) will catch it. The observation is used only to understand behavior; we still enforce stopping uploads before final verification, but the raw timeline is preserved for audit.

Stop the producer before trusting the cleanup ledger

Once the producers (uploader workers) have started, the controller must prevent any new upload attempts before verifying cleanup. In practice, this means signaling the upload threads to cease issuing new UploadPart calls, then joining them (waiting for completion). For example, in a script:

# Pseudocode for thread synchronization
for t in all_upload_threads:
    t.join(timeout=60)          # wait up to 60s
    if t.is_alive():
        raise Exception("Upload worker did not terminate")

This ensures that all started upload attempts have either finished successfully or errored out. We consider it an unrecoverable error (escalate) if any worker does not exit: it may indicate hanging or a logic bug. We do not use a fixed sleep, because that could end before all parts finish, violating correctness. Instead, the controller acts as a barrier: no further cleanup steps run until producers are fully quiescent. Only after all threads have joined do we proceed to abort and inspect. This strictly enforces AWS’s guidance that an abort should happen after parts complete.

At this point, every part upload has a final status. The ledger contains whether each UploadPart succeeded or failed, with timestamps. We trust these results when reconciling. Stopping the producers in this way separates the generation of data from the cleanup decision. We never race new parts into the system while cleaning up: once producers are stopped, the S3 state can be treated as static (except for the abort calls we make).

Enumerate uploads without losing a page or a same-key entry

Now we collect the S3-side inventory of multipart uploads in our test prefix. We use the S3 ListMultipartUploads API with a low MaxUploads (for example, 1 per page) to force pagination. For instance:

# First page
aws s3api list-multipart-uploads \
    --bucket example-bucket-refonte-mpu-lab-12345 \
    --prefix lab-run-20260930-01/ --max-uploads 1 \
    --query "Uploads[0].[Key,UploadId]" --output text
# --output example: "lab-run-.../data.bin   TARGETID"

If the result indicates "IsTruncated": true, we read the returned NextKeyMarker and NextUploadIdMarker and call again:

# Second page using markers from previous output
aws s3api list-multipart-uploads \
    --bucket example-bucket-refonte-mpu-lab-12345 \
    --prefix lab-run-20260930-01/ --max-uploads 1 \
    --key-marker lab-run-20260930-01/data.bin \
    --upload-id-marker TARGETID \
    --query "Uploads[0].[Key,UploadId]" --output text
# e.g.: "lab-run-.../data.bin   CONTROL1ID"

We repeat until IsTruncated is false. By using --prefix and paired markers (--key-marker, --upload-id-marker), we collect every upload entry in our range. In our fixture, the listing should produce three entries across pages: one for the target upload (at data.bin), one for the same-key control (also data.bin with a later initiation time), and one for the other-key control (data2.bin). We store all raw listing pages in the manifest. It is vital that we retain all pages and marker values; dropping pages (or misusing markers) can hide an upload. For example, if a script only read the first page, it might see only one data.bin upload and miss the other data.bin control, leading to an incomplete view. By following the documented pagination (with NextKeyMarker and NextUploadIdMarker), we ensure no upload is lost.

Test pagination with the smallest useful population

To prove our pagination code works, we deliberately set --max-uploads 1. With three total uploads in the fixture, this yields three API calls as shown above. If we had only taken one page (or used --max-uploads 2 and stopped prematurely), the target upload might or might not appear, depending on sort order. We verify that our loop sees all three uploads, matching the ledger’s count. As a sanity check, we can also run a truncated example: a buggy implementation that ignores pagination will quickly show a mismatch. For instance, calling with --max-uploads 2 would return only two of the three uploads, causing reconciliation against our ledger to fail (we’d detect an unseen upload). This exercise confirms that correct handling of IsTruncated, NextKeyMarker and NextUploadIdMarker is essential to a full inventory.

Reconcile parts against the exact target upload

With the upload listings in hand, we identify the target UploadId of interest (as recorded in our ledger, say $target_upload_id). We then invoke ListParts for that upload. Using a similar pagination strategy (MaxParts), we gather all parts still present under that UploadId. For example:

aws s3api list-parts \
    --bucket example-bucket-refonte-mpu-lab-12345 \
    --key lab-run-20260930-01/data.bin \
    --upload-id $target_upload_id --max-parts 1 \
    --query "Parts[0].[PartNumber,Size]" --output text

If IsTruncated is true, we repeat with the NextPartNumberMarker until done. Suppose we get output like:

1    6291456
2    5242880

This shows Part 1 and Part 2 (and no more). We then compare this list to the ledger’s record of parts we intended to upload for the target (two parts of 6291456 and 5242880 bytes). If they match exactly, there are no unexpected parts. If ListParts had returned only one part or returned extra parts not in the ledger, we flag an unresolved discrepancy. For example, if ListParts had returned just Part 1 and we expected two, perhaps the second part never arrived; or if it returned an additional Part 3 (unexpected), that would indicate data from another source. Any mismatch or error is recorded explicitly. We do not use ListParts to authorize deleting anything else; it’s simply a read. Importantly, we do this only for the target’s UploadId; we do not reconcile the control uploads here. The controls’ parts should still exist by design, but they are inspected separately for control purposes.

By the end of this step, we either have an exact match (all parts accounted for in ledger) or a concrete indication of leftover parts. The absence of parts in ListParts (empty result) would mean the upload has no parts left in S3; but we must interpret an empty list carefully (see next section). The reconciliation ensures no hidden parts remain under the target UploadId beyond what we know, confirming partial or full cleanup.

Classify terminal responses before closing the case

Now we interpret the S3 API responses to decide whether to close, retry, hold, or escalate. We consider each possible terminal result from our final list operations:

Observation

Interpretation / Required Checks

Action

Non-empty ListParts

One or more parts are still present under the target UploadId (maybe after abort). Indicates cleanup did not fully remove them. Check against ledger to see which remain.

Retry (abort again)

Empty ListParts

S3 returned a successful page with no parts. This suggests the upload has no data blocks left. If the upload still appears in ListMultipartUploads, we must retry abort. If the upload is gone from listings, this is evidence (but see NoSuchUpload).

Possibly Close if corroborated by next step; otherwise Retry.

ListParts → NoSuchUpload

The upload ID does not exist (it’s not found as a target of multipart). This can be normal after an abort: S3 indicates the upload record is gone. However, a NoSuchUpload could also happen if the upload was completed by a client before aborting. We must check: did our ledger record a completed upload? Did we see the sentinel object updated? Usually no, since we did not complete. If the sentinel is intact and our abort was sent, this likely means the abort cleared the upload.

If confirmed by sentinel control and absence from ListMultipartUploads, treat as Close. If uncertain (e.g. sentinel changed unexpectedly), Hold for review.

AccessDenied / Forbidden

We lack permission to list parts. This is a permission issue, not a storage state. The test cannot proceed until fixed.

Escalate (permission misconfig)

Timeout / Network Error

Transient failure querying S3. The clean decision is unknown.

Retry (bounded)

Target still in ListUploads

If after our abort attempts, ListMultipartUploads still shows the UploadId, S3 hasn’t finished aborting or abort failed.

Retry or Escalate after retries.

Unexpected UploadId

If we find an upload at our key with a different UploadId not in our ledger, it may indicate a new upload started inadvertently. Since our contract is only about the original UploadId, we do not abort the new one; but its presence violates our assumption of no new uploads.

Escalate (ambiguous: investigate source)

Crucially, absence of parts is not the whole story; we must correlate it with the abort vs completion distinction. For instance, a successful empty ListParts (above) might mean "all parts freed", but only a NoSuchUpload plus confirmation that no data appears at the key can really prove it. Therefore, we require multiple checks: the upload must not appear in ListMultipartUploads, ListParts must be empty or NoSuchUpload, the ledger must match, and the sentinel object must be unchanged. If any check fails or contradicts, we do not immediately close. Instead, we either retry aborting or hold the case. The table above summarizes how each result drives our decision.

Treat absence and its cause as separate questions

Note that an empty ListParts and a NoSuchUpload error both indicate "no parts found," but for different reasons. An empty list (with HTTP 200) means the upload exists but has zero parts currently. A NoSuchUpload (HTTP 404) means the upload record itself is gone. In both cases we must ask why: did the abort work, or did the upload complete in another way? For example, if NoSuchUpload occurs without us having aborted it (impossible here, but hypothetically), it could mean the upload was already completed or purged. In our controlled test, a NoSuchUpload after we called abort (and with no new CompleteMultipartUpload in the logs) generally means the abort was processed and removed the upload record. But to be safe, we verify the sentinel object remains and the completion history shows no CompleteMultipartUpload. In other words, absence of parts alone cannot tell the full story; we separate “upload gone” from “upload was aborted versus completed” and require our ledger to resolve the cause. Only when we reconcile that it must have been our abort and not a stray completion do we treat it as success. Otherwise we hold for manual review of the object and logs.

Retry cleanup within a bounded reconciliation loop

If initial verification finds residual parts or inconsistent signals, we retry the abort-and-verify sequence, but within fixed bounds. A possible pseudocode loop is:

max_attempts = 3
for attempt in range(max_attempts):
    # Stop producers and ensure quiescent (done earlier)
    # Re-list uploads
    uploads = list_all_multipart_uploads(bucket, prefix)
    if target_upload_id not in uploads:
        parts = list_parts(bucket, key, target_upload_id)
        if parts == [] or parts == "NoSuchUpload":
            result = "clean"
            break
    # If target still listed or parts remain, try abort again
    aws_abort_multipart_upload(bucket, key, target_upload_id)
    sleep(1)  # small backoff between retries
else:
    # Exceeded attempts without satisfying clean condition
    result = "incomplete"

Key points: we use the exact target_upload_id throughout; we do NOT grab a new UploadId if one appears. If a completely new UploadId at the same key shows up (meaning somehow another concurrent upload started), the script must escalate: by definition our contract is only to clean the original UploadId. Also, we enforce a final attempt limit (here 3 abort calls). After each abort, we re-run ListParts (and optionally ListMultipartUploads). If at any point ListParts is empty or returns NoSuchUpload and the upload is no longer listed, we break and close. If after max_attempts this never happens (perhaps parts persist due to some S3 glitch), we escalate for human investigation. This loop ensures we don’t silently give up; if parts remain, we try again, since AWS notes that multiple aborts may be needed. But it also prevents an infinite loop. Importantly, at no point do we switch our target to a new UploadId; doing so would risk cleaning the wrong upload. Any appearance of an unexpected UploadId is treated as a separate issue to be escalated, not silently swallowed.

Use lifecycle rules as a separate safety net

As a sanity check (not a substitute for our verification), we may also configure the S3 AbortIncompleteMultipartUpload lifecycle rule on the bucket. For example, we could apply a rule like:

{
  "Rules": [
    {
      "ID": "AbortStaleUploads",
      "Status": "Enabled",
      "Filter": { "Prefix": "" },
      "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
    }
  ]
}

Once applied with aws s3api put-bucket-lifecycle-configuration, running aws s3api get-bucket-lifecycle-configuration --bucket example-bucket-refonte-mpu-lab-12345 returns the configured rule. This ensures that any multipart upload older than 7 days is eventually aborted and its parts deleted automatically. However, this rule is asynchronous and time-based; it does not instantly clean up the specific upload we care about. It also only applies to incomplete uploads (as intended) and does not touch any completed objects. Thus, the lifecycle rule is merely a backup to reduce long-term orphaned parts, not a proof of immediate cleanup. For our one-hour test we expect it to have no effect. We keep it off or ignore it in our logic, because relying on DaysAfterInitiation in real time would be incorrect. In summary, lifecycle rules belong to long-term cost management, but do not replace the direct evidence needed for the immediate cleanup contract.

Do not confuse an upload rule with object expiration

It’s crucial to remember that the AbortIncompleteMultipartUpload lifecycle action only targets multipart upload parts, not objects. In particular, it will never delete the sentinel object at data.bin because that object is already complete. Likewise, an ordinary S3 expiration rule (for objects) is a separate mechanism and irrelevant here. Our cleanup contract concerns only the parts of the unfinished upload. In our disposable bucket, we show the current lifecycle rules (if any) and confirm that the sentinel object’s key is unaffected. The upload rule’s eligibility window (days after initiation) is not used as a deadline for our test. It’s simply another policy in the bucket’s config, unrelated to our synchronous check.

Recover from wrong-target or ambiguous cleanup

If during verification we discover that we aborted the wrong upload or the results are contradictory, we stop the automated process immediately. For instance, if we meant to abort TARGETID but evidence shows parts still under TARGETID, we must consider that a failure. We then preserve all our controller state (the full ledger of requests and responses). There is no “undo” for an abort: parts once freed are gone, and if necessary a new upload with the same key would get a fresh UploadId. Thus, if we aborted by mistake (say we used the control’s UploadId), that upload’s parts are lost and the target remains uncleared. In this case we would escalate, because our job was to clear a specific upload. The reviewer would examine the logs: check which UploadId got the abort calls, whether the sentinel object changed (it shouldn’t have), and whether either the target or control IDs appear unexpectedly in listings. Importantly, restarting the original producer would create a new UploadId, not resurrect the old one, so there is no automated rollback. The correct recovery often means performing the cleanup again from scratch on the right UploadId (perhaps after re-creating a fresh fixture). All of these steps should be documented in the evidence package if they occur.

Package evidence another operator can review

For accountability, we assemble a comprehensive, machine-readable evidence manifest and a summary matrix. An example JSON manifest (formatted here for illustration) might include:

{
  "bucket": "example-bucket-refonte-mpu-lab-12345",
  "prefix": "lab-run-20260930-01/",
  "region": "us-east-1",
  "timestamp": "2026-09-30T12:00:00Z",
  "sentinel": {
    "key": "lab-run-20260930-01/data.bin",
    "etag": "e3b0c44298fc1c14...",
    "size": 13
  },
  "uploads": [
    {
      "role": "target",
      "key": "lab-run-20260930-01/data.bin",
      "uploadId": "TARGETID",
      "parts_ledger": [
        {"PartNumber": 1, "Size": 6291456, "ETag": "\"etag1\""},
        {"PartNumber": 2, "Size": 5242880, "ETag": "\"etag2\""}
      ],
      "abort_requested": "2026-09-30T12:05:00Z",
      "terminal_status": "aborted",
      "listParts_result": [],
      "listMultipart_result": []
    },
    {
      "role": "same_key_control",
      "key": "lab-run-20260930-01/data.bin",
      "uploadId": "CONTROL1ID",
      "parts_ledger": [
        {"PartNumber": 1, "Size": 6291456, "ETag": "\"etag3\""}
      ],
      "abort_requested": null,
      "terminal_status": "in-progress",
      "listParts_result": [ /* should show 1 part if fetched / ]
    },
    {
      "role": "other_key_control",
      "key": "lab-run-20260930-01/data2.bin",
      "uploadId": "CONTROL2ID",
      "parts_ledger": [
        {"PartNumber": 1, "Size": 6291456, "ETag": "\"etag4\""}
      ],
      "abort_requested": null,
      "terminal_status": "in-progress",
      "listParts_result": [ / should show 1 part if fetched */ ]
    }
  ]
}

In the manifest, terminal_status reflects what the script determined (for target, "aborted"; for controls, "in-progress"). listParts_result and listMultipart_result contain the actual S3 responses (empty lists or NoSuchUpload for the target, listings for controls). We also capture pagination markers and raw JSON of each listing call (omitted here for brevity).

Alongside, we provide an acceptance matrix like:

Condition

Decision

Target UploadId no longer in ListMultipartUploads, ListParts empty or NoSuchUpload, and controls intact

Close

Residual parts found in ListParts

Retry

ListMultipart still shows target, or ListParts returns parts

Retry

NoSuchUpload but sentinel unexpectedly missing/changed

Hold/Escalate

Target UploadId appeared with different ID or unexpected result

Escalate

Any ListParts call AccessDenied/Timeout

Hold/Retry

This matrix explicitly ties evidence to our actions: we will close only when the upload’s absence is fully confirmed; retry if cleanup appears incomplete (known missing parts or still listed); hold if evidence is contradictory or incomplete; and escalate if we detect a wrong target or unexpected failure mode.

Assign ownership for abandoned transfer cleanup

Finally, ownership of this cleanup process should be clear. The team that originally produced the upload (e.g. the data-engineering or application team) is responsible for requesting abort and providing the controller with the right UploadId. The cleanup controller (e.g. platform reliability engineers) runs the verification routine and documents its outcome. A reviewer or auditor (perhaps a QA engineer or team lead) confirms the evidence before closure. Meanwhile, a separate cloud-cost or FinOps owner monitors any financial impact: as FinOps practices emphasize, cloud costs become a shared responsibility, and tagging or assigning cleanup tasks is part of that culture. In this model, if the cleanup script changes or if AWS SDK behavior updates, the owners should re-run and revalidate the tests as a form of regression check. But day-to-day, ownership is simply: producer owns the upload, ops owns the abort/verify process, and FinOps owns the cost accounts. This mirrors mature FinOps guidelines where technical teams and finance teams collaborate on cloud operations. For example, many organizations formalize such roles and insist on tagging or cost-center fields in cleanup logs (see FinOps best practices).

Build cloud-development habits around verified outcomes

Cleaning up abandoned uploads is a great project-based exercise for aspiring cloud engineers. Writing and running these verification scripts builds the same skills covered in a hands-on cloud curriculum: understanding AWS APIs, scripting reliably, and validating state. For instance, the Refonte Learning Cloud Development program (a 3-month, ~12–15 hours/week online course) includes guided projects that have students build and verify data pipelines on AWS. Such project-based data-pipeline practice reinforces the habit of writing automated tests and checks for every pipeline step. To learn more about verifying cloud workflows end-to-end, consider the Cloud Development program, which covers AWS S3, Lambda, Terraform, CI/CD and more through real-world projects and mentorship. By making every operation’s outcome explicit and automated, you ensure that “data left in S3” or “orphaned uploads” never slip through in production.