Amazon SQS acknowledgments can be misleading. AWS documentation warns that deleting a message requires the most recently received receipt handle, and that using an old handle “will succeed, but the message might not be deleted”. In other words, a successful DeleteMessage call does not prove the message is gone. Even after a correct delete, a Standard queue may redeliver a message if one replica failed to delete it. As a cloud application engineer, the question is: Which evidence can we trust to accept that work has truly been completed?
This article focuses on manual, non-Lambda SQS consumers. We trace a synthetic message (business key b-1) with two receives (attempts A and B) and two different handles (h-old, h-new). We show how a buggy adapter might ack with a stale handle, and how to repair it by binding each delete to its original attempt. A local Stubber-based test will verify exact parameters for each delete request. Finally, we discuss what remains uncertain: even a 200 OK from SQS only means the request was received, not that the message can’t reappear.
The outcome: Accept an acknowledgment only when it’s bound to the correct attempt and parameters; repair any adapter that mixes up handles; hold any unproven deletion claim when handles or effects are ambiguous; and reconcile any redelivered work via the application’s durable business ledger. (Local tests are run with boto3 1.43.18 / botocore 1.43.18 Stubber, no live AWS calls.)
Define the acknowledgment claim you can actually test
AWS SQS clients often treat DeleteMessage success as the final signal of completion. But by protocol, three things are distinct: the service response, the binding of a delete to the received attempt (via ReceiptHandle), and the actual removal of all copies of the message from the queue. A 200 OK only confirms the request was accepted, not that no copy remains. For standard queues, at-least-once delivery semantics guarantee a message may reappear even after a correct delete. We must separate what we can control locally from what the service “promises.”
This playbook assumes a manual Standard-queue consumer architecture (outside Lambda event-source mappings). We will not achieve global exactly-once processing here. Instead, we focus on binding each completion callback to its original receive attempt. Each attempt has a unique ID and the handle returned at that time. Our four decisions: ACCEPT the ack if we have the current handle for that attempt; REPAIR the local handle-tracking if our adapter would reuse an old handle; HOLD if we’re uncertain whether the delete actually took effect; and RECONCILE any redelivered business work according to the application’s state. We explicitly avoid implying that an SQS delete alone is an irrevocable guarantee.
For background on pipeline and consumer responsibilities, see how cloud-native pipeline operations frame end-to-end reliability.
Separate business keys, message IDs and receipt handles
When a producer sends a message to SQS, AWS assigns a MessageId (for example "m-1"). This ID is the same on every receive; it identifies the message on the queue but not a specific delivery attempt. By contrast, the ReceiptHandle (for example "h-old" or "h-new") is tied to one receive of that message. Every time you poll a Standard queue, even for the same message, SQS returns a fresh handle (a long opaque string) for that attempt.
Our example message includes a developer-defined business key (e.g. business_id="b-1" in the JSON body) that identifies the application’s unit of work. This key is outside SQS; it’s in the payload and is reusable across sends. The queue does not enforce uniqueness of this key. The MessageId is generated once per send and can be helpful for log correlation, but it’s not bound to the current receive. The ReceiptHandle is bound to one receive. To safely acknowledge, the application must use the latest handle for that attempt.
For example, suppose the consumer fetches a message, sees MessageId: m-1, extracts b-1 as the business key, and gets handle h-old. Later, the same message is received again (perhaps due to timeout) with handle h-new. We must keep {attempt=A, MessageId=m-1, ReceiptHandle=h-old, business_key=b-1} separate from {attempt=B, MessageId=m-1, ReceiptHandle=h-new, business_key=b-1}. Failure to do so is why naive adapters break. In particular, don’t assume you can delete by MessageId, and don’t reuse an old handle with the same MessageId. The docs warn: if you don’t use the most recent handle, “the message might not be deleted”.
This distinction is akin to owning a reservation at a hotel: the room number (MessageId) is fixed, but the keycard (ReceiptHandle) is only valid for your current stay (receive attempt). The next time you check in (re-receive), you get a new keycard even for the same room. If your bellhop uses the old keycard to try to check out (delete), the system will “succeed” his swipe (no error), but it might not register the check-out.
See the guide to streaming data integration for how continuous consumers manage stream processing state.
Build a network-free receive-and-delete fixture
We’ll simulate SQS calls locally with boto3 and botocore.stub.Stubber. This lets us enqueue fake responses and assert that our client calls exactly match expectations. First, we configure a client with dummy credentials and attach a Stubber:
import boto3, botocore
from botocore.stub import Stubber, ANY
print("Runtime:", sys.version.split()[0])
print("boto3", boto3.__version__, "botocore", botocore.__version__)
# Expected: boto3 1.43.18, botocore 1.43.18
# Create an SQS client with fake credentials
sqs = boto3.client('sqs',
region_name='us-west-2',
aws_access_key_id='fake',
aws_secret_access_key='fake')
# Attach Stubber to queue expected requests/responses
stubber = Stubber(sqs)
# Define synthetic queue URL and message
queue_url = 'https://sqs.us-west-2.amazonaws.com/123456789012/my-queue'
body = '... business_id="b-1" ...'
msg_id = 'm-1'
receipt_old = 'h-old'
receipt_new = 'h-new'We enable the stubber context, then add canned responses. First, two receives (A then B):# Step 1: Add ReceiveMessage response A
stubber.add_response('receive_message', {
'Messages': [{
'MessageId': msg_id,
'ReceiptHandle': receipt_old,
'Body': body
}]
}, expected_params={'QueueUrl': queue_url, 'MaxNumberOfMessages': 1})
# Step 2: Add ReceiveMessage response B
stubber.add_response('receive_message', {
'Messages': [{
'MessageId': msg_id,
'ReceiptHandle': receipt_new,
'Body': body
}]
}, expected_params={'QueueUrl': queue_url, 'MaxNumberOfMessages': 1})Then we queue DeleteMessage responses. We include two successes (for deletion with each handle) and will test them in sequence:# Expect a delete with the old handle, then a delete with the new handle
stubber.add_response(
'delete_message', {},
{'QueueUrl': queue_url, 'ReceiptHandle': receipt_old}
)
stubber.add_response(
'delete_message', {},
{'QueueUrl': queue_url, 'ReceiptHandle': receipt_new}
)
stubber.activate()
# Perform the stubbed calls
respA = sqs.receive_message(QueueUrl=queue_url, MaxNumberOfMessages=1)
respB = sqs.receive_message(QueueUrl=queue_url, MaxNumberOfMessages=1)
print("Received A:", respA['Messages'][0]['ReceiptHandle'])
print("Received B:", respB['Messages'][0]['ReceiptHandle'])This fixture uses canned responses, not a real SQS interaction. Each stubbed response is labelled and consumed exactly once, so we cannot infer real-world timing or invisible-window behavior from it. Any output from these calls is purely synthetic. In particular, seeing DeleteMessage accept h-old would only reflect our stubbed response; it is not evidence that the service actually deleted with an old handle.
A canned response is not an AWS observation
Even though our stubber returns success for the old-handle delete, remember: this is our scripted scenario. It simply demonstrates that the code will send that handle and accept a 200 OK. It does not prove that AWS would always succeed or delete in a real queue. We separate these provided responses from what an actual service experiment might show.
Replay two receive attempts for one synthetic message
With stub responses configured, our code simulates two receives in a row. This models (not enforces) what happens when the first attempt’s visibility expires or is released, and the queue hands the message again. In each receive, the client sees the same body and MessageId: m-1, but different handles.
attemptA = respA['Messages'][0]
attemptB = respB['Messages'][0]
print("Attempt A:", attemptA)
print("Attempt B:", attemptB)
We might log something like:
Attempt A: { MessageId='m-1', ReceiptHandle='h-old', Body='...b-1...' }
Attempt B: { MessageId='m-1', ReceiptHandle='h-new', Body='...b-1...' }
Notice: One MessageId can have multiple receive handles simultaneously. Here both attempts refer to MessageId m-1, but we have distinct "h-old" for attempt A and "h-new" for attempt B. Our local record keeps attempt_id='A' with h-old, and attempt_id='B' with h-new. Each tuple (attempt, handle, message ID, body) is immutable for the life of that callback.
We have modeled the timeline, but the stubber does not model time. We didn’t actually wait for visibility to expire. Instead, our code simply invoked two receives back-to-back. In a real queue, between these receives you might have released A’s visibility or waited out a timer. We are not testing SQS’s timing, only our client logic with two distinct responses.
Expose the cached-old-handle delete path
Now consider a faulty adapter that stores handles by MessageId. It might say: “when B is done, look up MessageId m-1 and use whatever handle we have”. In this bad case, B’s delete call would use the stale handle h-old. Let’s see what our stub does:
# Faulty adapter: attempt B completes but uses old handle
print("Faulty B delete with old handle:")
sqs.delete_message(QueueUrl=queue_url, ReceiptHandle=receipt_old)
print("Sent ReceiptHandle=h-old for B, got HTTP 200 OK (stubbed)")
Our stub returns success (empty body) for this call. We record that a DeleteMessage was issued with {ReceiptHandle='h-old'}. According to AWS DeleteMessage documentation, this could succeed (return 200) and yet leave the message undeleted. In a real system, we wouldn’t know. The docs explicitly warn: using an old handle may return success without deleting the message.
We also test the correct path for B:
# Correct adapter: attempt B uses its own new handle
print("Correct B delete with new handle:")
sqs.delete_message(QueueUrl=queue_url, ReceiptHandle=receipt_new)
print("Sent ReceiptHandle=h-new for B, got HTTP 200 OK (stubbed)")
Now the request used h-new as expected. We have no way to confirm side-effects here, but at least our parameters match the fresh receive.
Crucially, our stubbed demonstration shows how to catch the error of mixing up handles. We can also test a parameter-mismatch, like passing MessageId instead of ReceiptHandle, which Stubber would raise. This guards against typos or copying the wrong field. The stub’s parameter checking ensures we only treat the delete call as valid if the exact handle field is present. (If we accidentally wrote delete_message(QueueUrl=..., MessageId=msg_id), the stub would throw a ParameterMismatchError.)
By the end, call stubber.assert_no_pending_responses() to ensure we consumed all expected stubs. If any remain, our test or code missed a call. This helps ensure our test is complete and adversarial to prevent unnoticed mismatches.
Prevent late work from borrowing another attempt’s handle
Another bug pattern: suppose attempt A completes after attempt B started. A’s handler might naively look up “the current handle for MessageId m-1” and find B’s h-new (since B overwrote the cache). Then A might issue a delete with B’s handle, even though B hasn’t finished.
We simulate this: first we receive A, then B. Now A’s “complete” callback runs and tries to use the latest handle (h-new) to delete. Our adapter should not allow that. Instead of deleting with B’s handle, A should be rejected because B (a newer attempt) is still in flight.
# Faulty adapter: attempt A completes late and uses the newest handle (h-new)
print("Late A delete with B's handle (faulty):")
try:
sqs.delete_message(QueueUrl=queue_url, ReceiptHandle=receipt_new)
print("ERROR: A used handle h-new (belongs to B) for deletion")
except Exception as e:
print("Correctly blocked A using h-new (simulated by logic)")
Our fixed adapter will check the attempt ID before calling delete. If A’s attempt ID is older than the current authorized one for that message, it refuses to call SQS at all (or treats it as already superseded). In this single-process fixture, we can compare attempt IDs: A’s ID is “A” and B’s is “B”, so A is superseded. The adapter should log or raise an error instead of issuing the delete.
Thus, we reject superseded callbacks rather than letting them call DeleteMessage. This prevents the specific bug of “borrowing another attempt’s handle.” After this fix, every delete call is provably matched to its own receive. We don't claim this as a global SQS lock; it’s just local policy: “I will not delete if I’m not the current receiver.”
The newest handle is not a license for old work
It’s important to realize: even though we ensure “Use only your original handle”, it does not mean that an old attempt magically remains valid forever. If A had been the only one, h-old might become invalid after the visibility resets. We never assume an expired handle is valid just because we haven’t seen an error. We simply ensure no attempt uses another’s handle.
State what the repaired adapter does not guarantee
Our adapter now enforces that if it issues a delete, the ReceiptHandle is exactly the one from that attempt’s receive. The stubbed tests verify that by catching any ParameterMismatchError when we call the wrong handle. However, SQS behavior still leaves open possibilities:
• Visibility expiry: If we take too long before deleting, the message might become visible again. Even a correctly bound delete may race with the visibility window. (One could call ChangeMessageVisibility to hold it longer, but AWS won’t treat that as a persistent lock beyond the timeout.)
• External receivers: Our single-thread test knows no one else is polling, but in practice another consumer or retry mechanism might re-receive the message concurrently. We have no distributed locking. Our binding is local: it ensures our process won’t misuse handles, but it doesn’t stop someone else.
• Stale success: As noted, SQS itself might accept our delete request (200 OK) on an old handle and not remove all copies. We do not assume that success = deletion.
Put simply: our adapter’s correctness guarantees that if we see a 200 on DeleteMessage, it was with the right handle for that attempt. It does not guarantee the message is gone or that we won’t see it again. We rely on other mechanisms (visibility timeout, idempotent processing, a durable ledger keyed by b-1) to handle the rest. In short, we own our side of the API contract: we won’t erroneously mix handles. But AWS still owns eventual delivery semantics.
For context on maintaining API contracts in distributed systems, consider how teams approach contract ownership in design.
Make the Stubber assertions adversarial
Using botocore.stub.Stubber, we can force our code to use exact parameters. For example, we can employ ANY if some field is random, but we’ll assert the important ones. Below we illustrate such tests (in code comments) to be sure our delete calls are correct.
with Stubber(sqs) as stubber:
# Expect DeleteMessage with handle = h-old
stubber.add_response(
'delete_message', {},
{'QueueUrl': queue_url, 'ReceiptHandle': receipt_old}
)
stubber.activate()
# If code incorrectly uses a different handle or key, stubber would error.
sqs.delete_message(QueueUrl=queue_url, ReceiptHandle=receipt_old)
stubber.assert_no_pending_responses()
with Stubber(sqs) as stubber:
# Test invalid handle error path
stubber.add_client_error('delete_message', 'ReceiptHandleIsInvalid',
service_message='Invalid receipt handle',
expected_params={
'QueueUrl': queue_url,
'ReceiptHandle': 'invalid-handle'
})
stubber.activate()
try:
sqs.delete_message(QueueUrl=queue_url, ReceiptHandle='invalid-handle')
except botocore.exceptions.ClientError as err:
print("Received expected error code:", err.response['Error']['Code'])
stubber.assert_no_pending_responses()
# Parameter mismatch example
with Stubber(sqs) as stubber:
stubber.add_response(
'delete_message', {},
{'QueueUrl': queue_url, 'ReceiptHandle': receipt_old}
)
stubber.activate()
try:
# Missing or wrong parameter
sqs.delete_message(QueueUrl=queue_url, MessageId=msg_id)
except botocore.stub.ClientError as e:
print("Caught parameter error (mismatch):", type(e).__name__)Each add_response or add_client_error defines exact expected parameters. If our code sends anything else, the stubber raises an exception. This ensures we never accidentally send, say, the MessageId string in the ReceiptHandle field. In summary, our tests are adversarial: any deviation (wrong handle, missing handle, extra fields) will cause a clear failure in our test suite.
Assert the exact handle sent to DeleteMessage
The last snippet above catches even a parameter mismatch. If the code passed 'MessageId': 'm-1' instead of 'ReceiptHandle': '...', stubber will complain that the expected key is missing. This catches typos or logic bugs early. It underlines that the receipt handle string is the key field; every DeleteMessage must include it correctly.
Run a bounded service observation only in an owned queue
So far we’ve done local tests. Optionally, one could attempt a live experiment in a disposable, authorized Standard queue. This is just to observe behavior, not a substitute for the logic checks above. If we did: create a test queue, send one message, and try the same sequence (receive A, let invisibility lapse, receive B, call delete with A’s old handle, etc). We would:
• Ensure only our account and test client poll the queue.
• Keep B’s visibility open while testing A’s delete.
• Record AWS responses and new receives.
We’d stop after a finite number of receives (to bound time and costs). For each observation, we’d note:
• case_id, queue identity, message_id (m-1), business_key=b-1, attempt IDs, local state, handle digests (hash of actual handles for privacy), timestamps, request IDs, and response codes or errors.
• If a new receive happens, note its handle (digest) and time.
• Classify what we see: e.g. OBSERVED_REDELIVERY (the message reappeared within our window), NO_REDELIVERY_WITHIN_WINDOW, or INCONCLUSIVE (neither definitely saw nor timed out conclusively).
We would not loop indefinitely waiting for a redelivery (that risks uncontrollable delays). After a modest number of receives or a timeout, we conclude not observed in that window.
This bounded observation can tell us something about duplicate likelihood under specific conditions, but not anything mathematical. For example, if A’s old delete returned an error (ReceiptHandleIsInvalid) or succeeded, and then we never see the message again, it might suggest deletion, but it could still be a timing artifact. If we see it again, we know deletion didn’t fully stick. We only record what happened, careful not to over-interpret.
This observational trial is for demonstration. In production, we wouldn’t rely on it to fix the root issue. Instead, we stick to our local test’s logic and to idempotent business processing.
Interpret redelivery and empty polls without overclaiming
Suppose after our delete calls we continue polling. If the same message reappears with a new handle, we label OBSERVED_REDELIVERY and add a new attempt entry to our ledger (with a new handle and attempt ID). This provably shows the message still exists.
If we poll for some time and see nothing (queue empty or low count), we may tentatively note NO_REDELIVERY_WITHIN_WINDOW. But this is not proof of permanent deletion, just a bounded observation. The message might still be in flight elsewhere, or in some AZ copy. We avoid jumping to “message gone forever”. This is why we recommend a finite window and limited tries. Beyond that, declaring “deleted” is risky. (Amazon itself says a message might reappear even after delete.)
In any case, an empty poll or approximate queue count is weak evidence. Lack of duplicates within a few minutes doesn’t prove the absence of duplicates ever. Conversely, seeing the message once more is clear evidence it wasn’t fully deleted (or was re-sent).
No redelivery within a window is a bounded observation
So we might record something like:
observation_window=10s, redelivery_observed=NO_REDELIVERY_WITHIN_WINDOW
Record the local time of the last poll. But we do not change our acceptance decision purely on this. It’s just logged evidence. A more robust signal is the business ledger: if we process the message’s business key and commit it, we may mark it done internally anyway.
Reconcile message delivery with durable business state
Ultimately, the consumer’s authority to delete a message comes from application state, not SQS. Our code should check a durable business ledger keyed by b-1. Possible states:
• Committed: We’ve completed this business operation (e.g. updated a database) and should not repeat it. If this is the case before delete, then we can issue delete safely. If we discover after a redelivery that the work was already done, we can delete or ignore the duplicate.
• In Progress: We started the work but haven’t committed. This suggests we might want to finish it or roll back. We should not delete the message unless we commit. (If an old completion called delete erroneously, we may transition to a safe state or alert.)
• Unknown/Unstarted: This can happen if we lost track (e.g. a crash before writing to the ledger). In such case we might hold the delete (HOLD) until we know more, or we might attempt idempotent reprocessing.
We do not rely on redelivery to infer business state. For example, we won’t delete simply because we see a retry of a message (that’s just SQS’s protocol). Instead, on any receive or delete attempt, we consult the ledger. If the ledger says “done”, we can confidently delete; if it says “pending”, we proceed accordingly; if “none”, we might process and then delete.
For broader architecture on backend messaging responsibilities and business consistency, see backend message-processing responsibilities. The ledger is the source of truth, not the queue alone.
Capture evidence that supports a narrow acceptance decision
Each receive attempt should produce an audit record. For attempt X (A or B), include: case_id, message_id=m-1, business_key=b-1, attempt_id=X, local_state (e.g. STARTED, COMMITTED), handle_digest (hash of receipt handle, so we don’t log raw tokens), receive_time, completion_time, and if we called delete: delete_handle_digest, request_id (from AWS response), and response_status. Also note the ledger status (COMMITTED, IN_PROGRESS, UNKNOWN) and our decision (ACCEPT, REPAIR, HOLD, RECONCILE) and who owns it.
For example, after attempt B’s delete, we might have:
case_id=42, queue=..., message_id=m-1, business_key=b-1, attempt_id=B,
local_state=STARTED, handle_digest=sha256(h-new), receive_time=...,
completion_time=..., delete_handle_digest=sha256(h-new),
request_id=xyz, response_status=200, ledger_status=COMMITTED,
decision=ACCEPT, owner=our-service
This shows we used handle h-new for B and got 200, but notes the ledger said COMMITTED (so we should delete). We do not log the actual handle string or body in plaintext.
If an error or replay happens, we’ll record that too:
... delete_error=ReceiptHandleIsInvalid, observation_window=10s,
redelivery=OBSERVED
The evidence lets us review exactly what happened, but cannot by itself say “message is irrevocably gone.” It only shows, for each attempt, what was sent and what was returned.
Preserve attempt evidence without exposing live handles
We only put digest(h) in logs, not the raw handle. In our code example above, we print the stubbed handles (h-old, h-new) for clarity, but real logs should redact those. The queue URL can be partly masked too. Each piece of evidence is disjoint: we don’t hinge on one field proving something global. For instance, even if delete_message returns HTTP 200, we record it as a status, but we don’t “mark done” until the ledger says it’s safe.
Own the acknowledgment path and recovery decision
The core rule: never let an old or wrong handle be used. We enforce that in code and in tests. Once we have exactly matched request parameters (receipt handle from the same attempt), we consider the delete request correct. But we still verify local conditions (via the ledger). Then we make one of four decisions:
1. ACCEPT: The attempt’s work is done (ledger committed), and we just sent the correct DeleteMessage. In this case, we take the 200 response (or a handled invalid-handle error) as the endpoint of this attempt. We log acceptance.
2. REPAIR: We detected a binding bug (trying to use wrong handle). We correct our code to bind by attempt. We do not change any SQS state until it’s bound correctly.
3. HOLD: We are unsure if deletion happened (e.g. old handle succeeded, or we see an error code like ReceiptHandleIsInvalid). We refrain from marking this message done. We may retry deletion or reconciling logic later. The message stays in queue until cleared.
4. RECONCILE: We received a redelivery of a message we thought we already processed. Our responsibility is to ensure no duplicate side-effects. Based on the ledger: if the business effect was already committed, we can simply ack/delete. If not, we may re-run or roll back. The queue’s semantics don’t change our policy: redelivery by itself isn’t error; it triggers business-level idempotency checks.
Throughout, note that performance (throughput, latency) is secondary to correctness. We do not use observed latencies or counts as evidence of a delete. We also do not implement timeouts or retries here; those are system concerns. Our focus is purely semantic correctness: "did we send the right handle for the right attempt?" As a rule, do not interpret a speedy DeleteMessage response as a reason to stop your business logic, nor a slow response as a problem (unless it errors).
For more on balancing correctness and performance in backend systems, see performance checks beyond semantic correctness.
Build cloud workflows around explicit contracts
In summary, treat each SQS message like a contract: the message is “m-1 with b-1”, the receive attempt is a lease with a unique handle, and the effect is your business operation on b-1. We must explicitly bridge these: hold onto the handle for exactly that lease, verify with the ledger, and acknowledge only with exact parameters. Don’t assume the network or the service will prove anything for you beyond returning HTTP codes. Design your system so that if any doubt remains, the worst that happens is one duplicate processing (which your idempotency handles) or one message not getting deleted (which your periodic cleanup can address).
Building reliable cloud applications requires understanding these explicit contracts. For a broader foundation, explore Refonte Learning’s Cloud Development Program to master cloud architecture and operations.
