An Azure Blob lease break can look successful before write ownership has actually changed. The most dangerous shortcut is to see 202 Accepted, wait roughly the requested break period, and declare writer B the new owner. That conclusion skips the interval in which the lease can still be breaking and locked, and in which the existing lease holder A is documented as able to continue writing. Microsoft’s Lease Blob state table explicitly distinguishes Breaking from Broken, and its use-attempt table says a write carrying A can succeed while A’s lease is breaking.
This playbook is for an authorized storage or application team that needs evidence about one disposable Azure block blob. It separates four questions: did Azure acknowledge the break request; what lease state did the service subsequently report; which conditioned writes did the service accept; and what content was actually committed?
This harness has not been executed against Azure. Every result table below therefore presents documentation-derived expectations, not observed measurements. The harness is intended to be run against a real Azure Blob Storage endpoint before cloud qualification. An Azurite-only run may help debug the harness, but it must remain labelled emulator evidence.
The handoff ends in one of four decisions: ACCEPT proven ownership and content, WAIT while breaking remains observable, HOLD when evidence is missing or contradictory, or RECONCILE when content or ownership cannot safely be inferred.
Define the One-Blob Handoff Contract
The protected resource is exactly one small block blob at a predeclared path such as:
account: <owned-disposable-storage-account>
container: lease-handoff-lab
blob: writer-handoff.jsonThe test excludes production data, hierarchical namespace behavior, immutable retention, legal holds, snapshots as test targets, append/page blobs, container deletion, and writes to unrelated objects. The main handoff and the unlocked-gap control use the same owned blob path in separate object lifecycles: finish and clean up the main fixture before recreating that path as a fresh object for the destructive control.
There are three actors. Writer A initially owns the lease and is deliberately authorized to write the blob. Writer B is also authorized, but initially lacks the active lease identity. An operator is authorized to request the break. The lease therefore tests the storage concurrency boundary, not whether Azure RBAC rejects a participant. Azure documents a blob lease as an exclusive write/delete mechanism for the leased blob; it remains separate from authentication and authorization. Microsoft’s Lease Blob reference also states that data operations require authorization.
The acceptance vocabulary is intentionally strict:
Decision | Meaning |
ACCEPT | B acquisition is positively evidenced, B-conditioned write is accepted, stale A write is rejected, and committed content matches the accepted-write ledger. |
WAIT | Fresh properties still show the lease in breaking/locked; takeover is not complete. |
HOLD | A request outcome, state observation, API version, operation shape, or error is unknown or contradictory. |
RECONCILE | Storage ownership can be established, but committed content or a previously uncertain mutation must be resolved before another overwrite. |
Nothing here proves that process A stopped running. A lease constrains qualifying operations against this blob; it is not a process-kill mechanism, distributed leader-election proof, or fence around arbitrary remote side effects. That separation is consistent with the broader cloud-development architecture foundations: define the resource boundary before assigning a system-wide meaning to a storage primitive.
Pin the API and Build an Authorized Disposable Lab
For qualification, use the real Azure Blob Storage endpoint at <account>.blob.core.windows.net over HTTPS. Microsoft documents a separate emulator URI for Azurite, so emulator evidence and Azure-service evidence should never be merged into the same qualification row.
This lab explicitly requests Storage API version 2023-11-03. Do not merely record the SDK default: the Python BlobClient supports an api_version constructor argument, whose default otherwise follows the most recent service version compatible with that SDK. Record both the requested x-ms-version and the response x-ms-version; Lease Blob documents the latter as the service version used to execute the request.
A reproducible fixture can pin:
Python 3.13.7
azure-storage-blob==12.30.3
azure-core==1.41.0
azure-identity==1.25.3
Storage API: 2023-11-03
storage-operation retries: 0
max_single_put_size: 1048576azure-storage-blob 12.30.3 was published on PyPI by September 28, 2026; this proposed pinned lab environment has not been executed. Freeze the actual versions reported by the execution environment in the evidence bundle rather than calling a rolling SDK reference page proof of the package installed on the test host.
For the tiny overwrite, set a one-megabyte max_single_put_size and keep the fixture well below it. Microsoft’s current Python reference states that a blob at or below max_single_put_size is uploaded using one HTTP PUT; the harness must additionally inspect the emitted request and confirm that it was an ordinary blob PUT, rather than assuming the SDK made the expected choice.
Authenticate with an approved lab identity (preferably an approved Microsoft Entra credential or managed identity) and do not put account keys, bearer tokens, connection strings, or SAS query signatures into source or evidence. Microsoft recommends Entra ID/managed identities over Shared Key for Azure Storage authorization. Full RBAC setup is outside this lab; the prerequisite is simply that A, B, and the operator have the narrowly scoped permissions approved for this disposable blob.
Before the first command, record the account hostname, container, blob path, resource owner, test ticket/run ID, permitted actors, and cleanup authority. This discipline is the storage-specific counterpart of reproducible cloud infrastructure practices: the test should be reproducible without accidentally enlarging its ownership boundary.
Create an Independent Content and Request Ledger
Do not derive the expected answer from the same log being tested. Define the content sequence before execution.
Use canonical inert JSON such as:
{"purpose":"lease-handoff-lab","revision":3,"writer":"A"}Hash the exact bytes with SHA-256. The independent expected-content ledger for the main branch is:
Revision | Writer | Operation | Documentation-derived expectation | Expected committed revision afterward |
0 | fixture | initial Put Blob | accepted before leasing | 0 |
1 | A | Put Blob with A | accepted while leased | 1 |
2 | none | Put Blob without lease | rejected while leased | 1 |
3 | A | Put Blob bracketed by breaking observations | accepted if request truly occurs during breaking | 3 |
4 | A | Put Blob after observed broken | rejected | 3 |
5 | B | Put Blob after proven B acquisition | accepted | 5 |
6 | A | Put Blob while B owns lease | rejected | 5 |
This sequence follows Azure’s documented lease-state use table: A can write in Leased and Breaking; an unconditioned write is rejected while leased/breaking; A is rejected after Broken; and an unconditioned write can succeed once the resource is unlocked. The specific response code for the actual Put Blob operation is evaluated separately below.
Every storage request should produce a protected evidence row containing at least:
run_id
event_id
fixture_epoch
operation
actor_alias
request_started_utc
response_received_utc
http_method
blob_path
requested_api_version
returned_api_version
lease_alias_supplied # A, B, or omitted
lease_real_id # protected file only
payload_revision
payload_sha256
http_status
azure_error_code
x_ms_request_id
lease_state_before
lease_status_before
lease_state_after
lease_status_after
etag
request_shape
result_classification
notesThe reader-facing export removes lease_real_id, authorization headers, cookies, query credentials, tokens, and other secrets. It can retain hashes, aliases, request IDs, HTTP statuses, API versions, state labels, and inert fixture content.
A useful result classification is finer than success/failure: ACCEPTED_WRITE, EXPECTED_REJECTION, STATE_OBSERVED, BREAK_ACCEPTED_NOT_COMPLETE, INCONCLUSIVE_WINDOW, UNKNOWN_OUTCOME, and EVIDENCE_ERROR. UNKNOWN_OUTCOME always maps to HOLD until reconciled.
An ETag Is Not a Lease Ownership Identifier
Get Blob Properties returns both content metadata such as ETag and distinct lease fields including lease state and lease status. Microsoft documents x-ms-lease-state values including available, leased, expired, breaking, and broken, and x-ms-lease-status as locked or unlocked. Get Blob Properties also returns the service request ID and API version.
Keep these evidence dimensions separate. An ETag can help identify a content generation, but it does not say “B owns this lease.” Conversely, a generic leased observation identifies a lease state but not the private identity held by B.
Lease Blob itself does not modify the blob’s ETag or Last-Modified property, according to Microsoft’s REST reference. That is another reason not to conflate a content validator with a lease-holder identity.
Prove the Initial Lease A Controls the Write Path
Start only after revision 0 has been read back and its hash matches the ledger. Acquire A as an infinite lease (-1) so ordinary lease expiration does not compete with the planned 30-second break. Azure documents -1 as an infinite lease duration, while finite blob leases are 15–60 seconds.
Then run three controls in order:
A-conditioned overwrite: write revision 1 with A’s real lease ID. Require an accepted Put Blob response and read back revision 1.
Missing-ID overwrite: attempt revision 2 with exactly the same authenticated writer capability but omit the lease ID. Require rejection and verify revision 1 remains committed.
Unconditioned read: read the blob without a lease ID. Require success and revision 1
Microsoft’s Lease Blob remarks say an active lease ID must accompany Put Blob and other specified writes, while GET operations such as Get Blob and Get Blob Properties can succeed without it. The Put Blob REST specification is more operation-specific: if an existing blob has an active lease, an overwrite needs the valid lease ID; an invalid or missing applicable lease condition produces 412 Precondition Failed.
That makes the missing-ID write an important negative control. Merely succeeding with A would not prove that the lease was actually enforcing the write boundary; perhaps the request accidentally omitted the lease while the blob was not leased. The paired rejection demonstrates that the same resource was enforcing the active lease.
For every accepted write, immediately reconcile content by reading without a lease ID and comparing exact bytes or their SHA-256 hash with the independently prepared payload. An HTTP success says the write request succeeded; the read-back makes the committed content explicit in the acceptance record.
A Blob Lease Does Not Lock Every Kind of Access
Azure describes blob leases as pessimistic concurrency for write access, while also preserving read access under the documented rules. Microsoft’s Blob Storage concurrency guidance states that a client attempting to write a leased blob without the proper lease ID receives a precondition failure.
Do not translate that into “the blob is inaccessible.” The read-without-ID control is intentionally expected to pass. Nor should this experiment imply protection for another blob, a queue, a database row, an external API, or the writer process’s local state.
The lease similarly does not establish that A has stopped computing. It establishes whether a qualifying request to this one leased blob satisfies Azure’s lease condition. That narrow statement is strong because it is testable.
Request a Break Without Declaring the Handoff Complete
Capture a fresh properties observation immediately before the break request. It should show the expected initial lease state: for the designed fixture, leased and locked. Then have the authorized operator issue break_lease(lease_break_period=30).
A successful Lease Blob Break returns 202 Accepted; the response can also include x-ms-lease-time, an approximate number of seconds remaining. Azure explicitly says a break begins a period during which a new lease cannot yet be acquired, and that the blob may remain held longer than the proposed break period. The Python BlobLeaseClient.break_lease() wrapper likewise returns the approximate remaining interval rather than an ownership-transfer certificate.
Therefore record:
operation=Break Lease
requested_break_seconds=30
status=202 # expected, not observed here
x-ms-lease-time=<captured>
x-ms-request-id=<captured>
requested x-ms-version=2023-11-03
returned x-ms-version=<captured>Classify the event as BREAK_ACCEPTED_NOT_COMPLETE.
The next authoritative state comes from a fresh Get Blob Properties request, not from the local clock. If it reports breaking and locked, the decision is WAIT. Azure defines Breaking as a non-renewable state that remains locked until the break period expires, while Broken is the post-period state in which acquisition becomes allowed.
This distinction matters even in a much larger scalable cloud data-platform context: a storage state transition should not be replaced with a client-side sleep simply because higher-level pipelines would prefer a convenient handoff moment.
Test the Old Writer During the Breaking Interval
The key test is A’s conditioned overwrite while the Azure Blob lease is breaking.
The sequence must be:
Get Properties -> breaking/locked
Put Blob revision 3 with lease A
Get Properties -> breaking/locked
Read Blob -> revision 3Microsoft’s Lease Blob outcome table says “Write with (A)” succeeds in the Breaking (A) state. That documented behavior is the expected baseline; it is not an observed result for this unexecuted fixture.
Timing evidence matters. Keep both properties request IDs, the Put Blob request ID, request-start and response timestamps, the returned service versions, and the two state observations. To classify the intended test as reproduced, both bracketing observations should still show breaking/locked, the A-conditioned Put Blob must be accepted, and revision 3 must reconcile after the write.
If the post-write properties call already says broken/unlocked, the experiment crossed the transition boundary. Even if the Put Blob happened to succeed, the evidence no longer cleanly demonstrates which state Azure evaluated during that request. Mark it INCONCLUSIVE_WINDOW, safely finish that fixture, and repeat on a fresh disposable object lifecycle.
Now attempt B acquisition while a fresh properties response still shows breaking. Lease Blob’s operation table says acquisition by B during Breaking (A) fails with 409; the Blob Storage error catalogue includes LeaseIsBreakingAndCannotBeAcquired as a 409 Conflict. This is an independent negative control showing that a break acknowledgement has not yet opened the acquisition path.
Do not put B’s acquisition attempt before the A write if it will consume most of the remaining interval. The purpose is not to race Azure. The purpose is to create a defensible sequence whose timestamps show which state surrounded each test.
A Late Probe Cannot Prove an Earlier State
Suppose the test waits 35 seconds, observes broken, and then sends A’s write. A rejection proves only that stale A did not authorize that later request. It says nothing about whether A could have written during the earlier breaking interval.
That distinction prevents a common evidentiary mistake: using a correct post-transition result to make a historical claim about an unobserved state.
When the timing window is missed, repeat it. Do not backfill an illustrative breaking row, estimate the server transition from a stopwatch, or infer it from the break request’s initial time estimate.
Observe the Transition Without Trusting a Sleep
Use bounded polling against Get Blob Properties without supplying A or B. Azure allows Get Blob Properties without a lease ID on a leased blob, making it appropriate for this state observation.
A compact polling function is:
import time
def wait_for_broken(blob, deadline_seconds, trace_call):
deadline = time.monotonic() + deadline_seconds
while time.monotonic() < deadline:
props, ev = trace_call(
"Get Blob Properties",
lambda hooks: blob.get_blob_properties(hooks)
)
lease = props.lease
state = getattr(lease, "state", None)
status = getattr(lease, "status", None)
if str(state).lower().endswith("broken") and \
str(status).lower().endswith("unlocked"):
return props, ev
if not str(state).lower().endswith("breaking"):
raise RuntimeError(
f"HOLD: unexpected lease state={state}, status={status}"
)
time.sleep(1.0)
raise TimeoutError("HOLD: polling deadline expired before broken/unlocked")The deadline is an operator-selected bound, not Azure’s break-period estimate plus an assumed safety margin. A timeout means “we did not establish the required state within the observation window,” not “Azure must still be breaking.”
Likewise, a failed properties request does not prove anything about lease state. Record the transport/service failure and HOLD if the state is needed for the next mutation.
For state-changing requests, configure the storage client with automatic retries disabled in this diagnostic harness. The critical case is a transport timeout after Azure may already have processed the request: blindly issuing an unconditioned retry can create a second mutation while the first outcome is unknown. Instead, preserve UNKNOWN_OUTCOME, read current state/content, and reconcile before deciding whether another write is permissible.
A local sleep(30), a cached properties object, and the x-ms-lease-time returned by the initial break response are all useful timing aids. None substitutes for the fresh broken/unlocked observation.
Reject the Old Lease and Acquire the New One
Once fresh properties positively show broken and unlocked, send revision 4 as Put Blob with stale lease A. Do not omit the lease header. This is the direct test that A’s old credential no longer satisfies the write precondition.
For Put Blob specifically, Microsoft documents 412 Precondition Failed when an inappropriate lease ID accompanies the operation, for example, when the supplied ID does not correspond to an active lease. Record the actual status and Azure error code rather than manufacturing that expected value.
Next acquire B. A successful Lease Blob acquire is documented as 201 Created; the private lease ID returned to the client is the credential B will use for subsequent conditioned writes. Keep that GUID in protected local evidence and publish only alias B.
If the acquire response itself is lost after request transmission, stop. A later properties response saying leased cannot tell you whether this acquisition succeeded, whether another controlled action intervened, or what lease identity is active. Classify the acquisition UNKNOWN_OUTCOME and HOLD until ownership is reconciled.
With a positively captured B acquisition:
write revision 5 with B;
require acceptance and verify exact revision 5 content;
attempt revision 6 with A;
require rejection;
re-read and require content still at revision 5.
That sequence is the minimum positive/negative proof pair for Azure Blob writer handoff validation.
Leased Does Not Tell You Which Client Owns the Lease
Get Blob Properties exposes state and status but not the active secret lease GUID. A row showing leased/locked proves that some lease is in force, not that B possesses it.
B ownership therefore requires evidence from the acquisition transaction itself plus an operation that can only succeed when B supplies the matching private lease identity. The strongest scoped chain is:
Acquire B accepted
+
B-conditioned Put Blob accepted
+
A-conditioned Put Blob rejected
+
content == B's accepted revisionThat evidence proves the lease/write relationship for this blob during this test. It does not prove that B is the only running application instance anywhere else.
Expose the Unlocked Gap in a Separate Control
Do not demonstrate the unlocked gap inside a real takeover. Run it only after the main handoff has been accepted and safely cleaned up, then recreate the same owned path as a fresh object.
The control sequence is:
create fresh revision 100
acquire A
write revision 101 with A
request break
poll until broken/unlocked
write revision 102 WITHOUT a lease ID
read back revision 102
cleanupThe reason for this branch is explicit in Microsoft’s Lease Blob outcome table: “Write, no lease specified” is rejected in Leased and Breaking, but succeeds in Broken because the blob is unlocked.
That means “remove the stale lease header and retry” is not a safe recovery strategy. Once the lease is broken but before B has acquired its new lease, an otherwise authorized unconditioned writer can enter the gap.
The application rule should therefore be simple: a writer designed to require lease ownership must preserve that lease precondition across retries. If it loses confidence that its lease is valid, it stops mutating and reconciles; it does not degrade to an unconditioned write.
This is also where storage scope must remain explicit. The old process may still be alive, continue reading this blob, update another resource for which it has permission, emit a message, or call an external service. The lease protects the documented operations on the leased object, not an entire application transaction. Those are cloud-native application and pipeline boundaries that a single-object experiment cannot validate.
The gap control should be treated as destructive precisely because its success intentionally overwrites the blob while no lease exists. It belongs only on the fresh disposable lifecycle, never in a production takeover.
Interpret Errors for the Operation You Actually Sent
Lease errors should be stored as a pair:
HTTP status + Azure service error codeDo not reduce all lease conflicts to “409” or all stale writes to “412.”
There is a documentation nuance worth preserving. The broad Lease Blob “outcomes of use attempts” table lists, among other entries, a write using B against Leased (A) as failing with 409, while its Breaking (A) row for B says 412. The operation-specific Put Blob documentation, however, says that an active lease requires a valid lease ID and describes invalid or missing applicable lease IDs as 412 Precondition Failed.
For this harness the actual overwrite operation is Put Blob, so the Put Blob contract is the more specific baseline for interpreting the emitted overwrite. But the broader table’s differing code is not something to erase. Preserve it as a documentation discrepancy and record what the service actually returns for API version 2023-11-03.
Lease acquisition is a different operation. Azure’s Lease Blob table says an acquire attempt during breaking fails with 409, and the Azure Blob error catalogue maps LeaseIsBreakingAndCannotBeAcquired to 409 Conflict.
A compact fault matrix is therefore:
Operation and condition | Documentation-derived expectation | Decision if matched | If outcome is missing/different |
Put Blob, no ID while A active | reject; operation-specific baseline 412 | continue | HOLD |
Put Blob with A during cleanly bracketed breaking | accept | continue, reconcile content | HOLD/inconclusive |
Acquire B during breaking | reject; Lease Blob baseline 409 | continue | HOLD |
Put Blob with A after broken | reject; Put Blob baseline 412 | continue | HOLD |
Acquire B after broken | 201 | continue | HOLD if response uncertain |
Put Blob with B after acquisition | accept | continue | RECONCILE/HOLD |
Put Blob with A while B owns lease | reject | ACCEPT candidate | HOLD |
Unconditioned gap-control Put Blob after broken | accept | control reproduced | HOLD |
The core acceptance question is not whether a favorite status code appeared. It is which exact operation Azure accepted, under which lease condition and API version, and what bytes became committed.
Do Not Normalize Conflicting Evidence Into a Pass
If the service returns 409 for an overwrite where the operation-specific expectation was 412, store 409. Never “correct” it in the report.
First confirm the trace really emitted Put Blob rather than a lease request, block staging request, or unexpected SDK retry. Confirm x-ms-version, request URL shape, lease header presence, Azure error code, request ID, and installed package versions. An unexplained discrepancy is HOLD, not a cosmetically normalized pass.
The same rule applies when the SDK operation shape differs from the planned experiment. The Python client documents automatic chunking for upload_blob; for this tiny fixture, the trace must verify exactly one storage PUT with no comp=block/comp=blocklist sequence.
A minimal trace-aware harness can use the following complete core. It intentionally disables storage retries and sanitizes public evidence:
import hashlib
import json
import os
import time
from datetime import datetime, timezone
from urllib.parse import urlsplit
from azure.core.exceptions import HttpResponseError
from azure.identity import DefaultAzureCredential
from azure.storage.blob import BlobClient, BlobLeaseClient
API_VERSION = "2023-11-03"
ACCOUNT_URL = os.environ["AZURE_STORAGE_ACCOUNT_URL"]
CONTAINER = os.environ["AZURE_STORAGE_CONTAINER"]
BLOB = os.environ.get("AZURE_STORAGE_BLOB", "writer-handoff.json")
def utc():
return datetime.now(timezone.utc).isoformat()
def payload(revision: int, writer: str) -> bytes:
return json.dumps(
{"purpose": "lease-handoff-lab",
"revision": revision,
"writer": writer},
sort_keys=True,
separators=(",", ":"),
).encode()
def sha256(data: bytes) -> str:
return hashlib.sha256(data).hexdigest()
def new_blob() -> BlobClient:
# DefaultAzureCredential contains no embedded account key/SAS.
# Run only where its resolved identity is approved for this lab.
return BlobClient(
account_url=ACCOUNT_URL,
container_name=CONTAINER,
blob_name=BLOB,
credential=DefaultAzureCredential(),
api_version=API_VERSION,
max_single_put_size=1024 1024,
retry_total=0,
retry_connect=0,
retry_read=0,
retry_status=0,
)
def trace(op, actor, lease_alias, revision, fn):
ev = {
"operation": op,
"actor": actor,
"lease_alias": lease_alias or "omitted",
"payload_revision": revision,
"request_started_utc": utc(),
"requests": [],
}
def on_request(req):
r = req.http_request
u = urlsplit(r.url)
ev["requests"].append({
"method": r.method,
"path": u.path,
"query": u.query,
"requested_api_version": r.headers.get("x-ms-version"),
})
def on_response(resp):
r = resp.http_response
ev.update({
"http_status": r.status_code,
"x_ms_request_id": r.headers.get("x-ms-request-id"),
"returned_api_version": r.headers.get("x-ms-version"),
"x_ms_lease_time": r.headers.get("x-ms-lease-time"),
})
try:
result = fn(
raw_request_hook=on_request,
raw_response_hook=on_response,
)
ev["classification"] = "RESPONSE_RECEIVED"
return result, ev
except HttpResponseError as exc:
ev["http_status"] = getattr(exc, "status_code", None)
ev["azure_error_code"] = getattr(exc, "error_code", None)
response = getattr(exc, "response", None)
if response is not None:
ev["x_ms_request_id"] = response.headers.get("x-ms-request-id")
ev["returned_api_version"] = response.headers.get("x-ms-version")
ev["classification"] = "REJECTED_RESPONSE"
return exc, ev
except Exception as exc:
# A transport failure after send can have an unknown server outcome.
ev["classification"] = "UNKNOWN_OUTCOME"
ev["local_exception"] = type(exc).__name__
return exc, ev
finally:
ev["response_received_utc"] = utc()
def put(blob, revision, writer, lease_id=None, lease_alias=None):
data = payload(revision, writer)
result, ev = trace(
"Put Blob", writer, lease_alias, revision,
lambda *hooks: blob.upload_blob(
data, overwrite=True, lease=lease_id, hooks)
)
ev["payload_sha256"] = sha256(data)
# Qualification requires exactly one ordinary blob PUT.
if ev["classification"] == "RESPONSE_RECEIVED":
storage_puts = [
r for r in ev["requests"]
if r["method"] == "PUT"
]
if len(storage_puts) != 1 or "comp=block" in storage_puts[0]["query"] \
or "comp=blocklist" in storage_puts[0]["query"]:
ev["classification"] = "EVIDENCE_ERROR"
return result, ev
def props(blob):
return trace(
"Get Blob Properties", "observer", None, None,
lambda hooks: blob.get_blob_properties(**hooks)
)
def lease_values(p):
lease = p.lease
return str(lease.state).lower(), str(lease.status).lower()
def poll_broken(blob, seconds=75):
deadline = time.monotonic() + seconds
rows = []
while time.monotonic() < deadline:
p, ev = props(blob)
rows.append(ev)
if isinstance(p, Exception):
raise RuntimeError("HOLD: properties outcome unavailable")
state, status = lease_values(p)
if state.endswith("broken") and status.endswith("unlocked"):
return p, rows
if not state.endswith("breaking"):
raise RuntimeError(
f"HOLD: unexpected {state=}, {status=}")
time.sleep(1)
raise TimeoutError("HOLD: broken/unlocked not observed by deadline")
def read_bytes(blob):
result, ev = trace(
"Get Blob", "observer", None, None,
lambda hooks: blob.download_blob(hooks).readall()
)
return result, ev
def assert_content(blob, revision, writer):
actual, ev = read_bytes(blob)
expected = payload(revision, writer)
if isinstance(actual, Exception) or actual != expected:
raise RuntimeError("RECONCILE: blob content != accepted ledger")
return evThe sequence driver should use those functions in the order specified in the surrounding playbook: create revision 0; acquire infinite A; write/reconcile 1; reject 2; break for 30 seconds; bracket the 3 probe; reject premature B acquire; poll; reject 4; acquire B; write/reconcile 5; reject 6; and finally reconcile revision 5.
For lease operations, use BlobLeaseClient(blob), retain .id only in protected process memory/local evidence, and invoke acquire(lease_duration=-1) or break_lease(lease_break_period=30). Those are the callable operations documented by the Python SDK.
Reconcile Every Accepted Write With the Blob Content
The accepted-write ledger, not attempted requests, defines expected content.
After every accepted Put Blob, perform an authorized unconditioned read and compare the exact canonical payload or hash. For the normal main sequence the expected committed progression is:
0 -> 1 -> 3 -> 5Revisions 2, 4, and 6 are attempted negative controls and must never appear in the committed ledger if their writes were rejected.
An UNKNOWN_OUTCOME is different from a rejection. If revision 5 timed out after the request left the client, do not write “B write failed.” Read the blob. If it contains revision 5, that is strong evidence the mutation occurred, but the ownership decision still needs the surrounding B-acquisition and lease-condition evidence. If it contains an earlier revision, continue investigating request IDs and state rather than assuming the timed-out write was never processed.
The final normal-path acceptance record should therefore connect:
B acquisition response
-> B-conditioned revision 5 accepted
-> exact revision 5 read back
-> stale A revision 6 rejected
-> revision 5 still presentProperties and content answer different questions. broken/unlocked establishes state at an observation point; leased/locked later establishes that a lease exists; the B acquire record ties that lease to B; and the B-conditioned Put Blob plus read-back establishes the successful protected write.
None of these observations establishes that process A has stopped or that unrelated application effects have quiesced.
Recover a Partial Handoff Without Overwriting Newer Data
When a handoff stops halfway, speculative retrying is the wrong first move. Freeze mutations and determine what Azure and the content ledger already establish.
A recovery procedure is:
Preserve the original request rows, exceptions, request IDs, API versions, timings, and hashes.
Read current properties without a stale lease condition.
Read current blob content without a lease ID.
Compare the content with every known accepted or potentially accepted revision.
Determine whether an active owner is positively known from its acquisition evidence and private lease identity.
Perform a new write only after ownership and intended content have both been reconciled.
For example, suppose B’s acquisition was captured successfully but B’s revision 5 Put Blob hit a client transport timeout. Do not automatically repeat revision 5. First read the blob. If revision 5 is already there, an unconditional retry would be redundant and could become dangerous if later application revisions exist. If revision 3 remains, B may decide (under the application’s own recovery contract) to retry using the confirmed B lease.
If acquisition itself has an uncertain outcome, an observation of leased is insufficient to manufacture B ownership. HOLD until the private ownership evidence can be reconciled. Depending on the controlled lab state, that may mean proving the known B lease ID through a conditioned operation or abandoning that fixture and creating a fresh test object.
Cleanup follows the same certainty rule. When B is positively known to own the lease, release B and confirm the release before deleting the owned blob. Do not run a broad account/container cleanup routine as part of this test. If ownership is uncertain, leave the disposable resource for the authorized operator rather than turning cleanup into another ambiguous state-changing operation.
A Remembered Pre-Break Payload Is Not a Recovery Authority
Revision 1 may have been the last payload remembered before the break, but revision 3 can legitimately have been accepted while A was still in breaking. Azure’s documented use table explicitly permits that A-conditioned breaking-state write.
Therefore recovery must not blindly restore revision 1.
The recovery authority is the reconciled content history: accepted write evidence, current bytes, and the application’s approved logical revision. If current content is newer than the remembered pre-break image, preserve it until its provenance is resolved. Once B is the confirmed owner, B can apply the approved next revision, conditionally and deliberately, rather than automatically restoring a remembered old copy.
This is where the distinction between HOLD and RECONCILE matters. HOLD means the system lacks enough trustworthy evidence to act. RECONCILE means evidence shows a content/ownership situation that must be resolved before another mutation.
Apply the ACCEPT, WAIT, HOLD, and RECONCILE Matrix
The final decision should be mechanical enough that another reviewer can reach the same result.
Evidence state | Decision |
Fresh properties show breaking/locked | WAIT |
Break returned 202, but no fresh post-break state exists | WAIT or HOLD, never ACCEPT |
Request/response lost for a potentially mutating operation | HOLD |
Breaking probe crossed the state transition | HOLD that reproduction claim; repeat fresh fixture |
broken/unlocked observed, but B acquisition not proven | HOLD |
B ownership proven, but blob content conflicts with accepted/unknown writes | RECONCILE |
B acquisition captured, B-conditioned write accepted/read back, A-conditioned write rejected, final content correct | ACCEPT |
Actual HTTP/error/API/request shape contradicts the documented baseline and remains unexplained | HOLD |
The ACCEPT statement should be deliberately narrow:
For this identified Azure block blob, API version, request sequence, and evidence run, the observed handoff is accepted: stale lease A no longer authorized the tested overwrite, lease B was positively acquired, a B-conditioned Put Blob was accepted, an A-conditioned Put Blob was rejected, and final content matched the accepted-write ledger.
That statement does not qualify another blob, larger/multipart uploads, changed retry policies, another API version, an emulator, or application side effects outside the blob.
Requalify when the service API version changes, SDK/retry behavior changes materially, the upload crosses the single-Put threshold, the application begins staging blocks separately, credential/authorization behavior changes, or the design expands from one resource to multiple resources. Larger Azure analytics workflows built above storage are examples of systems that may depend on storage without being validated by this one-blob acceptance test.
The storage owner should own lease-state and request evidence; the application owner should own logical-content reconciliation and confirmation that external writer behavior is safe. A single Azure lease cannot merge those responsibilities.
Build Cloud Development Skills Around Explicit State Contracts
The engineering lesson in an Azure Blob lease break is not “wait 30 seconds.” It is to distinguish acknowledgement, state, ownership, accepted mutation, and committed content.
A 202 Accepted proves that Azure accepted the break request. It does not prove the break period has ended. breaking/locked means WAIT. An unknown state-changing request means HOLD. A confirmed B owner with uncertain content means RECONCILE. Only the complete evidence chain (observed transition, stale-A rejection, captured B acquisition, B-conditioned accepted write, and reconciled final bytes) supports ACCEPT for this protected blob. Azure’s own Lease Blob model makes the critical distinction: A’s write is documented as possible while the lease is breaking, whereas acquisition becomes available after the lease reaches the broken/unlocked state.
Engineers developing this kind of operational judgment need more than syntax: they need explicit resource boundaries, reproducible environments, security-aware cloud access, and testable state contracts. Refonte Learning’s Cloud Development Program lists cloud architecture, Docker and Kubernetes, Infrastructure as Code, security, performance, DevOps, and a cloud-application capstone in a three-month program requiring 12–15 hours per week; its inspected curriculum does not establish that this specific Azure Blob lease-break lab is taught.
For this handoff, the operational rule remains narrower and more important: never declare a new writer from elapsed time or a break acknowledgement alone. Prove the storage state, prove the owner through its conditioned operation, and reconcile the bytes before proceeding.
