Linux systems engineer tracing an inherited flock lock across parent and child processes

Why Your Child Process Keeps the flock Lock Alive

Fri, Oct 9, 2026

The problem starts with a blocked worker after its parent process closed the lock file descriptor. Did the forked child inherit the lock and keep it held? We will not assume the child was at fault; instead we design a controlled experiment to separate process liveness from lock ownership. By orchestrating a parent that holds a flock lock and a child that reports state over a socket pair, we observe exactly when each lock owner closes its descriptor. This ordered experiment shows precisely who still holds the lock at each step. In our setup (single-threaded CPython 3.12 on Linux 6.18 x86_64, overlayfs), the parent never exits, ensuring the child remains alive. Two protocols are compared: one where the child keeps its inherited lock descriptor, and one where the child deliberately closes it early. Each contender attempt opens the file fresh (new open file description) so that locks truly contend as documented by the kernel.

We record 9 lock attempts across these protocols, verifying device/inode identity each time to ensure the same file was targeted. The evidence shows: if the child retains its descriptor, closing only the parent’s copy does not free the lock; if the child closes its copy (before the parent’s final work), then after the parent releases its own copy the lock can be acquired even while the child is running. We use fcntl.flock(..., LOCK_EX | LOCK_NB) for non-blocking tests. When an incompatible lock prevents acquisition, it raises OSError with errno EAGAIN/EWOULDBLOCK. Each acquisition or block is captured in a JSON report. This page presents the full fixture code, the measured outcomes, and an explicit ownership policy. It explains why the child’s inherited descriptor kept the lock, why FD_CLOEXEC didn’t help (fork ignores that flag), and when to close unwanted inherited descriptors.

State which processes should own the lock

Consider a local maintenance job that’s been waiting on a file lock. Its parent process has completed the protected work and closed its file descriptor (fd), but the forked child is still alive. Should the lock now be free or kept? Two ownership models emerge: either the child intentionally shares the lock for legitimate work, or it accidentally retains it and should let it go. We pose this question precisely without assuming the parent exited; it only closed its fd. We interpret an ACQUIRED result (fresh lock taken) as evidence that at that moment the lock was available, not as a guarantee of future access.

The Linux docs clarify that a flock lock is tied to an open file description (OFD), not merely the file name. All duplicate fds (from fork/dup) refer to the same OFD (and thus the same lock). The lock is only released by an explicit LOCK_UN on one of the duplicates or when all those duplicates are closed. In other words, if the child still has a copy of the holder descriptor open, the lock remains held despite the parent’s close. This is foundational: closing one of multiple duplicates does not release the lock. Moreover, an independent open (same pathname but a fresh open() call) creates a new OFD, which can contend for the lock. We use this principle in our test.

For our purposes, both parent and child started with a non-inheritable lock fd in the parent (so it wouldn’t survive exec) and then forked. The child inherits that descriptor by fork (because fork duplicates all fds into the child). If our policy is shared ownership, the child would hold its copy until it’s truly done; if not needed, the child should close its copy early. This experiment forces us to adopt an explicit policy after observing which descriptor owners actually keep the lock. We will see that mere “process is alive” is not enough; a process must drop its fd to free the lock.

Pin the Linux filesystem and process model

Environment. We ran the fixture on Linux 6.18.44 (glibc 2.39, x86_64) with Python 3.12.14. The underlying directory was an overlayfs (recorded via stat -f); we avoided any NFS or SMB mounts because flock behaves differently there. The script is a single-threaded Python process. In a multi-threaded program, the child can safely call only async-signal-safe functions after fork() and before execve(). We also require os.pidfd_open (Linux 5.3+) to wait for child termination reliably. A brief stat of the temp directory confirmed “overlayfs”, and the test aborts if a network filesystem is detected (see require in code).

  • Python version: 3.12.14 (the code’s report["python"]). We set the parent fd inheritable flag to False (os.set_inheritable(fd, False)), which only matters for exec (no exec() is used here).

  • Filesystems: We use only regular files on a local filesystem (overlayfs in test). We exclude NFS/SMB because older kernels wouldn’t lock over NFS, and SMB semantics changed in Linux 5.5 (flock becomes a mandatory lock over CIFS).

  • Processes: The parent and child are plain processes (no threads). fork() duplicates the parent’s file descriptors into the child. Each test run opens a lockfile with os.open(...), fcntl.flock(), etc. The code uses a socketpair handshake with 3-second timeouts (TIMEOUT=3.0) to synchronize parent and child.

For readers, familiarity with general Linux administration foundations helps; things like file permissions, processes, and commands (bash, stat, etc.) are assumed. In particular, note that replacing a file path (rename/unlink) does not change existing open descriptors (so the path can be removed and the fd still points to the original content). Similarly, FD_CLOEXEC (close-on-exec) only affects execve, not fork. Throughout, the parent is never terminated prematurely; we simply test what happens when it closes its copy of the lock file’s fd.

Separate a path, a descriptor and an open file description

It helps to be precise: a path is the pathname naming a file in the filesystem. A file descriptor (FD) is a small integer handle in a process, returned by open(). An open file description (OFD) is the kernel’s internal record for that open(), holding the device/inode, file offset, status flags, and flock locks. Every call to open() creates a new OFD; a file descriptor is just a reference to an OFD. If you rename or delete the path, existing FDs still refer to the original OFD and file contents. Crucially, fork() copies the parent’s FD table into the child, so the child’s FD refers to the same OFD as the parent’s corresponding FD. By contrast, if the same process does a fresh open() on the same file, it creates a new OFD (independent of the first). We keep these distinctions clear: path = name, descriptor = handle (per-process), open file description = the kernel’s entry. Our procedure uses both fork and fresh open to tease apart these relationships.

Run the complete fixture against private regular files

Below is the full driver script flock_fixture.py used for testing. It sets up its own temporary directory, creates the lockfile, and orchestrates the parent/child protocol for the two cases. The code is shown in full so the test is reproducible. Note how the two run_case protocols differ only in the early_close flag: False (child retains descriptor) vs True (child closes before READY). Each case records events and independent lock attempts (via fresh os.open) on its own file, with nine attempts across the two cases. We use a three-second deadline on all socket handshakes (see TIMEOUT) so a hung child will trigger a timeout error.

"""Linux flock/fork descriptor lifetime; parent remains alive, no exec in tested children."""
import errno
import fcntl
import json
import os
from pathlib import Path
import platform
import selectors
import signal
import socket
import stat
import subprocess
import tempfile
import time
import hashlib

TIMEOUT = 3.0

def require(condition, message):
    if not condition:
        raise RuntimeError(message)

def identity(fd):
    value = os.fstat(fd)
    return {"device": value.st_dev, "inode": value.st_ino,
            "regular_file": stat.S_ISREG(value.st_mode)}

def send(sock, value):
    payload = json.dumps(value, separators=(",", ":")).encode() + b"\n"
    require(len(payload) < 4096, "protocol message too large")
    sock.sendall(payload)

def receive(sock):
    deadline = time.monotonic() + TIMEOUT
    result = bytearray()
    while not result.endswith(b"\n"):
        remaining = deadline - time.monotonic()
        if remaining <= 0:
            raise TimeoutError("handshake deadline expired")
        sock.settimeout(remaining)
        chunk = sock.recv(1)
        if not chunk:
            raise RuntimeError("protocol EOF before complete message")
        result.extend(chunk)
        require(len(result) < 4096, "protocol response too large")
    sock.settimeout(TIMEOUT)
    message = json.loads(result)
    if message.get("event") == "ERROR":
        raise RuntimeError("child error: " + message.get("error", "unspecified"))
    return message

def wait_ready(pidfd, timeout):
    with selectors.DefaultSelector() as selector:
        selector.register(pidfd, selectors.EVENT_READ)
        return bool(selector.select(timeout))

def child_protocol(child_sock, parent_sock, holder_fd, early_close):
    parent_sock.close()
    child_sock.settimeout(TIMEOUT)
    descriptor = holder_fd
    exit_code = 0
    try:
        inherited = identity(descriptor)
        inheritable = os.get_inheritable(descriptor)
        if early_close:
            os.close(descriptor)
            descriptor = None
        send(child_sock, {"event": "READY", "pid": os.getpid(),
                         "parent_pid": os.getppid(), "identity": inherited,
                         "inheritable_flag_at_fork": inheritable,
                         "lock_descriptor_open": descriptor is not None})
        while True:
            message = receive(child_sock)
            command = message.get("command")
            if command == "PING":
                send(child_sock, {"event": "ALIVE", "pid": os.getpid(),
                                  "lock_descriptor_open": descriptor is not None})
            elif command == "CLOSE":
                require(descriptor is not None, "descriptor already closed")
                os.close(descriptor)
                descriptor = None
                send(child_sock, {"event": "CLOSED", "pid": os.getpid(),
                                  "lock_descriptor_open": False})
            elif command == "EXIT":
                require(descriptor is None, "EXIT requested while child still holds descriptor")
                send(child_sock, {"event": "BYE", "pid": os.getpid()})
                break
            else:
                raise RuntimeError("unknown parent command")
    except Exception as exc:
        exit_code = 1
        try:
            send(child_sock, {"event": "ERROR", "pid": os.getpid(),
                              "error": type(exc).__name__ + ": " + str(exc)})
        except Exception:
            pass
    finally:
        if descriptor is not None:
            os.close(descriptor)
        child_sock.close()
    os._exit(exit_code)

def run_case(name, early_close, directory, report):
    entry = {"status": "HOLD", "parent_pid": os.getpid(), "early_child_close": early_close,
             "events": [], "contender_attempts": []}
    report["cases"][name] = entry
    path = directory / (name + ".lock")
    holder_fd = None
    parent_sock = child_sock = None
    pid = pidfd = None
    reaped = False

    def event(operation, values):
        entry["events"].append({"sequence": len(entry["events"]) + 1,
                               "operation": operation, values})

    def contender(label, expected):
        trial = {"label": label, "status": "HOLD", "expected": expected,
                 "pid": os.getpid(), "open_method": "fresh independent os.open"}
        entry["contender_attempts"].append(trial)
        independent_fd = None
        try:
            independent_fd = os.open(path, os.O_RDWR)
            trial["identity"] = identity(independent_fd)
            require(trial["identity"] == entry["parent_identity"], "lock path identity changed")
            trial["inheritable_flag"] = os.get_inheritable(independent_fd)
            try:
                fcntl.flock(independent_fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
            except OSError as exc:
                trial["errno"] = exc.errno
                if exc.errno not in {errno.EAGAIN, errno.EWOULDBLOCK}:
                    raise
                trial["status"] = "BLOCKED"
            else:
                trial["status"] = "ACQUIRED"
            require(trial["status"] == expected, "unexpected contender result: " + label)
            event("contender", label=label, status=trial["status"], identity=trial["identity"])
        except Exception as exc:
            trial["error"] = {"type": type(exc).__name__, "message": str(exc)}
            raise
        finally:
            if independent_fd is not None:
                os.close(independent_fd)
                trial["descriptor_closed"] = True

    def ping(expected_open):
        send(parent_sock, {"command": "PING"})
        message = receive(parent_sock)
        require(message.get("event") == "ALIVE" and message.get("pid") == pid,
                "wrong child liveness response")
        require(message.get("lock_descriptor_open") is expected_open, "wrong child descriptor state")
        event("child_alive", message)

    try:
        holder_fd = os.open(path, os.O_CREAT | os.O_EXCL | os.O_RDWR, 0o600)
        os.set_inheritable(holder_fd, False)
        entry["parent_identity"] = identity(holder_fd)
        require(entry["parent_identity"]["regular_file"], "lock target is not a regular file")
        entry["parent_inheritable_flag"] = os.get_inheritable(holder_fd)
        require(entry["parent_inheritable_flag"] is False, "fixture requires non-inheritable flag")
        contender("unlocked_baseline", "ACQUIRED")
        fcntl.flock(holder_fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
        event("parent_acquired", identity=identity(holder_fd))
        contender("parent_held_baseline", "BLOCKED")
        parent_sock, child_sock = socket.socketpair()
        parent_sock.settimeout(TIMEOUT)
        pid = os.fork()
        if pid == 0:
            child_protocol(child_sock, parent_sock, holder_fd, early_close)
        child_sock.close()
        child_sock = None
        pidfd = os.pidfd_open(pid)
        entry["child_pid"] = pid
        message = receive(parent_sock)
        require(message.get("event") == "READY" and message.get("pid") == pid,
                "wrong READY identity")
        require(message.get("parent_pid") == os.getpid(), "unexpected parent changed")
        require(message.get("identity") == entry["parent_identity"], "fork identity mismatch")
        require(message.get("inheritable_flag_at_fork") is False,
                "non-inheritable flag unexpectedly changed")
        require(message.get("lock_descriptor_open") is (not early_close), "incorrect READY state")
        entry["child_ready"] = message
        event("child_ready", message)
        contender("after_ready_parent_still_holds", "BLOCKED")
        event("controlled_parent_work_boundary_complete")
        os.close(holder_fd)
        holder_fd = None
        event("parent_closed_own_descriptor_without_LOCK_UN")
        ping(not early_close)
        contender("after_parent_close_child_alive", "ACQUIRED" if early_close else "BLOCKED")
        ping(not early_close)
        if not early_close:
            send(parent_sock, {"command": "CLOSE"})
            message = receive(parent_sock)
            require(message.get("event") == "CLOSED" and message.get("pid") == pid,
                    "missing child CLOSE acknowledgement")
            require(message.get("lock_descriptor_open") is False, "child did not close descriptor")
            event("child_closed_ack", message)
            ping(False)
            contender("after_child_close_child_alive", "ACQUIRED")
            ping(False)
        send(parent_sock, {"command": "EXIT"})
        message = receive(parent_sock)
        require(message.get("event") == "BYE" and message.get("pid") == pid, "missing BYE")
        require(wait_ready(pidfd, TIMEOUT), "child did not exit before deadline")
        waited_pid, status = os.waitpid(pid, 0)
        reaped = True
        entry["waitpid"] = {"pid": waited_pid, "exit_code": os.waitstatus_to_exitcode(status)}
        require(waited_pid == pid and os.waitstatus_to_exitcode(status) == 0, "child exit failure")
        event("child_reaped", entry["waitpid"])
        require(os.getpid() == entry["parent_pid"], "parent process changed")
        entry["status"] = "PASS"
    except Exception as exc:
        entry["error"] = {"type": type(exc).__name__, "message": str(exc)}
        raise
    finally:
        if holder_fd is not None:
            os.close(holder_fd)
        for endpoint in (parent_sock, child_sock):
            if endpoint is not None:
                endpoint.close()
        if pid is not None and pid > 0 and not reaped:
            try:
                os.kill(pid, signal.SIGKILL)
                entry["failure_cleanup_signal"] = "SIGKILL"
            except ProcessLookupError:
                pass
            if pidfd is None:
                try:
                    pidfd = os.pidfd_open(pid)
                except ProcessLookupError:
                    pass
            if pidfd is not None and wait_ready(pidfd, TIMEOUT):
                waited_pid, status = os.waitpid(pid, 0)
                reaped = True
                entry["failure_cleanup_reaped"] = waited_pid
            else:
                waited_pid, status = os.waitpid(pid, os.WNOHANG)
                reaped = waited_pid == pid
                entry["failure_cleanup_reaped"] = waited_pid
            require(reaped, "HOLD: owned child did not exit within cleanup deadline")
        if pidfd is not None:
            os.close(pidfd)
        if path.exists() and (pid is None or reaped):
            path.unlink()
            entry["owned_lock_file_removed"] = True

def main():
    directory = Path(tempfile.mkdtemp(prefix="flock-inheritance-", dir=Path(__file__).parent))
    report = {"status": "HOLD", "scope": "Linux local filesystem, fork only; original parent remains alive",
              "python": platform.python_version(), "platform": platform.platform(),
              "parent_pid": os.getpid(), "timeout_seconds": TIMEOUT, "cases": {},
              "source_sha256": hashlib.sha256(Path(__file__).read_bytes()).hexdigest()}
    try:
        require(platform.system() == "Linux", "fixture requires Linux")
        require(hasattr(os, "fork") and hasattr(os, "pidfd_open"), "fork and pidfd_open required")
        report["filesystem_type"] = subprocess.check_output(
            ["stat", "-f", "-c", "%T", str(directory)], text=True, timeout=TIMEOUT).strip()
        require(report["filesystem_type"] not in {"nfs", "nfs4", "cifs", "smb2"},
                "this fixture excludes NFS/SMB mounts")
        run_case("child_retains_descriptor", False, directory, report)
        run_case("child_closes_before_ready", True, directory, report)
        require(set(report["cases"]) == {"child_retains_descriptor", "child_closes_before_ready"},
                "case inventory incomplete")
        require(all(case["status"] == "PASS" for case in report["cases"].values()), "case not passed")
        require(sum(len(case["contender_attempts"]) for case in report["cases"].values()) == 9,
                "contender attempt inventory incomplete")
        report["status"] = "PASS_TWO_PROTOCOLS_NINE_ATTEMPTS"
    except Exception as exc:
        report["error"] = {"type": type(exc).__name__, "message": str(exc)}
        raise
    finally:
        result_path = directory / "results.json"
        result_path.write_text(json.dumps(report, indent=2) + "\n")
        print(json.dumps({"status": report["status"], "report": str(result_path)}, indent=2))

if name == "__main__":
    main()

This driver creates two cases, each with ordered events and either four or five contention attempts, for a total of nine attempts. The socketpair channel is used to synchronize parent and child: the child first sends a READY event (with PID, PPID, file identity, and descriptor state), then we ping with ALIVE, send CLOSE or EXIT commands, and observe CLOSED or BYE events. The three-second deadline (TIMEOUT) bounds each handshake so failures time out. The expected results.json status is PASS_TWO_PROTOCOLS_NINE_ATTEMPTS. The table below summarizes the labels and expected result of each lock attempt (lock conflicts raise OSError with errno EAGAIN/EWOULDBLOCK). The parent remains alive throughout; no exec() is done and no explicit LOCK_UN is issued in this test.

Validate the unlocked and parent-held baselines

Before testing inheritance, we confirm the environment’s lock behavior. First, with no lock held, an independent open should succeed. We perform a fresh os.open(path, O_RDWR) and attempt flock(LOCK_EX|LOCK_NB): this must return ACQUIRED. Next, the parent acquires an exclusive flock on its own FD (non-blocking). We then try another fresh open: now the lock is held by the parent’s OFD, so a new OFD must be denied. Indeed, flock(...LOCK_EX|LOCK_NB) should block (we see OSError with errno for lock failure). If either of these baselines fails (e.g. first attempt is blocked or second is acquired), the test is invalid because it contradicts basic flock rules. In our run, the “unlocked baseline” was always ACQUIRED and the “parent-held baseline” always BLOCKED, as expected.

The reason for fresh opens is fundamental: each call to open() returns a new open file description. The Linux manual clarifies that separate opens on the same file create independent OFDs, so locks on one can block the other. In other words, even in the same process, two distinct opens contend for an exclusive flock just like separate processes would. If we had tried to reuse the inherited descriptor (or a dup) for contention, we would not be testing an independent contender: duplicates share the lock. Our use of a fresh open in the parent process ensures each contender truly contends for the lock, just as if it were another process.

Why every contender must use a fresh open

It’s critical that each “contender” attempt calls os.open() anew. The man page warns that duplicating or forking an existing descriptor does not test a real contention, since they refer to the same OFD. By using a fresh independent os.open(), each contender gets a distinct OFD with the same file path. This matches flock’s documented scope: independent OFDs may block each other. Calling flock on the child’s inherited descriptor (or a dup of the parent’s) would merely invoke the same lock and not produce a BLOCKED result. We deliberately keep all lock attempts in the original parent process to keep the PID constant; “same PID, different OFD” is our contention model. No third process is introduced. Thus the parent process itself acts as both holder and all nine contenders via separate opens.

Let the forked child retain the inherited descriptor

We now follow the child_retains_descriptor protocol. After the baselines, the parent has applied flock(LOCK_EX) on holder_fd. Next, we create a socketpair() and fork. In the child (early_close=False), the inherited descriptor is still open. The child immediately sends a READY event with JSON containing:

  • pid (child’s PID) and parent_pid (must equal parent),

  • identity (the parent's parent_identity from os.fstat(): same device and inode),

  • inheritable_flag_at_fork (should be False as we set),

  • lock_descriptor_open (should be True here since we did not close it).

Upon receiving this READY message, the parent checks that the child’s reported file identity matches its own and that the non-inheritable flag remained False (it is only relevant for exec). We effectively tie the child’s FD to the same OFD as the parent’s FD, as Linux fork semantics dictate.

With child READY confirmed, we run the next contender attempt (“after_ready_parent_still_holds”), which again must block since the parent’s lock is still active. The child is still alive and has not yet been told to release anything. All these handshake checks ensure we know exactly what state the child’s descriptor is in. In summary, in this phase the child process is actively holding (inherited) the lock descriptor, and all lock tests reflect that.

Use READY as a local test barrier

The READY event is our own synchronization mechanism, not to be confused with any systemd or service notification. We ensure the parent does not close its descriptor until it sees the child’s READY payload. This is analogous to but distinct from a systemd “READY=1” signal: here, the child explicitly reports its current state. In other words, READY is just an application-defined checkpoint. The parent validates the child’s message carefully (matching PIDs, file identity, descriptor-open flag) before proceeding. (See the systemd startup readiness validation article for how such handshakes coordinate producer and consumer; here we borrow the concept of a readiness barrier, but with a different intent and a direct socket protocol.)

Close only the parent descriptor and probe again

Next, the parent completes its controlled work boundary. It invokes os.close(holder_fd) on the lock fd without issuing a LOCK_UN. This action logs “parent_closed_own_descriptor_without_LOCK_UN”. Now only the child’s copy of the descriptor remains open (in the child process). We immediately PING the child (sending a {"command": "PING"}) and verify an ALIVE response confirming the child is still running and still holds the descriptor (lock_descriptor_open=True).

After confirming liveness, we attempt the next lock (“after_parent_close_child_alive”) with a fresh open. In our environment this probe returned BLOCKED when the child had its descriptor open. The Python flock call produced errno 11, which corresponds to EAGAIN/EWOULDBLOCK on Linux (we interpret either EAGAIN or EWOULDBLOCK as blocked). This outcome shows the lock remained held by the child’s inherited copy. Importantly, we did not do a LOCK_UN; merely closing the parent’s FD does not free the lock because a duplicate remains. This is exactly what the flock(2) page described: the lock stays until all duplicates are closed.

At this point the parent has closed its own handle, but the child still has its handle open. The lock is still blocking other open attempts. We emphasize again: we only closed the parent’s descriptor. We did not delete the file or exit the parent. In this scenario, the lock must remain held because the child’s duplicate descriptor is still open. The syscall did not return EACCES or any other error, only the expected EAGAIN indicating contention. This precise state (“parent closed, child alive, still blocked”) will contrast with the next step.

Release the child descriptor while the child stays alive

Now we instruct the child to close its unneeded descriptor. The parent sends {"command": "CLOSE"} over the socket. The child closes its descriptor inside child_protocol and replies with a {"event": "CLOSED", "pid": child_pid} message. The parent verifies this CLOSED event and that lock_descriptor_open is now False. Another PING confirms the child remains alive (alive but with no descriptor).

Finally, we make one more lock attempt with a fresh open. This time it succeeds (status ACQUIRED). The transition of this attempt from BLOCKED to ACQUIRED is the key evidence: once all inherited descriptors are closed, the lock is free. The child process is still running (we’ll exit it next), but it no longer holds the lock. The events captured (“child_closed_ack” then a successful “after_child_close_child_alive” lock) show the lock became available exactly when the last duplicate descriptor went away.

Distinguish process liveness from lock ownership

One must not confuse a process simply existing with it owning a lock. Our experiment shows a running child can relinquish its handle and free the lock before it exits. Conversely, a parent closing its own handle does not mean “no one holds the lock” if a child still has it. This is akin to how a systemd service might be “up” (running) but not ready (hasn’t sent READY=1) in a Type=notify service. Here, the child’s READY and CLOSED messages are explicit signals of lock state. Without such a signal, the lock’s fate is unclear. In practice, an administrator should not assume an inherited descriptor was closed just because the process is still listed; one must check or synchronize explicitly.

Compare a child that closes its unneeded descriptor before READY

In the child_closes_before_ready protocol, the child immediately closes its inherited descriptor before sending READY. The child’s READY event still reports the same file identity, and then it sets lock_descriptor_open: False right away. In effect, the child gives up the lock immediately. The parent, still holding the lock, sees this: the next contender (“after_ready_parent_still_holds”) is still BLOCKED, because the parent hasn’t closed its own FD yet.

Then the parent closes its FD at its boundary. Now, since the child had already closed its copy, no duplicate remains. The next lock attempt (“parent closed, child alive”) is ACQUIRED immediately, even though the child process is still alive. This confirms that with an intentional early close, the child effectively transfers full ownership of the file to the parent’s schedule. The child’s closure occurred while the parent still held the lock, so it never blocked other attempts. Only after the parent released its lock do new contenders see it as free. This protocol validates a proactive policy: if the child does not need the lock for its work, it should close its copy before the parent finishes.

Reconcile all nine attempts and their event order

Putting it all together, we observe these outcomes for the two protocols:

Protocol

Contender attempt

Observed result

Child retains descriptor

Unlocked baseline

ACQUIRED

Child retains descriptor

Parent-held baseline

BLOCKED

Child retains descriptor

After READY, parent still holds

BLOCKED

Child retains descriptor

Parent closed, child alive with descriptor

BLOCKED

Child retains descriptor

Child closed, child alive without descriptor

ACQUIRED

Child closes before READY

Unlocked baseline

ACQUIRED

Child closes before READY

Parent-held baseline

BLOCKED

Child closes before READY

After READY, parent still holds

BLOCKED

Child closes before READY

Parent closed, child still alive

ACQUIRED

All listed events were emitted in sequence in each case, including the child’s READY/BYE and the parent’s ping/alive handshakes. Both child PIDs were reaped with exit code 0, and the final status of each case was PASS. After both protocols, the lock file was unlinked (cleanup) because no process held it. Importantly, each contender attempt targeted the same recorded file: the device and inode from os.fstat() never changed between attempts. However, device/inode equality alone does not prove the attempts used the same OFD; we know they were fresh opens. We establish the descriptor provenance by the controlled fork/open sequence, just as the rsync and hard-linked backup identity article distinguished identical inodes from identical content.

In summary, with the child retaining its descriptor, the second-to-last attempt stayed BLOCKED, only switching to ACQUIRED after the child explicitly closed its fd. With the child closing early, the lock became free immediately after the parent’s close. This full matrix of 9 attempts, with their ACQUIRED/BLOCKED outcomes, underlines exactly how lock availability depends on who still holds an open descriptor.

Use file identity without overclaiming descriptor identity

Throughout, we identified the file by its (device, inode) pair to ensure every attempt was on the same file. Yet we never assumed “same inode means same descriptor”; we relied on the known sequence of opens and forks. The rsync and hard-linked backup identity article showed that identical inode numbers can hide complex aliasing issues. Here, we similarly do not overclaim: matching device/inode means the same file target, but not the same open instance. The key is that our code explicitly orchestrated separate opens and a fork so we know which FD belonged to which process, rather than inferring anything from the inode alone.

Explain close-on-exec and explicit unlock within their scopes

One might wonder about FD_CLOEXEC: we set the parent’s fd to non-inheritable (inheritable=False) as per PEP 446, but this only affects what happens on an exec(). The proposal notes that “close-on-exec flag has no effect on fork(): all file descriptors are inherited by the child process”. Since we never do an exec in these children, the child naturally inherits the non-exec flag without closing the fd. In other words, FD_CLOEXEC only prevents descriptor inheritance across a new program execution; it does not prevent inheritance on fork.

Also note: we intentionally never use flock(..., LOCK_UN) to release the lock early. If we had applied LOCK_UN on the shared OFD, it would remove the lock for both processes instantly. That would defeat our negative-control test, because the parent closing its descriptor alone (without explicit unlock) is what we want to examine. Our findings do not imply every subprocess inherits locks inappropriately; rather, they show that inherited locks survive forks. FD_CLOEXEC being set is irrelevant without an exec, and adding a subprocess call (with close_fds, etc.) is a different scenario not in scope here.

Contain protocol errors and clean up only owned resources

Any unexpected error (wrong identity, EOF, timeout, etc.) in the handshake is treated as a protocol failure (HOLD). The script catches exceptions, logs the error event, and then cleans up only what the parent “owns”: it closes the parent’s descriptors, kills the owned child if needed, waits with pidfd_open for it to die, and only then removes the lock file. A socket timeout or runtime exception leads to an immediate HOLD status in results. The timeouts (3 seconds) are just a safety bound for our tests, not an absolute guarantee that a real process can always be killed quickly in production.

Reject a READY message to check the failure path

We also verified the failure cleanup logic by forcing the first READY to fail. In a separate run (not the normal two-protocol result), we monkey-patched the receive() function to change the first "READY" into "REJECTED_READY". The parent’s require triggers a RuntimeError, which we catch and interpret as expected. The result for that single case (bad_ready) is status HOLD, with the child killed and reaped. The relevant verification code is below. This ensures our code handles a STARTUP handshake failure correctly (it reaps the child and deletes the lock file). Note that we did not exhaustively test every possible cleanup failure, but this injected fault shows the failure path works as intended.

import importlib.util, json, tempfile
from pathlib import Path
spec = importlib.util.spec_from_file_location("flock_lab", "flock_fixture.py")
lab = importlib.util.module_from_spec(spec)
spec.loader.exec_module(lab)
original_receive = lab.receive
def reject_ready(sock):
    message = original_receive(sock)
    if message.get("event") == "READY":
        return {**message, "event": "REJECTED_READY"}
    return message
lab.receive = reject_ready
directory = Path(tempfile.mkdtemp(prefix="flock-failure-check-"))
report = {"status": "HOLD", "cases": {}}
try:
    try:
        lab.run_case("bad_ready", False, directory, report)
    except RuntimeError as error:
        lab.require(str(error) == "wrong READY identity", "Wrong controlled failure")
    else:
        raise RuntimeError("Controlled fault did not fail")
    case = report["cases"]["bad_ready"]
    lab.require(case["status"] == "HOLD", "Failure was not held")
    lab.require(case.get("failure_cleanup_reaped") == case.get("child_pid"),
                "Owned child not reaped")
    lab.require(case.get("owned_lock_file_removed") is True, "Owned lockfile not removed")
    report["status"] = "PASS_FAILURE_HOLD_AND_REAP"
finally:
    result_path = directory / "results.json"
    result_path.write_text(json.dumps(report, indent=2) + "\n")
    print(result_path, report["status"])

Choose an explicit descriptor-ownership policy

Based on these results, we adopt a clear policy in our system:

Situation

Action

Shared work intended (child legitimately needs lock)

Keep lock descriptor open in both parent and child until all are done (document this shared ownership and expected lifetime).

Child does not need to write after parent

Child should close its inherited descriptor before parent finishes (release lock as soon as possible).

State is ambiguous

Hold further progress and inspect which processes still hold open handles (don’t assume ownership).

Unexpected lock acquisition

Flag this as an error (the assumption of exclusive access is violated).

In all cases, the parent holds its descriptor until its own protected work is complete (the “controlled parent work boundary”). We never suggest deleting lock files by hand or sending signals to arbitrary PIDs to resolve this; instead, it should be handled via process protocols. A mere pathname check (e.g. presence of .lock) or age-based heuristic is not sufficient. If after proper coordination the lock appears unexpectedly free or taken, treat it as a failure to meet the mutual-exclusion requirement. In our scenario, we saw that only when all known owners closed the descriptor did the lock truly become available.

Assign reviewers and revalidation triggers

To ensure this result is integrated into practice, assign roles as follows:

  • Worker maintainer (development team): owns the descriptor-ownership policy (deciding when children should keep or close locks).

  • Platform reviewer (OS/kernel specialist): verifies assumptions about the process model and filesystem behavior (e.g. after kernel updates or filesystem changes, check if fork or flock semantics changed).

  • QA/test engineer: reproduces the nine-attempt matrix and checks the failure path ledger and events.

  • Operations lead: retains the collected evidence (logs/results) from runs to audit compliance and support troubleshooting.

Revalidation should be scheduled whenever relevant system components change: for example, if a new Python/POSIX fork/exec model emerges, if Linux or glibc is updated, or if the filesystem driver changes (since overlayfs or a different filesystem type might behave differently with flock). These tests could be automated as part of continuous integration; see system administration automation practices for guidance on incorporating such operational checks into DevOps workflows.

Develop this troubleshooting discipline with Refonte Learning

This case study illustrates an evidence-based troubleshooting approach in the context of Linux systems, which is exactly what Refonte Learning’s System Administration Program teaches. The program (10–12 hours/week commitment) covers Windows and Linux administration, networking, security, and command-line tools, providing hands-on capstone projects. On completion one earns both a Training Certificate and an Internship Certificate. Course topics explicitly include Linux file management, user administration, and networking.

For aspiring system administrators, this example shows how to connect low-level kernel behavior to an operational policy. Knowing that “forked child inherits flock locks” and that FD_CLOEXEC doesn’t apply on fork is valuable in root-cause analysis. Refonte’s curriculum, led by experienced instructors like Alice Smith, builds such foundational knowledge. By joining the System Administration Program, learners can further practice evidence-driven troubleshooting by writing their own tests, analyzing logs, and developing clear ownership policies. These skills are critical in real-world IT operations. We invite you to explore the program to deepen your expertise, but remember: mastering the craft of system administration always depends on rigorous testing and documentation, as we’ve demonstrated here.