Cloud security engineers reviewing security group rules and network connection activity on office monitors

An AWS Security Group Rule Is Gone. Why Is the Connection Still Open?

Wed, Oct 7, 2026

Cloud incident responders often encounter situations where a permissions change doesn’t have the expected effect. In this controlled test, a narrow ingress rule (only allowing a specific client IP on TCP/18443) will be removed via the AWS API, and we ask: Can the original TCP session still exchange messages while fresh connection attempts fail? AWS documentation on security group connection tracking suggests yes for tracked flows.

We will implement a strict playbook: enroll two test EC2 instances in a private VPC (non-default), configure the source-/32 rule, open a long-running client session (the “old” session), then revoke the rule by its SecurityGroupRuleId. We then attempt new connections (ensuring the old session does not reconnect) and observe whether new TCP handshakes time out while the old session still receives valid replies. We carefully capture the rule state before and after revocation (via describe-security-group-rules), mark the client phase, and close only the specific fixture session through the private local administrative queue.

Finally, we restore the narrow rule (using authorize-security-group-ingress) and confirm a new connection resumes application-layer exchange. This lets us prove exactly what we see and what we do not observe, instead of assuming blanket instance isolation.

Throughout, we distinguish documented expected behavior (tracked vs untracked flows, additive SG rules, canonicalization and small delays in AWS rule changes) from the outcomes this experiment will actually measure. We do not claim every session is terminated or the instance is fully isolated; that would require additional controls (e.g. NACLs to immediately cut connections). The goals here are limited, reproducible checks: old session alive? new blocked? closure scoped? rule restored?

The evidence we gather (session UUID, network tuple, rule ID, JSON logs, error codes) will drive specific conclusions rather than broad generalizations. No AWS integration was executed for this article. The AWS outcomes below are expectations and decision criteria, not measured results; a local loopback protocol check does not establish AWS rule behavior.

Define the claim you are about to test

We start by enumerating the claims to test versus the trust boundaries we accept:

  • Rule absent: First, the security-group rule was revoked. We ensure this by capturing the SecurityGroupRuleId ahead of time and confirming its absence after the API call. We accept a small propagation delay; AWS documents that rule changes “are applied as quickly as possible” but may not be instantaneous. We will explicitly run describe-security-group-rules to verify removal before proceeding.

  • New admission blocked: The revocation should block new connection attempts that match the rule (same protocol/port/CIDR). We test this by repeatedly attempting fresh TCP connects from unused source ports while the rule is absent. If these new attempts consistently time out (while control connects on a separate port still succeed), we conclude new admission was blocked for those probes during the observed interval. Note that alternate paths (different rules or multiple SGs) could still allow traffic, so we carefully inspect the aggregate permission set to ensure only the narrow rule was granting access on our test port.

  • Old session active: The original tracked TCP session might continue to exchange packets even after revocation. We keep the old client socket open (it never attempts to reconnect on failure). If its application-level ACKs and replies continue after rule removal (until we explicitly close it), that demonstrates an active identified session, consistent with documented connection tracking. We will verify matching session IDs and nonces in each round to prove it’s the same session. We do not trust passive disconnections or retries; we drive the session until either an error or explicit closure by admin.

  • Instance isolated: We do not claim isolation of the entire instance from the network. AWS stateful security groups allow tracked flows to persist until timeouts. AWS describes a stateless network ACL as a separately scoped way to interrupt existing traffic. We will stop at closing the one fixture session; we will not sweep or reset the instance’s network. Thus we cannot prove or disprove that every session ended (nor is that our goal).

Trust assumptions: we assume exclusive control of the two test instances and their VPC; no other admin or automation is changing rules or interfaces during the test. We operate with full AWS CLI v2 (no dry-run) and Python 3.12 under a role/user authorized for these operations. We log rule IDs, timestamps, command outputs and all TCP-layer events. The experiment stops if any command fails (unauthorized, missing resource, etc.) or if baseline prerequisites are unmet.

Separate tracked sessions from new admission

To understand the outcome, recall how AWS tracks connections. If an incoming TCP flow was individually allowed by the security group (e.g. source IP/port, dest IP/port match a rule), AWS will statefully track that connection. After the handshake, return traffic is automatically allowed even if only one direction is explicitly permitted. If you change or remove the rule afterwards, the tracked session normally continues until it naturally ends or times out.

In contrast, broad reciprocal rules can produce an untracked flow: AWS describes a TCP or UDP allow from all addresses, with a corresponding all-address rule in the other direction admitting response traffic on any port. Unless the flow is automatically tracked, removing the rule that enables it interrupts that flow. Our experiment uses the narrow case (a /32 source rule to one port), which falls in the tracked category.

We must also account for effective permissions from all security groups on the instance. AWS aggregates rules across all SGs attached to an interface. That means if any SG still permits the flow (inbound on dest port 18443, from our client IP/port), then new packets could succeed despite our one-rule revocation. Therefore we first enumerate every SG on the target ENI and collect all ingress rules.

We then identify that exactly one rule permits tcp:18443 from our client /32; no other rule (including rules in other SGs) allows that combination. Removing this rule thus closes the only (narrow) admission path for new flows. If there were a wider rule (e.g. 0.0.0.0/0) on another SG, then new connections might still get through (the broad reciprocal-rule exception can also change whether a flow is tracked, as discussed below).

We also note that unrelated factors like network path, NACLs, or client behavior could end a session. For example, TCP idle timeouts in the OS or VPC can eventually expire a connection (we do not rely on that, and no universal survival duration follows from this experiment). We ensure the server and client sockets stay active (no close()) and keep application heartbeats going rather than deliberately leaving the connection idle.

We also verify that no other process sends RST/FIN. In short: if we see an ACK from the server after revocation, it must be due to the existing socket, not a new one.

For completeness, AWS’s official narrow-case example describes an SSH connection admitted by a narrow inbound rule. When that rule is changed to stop admitting new connections from the current client address, the existing connection is not interrupted because it is tracked. That is precisely our hypothesis to test.

(We also note stateless NACLs: a Network ACL is stateless, and AWS describes a blocking NACL as a way to interrupt existing traffic. But NACLs are out of scope here. We do not mutate any NACL; all testing is done purely at the SG level, with the other path components held unchanged.)

For relevant background, see AWS network security foundations in Refonte’s learning materials. AWS documents that security group rules are stateful and additive, so revoking one rule does not imply that traffic stops when other rules or tracked state permit it. We will rely on documented behavior rather than assume it.

AWS network security foundations

AWS Security Groups are allow-only (no deny rules) and rules are aggregated across all associated groups. When a rule is removed or modified, AWS “automatically applies” the change to all instances, but any existing tracked connections can keep flowing until they timeout. This is why removing a rule does not guarantee new traffic cannot sneak through via another rule or stateful allowance. In our test, we explicitly list all rules on the instance’s ENIs and verify only the single targeted CIDR/port combination gave access.

That means once we revoke it, the aggregate effective rules no longer provide the selected new-admission permission; existing tracked state is a separate question.

Statefulness vs stateless alternatives

AWS EC2 SGs are stateful. Once a session is accepted, the return packets do not need an explicit rule. AWS user documentation notes that to drop an existing connection immediately you can use a separately scoped stateless NACL. We do not use those here. Instead, we will close the old session from the end using a local admin mechanism inside the server, as a controlled final step.

That way we see exactly how a cooperative endpoint closure works without requiring AWS to end the session. This emphasizes the limits of SG changes: revoking the narrow rule does not itself prove an immediate break in the tracked session.

Freeze the lab boundary and management path

We fix every environmental detail except the single targeted rule. Two Linux EC2 instances (client and server) are in one private subnet of a non-default VPC, each with just a private IPv4 address. Each has exactly one Elastic Network Interface (ENI) with one security group. We record the AWS account and region, instance AMI IDs, OS versions, Python 3.12 and AWS CLI v2 versions. In particular, we do not alter routing, enable flow logs, modify NACLs, touch other SGs, or enable public/Internet connectivity. No load balancers, NAT instances, or other middleboxes are involved.

  • We choose two TCP ports: 18443 as the test port and 18444 as an independent control port. Both will have simple JSON-based test services (below). Neither port is standard (no TLS implied), and no other SG rule will admit any traffic to 18443 except our one narrow /32 rule. A separate control rule (source-/32 to port 18444 on the same SG) will allow baseline monitoring connections. This keeps control traffic fully independent.

  • The same single SG must allow the old session and control connections. We therefore create two inbound rules: one ingress rule with IpProtocol=tcp, FromPort=ToPort=18443 and CidrIpv4 set to the client’s private /32; another ingress rule for protocol=tcp, port=18444, Cidr=client/32. No other rule grants those combinations.

  • We record sender egress and leave the verified outbound rules unchanged. A default allow-all rule is recorded only if it is actually present; the fixture does not replace egress policy.

  • All AWS CLI commands set the Region via $AWS_REGION and use credentials already authorized for the required operations. We presume authorized management access to both instances already exists and has been independently verified. We do not teach provisioning or record credentials here. Where necessary, diagnose an unmanaged Systems Manager instance before beginning this experiment, rather than repairing the management path during the test.

  • We ensure no host firewall (iptables) is blocking the chosen ports. We do not enable any persistent VPC Flow Logs for this test (that’s a separate Refonte blog topic).

  • If any prerequisites are missing (e.g. unable to contact AWS API, EC2 CLI calls fail, or the server application is not listening), we abort as BASELINE_INVALID.

Keep control traffic independent

On the server, we run two listener processes using the Python fixture: one bound to port 18443 (the target) and one to 18444 (the control). Both use the same protocol and framing, but port 18444 remains untouched by revocation. On the client side, we will open the “old” connection to 18443 and a probe to 18444 before touching any SG rule. If the control port ever fails, the environment is faulty and we stop.

We use a fresh working directory on each host for logs, with mode 0700 (create the parent run directories before starting the fixture). We start the old-hold client on 18443 and use fresh probes for the target and independently admitted control port 18444. All probes explicitly bind unused source ports: for example, old session uses port 41400 on client; fresh probes use 41501+ for 18443 and 42501+ for 18444.

We must not reuse a socket or source-port from a prior run, because TCP TIME_WAIT could mislead (we want errors to mean real blockage, not local reuse). On each run we document the chosen source ports.

Before we revoke anything, we verify that both server ports accept connections and exchange messages correctly. We do this by running in parallel: Set RUN, SERVER_IP, CLIENT_IP, SERVER_RUN and CLIENT_RUN from the verified manifest. Each run directory belongs to its respective host, must be newly created with mode 0700, and must not overwrite prior evidence. Initialize phase.txt to baseline before starting hold; old.json must not already exist. Run each long-lived process in its own managed terminal, and capture stdout, stderr and process identity separately.

# On the receiver, in separate managed terminals:

python3 socket_lab.py serve "$SERVER_IP" 18443 "$RUN" \
    "$SERVER_RUN/admin"
python3 socket_lab.py serve "$SERVER_IP" 18444 "$RUN" -
# On the sender, start hold separately from the probes:

python3 socket_lab.py hold "$SERVER_IP" 18443 "$RUN" \

    "$CLIENT_IP" 41400 "$CLIENT_RUN/phase.txt" \

    "$CLIENT_RUN/old.json"

python3 socket_lab.py probe "$SERVER_IP" 18443 "$RUN" \

    "$CLIENT_IP" 41501

python3 socket_lab.py probe "$SERVER_IP" 18444 "$RUN" \

    "$CLIENT_IP" 42501

The first client command (hold) creates the initial session (it will continuously send messages until closed). The two probe commands attempt a single round-trip (fresh-target, fresh-control). All output (JSON events) is logged separately on each host for review. We expect: the hold client to get a session hello from the server and then match ACKs per second; the fresh probes should see a successful ACK/response (“ack”) and then exit; control (18444) should similarly succeed.

If any connect or protocol stage fails here, the baseline is invalid. Both listeners must report ready, and the private administrative queue must report admin_ready, before clients run; otherwise we abort.

Record the effective rules before touching one

Before making any change, we capture the current configuration and caller identity. Using aws sts get-caller-identity we record the AWS account/role. We note our CLI and Python versions at runtime. Then we list each instance’s ENI and all attached SGs, and retrieve their rules with describe-security-group-rules. For example:

aws ec2 describe-network-interfaces \

    --network-interface-ids "$RECEIVER_ENI" \

    --region "$AWS_REGION" --output json > eni.json

aws ec2 describe-security-group-rules \

    --filters "Name=group-id,Values=$TARGET_SG" \

    --region "$AWS_REGION" --output json > rules.json

We also gather any other SG attached to the ENI similarly, to have the full JSON of inbound rules. We do not use pagination limits (--no-paginate or --max-items) so as to capture everything. If any group or rule appears unexpectedly, we investigate. Repeat this complete retrieval for every group on both actual ENIs, including sender egress. Save each output under its own verified per-run path; the example filenames are not permission to overwrite the baseline during readback.

Next, from the combined JSON we parse out exactly one matching ingress rule with IpProtocol == "tcp", FromPort == 18443, ToPort == 18443, and CidrIpv4 == "<client_ip>/32", with IsEgress=false. We save its SecurityGroupRuleId (a value like sgr-abcdef0123456789). At the same time, we examine all other rules in that group and across attached groups to ensure no alternative rule would also allow 18443 from that client (for example, a broader CIDR, a group reference, etc.).

We should see one and only one such rule; any additional matching rule would complicate our test. We record the entire rule object (including any description or tags) into a JSON snippet to use later for restoration. For instance, a saved restore document might look like: This is an illustrative reconstruction shape, not measured AWS output. Replace the example CIDR with the verified private sender /32 and retain only its saved description, omitting Description when absent.

[{"IpProtocol":"tcp","FromPort":18443,"ToPort":18443,
  "IpRanges":[{"CidrIp":"203.0.113.10/32","Description":"Client SSH rule"}]}]

This step certifies the “before” state: who we are, the attached SGs and rules, and precisely which rule we will revoke. The risk of any mismatch (account/region, partial describe, wrong rule) invalidates the baseline and we stop.

Capture the one rule and every possible alternate allow

The following hypothetical JSON illustrates selected rule fields, not a complete baseline or a measured AWS response. The illustrative address and abbreviated identifiers must be replaced by values from the authorized run. The unchanged control admission must also appear in the complete inventory.

{

  "SecurityGroupRules": [

    {

      "SecurityGroupRuleId": "sgr-abc123",

      "GroupId": "sg-0abcd12345",

      "IsEgress": false,

      "IpProtocol": "tcp",

      "FromPort": 18443,

      "ToPort": 18443,

      "CidrIpv4": "203.0.113.10/32",

      "Description": "Allow from client"

    },

    {

      "SecurityGroupRuleId": "sgr-def456",

      "GroupId": "sg-0abcd12345",

      "IsEgress": true,

      "IpProtocol": "-1", "FromPort": -1, "ToPort": -1,

      "CidrIpv4": "0.0.0.0/0"

    }

  ]

}

In this illustrative fragment, sgr-abc123 is the selected ingress rule. Only the complete aggregate inventory can establish that no alternate rule admits the target connection. We extract its ID (sgr-abc123) into a variable $RULE_ID for the revoke command, and store its IpPermission details in restore-permissions.json (removing fields like the RuleId that the API will assign anew). We also note any tag or description that should be restored later.

We will later repeat the complete group-filtered rule retrieval and check for the saved rule ID to establish presence or absence. This preparatory inventory is critical: AWS adds a unique ID for each rule (introduced in 2021), and we rely on it for precise changes.

Build a session-aware application fixture

We need a transparent TCP service that assigns each connection a unique session ID and echoes each message back for matching. We use the provided socket_lab.py (below) which implements exactly that with JSON-over-TCP framing. Each client message includes {"run":RUN,"sid":SID,"seq":N,"phase":PHASE,"nonce":RAND}; the server responds with {"type":"ack", ... same fields ...}. Key points:

  • On accept of a new connection, the server creates a random sid (UUID) for that session, stores the conn object, and immediately sends a “hello” message containing run and sid. The client verifies the hello and saves sid. All further messages in that session must carry that same sid.

  • Each send() uses sock.settimeout(3) and sock.sendall(...). In Python, socket.sendall() blocks until all bytes are sent or error. On success it returns None; on error it raises an exception. We rely on this: sendall() completion is only send-side success, not proof that the remote application processed the line. We serialize to JSON (no whitespace) and append "\n" as a delimiter.

  • The receive() function reads one byte at a time (sock.recv(1)) until a newline, assembling into a JSON object. If the complete frame is not received within the three-second deadline, it raises TimeoutError("frame deadline"). If the peer closes (recv returns empty), it raises EOFError("peer closed"). Both cases terminate the loop.

  • The client loop sends a new message every second (seq++, new nonce). For mode hold, it reads the current phase from a file so we can insert a marker; for mode probe, it sends one message and exits. Each recv on the client expects exactly {"type":"ack", ...} matching what it sent. Any mismatch raises an error and exits.

  • On a shutdown request (mode close), the trusted operator’s local queue request asks the server to invoke conn.shutdown(socket.SHUT_RDWR) on the socket, causing the server to close it. The client will then see an EOF or error on its next recv.

Save the complete standard-library fixture below as socket_lab.py. It uses bounded newline-delimited JSON, thread-safe JSONL event output and a private local administrative request directory. The administrative directory is created inside the receiver’s private run directory; it exposes no administrative TCP endpoint. Capture each process’s stdout and stderr in its own evidence file. The old client deliberately has no reconnect loop.

import datetime

import json

import secrets

import socket

import sys

import threading

import time

import uuid

from pathlib import Path



PRINT_LOCK = threading.Lock()

SESSIONS_LOCK = threading.Lock()

SESSIONS = {}



def emit(event, **data):

    row = dict(

        event=event,

        utc=datetime.datetime.now(datetime.timezone.utc).isoformat(),

        monotonic=time.monotonic(),

        **data,

    )

    with PRINT_LOCK:

        print(json.dumps(row, sort_keys=True), flush=True)



def send(sock, obj):

    sock.settimeout(3)

    sock.sendall((json.dumps(obj, separators=(",", ":")) + "\n").encode())



def receive(sock):

    end, data = time.monotonic() + 3, bytearray()

    while len(data) < 4096:

        remaining = end - time.monotonic()

        if remaining <= 0:

            raise TimeoutError("frame deadline")

        sock.settimeout(remaining)

        part = sock.recv(1)

        if not part:

            raise EOFError("peer closed")

        if part == b"\n":

            return json.loads(data)

        data.extend(part)

    raise ValueError("frame exceeds limit")



def serve(bind, port, run, control):

    def worker(conn, peer):

        sid = uuid.uuid4().hex

        with SESSIONS_LOCK:

            SESSIONS[sid] = conn

        emit("accepted", run=run, sid=sid, peer=peer,

             local=conn.getsockname())

        try:

            send(conn, dict(type="hello", run=run, sid=sid,

                            peer=peer, local=conn.getsockname()))

            last_seq = -1

            while True:

                msg = receive(conn)

                if (set(msg) != {"run", "sid", "seq", "phase", "nonce"}

                    or msg["run"] != run or msg["sid"] != sid

                    or type(msg["seq"]) is not int

                    or msg["seq"] <= last_seq):

                    raise ValueError("invalid request identity")

                last_seq = msg["seq"]

                emit("received", **msg)

                send(conn, dict(type="ack", **msg))

        except Exception as exc:

            emit("session_end", run=run, sid=sid,

                 error=type(exc).__name__)

        finally:

            with SESSIONS_LOCK:

                SESSIONS.pop(sid, None)

            conn.close()



    def admin():

        directory = Path(control)

        directory.mkdir(mode=0o700)

        emit("admin_ready", run=run, path=control)

        while True:

            for path in directory.glob("*.request.json"):

                request_id = path.name.removesuffix(".request.json")

                try:

                    request = json.loads(path.read_text())

                    if request.get("run") != run:

                        raise ValueError("wrong run")

                    sid = request["sid"]

                    with SESSIONS_LOCK:

                        target = SESSIONS.get(sid)

                    if target is None:

                        raise ValueError("unknown session")

                    target.shutdown(socket.SHUT_RDWR)

                    target.close()

                    emit("admin_closed", run=run, sid=sid,

                         request_id=request_id)

                    result = dict(ok=True, run=run, sid=sid,

                                  request_id=request_id)

                except Exception as exc:

                    result = dict(ok=False, error=type(exc).__name__,

                                  request_id=request_id)

                temporary = directory / (request_id + ".response.tmp")

                temporary.write_text(json.dumps(result))

                temporary.replace(

                    directory / (request_id + ".response.json")

                )

                path.unlink()

            time.sleep(0.1)



    if control != "-":

        threading.Thread(target=admin, daemon=True).start()

    with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as listener:

        listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)

        listener.bind((bind, int(port)))

        listener.listen()

        emit("ready", run=run, port=int(port))

        while True:

            conn, peer = listener.accept()

            threading.Thread(target=worker, args=(conn, peer),

                             daemon=True).start()



def client(mode, host, port, run, source, source_port, *files):

    sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)

    sid, stage = None, "connect"

    try:

        sock.settimeout(3)

        sock.bind((source, int(source_port)))

        sock.connect((host, int(port)))

        stage = "hello"

        hello = receive(sock)

        if (hello["type"] != "hello" or hello["run"] != run

            or hello["peer"] != list(sock.getsockname())

            or hello["local"] != list(sock.getpeername())):

            raise ValueError("hello mismatch")

        sid = hello["sid"]

        emit("connected", **hello)

        if mode == "hold":

            with open(files[1], "x") as state:

                json.dump(hello, state)

        seq = 0

        while True:

            phase = (Path(files[0]).read_text().strip()

                     if mode == "hold" else "fresh")

            msg = dict(run=run, sid=sid, seq=seq, phase=phase,

                       nonce=secrets.token_hex(16))

            stage = "send"

            send(sock, msg)

            stage = "ack"

            if receive(sock) != dict(type="ack", **msg):

                raise ValueError("ack mismatch")

            emit("matched_ack", **msg)

            if mode == "probe":

                return 0

            seq += 1

            time.sleep(1)

    except Exception as exc:

        emit("client_end", run=run, sid=sid, stage=stage,

             source=source, source_port=int(source_port), port=int(port),

             error=type(exc).__name__)

        return 2

    finally:

        sock.close()



def close_session(path, run, sid):

    directory, request_id = Path(path), uuid.uuid4().hex

    temporary = directory / (request_id + ".request.tmp")

    with temporary.open("x") as handle:

        json.dump(dict(run=run, sid=sid), handle)

    temporary.replace(directory / (request_id + ".request.json"))

    response = directory / (request_id + ".response.json")

    end = time.monotonic() + 6

    while time.monotonic() < end:

        if response.exists():

            result = json.loads(response.read_text())

            emit("close_result", **result)

            return 0 if result.get("ok") is True else 2

        time.sleep(0.1)

    raise TimeoutError("local close response deadline")



if name == "__main__":

    mode, *args = sys.argv[1:]

    if mode == "serve":

        serve(*args)

    elif mode in ("hold", "probe"):

        sys.exit(client(mode, *args))

    elif mode == "close":

        sys.exit(close_session(*args))

    else:

        raise SystemExit("mode: serve, hold, probe, close")

With this fixture, each message echo includes the same run and sid. The server logs received, while the client logs matched_ack only after checking the entire reply. A hypothetical client record might contain {"event":"matched_ack","run":"R1","sid":"xxxx","seq":0,"phase":"baseline","nonce":"..."}. Match those request fields to the server receipt. Session UUID and network tuple establish connection identity; phase, sequence and nonce bind each challenge to the relevant observation. TCP buffers may join or split writes, so receive reconstructs frames byte by byte.

Bind every ACK to a session, phase and nonce

Each TCP chunk belongs to exactly one open socket, and we assign it a unique sid. The client never changes its sid or reconnects; if a send/recv fails, the client exits with an error event. The server tracks each socket in SESSIONS[sid]. Our protocol ensures each application-layer message (JSONL) is bounded to a short newline-delimited frame, with a strict three-second receive deadline. The server’s response echoes back the entire JSON message with "type":"ack" prepended, so the client can assert it was not corrupted or replayed.

All of these provide evidence of actual application-level processing, not just TCP-level SYN-ACK.

Since TCP does not align to JSON messages, the framing loop accumulates bytes until \n. If the session is terminated (e.g. shutdown closes the socket), the other side’s receive() will get EOF and raise an exception (emitting 'session_end' or 'client_end' events). By capturing both sides’ logs, we can correlate one server-close with one client-EOF.

All this ensures our evidence includes session identity (sid and network tuple), message sequence and nonces. This will let us prove exactly which data was sent when, and confirm it was processed by the app (not just TCP-layer). As a result, a matching received record on the server and matched_ack on the client means the packet passed the security group and the server app replied to it (it’s not enough to hear the client’s OS-level success).

Establish the baseline without reconnects

We now run the commands listed above in separate managed terminal sessions on the appropriate hosts, with a fresh phase.txt file containing “baseline” before the hold client starts. We expect:

  • Both server listeners print ready, and the target listener’s private administrative queue prints admin_ready.

  • The client hold should print one "connected" event with type="hello" and then a "matched_ack" for seq=0 with phase="baseline".

  • Both probe clients should print one "matched_ack" (seq=0, phase="fresh").

  • The server should show matching received events for the hold client and both fresh probes, with each receipt matching its client’s matched_ack fields.

  • The old client’s sid in old.json must match the server’s accepted sid. The network addresses in the logs should match the expected client and server private IPs and chosen ports.

If any of these conditions fails (no hello, wrong sid, missing ack, stale phase), we raise BASELINE_INVALID and inspect the error (perhaps an SG rule was missing or our listener wasn’t running). Once verified, the state is stable: the “old” session remains open and active, fresh sessions can be created, control port is fine, and the network tuple is known.

At this point, we have established the history (old.json, server logs) and state (open sockets and session ID). We will now insert a marker to delineate phases: typically by writing a new string to phase.txt. The client’s loop will read that and start sending with an after-readback phase containing a unique operator event token. We do not update the file until after the revoke/readback steps, so the accepted post-readback evidence contains that marker and a newly generated nonce.

An older buffered reply is not a new post-marker challenge. We use rename or similar to atomically replace phase.txt. This ensures a clean handoff between before/after without trusting the clocks on both hosts.

Revoke the captured rule and verify its absence

Now we perform the minimal mutation: remove the single ingress rule we identified. Since we saved the exact SecurityGroupRuleId ($RULE_ID), we do:

aws ec2 revoke-security-group-ingress \

    --group-id "$TARGET_SG" \

    --security-group-rule-ids "$RULE_ID" \

    --region "$AWS_REGION" \

    --output json > revoke-output.json

We capture the exit status and complete JSON output in revoke-output.json. If the CLI call fails with a nonzero exit code, we stop dependent steps. The AWS CLI revoke reference contains both examples described as producing no output and response fields such as Return; retain the actual response from the installed version rather than assuming an example is a measured result. A zero exit status is not a substitute for configuration readback. DryRunOperation indicates an authorization check without performing the mutation, not successful revocation.

Now we read back the actual state before proceeding. We run the same describe commands as before (same filters) to get updated JSON. We must see that sgr-abc123 (our $RULE_ID) is now absent from the results. AWS docs explicitly recommend verifying by describe after revoke. We also confirm no other rules changed unexpectedly (by comparing against our saved baseline). Only after complete successful readback do we atomically replace the client’s phase.txt with after-readback plus the unique operator event token.

The hold client continues to exchange framed messages. Qualifying post-readback challenges must read the new marker and generate their nonces afterward. This marker will also tag any further old-session messages as belonging after revocation.

(If for some reason the revoke call returned success but the describe still shows the rule ID present, or shows a permission we didn’t expect, we classify RULE_CHANGE_UNVERIFIED and abort. A readback discrepancy remains unresolved. AWS separately documents property-mismatch behavior that differs between default and non-default VPCs; this contract uses an exact rule ID in a non-default VPC. We will limit our observation window but keep in mind AWS does not promise atomic immediacy. If needed, we would wait a moment and re-describe to confirm absence.)

Compare the old socket with genuinely fresh connections

With the rule gone and phase marked, we continue the experiment. The old session (mode=hold) automatically sends a new message each second. We match the server’s received records to the old client’s matched_ack records. Simultaneously, we launch new probe attempts to the target port (18443) from unused source ports: e.g. 41502, 41503, etc. Each probe is paired with a control probe to port 18444. For each pair:

  • If a fresh probe to 18443 connects, completes its hello and gets a matching ACK, probe mode immediately closes that socket and records successful new admission at that point. A completed handshake without a valid application exchange is still not a connect-stage timeout.

  • If the probe to 18443 fails at the connect stage (times out at TCP SYN/ACK level), we record that as a timeout.

  • A local bind error, a wrong hello or ACK, and a failure after a completed handshake must be recorded separately. None is the intended connect-stage timeout. A protocol mismatch could mean, for example, that the wrong service or address was reached.

  • We require the paired fresh control on 18444 to succeed for the attempt to be valid: if control also fails, our test environment might be compromised or the server down. We would classify such a case as OBSERVATION_INCONCLUSIVE.

We continue probing for up to 60 seconds or until we see three consecutive target-probe timeouts with the control port always succeeding in between. We choose a 3-second operation timeout per attempt, and reserve enough of the overall budget for both probes and their connect, hello and ACK operations. If the required three-pair pattern never appears within 60 seconds (or the old session ends early), we halt as inconclusive.

If we do get repeated timeouts on 18443 while 18444 remains fine, and the old session continues to get acks, that is evidence that new admission is blocked while the old session is still active.

This waiting accounts for the documented possibility of a short propagation delay. We do not immediately conclude on a single failed probe, because AWS documents possible propagation delay. By requiring three timeouts over multiple source ports, we reduce the chance that a transient network glitch fooled us. Throughout, we log exact timestamps of each attempt to correlate with system logs.

At the end of this period we have one of these cases:

  • If the old session is still fully responding (matched_ack continues) and three consecutive fresh-target attempts timed out during the qualifying interval while controls passed, we classify NEW_ADMISSION_BLOCKED_OLD_SESSION_ACTIVE. The evidence will show identical sid and session info pre- and post-marker for old messages, and all new probes failing TCP handshake.

  • If a fresh attempt succeeds before the qualifying interval, new admission was available at that point. Preserve the success as transitional evidence rather than discarding it or assuming it proves an alternate rule. It resets the consecutive-timeout count. Persistent success through the observation budget requires alternate-rule, identity and path checks and an OBSERVATION_INCONCLUSIVE result, not a blocked-admission conclusion.

  • If controls fail or something else goes wrong, we label OBSERVATION_INCONCLUSIVE.

This yields objective outcomes from our bounded experiment. We do not interpret a timeout as definitely a permission denial (it could be dropped for other reasons), but by isolating all other variables (control port ok, old session ok, readback confirmed only rule gone) we support a limited interpretation consistent with the intended admission change, not a universal assertion about the cause of every timeout. We avoid logical leaps: if old session ended on its own, we won’t call it “SG cut it off”; it would simply become OBSERVATION_INCONCLUSIVE because the event doesn’t match our expected evidence for either state.

Treat propagation as an observation window

To spell it out: our local policy is a 60-second overall window, with each socket operation timed at 3s. We consider a consistent pattern of failed new connections with successful controls as evidence. However, we note explicitly that AWS does not guarantee a cutoff at any fixed time after revoke-security-group-ingress. The API docs even say a small delay may happen. So we label our findings based on what we observed, not an absolute rule.

If within our window, the required pattern never appears, we label OBSERVATION_INCONCLUSIVE. The 60-second window is a local acceptance policy, not an AWS propagation guarantee. Preserve early successes and the actual chronology; this deadline does not predict what would happen later.

Reconcile outcomes instead of declaring isolation

The table defines expected branches and the evidence required to classify a future authorized AWS run. These are acceptance criteria, not recorded AWS results.

State to classify

Required evidence

Classification

Baseline incomplete

Missing expected initial ACKs, incomplete rule listing, or control port fail

BASELINE_INVALID

Rule state unproved

API error, describe shows rule still present, or group identity mismatch

RULE_CHANGE_UNVERIFIED

Old ACKs continue; fresh target fails; controls pass

Same sid and tuple; new post-marker nonces; complete readback; three paired connect-stage timeouts; fresh controls and intervening old ACKs

NEW_ADMISSION_BLOCKED_OLD_SESSION_ACTIVE

Controls fail or observations conflict

Any stage error (bind fail, wrong identity) or mixed results

OBSERVATION_INCONCLUSIVE

Exact local closure observed (after admin close)

Affirmative response for the exact run and sid; server admin_closed; old-client EOF or reset within the window

FIXTURE_SESSION_CLOSED

Narrow admission recreated (fresh ACK after restore)

New rule ID returned by authorize, describe shows it, fresh ACK on 18443

NARROW_RULE_RESTORED

We never claim INSTANCE_ISOLATED. Even if old connections dropped, that alone doesn’t prove our SG mutation caused it; it might have timed out. Also, a single closed session doesn’t mean no others exist. Our goal is only to conclude about the tracked session vs new admission given the observed evidence. The limits of missing VPC Flow Logs as evidence illustrate the same boundary: the available observation must support the specific conclusion being claimed, not a broader isolation claim.

Notice the nuance: the critical path is the third row. Only when all conditions align (rule gone in AWS, old session acks still match sid, fresh-target times out thrice, fresh-control always ok) do we say new admission blocked and old session active. If any control fails or error in probing, we play it safe with INCONCLUSIVE. We keep all raw logs (phase markers, JSONL events, AWS outputs) so an incident reviewer can retrace steps.

If the old session ended on its own without an admin_closed event, we also say INCONCLUSIVE; it might have ended by timeout or error not related to our SG change.

This matches AWS’s own caveats: security group changes can result in “untracked” flows being immediately dropped or “tracked” ones persisting, but there’s no promise of a universal deadline. Our test is an acceptance policy: did things behave as documented for the narrow tracked case? If yes, classify as NEW_ADMISSION_BLOCKED_OLD_SESSION_ACTIVE; if not, classify appropriately.

Close exactly the controlled fixture session

To prove we can end the original session at will, we instruct the server to shut down that socket. We take the sid from the old client’s saved old.json and issue:

python3 socket_lab.py close "$SERVER_RUN/admin" "$RUN" "$SID"

This creates a request file in the receiver’s private administrative queue. The server’s admin thread should log admin_closed for that run and sid and atomically publish an affirmative response containing the request identity. Python socket shutdown and receive behavior explain the endpoint mechanics; the queue and its event names belong to this fixture, not an AWS API. An empty receive raises EOFError in the fixture and produces client_end. Capture the response and both endpoint events.

We require that within our observation window, we see the affirmative response for the correct run and sid, exactly one admin_closed, and a client-side socket error (EOF or reset) corresponding to that sid. If we get “unknown session” or timeout from the admin, or no client EOF, we consider this step failed (but still evidence). Meanwhile, we keep the server listening and allow fresh control connections to still work (since we only closed one socket object).

The target port 18443 remains gated by the missing rule, so new inbound to 18443 should still time out (we test a quick probe just to confirm). The control port (18444) should continue to allow connections unaffected. Buffered pre-close ACKs do not prove continued processing after closure.

Distinguish endpoint cooperation from incident containment

It’s worth emphasizing: this shutdown is done by the trusted server-side admin thread, not by AWS. In a real incident with a compromised host, you might not be able to run this script. So this action is not general containment; it’s just a way to test that we can end the session if we have local control. The matched pre-close challenges establish the old session’s activity; the administrative response and endpoint events establish its scoped closure.

However, it does not prove that every session (or all unknown sessions) would respond the same, nor that a malicious endpoint would allow closure. A comprehensive isolation would require network-layer controls beyond the SG and trusted cooperation. We are closing exactly one fixture session as a controlled experiment, not quarantining the instance completely. Broader AWS security incident-response preparation addresses a wider scope, but its generic containment guidance still requires separate authorization and evidence.

Restore the narrow rule and create a new session

Finally, we bring the configuration back to its original state (except the old session remains closed). We carefully reconstruct the save of $RULE_ID into the AWS CLI command:

aws ec2 authorize-security-group-ingress \

    --group-id "$TARGET_SG" \

    --ip-permissions file://restore-permissions.json \

    --region "$AWS_REGION" --output json > authorize-output.json

Here restore-permissions.json contains the rule array we built earlier (protocol, from/to ports 18443, client/32, and description if any). AWS will canonicalize the CIDR if needed (though a /32 is canonical). The CLI will return a new SecurityGroupRuleId for this added rule. We capture that and confirm it appears in the new describe output. Never assume the removed identifier is reused; record the returned rule identity.

We then repeat our fresh-target probe on port 18443 with a new source port. We expect this probe to succeed (eventually returning a matched ACK) now that the rule is back. If it does, we record NARROW_RULE_RESTORED and note the new sid/tuple from that connection. If it still fails, we record failure evidence.

We also ensure the rule readback shows the correct properties (protocol, port, CIDR, description). This completes the cycle: narrow rule removed, old kept alive, new blocked; then old closed, rule re-added, new allowed again. The old socket does not resurrect; it should remain closed. The evidence supports only restored narrow admission, not a rollback of connection history. Recheck current state before restoration, reconcile concurrent changes, and restore saved rule tags explicitly when present. Give fresh recovery probes a bounded observation window for propagation; do not broaden the allow rule to obtain a pass.

(If an error occurs on authorize or describe, we would classify RECOVERY_UNVERIFIED. We do not proceed further beyond this exercise.)

Set acceptance criteria and rerun the same contract

Our acceptance criteria for a successful test run are:

  • Baseline validated: initial exchange succeeded, and we captured session info.

  • Rule changed/verified: revoke call succeeded and describe confirms rule gone.

  • Old-session proof: matched_ack events with identical sid before and after revocation marker.

  • Fresh-control checks: all probes on port 18444 succeed unchanged.

  • New-target block: at least three post-marker probes to 18443 time out at TCP, with no control port failure.

  • Session close: admin_closed and client EOF observed.

  • Rule restore: new rule ID returned, describe shows it, new-target probe gets ack.

If all these are met, we declare the test passed with classification NEW_ADMISSION_BLOCKED_OLD_SESSION_ACTIVE (then FIXTURE_SESSION_CLOSED and NARROW_RULE_RESTORED). If only some parts work, we keep their partial evidence with appropriate classifications (for example, a success interrupts the required consecutive-timeout interval). Each run uses new RUN, new phase.txt, new source-port seeds; we never append to old logs. The contract is fixed, so after one successful run, we could stop (or rerun from scratch if needed for confidence).

We do not broaden the test to different settings (like varying timeouts or adding NACLs); we stick to the defined scope.

Assign ownership to the evidence and recovery decision

Finally, we label who owns what in this scenario. The Cloud/Network owner provided the security group and its rules; the snapshot of groups and rules (including the target SG ID) belongs to their account. The Application owner (or incident response lead) owns the session identity (SID) and the decision to close it via a trusted admin. The Incident lead is responsible for making the call on what to do (e.g. quarantine vs restore vs escalate), using this evidence.

The Reviewer or auditor ensures we preserved all original evidence: the rule inventory, describe JSONs, the phase marker, all logs of sent/received JSON (client/server). These artifacts are owned by the incident response team and should be archived.

We do not delete anything or purge logs until the decision is finalized and documented. The only disposable items are the test resources (the EC2 instances and SG can be torn down once the exercise and any formal report are complete). Any final cleanup (e.g. removing the test rule after restore) should be done only after preserving evidence, and ideally via separate commands after this log proves success.

This process of recording cloud security architecture decisions makes ownership and review explicit. We capture which account and role ran the experiment, what changes were made, how the session behaved, and who confirmed closure and recovery. In a real incident, such records would be part of the incident postmortem.

Practice cloud security decisions with Refonte Learning

This exercise embodies the principle of “test the control, don’t just assume it works.” As cloud security professionals, developing the habit of experimenting to verify permission changes is crucial. If you found this scenario enlightening, consider exploring hands-on curricula like Refonte Learning’s Cloud Security Engineer Essentials program, which covers IAM, incident-response planning and cloud security, with practical projects. Its FAQ names AWS Security Hub, Azure Security Center, Google Cloud Security Command Center and Splunk.

(Learners should verify on the current course syllabus whether specific labs like this are included.) Such programs emphasize not just theory but evidence-based practice. The next step is to practice designing similar tests: pick a cloud control, form a hypothesis of its effect, and write a concise runbook with verifiable checkpoints. Using structured playbooks and evidence logs, rather than guesswork, will improve any security response.