Cloud engineer reviewing terminal output and cloud storage object listings on a laptop and external monitor.

Why AWS CLI Text Output Can Choose the Wrong Largest Object article2

Mon, Oct 5, 2026

Cloud inventory scripts often use AWS CLI queries (JMESPath) to pick items such as “the largest object” by size. However, AWS CLI’s behavior differs between output modes: in text mode the CLI applies the query to each page of results separately, whereas in JSON mode it assembles the full result before querying. This means a --query filter like sort_by(Contents,&Size)[-1].[Key,Size] run with --output text can yield one maximum per page, not the true global maximum, misleading downstream wrappers like head or tail.

In this playbook we enforce a global selection contract: we declare and verify the complete set of objects (our manifest) and ensure our script’s result indeed comes from the full population.

We demonstrate with seven inert test objects (01.bin…07.bin of sizes 10, 70, 20, 90, 30, 80, 40 bytes respectively) in a dedicated S3 prefix. The unique global maximum is 04.bin (90 bytes). We capture both raw and reduced outputs in JSON and text modes, compare them, and document the exact AWS CLI calls and environment used. AWS documentation on output filtering, CLI pagination and S3 object listing explains the behavior of --query and pagination.

Code and manifest files are provided for repeatability, and we answer four success questions (did the CLI run, was the population complete, was query scope global, and did downstream processing preserve the correct result). Ultimately we decide to ACCEPT, REPAIR, RERUN or HOLD based on evidence. The result is a precise audit ensuring the chosen “largest” object truly reflects the entire prefix, not just one page.

Declare the global object-selection contract

We define “largest” strictly as the object with the maximum Size in our fixed prefix, not the first/last key by name or modification time. We have pre-authored an immutable manifest of seven keys with exact byte sizes:

Key

Size (bytes)

01.bin

10

02.bin

70

03.bin

20

04.bin

90

05.bin

30

06.bin

80

07.bin

40

In this list, the unique global maximum is 04.bin at 90. No two objects share that size. We will reconcile the CLI listing to this declared manifest; in other words, verify the output contains exactly these seven keys and sizes. We reconcile an explicitly declared S3 object population by checking the listing against our manifest. (This is the same reconciliation principle used in partial-delete audits: do not trust the CLI output until it exactly matches the known set.) Only after confirming completeness do we accept the result. No delete or mutation occurs here; this is a read-only audit.

The AWS S3 API (for general-purpose buckets) guarantees results are sorted lexicographically by key, not by size, so the natural order will not place the largest object at an extremity. This justifies our emphasis on full enumeration rather than relying on ordering. We will only accept a “largest” key once we have passed both an exact-population check and a selection check. The manifest (above) is entirely independent of any CLI output; we will not reverse-engineer it from the listing.

Pin the CLI, account and owned S3 prefix

We record all relevant AWS settings to bind our evidence to a specific environment. For example:

$ aws --version
aws-cli/2.36.44 Python/3.9.5 Linux/5.15.0   # example CLI build
$ aws sts get-caller-identity
{
    "UserId": "AROABCDEFGHIJKLMN:role-session",
    "Account": "123456789012",
    "Arn": "arn:aws:iam::123456789012:role/MyRole"
}
$ echo $AWS_REGION
us-west-2
$ echo $AWS_PROFILE
default

We choose an owned, general-purpose bucket (not a directory or access-point bucket) and a dedicated prefix. Let’s say BUCKET=my-demo-bucket and PREFIX=cli-pagination-test-20261005/ (a unique run ID). We will use --expected-bucket-owner 123456789012 on each s3api call to ensure the bucket is indeed ours. AWS S3 docs note that if the provided owner ID doesn’t match, the request fails with 403, so this explicitly ties the query to our bucket.

We make no account config changes; we do not disable TLS or print credentials. All setup writes go only to this inert prefix; measurement commands are only reads of that prefix. In short, we bind evidence to the actual cloud resource by specifying account, bucket and region.

The prefix is frozen for this run (no other processes may write to it), so that our listing reflects only our seven test objects. Minimal IAM permissions are needed: the setup needs s3:PutObject on this prefix and s3:ListBucket on it, and the measurement needs only s3:ListBucket.

Keep setup permission separate from measurement permission

To avoid any conflation of steps, we distinguish roles: setup has permission to upload our test objects (and nothing else in the bucket), while the measurement phase only needs read (list) permission on that prefix. We do not assume any automated cleanup; we expect the engineer to delete these inert objects when done. We do not attempt any atomic snapshot or bucket deletion; this is a controlled lab, not production. By separating write (upload) actions from the listing step, we ensure the listing truly observes the final object set.

Build the seven-object manifest before uploading

We create exactly seven local files of the specified sizes and an expected-manifest file before any listing. For example:

$ mkdir -p /tmp/s3test; cd /tmp/s3test
$ echo -n > manifest.json
$ echo '[' >> manifest.json
$ first=true
$ sizes=(10 70 20 90 30 80 40)
$ for i in {1..7}; do
    key=$(printf "%02d.bin" "$i")
    size=${sizes[$((i-1))]}
    head -c $size /dev/zero > "$key"    # create a file of given size
    if $first; then first=false; else echo ',' >> manifest.json; fi
    echo "{\"Key\":\"$key\",\"Size\":$size}" >> manifest.json
  done
$ echo ']' >> manifest.json

This yields a valid manifest.json containing our seven objects and sizes (in the table above). We then upload them and record outcomes:

$ for key in 01.bin 02.bin 03.bin 04.bin 05.bin 06.bin 07.bin; do
    aws s3 cp "$key" "s3://$BUCKET/$PREFIX$key" \
        --expected-bucket-owner 123456789012 \
        --region us-west-2
done

Expected output (example) from aws s3 cp:

upload: ./01.bin to s3://my-demo-bucket/cli-pagination-test-20261005/01.bin
upload: ./02.bin to s3://my-demo-bucket/cli-pagination-test-20261005/02.bin
...

Each line confirms the file was sent. We then verify all uploads succeeded: for instance, by running aws s3api list-objects-v2 --bucket $BUCKET --prefix $PREFIX --output text and checking for exactly 7 keys. If any upload had failed or an unexpected object appeared, we abort. Crucially, we do not generate the manifest from the CLI; it is our source of truth.

Make the winner independent of key ordering

Notice that the sizes (10, 70, 20, 90, 30, 80, 40) are arranged so the largest object (04.bin) is not last by key. Lexicographically the keys go 01,02,…,07, and 07.bin (40 bytes) is smaller than 04.bin (90). This ensures a naive wrapper like tail -n1 on the text output (which would pick the last page’s max) would be wrong, since the true max (90) is in the middle.

A quick local test confirms JMESPath’s sort_by(Contents,&Size)[-1] on the full list picks 04.bin/90, whereas on page-sized chunks it would pick 02.bin (70), 04.bin (90), 06.bin (80), 07.bin (40) as seen later. We establish the manifest so that our chosen “winner” is independent of key order, thereby exposing exactly this pitfall if one only looked at first/last lines.

Capture a complete JSON listing without reduction

Next, we retrieve the full listing in JSON without filtering. Using AWS CLI and the same page size (2) but no --max-items or query, we do:

$ aws s3api list-objects-v2 \
    --bucket "$BUCKET" --prefix "$PREFIX" --page-size 2 \
    --output json > full-list.json

This command (with --no-cli-pager implied) will make multiple API calls (page by page) but then output one combined JSON structure. AWS documentation confirms that with JSON/YAML output, the CLI “completely process[es] [the] output as a single structure before the --query filter is applied”. The result in full-list.json should contain all 7 objects. We then verify it matches our manifest:

import json, sys
expected = {
    item["Key"]: item["Size"]
    for item in json.load(open("manifest.json"))
}
data = json.load(open("full-list.json"))
contents = data.get("Contents") or []
result = {item["Key"]: item["Size"] for item in contents}
if expected != result:
    print("ERROR: Listing does not match manifest", file=sys.stderr)
    sys.exit(1)

This Python check confirms the CLI saw exactly our seven keys with the right sizes. A successful match here means the entire population was retrieved as a single JSON before any filtering. We do not proceed to reduction until this gate passes. The log should note that 7 items were found.

Run the same reduction in text and JSON modes

With the full listing in place, we apply our key query in both output modes. The JMESPath query we use is sort_by(Contents, &Size)[-1].[Key,Size], which sorts objects by Size ascending and picks the last element (the largest). According to the JMESPath specification, sort_by sorts ascending and preserves stability, so [-1] yields the maximum by size.

  • JSON mode: We run

$ aws s3api list-objects-v2 \
    --bucket "$BUCKET" --prefix "$PREFIX" --page-size 2 \
    --query "sort_by(Contents,&Size)[-1].[Key,Size]" --output json

The CLI assembles all pages and then applies the query once. The expected output is a single JSON array containing ["04.bin",90] (the global max). For example:

[
    [
        "04.bin",
        90
    ]
]

This correctly identifies 04.bin/90 from all 7. AWS documentation confirms that JSON output applies the query globally, so it gives one result line.

  • Text mode: We run the same command with --output text:

$ aws s3api list-objects-v2 \
    --bucket "$BUCKET" --prefix "$PREFIX" --page-size 2 \
    --query "sort_by(Contents,&Size)[-1].[Key,Size]" --output text

Because of --output text, AWS CLI applies the query to each page separately. With our 7 objects split into pages of 2, this yields four lines (one per page). In our case the output is:

02.bin    70
04.bin    90
06.bin    80
07.bin    40

Each line is the page’s local maximum by size. Note that one of these lines (04/90) is the true max, but the others are not. (In general the text-mode output will include the top item from each page.) This disparity is expected per AWS docs: text output “runs the query once on each page of the output”. In summary, JSON returns one correct global winner, while text returns multiple page-local winners.

Expose the shell wrapper that keeps the wrong line

Now consider what happens if a script blindly takes only the first or last line of the text-mode output. Save the text query result:

$ aws s3api list-objects-v2 \
    --bucket $BUCKET --prefix $PREFIX --page-size 2 \
    --query "sort_by(Contents,&Size)[-1].[Key,Size]" \
    --output text > result.txt

The file result.txt contains the four lines shown above. For example:

$ head -n1 result.txt
02.bin    70
$ tail -n1 result.txt
07.bin    40

Using head -n1 would pick 02.bin/70 as if it were the largest, and tail -n1 would pick 07.bin/40. Neither is the true global max. Importantly, neither wrapper establishes a global ordering by Size. This scenario highlights the danger: a successful CLI invocation (no errors) can still produce an incorrect final choice.

Indeed, one should recall that success responses do not settle every caller obligation (a lesson from SQS deletion validation); we must carefully check the logic, not just the exit status. In practice, this means avoiding such wrappers or using JSON output so that the query sees all data.

Vary page size while keeping the population fixed

We now test other --page-size values (while keeping the same seven-object prefix) to confirm behavior. The JSON-mode query should always yield 04.bin/90 (the global maximum) regardless of page size. The text-mode query yields different numbers of page results: fewer pages if larger page-size. For example:

  1. Page-size=1: Each object is its own page. Text query emits 7 lines, one per object (the object itself). The global max is still 04/90, but text output has all seven lines (01/10, 02/70, 03/20, 04/90, 05/30, 06/80, 07/40).

  2. Page-size=2: (our baseline above) 4 pages: text output lines 02/70, 04/90, 06/80, 07/40.

  3. Page-size=3: pages of 3,3,1. The text output lines would be 02/70, 04/90, 07/40.

  4. Page-size>=7: one page with all 7. Text output yields one line 04/90. (Any request ≥7 returns only 7 items because the bucket has 7 total.)

The following table summarizes:

Requested --page-size

# of Pages

Page-local maxima (Key/Size)

1

7

01/10, 02/70, 03/20, 04/90, 05/30, 06/80, 07/40

2

4

02/70, 04/90, 06/80, 07/40

3

3

02/70, 04/90, 07/40

≥7

1

04/90

In all cases, the JSON output would return 04.bin/90 once, as it considers all pages before sorting. The text output simply reflects more or fewer lines based on how many pages were used. This matches generic AWS guidance: changing --page-size only affects how many calls are made and how results are split, not which items appear.

Separate requested maximum page size from observed boundaries

It’s important to distinguish the requested page size from what S3 actually returns. S3 will never return more items than requested per page. For example, if we request 3, the first two pages may have 3 items, but the final page will have 1 (since only 7 total). AWS docs note “KeyCount will always be ≤ MaxKeys”; similarly, each page’s contents count is ≤ --page-size.

The observed page boundaries therefore depend on the total population. If we request 10 (larger than 7), S3 simply returns all 7 in one page. We do not assume a fixed number of pages; we always compute the full union of returned objects. In other words, the list of pages adapts to the population (and possible truncation at the end), in line with the API: “If you ask for N keys, your result will include N keys or fewer”.

Our experiment confirms these limits: in every case the JSON mode still found 7 objects and the same global max, while text mode had the per-page maxima shown above.

Test a truncated JSON result as a negative control

So far we always retrieved the full list. As a negative control, we artificially truncate the JSON output using --max-items. For instance:

$ aws s3api list-objects-v2 \
    --bucket "$BUCKET" --prefix "$PREFIX" \
    --page-size 2 --max-items 2 \
    --output json > truncated.json

With --max-items 2, AWS CLI will only return up to 2 items and include a NextToken in the output. For example, truncated.json might show only 2 object entries (say 01.bin and 02.bin) and a NextToken. The query on this output (if applied) would yield whichever of those 2 is larger, but this is not the true max of the full set. We check using our manifest gate and see only 2 keys returned (not 7), so we fail the manifest check. In other words, we detect this as an incomplete population: the JSON output has been truncated by the client.

We should not then try to continue fetching (e.g. by passing the NextToken as a starting-token) within the same logical query, because the CLI has already applied the cut-off. We simply conclude the result is invalid for global selection.

This example shows why JSON output alone is not enough; if someone uses --max-items (or inadvertently --no-paginate), the query scope is limited. In that case, our playbook would recommend RERUN without the truncation.

Do not confuse --no-cli-pager with --no-paginate

A common misunderstanding is to mix up --no-paginate with the CLI’s output pager. In AWS CLI v2, --no-cli-pager simply disables sending output to a pager (like less), it does not affect which data is retrieved. In contrast, --no-paginate or --max-items affect the AWS API calls: they can stop the retrieval of further pages.

For example, disabling CLI pagination (--no-paginate) on an S3 list means you get only the first page (default 1000 items). In our case with 7 objects, disabling pagination would still return all 7 (since default page size ≥7). But in larger buckets it would not.

In any case, remember: --no-cli-pager is unrelated to data completeness: it only changes how the output is displayed, whereas --no-paginate stops at the first page of data.

Reduce only after validating the assembled population

We now perform the reduction step only after confirming the manifest. As a final check, we run a reliable consumer on full-list.json to output the largest object, or “NO_MATCH” if empty. For example:

import json, sys
data = json.load(open("full-list.json"))
contents = data.get("Contents") or []
if not contents:
    print("NO_MATCH")
    sys.exit(1)
# Verify manifest match (again)
expected = {
    item["Key"]: item["Size"]
    for item in json.load(open("manifest.json"))
}
result = {c["Key"]: c["Size"] for c in contents}
if (
    set(result.keys()) != set(expected.keys())
    or any(result[k] != expected[k] for k in result)
):
    print("ERROR: Inventory mismatch", file=sys.stderr)
    sys.exit(1)
# Select largest
max_obj = max(contents, key=lambda x: x["Size"])
print(max_obj["Key"], max_obj["Size"])

This code uses only the full, validated data. It never trusts the text “None” string as a key (if Contents were missing, we handle it by or []). It prints 04.bin 90 in this run. If Contents had been empty, it would print “NO_MATCH” and exit, signaling an empty result set. This fully decouples reduction logic from listing logic: we only compute the maximum after we are sure we have the intended population.

Capture a reproducible query-scope ledger

Throughout these steps, we keep a detailed log. Each run is tagged with a unique run ID and timestamp. The log records the CLI version, account, region, bucket and prefix, and the exact commands run (including --output, --query, --page-size, etc.). We also capture stdout, stderr, and the exit status of each command. For instance:

$ run_id=$(date -u +"%Y%m%dT%H%M%SZ")
$ aws --version 2> /tmp/log
$ echo \
    "RunID: $run_id, Bucket: $BUCKET, Prefix: $PREFIX, CLI: $(aws --version 2>&1)" \
    >> /tmp/log
$ aws s3api list-objects-v2 \
    --bucket $BUCKET --prefix $PREFIX --page-size 2 \
    --output json > full-list.json 2>> /tmp/log
$ sha256sum full-list.json >> /tmp/log
$ aws s3api list-objects-v2 \
    --bucket $BUCKET --prefix $PREFIX --page-size 2 \
    --query "sort_by(Contents,&Size)[-1].[Key,Size]" \
    --output text > text-result.txt 2>> /tmp/log
$ echo "Text-mode result: $(cat text-result.txt | sed 's/^/    /')" >> /tmp/log
$ aws s3api list-objects-v2 \
    --bucket $BUCKET --prefix $PREFIX --page-size 2 \
    --query "sort_by(Contents,&Size)[-1].[Key,Size]" \
    --output json > json-result.json 2>> /tmp/log
$ echo "JSON-mode result: $(jq -c . json-result.json)" >> /tmp/log
$ echo "Script exit status: $?" >> /tmp/log

This preserves the raw outputs and digests. We avoid sharing sensitive debug in the report, but an engineer’s log will include the terminal output. We also follow good scripting practice (e.g. set -o pipefail) to preserve CLI failures in shell logging (a common trap).

Distinguish four success questions

Our log/ledger must explicitly answer four questions:

  1. Did the command finish successfully? Check each exit code. If any CLI call failed (non-zero), we abort (HOLD).

  2. Was the intended population retrieved? Confirm the listing matches the manifest (7 objects). If not (as with --max-items), we do not trust the result.

  3. Was the expression evaluated at the intended scope? Confirm whether we used JSON (global) or text (page-local). If text mode was used, we know the query was per-page.

  4. Did the downstream consumer preserve the correct result? For example, did we take head, tail, or otherwise pick the wrong line from text output? This is a manual check.

A “yes” on 1 and 2 is necessary to accept the data; a “yes” on 3 ensures global correctness; a “yes” on 4 ensures we did not introduce an error in the final step. We do not assume any one “yes” implies the others.

We also ensure that any stderr output (such as AWS debug or errors) is logged separately. We do not rely on partial outputs; every filter result is checked against our ledger.

Handle empty populations, ties and changing writers

Our baseline manifest uses unique sizes. But we consider controls:

  • Empty prefix: If we had no objects in the prefix, the listing would show Contents missing or empty. Our reduction code would then print “NO_MATCH”. That is the correct behavior for an empty population. A missing Contents should not cause an unhandled error.

  • Ties: In our test there are no ties. If two objects shared the maximum size, our query [-1] would arbitrarily pick one (based on stable sort). This is not a robust rule. A true audit should either define a tie-break (e.g. smallest key lexicographically) or return all tied keys for manual decision. In practice, any tie in size should be treated as ambiguous, and the script should HOLD for manual review rather than blindly choosing one. (JMESPath’s behavior under ties is an internal detail, so it’s safer not to rely on it.)

  • Concurrent writes: We assume a writer-free window. Our static lab disabled any background writes to the prefix. In a dynamic environment, S3 is strongly consistent for new objects, but if concurrent uploads/deletions were happening, our pre-post inventory comparison would not be an atomic snapshot. We must HOLD if we cannot guarantee no changes occurred during the listing. (Achieving a transactional snapshot in S3 would require different tools, which is beyond this audit’s scope.)

In summary, as long as our seven objects remain unchanged during the test, and no two are tied for largest, our logic stands. Any deviation (unexpected keys, equal sizes, or active writers) invalidates the single-run result and must be treated cautiously.

Choose accept, repair, rerun or hold

We now consolidate the possible outcomes in a decision matrix:

  • Accept (no change needed): JSON query on complete list gave the correct result (04.bin 90), and our manifest check passed. The wrapper (if any) also captured that line properly. In this case the selection is globally correct.

  • Repair query/output: If a text-mode query was used and a wrapper took a wrong line, we accept the listing but repair the script (e.g. switch to JSON output or collect all lines). The inventory is correct but our use of it was wrong. We report this as a query logic fix (e.g. “use --output json or remove head/tail”).

  • Rerun with full population: If --max-items or --no-paginate caused truncation, we must rerun without those limits. The query itself isn’t broken, just the input was incomplete.

  • Hold (incomplete seed or list): If the initial manifest upload didn’t complete, or if the listing misses expected objects (by error or writers), we do not trust the result. This is a HOLD condition until environmental issues are resolved.

  • Hold (ambiguous/tied values): If two objects tied for max size, or if bucket consistency is in doubt, we HOLD. No automatic resolution is safe without a tie-break policy.

  • Hold (unexpected population): If extra objects appear or an owner check fails, that is a potential security or bucket issue, so we hold.

Importantly, we do not immediately act on the selected object (for example, by deleting or re-encrypting it). First, we produce a corrected report. Any downstream action (delete, migrate, etc.) must have its own approval and exact identity. In other words, we always separate object selection from retention authority.

The report is reconciled and repaired independently; any later mutation of the “winner” would follow a separate workflow (as outlined in S3 object-hold/release processes).

Assign the inventory owner and regression triggers

Finally, we assign roles and maintenance. The inventory owner (e.g. cloud ops lead) is responsible for the manifest and verifying runs periodically. The CLI-script author or DevOps engineer who wrote the query pipeline must be involved, and another reviewer should validate the results. We record the exact AWS CLI version in our ledger; whenever the AWS CLI or JMESPath library is updated, we should re-run this test to catch any changes in filtering behavior.

Likewise, if our tooling (e.g. shell, jq/Python) or target bucket type changes (e.g. switching to a directory bucket or an access point) this lab should be rerun. We preserve manifest.json under version control so it is never regenerated by the script.

For broader production use, note that real inventories (with frequent updates or many objects) might require stricter consistency controls or pagination logic (e.g. polling until no continuation token). That is beyond our static lab, but should be documented: this playbook assumes a fixed set of objects and S3’s current consistency model.

Build the cloud foundations behind reliable automation

Robust command-line automation ties together cloud identity, precise queries, and clear validation. Practices like explicit owner checks, exact manifests, and separate reporting-versus-action paths are foundational to reliable DevOps.

For engineers interested in deepening their cloud skills, consider the Refonte Cloud Development Program. It is a 3-month, part-time (12–15 hours per week) training covering major cloud platforms and tools (AWS/Azure/GCP, Docker, Kubernetes, Terraform, etc.) with no prior experience required. The curriculum includes cloud architecture, infrastructure-as-code, security and more, and participants earn a certificate upon completion.

Evaluating this program’s current curriculum and prerequisites can help further your understanding of topics like AWS CLI, automation, and cloud governance.