Suggested for an image showing an engineer at work: DevOps engineer reviewing Bash pipeline exit codes and tee log output on a workstation.

Keep tee From Turning a Failed CI Test Green

Mon, Sep 28, 2026

A test command can print an unmistakable failure, stream that output through tee, and still hand CI a zero exit status. That is not a contradiction. Output and process status are separate channels. tee can faithfully display the producer's diagnostics while Bash, under its default pipeline rule, reports only the last command's status. When tee succeeds, a producer that exits 7 can therefore appear green to the caller. GNU Bash documents this last-command rule, and documents pipefail as a different aggregate rule rather than a record of every component status. GNU Bash 5.3 pipeline semantics are the fixed reference family for this playbook.

The acceptance target here is deliberately narrow: a required test producer and a required GNU tee log sink must both succeed. The wrapper must preserve a nonzero producer status, independently detect sink failure, emit both component statuses, and return a caller-visible nonzero result whenever either required component fails. Four cases prove that contract: producer/sink success, producer failure only, sink failure only, and both failing. The decision after evidence collection is equally explicit: accept the wrapper, repair it, rerun an inconclusive test, or hold the release. A green result proves only this wrapper contract for the declared commands and environment; it does not prove test discovery completeness, durable log storage, deployment correctness, or application correctness.

Define the wrapper’s success contract

Start with the release decision, not with shell folklore. The wrapper has two required components: the producer, which represents the test command, and the sink, which represents the file write performed by GNU tee. Success means producer status = 0 and sink status = 0. Any other vector is a failure of the wrapper's declared gate, even when output is visible in the CI console.

Treat the two evidence forms as separate acceptance dimensions rather than redundant signals. A component vector explains which required command failed; the process status exposed to the caller determines whether the outer gate can stop. A wrapper that records the vector but exits zero is diagnostically interesting but release-unsafe. A wrapper that returns nonzero but discards the vector may stop the release, yet still make triage ambiguous when producer and sink can fail independently.

The wrapper must therefore expose two forms of evidence. First, it must retain the component vector (producer_rc, sink_rc) for diagnostics. Second, it must return one process exit status to its caller. This playbook proposes an application policy: if the producer is nonzero, return the producer status; otherwise, if the sink is nonzero, return the sink status; otherwise return zero. That precedence is not Bash's built-in pipefail rule. Bash with pipefail returns the rightmost nonzero component, which means a simultaneous producer failure and tee failure can aggregate to the sink's code instead of the producer's.

This contract belongs at the execution boundary before broader CI/CD pipeline boundaries such as image publication or deployment. Refonte Learning's Kubernetes CI/CD article describes those wider build, test and deployment stages; this playbook intentionally stops at the shell wrapper that transports test status through a logging pipeline.

Acceptance therefore asks two questions separately: did the test producer pass, and did the required diagnostic sink succeed? Do not collapse them into “the pipeline failed.” If the test failed, engineers need the producer's status. If the log sink failed, engineers need that evidence too. If invocation evidence is missing, the result may be inconclusive even when an exit code exists.

Record the actual shell invocation

Shell reliability work begins by identifying what actually ran. Record the operating system, Bash and coreutils versions, executable paths, script arguments, shell options, SHELLOPTS, file modes, fixture revision, working directory, and exact producer/sink command. Do this before interpreting a CI color.

The research baseline is September 26, 2026. The technical source claims used here were rechecked on September 28, 2026; no material discrepancy was found for the behavior relied upon. The disposable local execution used for this article was performed on September 28 and named local-lab-2026-09-28/de97c1f7f0ea. It ran Debian GNU/Linux 13 (trixie), Linux kernel 6.18.44, /usr/bin/bash version 5.2.37(1)-release, and /usr/bin/tee from GNU coreutils 9.7. The Bash 5.3 manual, edition 5.3 updated May 18, 2025, remains the reference family; the observed runtime was not Bash 5.3, so the local result is evidence for the named build, not a fabricated 5.3 execution. The official manual identifies itself as Bash 5.3 and gives that update date.

A minimal manifest command set is:

printf 'fixture_revision=%s\n' 'refonte-tee-lab-v1'
printf 'pwd=%s\n' "$PWD"
uname -a
cat /etc/os-release
printf 'bash_path=%s\n' "$(command -v bash)"
bash --version | head -n 1
printf 'tee_path=%s\n' "$(command -v tee)"
tee --version | head -n 1
printf 'SHELLOPTS=%s\n' "$SHELLOPTS"
printf 'exported_SHELLOPTS=%s\n' "$(env | sed -n 's/^SHELLOPTS=//p')"
set -o
stat -c '%A %a %U:%G %n' . producer.sh wrapper.sh bad-sink FIXTURE_REVISION

In the observed local lab, the work directory was mode 0700; producer.sh and wrapper.sh were 0755; bad-sink was an existing directory at 0755; and FIXTURE_REVISION was 0644. The outer collection shell reported SHELLOPTS=braceexpand:errexit:hashall:interactive-comments:nounset, with no exported SHELLOPTS environment entry; set -o showed errexit=on, nounset=on, and pipefail=off. A fresh child started by bash ./option-probe.sh reported errexit=off, nounset=off, and pipefail=off before that script changed any options.

This is the kind of provenance needed to reconcile a local result with a CI result rather than assuming that the two shells shared settings. The lab process happened to run as root, which is exactly why the negative sink is a directory rather than a “non-writable” file. A permission-denial fixture can accidentally succeed under elevated privileges; opening an existing directory as a regular tee output file fails for the intended type reason instead.

For the observed producer-failure/sink-success case, the caller invocation was:

bash ./wrapper.sh P7_S0 7 refonte-P7_S0-20260928 good-fail.log

The wrapper's declared producer and sink arguments corresponded to:

bash ./producer.sh 7 refonte-P7_S0-20260928 2>&1 | tee good-fail.log

For the both-failed case, the sink operand was bad-sink. Recording those arguments matters because a different path, option, redirection order, or interpreter can create a different contract even when the log text looks similar.

Do not infer child-shell options from the parent

bash script.sh and ./script.sh are not identical invocation statements. The first explicitly runs the bash found by command lookup and treats script.sh as its script argument; the second asks the operating system to execute the file, which then uses its shebang, such as #!/usr/bin/env bash, to locate an interpreter. Record which form CI used.

Do not say that an inner Bash automatically inherits the parent's set -e or set -o pipefail. In the observed lab, bash ./option-probe.sh started with errexit=off and pipefail=off even though the invoking shell had errexit=on, because SHELLOPTS was not exported. A separate controlled run that explicitly exported SHELLOPTS with pipefail enabled caused the child Bash to start with pipefail=on. The operational rule is therefore to inspect the child configuration, not infer it from the parent. The GNU Bash variables documentation, including SHELLOPTS and PIPESTATUS, is the reference page for these shell variables.

Build a producer and sink that fail predictably

The fixture must fail only for reasons it owns. It needs no network, credentials, deployment target, container runtime, or privileged filesystem mutation. That makes a rerun safe when evidence is inconclusive and keeps the laboratory independent from security gates that depend on correct execution status, which are a separate concern.

Create an owned disposable directory and the producer:

work=$(mktemp -d "${TMPDIR:-/tmp}/refonte-tee-lab.XXXXXX")
cd "$work"
printf '%s\n' 'fixture_revision=refonte-tee-lab-v1' > FIXTURE_REVISION

cat > producer.sh <<'BASH'
#!/usr/bin/env bash
set -u

if (( $# != 2 )); then
  printf 'usage: %s EXIT_STATUS RUN_MARKER\n' "$0" >&2
  exit 64
fi

status=$1
marker=$2
case $status in
  0|7) ;;
  *)
    printf 'producer: unsupported status=%s\n' "$status" >&2
    exit 64
    ;;
esac

printf 'RUN_MARKER=%s\n' "$marker"
printf 'DIAGNOSTIC_MARKER=%s status=%s\n' "$marker" "$status" >&2
exit "$status"
BASH

chmod 0755 producer.sh
mkdir bad-sink

The producer writes a unique run marker to stdout, writes a diagnostic marker to stderr, and exits exactly 0 or 7. It performs no deployment and no external write. A regular path such as good.log is the good sink. The existing directory bad-sink is the intentional failing sink.

GNU tee copies standard input to standard output and to the files named as operands. Its documented output-error behavior includes reporting failures on non-pipe outputs and returning a nonzero status for an output error. See the GNU Coreutils tee invocation manual.

The four cases are:

Case

Producer argument

Sink operand

Required vector shape

Wrapper result

P0_S0

0

writable file

(0, 0)

0

P7_S0

7

writable file

(7, 0)

7

P0_SF

0

existing directory

(0, nonzero)

sink status

P7_SF

7

existing directory

(7, nonzero)

7

On local-lab-2026-09-28/de97c1f7f0ea, GNU tee 9.7 returned 1 for the directory sink, so the observed vectors were (0,0), (7,0), (0,1), and (7,1). Treat 1 as an observed build result, not as the application contract. The contract only requires the sink failure to be nonzero.

Reproduce the default last-command false green

Make the masking defect concrete before repairing it. With a valid writable sink, run exactly the shape that causes trouble:

log=default.log
bash producer.sh 7 2>&1 | tee "$log"

Under Bash's default pipeline semantics, the pipeline's status is the status of its last command, unless pipefail is enabled. Bash also waits for a foreground pipeline to complete before returning its status. Therefore, if the producer exits 7 but tee successfully writes the log and exits 0, the aggregate pipeline status is 0. That is documented behavior, not a new feature or a CI-specific bug.

Do not try to collect PIPESTATUS and $? from the same pipeline by inserting assignments in an ad hoc order. Keep aggregate-status demonstrations in separate shell processes so they cannot corrupt the component vector you intend to study. For example:

set +e
bash -c '
  set +o pipefail
  bash ./producer.sh 7 refonte-default-demo 2>&1 | tee default.log
'
default_rc=$?

bash -c '
  set -o pipefail
  bash ./producer.sh 7 refonte-pipefail-demo 2>&1 | tee pipefail.log
'
pipefail_rc=$?
set -e

printf 'default_rc=%d pipefail_rc=%d\n' "$default_rc" "$pipefail_rc"

The named local build observed default_rc=0 and pipefail_rc=7. This is an actual execution result for Bash 5.2.37/coreutils 9.7, and it agrees with the Bash 5.3 manual semantics cited above.

Separate output visibility from execution success

Because 2>&1 is attached to the producer before the pipe, both its stdout marker and stderr diagnostic flow into tee. Seeing DIAGNOSTIC_MARKER=... status=7 in console output or in a successful log file proves that bytes were emitted and copied through that path. It does not change the producer's exit status, and it does not force the default pipeline aggregate to become nonzero. Bash documents 2>&1 | as the behavior represented by the |& shorthand, including the relationship between the redirection and pipe.

The converse matters too. A green step is not proof that all expected tests were discovered, and a marker in a log is not proof that the log is complete or durably stored. This wrapper acceptance test establishes only the declared process-status and sink-check contract. Keep those evidentiary boundaries explicit when deciding whether a release gate is trustworthy.

Apply pipefail and inspect what it does not encode

set -o pipefail repairs one important class of false green: a nonzero command earlier in a pipeline can now make the aggregate pipeline nonzero. Bash defines the aggregate as the value of the rightmost command that exits nonzero, or zero when all components succeed. GNU Bash's pipeline documentation and the set builtin documentation both specify that rule.

For the four-case fixture, let S be the nonzero tee status:

Case

Default aggregate

pipefail aggregate

Information retained by aggregate alone

P0_S0

0

0

both appear successful

P7_S0

0

7

failure detected; producer code happens to survive

P0_SF

S

S

sink failure detected

P7_SF

S

S

failure detected, but producer 7 is not preserved

The local build, where S=1, observed the both-failed pipefail aggregate as 1, not 7. That is exactly what “rightmost nonzero” predicts. pipefail is therefore a useful aggregate failure detector, but it is not a two-element diagnostic record and it is not the producer-first precedence policy required here.

Use pipefail where its aggregate rule is the desired shell contract. Do not describe it as equivalent to “return the test's exit code.” When the sink can also fail, those are different policies.

Capture PIPESTATUS before another command replaces it

Bash provides PIPESTATUS, an array containing the statuses from the most recently executed foreground pipeline, including a pipeline containing a single command. That last detail explains why evidence disappears so easily: after another simple command runs, that command is itself the most recent foreground pipeline. The GNU Bash variables documentation for PIPESTATUS is the reference definition.

The practical rule is to copy the complete array immediately. The first simple command after the pipeline must be:

bash ./producer.sh "$producer_status" "$marker" 2>&1 | tee "$log"
statuses=("${PIPESTATUS[@]}")

Do not put echo, a test, another helper, or rc=$? between those lines. The array assignment works because Bash expands ${PIPESTATUS[@]} from the preceding pipeline while executing that first assignment; the copied statuses array can then be used by later commands.

A compact collector is:

#!/usr/bin/env bash
set +e

case_id=$1
producer_status=$2
marker=$3
log=$4

bash ./producer.sh "$producer_status" "$marker" 2>&1 | tee "$log"
statuses=("${PIPESTATUS[@]}")

if (( ${#statuses[@]} != 2 )); then
  printf 'capture error: expected 2 statuses, got %d\n' "${#statuses[@]}" >&2
  exit 70
fi

printf 'CAPTURE case=%s producer_rc=%d sink_rc=%d\n' \
  "$case_id" "${statuses[0]}" "${statuses[1]}"

This collector is diagnostic only; the accepted wrapper later applies release policy.

Prove the evidence-capture bug with a negative test

The deliberately broken form is:

bash ./producer.sh 7 refonte-broken-capture 2>&1 | tee broken.log
rc=$?
statuses=("${PIPESTATUS[@]}")
printf 'aggregate=%d vector=(%s)\n' "$rc" "${statuses[*]}"

The rc=$? line is itself a simple command. By the time the following line expands PIPESTATUS, the original two-command vector has been replaced. On the named local build with default pipefail off, the producer failed, tee succeeded, rc captured 0, and the later vector was (0) instead of (7 0).

A proper negative test therefore asserts both vector length and component identity. For P7_S0, immediate capture must produce length two with producer 7 and sink 0. A capture after an intervening assignment or echo must not be accepted as evidence for that pipeline. Merely checking “some aggregate was nonzero” is insufficient because it cannot establish which required component failed, and in the masking case it may be zero anyway.

Write an explicit failure-propagation wrapper

The accepted wrapper should be a standalone process with its own declared behavior. It must allow the pipeline to finish, capture PIPESTATUS immediately, log both statuses, and then choose its caller-visible code according to producer-first precedence. It should not depend on whatever errexit or pipefail state an outer CI shell happens to have.

#!/usr/bin/env bash
set -u

if (( $# != 4 )); then
  printf 'usage: %s CASE_ID PRODUCER_STATUS RUN_MARKER LOG_PATH\n' "$0" >&2
  exit 64
fi

case_id=$1
producer_status_arg=$2
run_marker=$3
log=$4

producer_cmd=(bash ./producer.sh "$producer_status_arg" "$run_marker")
sink_cmd=(tee "$log")

printf 'WRAPPER case=%s producer_argv=' "$case_id" >&2
printf '%q ' "${producer_cmd[@]}" >&2
printf 'sink_argv=' >&2
printf '%q ' "${sink_cmd[@]}" >&2
printf '\n' >&2

# Intentionally permit the required pipeline to return so its vector can be copied.
set +e
"${producer_cmd[@]}" 2>&1 | "${sink_cmd[@]}"
statuses=("${PIPESTATUS[@]}")

if (( ${#statuses[@]} != 2 )); then
  printf 'WRAPPER case=%s internal_error=status_vector_length actual=%d\n' \
    "$case_id" "${#statuses[@]}" >&2
  exit 70
fi

producer_rc=${statuses[0]}
sink_rc=${statuses[1]}
printf 'WRAPPER case=%s producer_rc=%d sink_rc=%d\n' \
  "$case_id" "$producer_rc" "$sink_rc" >&2

if (( producer_rc != 0 )); then
  exit "$producer_rc"
fi
if (( sink_rc != 0 )); then
  exit "$sink_rc"
fi
exit 0

set +e is deliberate here, not an excuse to ignore failures. After the pipeline, the script does only vector validation, diagnostic reporting, explicit status selection, and exit. There is no broad || true, no continue-on-error, and no final exit 0 that can erase a required failure. The caller must inspect the wrapper process's actual exit status.

The precedence is an application contract:

producer_rc != 0  -> wrapper_rc = producer_rc
else sink_rc != 0 -> wrapper_rc = sink_rc
else               -> wrapper_rc = 0

Bash does not supply that producer-first rule through pipefail; the wrapper implements it after obtaining the complete vector.

Audit the errexit contexts that change control flow

set -e is not universal exception handling. The Bash set documentation describes several contexts where a nonzero status does not trigger immediate shell exit: tests used by if, command lists following while or until, non-final commands in && and || lists, non-final pipeline commands subject to the pipefail state, and statuses inverted by !. See the GNU Bash set builtin documentation.

Minimal examples show why context matters:

# Standalone failure: errexit can terminate the script.
set -e
false
echo 'not reached'

# if-test context: nonzero is used as a condition.
set -e
if false; then
  echo 'not reached'
fi
echo 'still running'

# Non-final command of an AND list: failure controls the list.
set -e
false && echo 'not reached'
echo 'still running'

# Negation context: status is deliberately inverted.
set -e
! false
echo 'still running'

These are not edge-case curiosities. They are why the statement “we use set -e, therefore every failing test stops CI” is too strong. Bash explicitly conditions errexit behavior on syntactic context, and its manual further notes effects when compound commands or functions execute in contexts where -e is being ignored.

Pipelines add another dimension. Without pipefail, producer | tee can aggregate to zero when only the producer fails. With pipefail, the same standalone pipeline can aggregate nonzero and interact with errexit. But place the pipeline inside an if test and the if context changes how errexit applies. This is why a laboratory that wraps the whole experiment in a conditional can accidentally test a different control-flow context from the production script.

Keep deliberate status collection narrowly scoped

The accepted wrapper intentionally suppresses automatic exit long enough to capture the vector. That is acceptable only because the suppression is paired immediately with explicit validation and a nonzero failure path. Do not generalize it to a large script that continues performing unrelated work after failures.

A useful review question is: “What commands can execute after the required pipeline fails?” In this wrapper the answer should be limited to copying PIPESTATUS, checking vector shape, printing the two statuses, choosing precedence, and exiting. If deployment, publication, mutation, or unrelated test commands occur in that window, the scope is too broad and the wrapper should be refactored.

The important distinction is between controlled observation and failure suppression. The former lets the failed pipeline return so the wrapper can inspect its components and then enforces policy. The latter simply continues without restoring a meaningful nonzero outcome. Only the first pattern belongs on the accepted path.

Validate log-sink failure independently

A test wrapper that preserves producer failure but ignores logging failure is only half repaired when the log sink is required. Exercise P0_SF first: producer success removes test failure from the equation, so any nonzero sink status is independently attributable to the intentional directory operand.

marker='refonte-P0_SF-local'
bash ./producer.sh 0 "$marker" 2>&1 | tee bad-sink
statuses=("${PIPESTATUS[@]}")
printf 'producer=%d sink=%d\n' "${statuses[0]}" "${statuses[1]}"

On the named local build, GNU tee printed:

tee: bad-sink: Is a directory

The immediate vector was (0 1). The accepted wrapper returned 1. That is consistent with GNU tee's documented behavior of diagnosing non-pipe output errors and returning failure when an output error occurs.

Then P7_SF observed (7 1), logged both values, and returned 7 because the proposed contract gives producer failure precedence.

This is also where pipefail and explicit policy visibly diverge. In a separate aggregate demonstration with both components failing, the local pipefail result was 1, the rightmost nonzero status. The explicit wrapper returned 7, preserving the producer while still recording sink_rc=1. That difference is intentional and should be part of code review, not left as an accidental property of shell options. Bash's documented rule supports the expected aggregate distinction.

Do not swallow the tee error merely because the test itself passed. If usable diagnostic capture is a declared release requirement, sink failure makes the gate fail. Equally, do not overclaim what sink success means: a zero tee status and expected markers show that this invocation wrote successfully according to the tested command, not that remote retention, later upload, replication, or archival durability has been established.

Check the GitHub Actions boundary separately

GitHub Actions adds another process boundary. Current GitHub workflow-syntax documentation says that on Linux/macOS an unspecified shell uses the internal template bash -e {0} when Bash is available, whereas an explicitly selected shell: bash uses bash --noprofile --norc -eo pipefail {0}. GitHub explicitly notes that the unspecified and explicit Bash forms run different commands. GitHub Actions workflow syntax documents those templates.

The same documentation says the runner executes a temporary file containing the run commands, and documents fail-fast behavior for the built-in Bash shell, including -o pipefail when shell: bash is explicitly specified.

That distinction is exactly why “GitHub uses Bash” is not enough evidence. Pin the shell you intend to test. The following is an unexecuted integration example; no GitHub repository or hosted runner was used for this article, so there is no invented runner build, job identifier, temporary-script path, or observed GitHub result to report.

name: tee-wrapper-regression

on:
  workflow_dispatch:

jobs:
  wrapper-regression:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Record runner and shell boundary
        shell: bash
        run: |
          printf 'runner_os=%s\n' "$RUNNER_OS"
          printf 'runner_arch=%s\n' "$RUNNER_ARCH"
          printf 'image_os=%s\n' "${ImageOS-unknown}"
          printf 'image_version=%s\n' "${ImageVersion-unknown}"
          command -v bash
          bash --version | head -n 1
          command -v tee
          tee --version | head -n 1
          set -o
          printf 'SHELLOPTS=%s\n' "$SHELLOPTS"

      - name: Run owned local wrapper regression
        shell: bash
        run: bash ./regression.sh

The local fixture remains independently runnable; Actions is only a secondary integration boundary. This avoids turning the article into a tutorial on the wider DevOps toolchain, which covers broader tooling concerns.

GitHub also documents the success/failure role of process exit status: zero maps to success and nonzero to failure for actions. GitHub's exit-code guidance makes that mapping explicit.

Do not let CI configuration erase the repaired status

A correct child wrapper is insufficient if an outer layer deliberately converts failure to success. Inspect continue-on-error, custom shell templates, parent scripts, || true, unconditional exit 0, and any command that runs after the wrapper and becomes the parent script's final status. GitHub documents the step-level continuation setting and separately documents shell exit-code handling; a review must therefore verify that failure is not deliberately tolerated at the outer boundary.

The accepted boundary is simple:

bash ./regression.sh

The regression script must return its final status unchanged to the Actions shell, and that shell must return failure to the runner when the regression fails. GitHub's documented Bash shell handling ties the script's exit status to step success/failure unless workflow configuration intentionally changes the treatment.

Do not infer that the inner bash ./regression.sh inherited the Actions shell's -e or pipefail; the inner script must own its options and checks. Likewise, do not claim an “actual GitHub shell command” unless a run recorded it. The documented template bash --noprofile --norc -eo pipefail {0} is evidence about GitHub's configured explicit-Bash behavior, not a transcript of this unexecuted example.

Package a reproducible wrapper regression test

A regression test should turn the four-case matrix into assertions, not screenshots. It should create fresh sinks, generate a unique run prefix, call the wrapper as a separate process, inspect caller exit status, verify producer and diagnostic markers, and retain per-case stdout/stderr. It should also record the environment manifest so a later shell, runner, or coreutils change can be reconciled.

#!/usr/bin/env bash
set -u

failures=0
run_id="refonte-$(date -u +%Y%m%dT%H%M%SZ)-$$"
evidence="evidence-${run_id}"
mkdir "$evidence"
bad_sink="$evidence/bad-sink"
mkdir "$bad_sink"
good0="$evidence/good-0.log"
good7="$evidence/good-7.log"
: >"$good0"
: >"$good7"

check_case() {
  local case_id=$1 producer_status=$2 sink=$3 expected_caller=$4
  local marker="${run_id}-${case_id}"
  local out="$evidence/${case_id}.out"
  local err="$evidence/${case_id}.err"
  local rc

  set +e
  bash ./wrapper.sh "$case_id" "$producer_status" "$marker" "$sink" \
    >"$out" 2>"$err"
  rc=$?
  set -e

  if (( rc != expected_caller )); then
    printf 'FAIL %s caller expected=%d actual=%d\n' \
      "$case_id" "$expected_caller" "$rc" >&2
    failures=$((failures + 1))
  fi

  if ! grep -Fq "RUN_MARKER=$marker" "$out"; then
    printf 'FAIL %s missing stdout marker\n' "$case_id" >&2
    failures=$((failures + 1))
  fi

  if ! grep -Fq \
      "DIAGNOSTIC_MARKER=$marker status=$producer_status" "$out"; then
    printf 'FAIL %s missing diagnostic marker\n' "$case_id" >&2
    failures=$((failures + 1))
  fi

  case $case_id in
    P0_S0) expected_vector='producer_rc=0 sink_rc=0' ;;
    P7_S0) expected_vector='producer_rc=7 sink_rc=0' ;;
    P0_SF) expected_vector='producer_rc=0 sink_rc=' ;;
    P7_SF) expected_vector='producer_rc=7 sink_rc=' ;;
  esac

  if ! grep -Fq "$expected_vector" "$err"; then
    printf 'FAIL %s status vector evidence mismatch\n' "$case_id" >&2
    failures=$((failures + 1))
  fi

  if [[ $case_id == *_SF ]] && grep -Fq 'sink_rc=0' "$err"; then
    printf 'FAIL %s sink unexpectedly succeeded\n' "$case_id" >&2
    failures=$((failures + 1))
  fi
}

check_case P0_S0 0 "$good0" 0
check_case P7_S0 7 "$good7" 7

# Discover this build's nonzero tee status rather than hard-coding 1.
set +e
bash ./wrapper.sh PROBE_SF 0 "${run_id}-PROBE_SF" "$bad_sink" \
  >"$evidence/PROBE_SF.out" 2>"$evidence/PROBE_SF.err"
sink_caller=$?
set -e

if (( sink_caller == 0 )); then
  printf 'FAIL probe: directory sink returned zero\n' >&2
  exit 1
fi

check_case P0_SF 0 "$bad_sink" "$sink_caller"
check_case P7_SF 7 "$bad_sink" 7

if (( failures != 0 )); then
  printf 'REGRESSION FAIL failures=%d\n' "$failures" >&2
  exit 1
fi

printf 'REGRESSION PASS run_id=%s\n' "$run_id"
exit 0

This script distinguishes expected from observed values. Producer status is specified exactly. Sink failure is required to be nonzero. The build-specific sink code is probed and then used for the sink-only caller assertion rather than pretending that every GNU tee build or failure class must use the exact code observed here.

The final local fixture was executed with the explicit wrapper shown above. The four cases returned caller statuses 0, 7, 1, and 7 for P0_S0, P7_S0, P0_SF, and P7_SF, respectively. The sink-failure cases recorded tee: .../bad-sink: Is a directory, and the wrapper recorded (0,1) and (7,1). The regression harness completed with status 0. Those are local observations from the named disposable environment, not GitHub Actions observations.

For stronger evidence, parse and assert the complete wrapper status line rather than relying only on substrings. The manifest should travel with the case evidence:

fixture_revision
run_id
OS and kernel
bash path and version
tee path and coreutils version
working directory
file modes
outer set -o state
outer SHELLOPTS
whether SHELLOPTS was exported
wrapper invocation argv
producer argv
tee argv
producer_rc
sink_rc
wrapper_rc
caller_rc

Package the manifest and evidence directory with the same run ID, but do not confuse retention with acceptance. Cleanup should delete only the owned disposable fixture directory created by mktemp; after evidence has been copied where policy requires, leave the fixture directory and remove exactly the recorded path:

cd /
printf 'removing owned fixture: %q\n' "$work"
rm -rf -- "$work"

Never generalize that cleanup into a broad wildcard or parent-directory deletion.

An evidence bundle should make it possible to answer: which revision ran, which executable paths were resolved, which options were active in the relevant process, which exact arguments were passed, what each pipeline component returned, and what the caller observed. If one of those fields is absent after a runner or shell change, classify the comparison as incomplete rather than filling the gap from memory. The purpose of the fixture is not to produce a decorative CI artifact; it is to make status propagation falsifiable.

Decide whether to accept, repair, rerun or hold

The evidence should drive one of four decisions. “Rerun” is not a generic response to red CI; it is appropriate when the previous result is inconclusive and repeating the operation is safe. This fixture is intentionally side-effect free, so rerunning it is safe. A real test suite may have different repeatability constraints and requires its own authority.

Evidence

Decision

Primary owner

Release meaning

Four cases produce correct vectors; caller results follow producer-first precedence; invocation and environment are recorded

Accept wrapper

Wrapper maintainer

Wrapper contract is qualified for the declared build and tested invocation

P7_S0 returns 0, vector is lost, sink failure is ignored, or parent maps nonzero to zero

Repair

Wrapper or CI owner

Gate is not trustworthy; prior green must not be treated as wrapper-qualification evidence

Status vector or invocation record is missing/corrupted, runner terminates unexpectedly, or test does not complete

Rerun inconclusive test

Test-infrastructure owner

No defensible pass/fail conclusion yet; rerun only when the operation is safe to repeat

Producer returns nonzero on a conclusive run, required sink fails, or wrapper/CI contract remains unqualified for the release

Hold release

Release authority with owning engineer

Required gate has not been satisfied

A crucial distinction is between a failed test and a failed gate mechanism.

P7_S0 with correctly captured (7,0) is a conclusive producer failure. The wrapper worked; the test did not. Hold the release until the test defect, product defect, or intended expected behavior is resolved under the project's normal authority.

P0_SF is different. The producer passed, but the required evidence sink failed. If successful diagnostic capture is part of the acceptance contract, the release gate remains unsatisfied even though the test itself returned zero.

A default-pipeline P7_S0 that appears green is different again. The producer failed, but the wrapper reported the successful tee result. That is a transport defect in the gate. Repair the wrapper and treat the earlier green as unqualified because the mechanism did not preserve producer failure according to the intended contract. Bash's default last-command rule explains the mechanism; it does not make such a wrapper acceptable for a producer-required release gate.

The “rerun” branch belongs only to inconclusive evidence. Examples include a missing status vector, a truncated execution record, a runner ending before the wrapper returned, or uncertainty about which shell actually interpreted the script. Do not use rerun to erase a conclusive producer failure. Conversely, do not call a test “failed” merely because its mandatory logging sink failed; record the producer and evidence-path outcomes separately.

Historical production test results deserve conservative handling. Once you discover that a wrapper could mask producer failures, you cannot retroactively infer which old green runs contained hidden failures unless independent evidence preserved the producer statuses. Console text may help investigation, but marker presence alone does not recreate a trustworthy exit contract. The release owner should hold affected decisions until sufficient independent evidence exists or a safe rerun produces a conclusive result.

Record producer failure and evidence failure separately

Use separate fields in incident or release records:

producer_rc
sink_rc
wrapper_rc
caller_rc
fixture_revision
invocation
environment_manifest
evidence_complete

Avoid a single label such as “pipeline error.” It destroys the distinction needed for ownership and recovery.

For example, (7,0) means the test failed but log writing succeeded. (0,1) on the observed GNU tee build means the test passed but the required sink failed. (7,1) means both failed; the accepted wrapper returns 7 while retaining sink_rc=1.

Missing PIPESTATUS, missing invocation evidence, or a parent that overwrote the wrapper's status is different again: the test result may be inconclusive even when some output exists. Preserve those distinctions all the way into the release record.

Assign ownership for the gate and the diagnostic path

Three owners should be named even when one person fills multiple roles. The test author owns what producer status 0 or 7 means and whether a test rerun is semantically safe. The wrapper maintainer owns the process contract: exact invocation, immediate PIPESTATUS capture, producer-first precedence, and sink enforcement. The CI configuration owner owns the outer shell, step continuation policy, runner changes, and preservation of the wrapper's final status.

This ownership model complements broader DevSecOps tool responsibilities without turning shell acceptance into a product ranking. Refonte's contextual article groups CI/CD and security tooling by function; the narrower lesson here is that a tool's presence does not qualify the process-status boundary around it.

Require regression review after changes to Bash, coreutils, runner image, shebang, script invocation form, shell template, wrapper code, or sink semantics. A version change is not automatically a failure, but it is a trigger to rerun the owned four-case matrix and compare evidence.

A good review record should answer two questions independently. First: did the implementation continue to conform to the declared producer/sink contract? Second: is the evidence sufficient to know that the same shell and invocation boundary were exercised? Passing the first while losing the second is not enough to extrapolate qualification to an unknown runner configuration.

If installed behavior disagrees with the Bash 5.3 reference semantics or GNU tee documentation relevant to the contract, record the discrepancy and hold acceptance for that build rather than selecting whichever result is convenient. Bash's current reference pages document the pipeline and errexit rules used here, while GNU coreutils documents the sink's output-error contract.

Strengthen DevOps practice with executable failure contracts

Reliable automation is built from executable contracts that survive composition. A test command, a pipe, a log copier, a wrapper process, and a CI runner each have their own status boundary. The practical skill is to state which failures are required, capture evidence before it is overwritten, and make the final caller status match the release policy.

The same discipline applies when teams add automated database-change checks: the higher-level check is only as trustworthy as the execution boundary that transports its result. That contextual Refonte article discusses broader database automation; it does not establish the Bash semantics in this playbook.

For this fixture, acceptance is deliberately modest. The wrapper is acceptable only when all four local cases pass, both statuses are logged, producer status wins when nonzero, sink-only failure remains nonzero, and the caller sees the wrapper's exit unchanged. The evidence should also identify the tested shell, tee build, paths, options, fixture revision, arguments, and file conditions.

That answers the operational question directly: the naïve producer | tee wrapper does not preserve a failing producer when tee succeeds under Bash's default pipeline rule. pipefail detects such failure but, when both components fail, follows Bash's rightmost-nonzero rule rather than preserving the producer. An accepted wrapper therefore needs immediate PIPESTATUS capture and an explicit producer-first policy while independently treating a required tee sink failure as nonzero.

A green accepted wrapper means only that its declared commands and required sink checks passed under the qualified invocation. It does not establish complete test discovery, artifact trust, successful deployment, or correctness of a deployed system.

Refonte Learning's DevOps Engineering program lists a three-month format at 12–14 hours per week and includes Linux fundamentals and scripting plus CI/CD in its published curriculum. Those confirmed areas align with practicing small, verifiable shell and pipeline exercises such as this one; the program page is not evidence for the Bash or GNU tee behavior described above.