Backend engineer debugging Redis Lua script errors and retained writes across multiple monitors

Redis Stopped the Script. Which Writes Stayed?

Wed, Sep 30, 2026

When a Redis Lua script hits a runtime error, engineers often assume everything was undone or never executed. In fact, Redis guarantees atomic visibility (no other client sees partial effects during execution), but it does not automatically roll back earlier writes. This can confuse engineers expecting SQL-style rollback. Backend engineers and maintainers need to verify which writes occurred before the error, and then decide how to handle the partial state.

We will examine a controlled fixture where a script does INCR then LPUSH, with the LPUSH failing on purpose. We trace the exact key changes, compare variants with redis.call vs redis.pcall, test a retry, and test adding a guard with TYPE. Each case has a known starting state, a fixed script, and independent key reads.

We use published Redis documentation and evidence expectations at every step. The outcome is not a “cache vs database” treatise but a precise state-accuracy check.

(For broader context, this kind of disciplined state audit complements the broader backend readiness skills needed for production reliability.) Ultimately we issue one of ACCEPT, REPAIR, HOLD or RECONCILE based on the exact observed state.

Separate a script error from a no-effect guarantee

A Redis Lua script executes atomically and blocks other commands; no other commands interleave, but atomic visibility is not the same as “transactional rollback”. In our scenario, the script has three sequential commands: INCR, LPUSH, then SET. We intentionally arrange for LPUSH to throw a WRONGTYPE error after INCR. This poses three distinct claims to examine:

  • No interleaving: Redis ensures no other client can see intermediate state. During the script, the database is effectively locked, so we do not worry about races. However, this lock alone does not roll back partial writes.

  • Script termination on error: By default, redis.call('LPUSH', ...) will throw an error and stop the script. Any writes before the error remain in the database. So the script execution will not proceed to the SET if LPUSH fails.

  • Rollback myth: Redis does not silently undo writes if a script fails. (By contrast, some database transactions might.) As a historical point, Redis transaction documentation explicitly notes that a transaction error does not trigger rollback; all commands still execute. For Lua scripts, an error just aborts remaining commands; there is no built-in undo.

Our task: detect exactly which keys were changed by the time of error. That allows the caller to decide next steps. We test each variant under a known initial state, then read all keys to collect evidence. If any outcome is ambiguous (e.g. network timeout, concurrent writes, or undefined state), we must HOLD and gather more info rather than blindly retry or accept.

This narrow experiment (no clusters, single node) is like a minimal reproducible scenario. Think of it like setting up a single-component integration test, similar in spirit to thorough database integration practices that isolate one subsystem. We will explicitly enumerate server/client versions, keys, arguments, and script content so readers can trust the precise behavior.

Record the execution path before touching state

Before running any scripts, we must document the environment and script code exactly. Assume a standalone Redis server (e.g. Redis 8.x on localhost:6379, database 0, default ACL user, no cluster or AOF interference).

Note that debugging mode is off; we are not running a SCRIPT DEBUG YES forked session. (In asynchronous debug mode, Redis would fork and roll back changes, but we do not use that, because it would discard the very effects we need to observe.) Our execution will use normal EVAL (atomic, blocking) or redis-cli --eval in synchronous mode.

We pick a unique prefix for our keys, e.g. fixtest:. The three keys are:

  • fixtest:counter (initial value "0", string that represents an integer)

  • fixtest:ledger (initial value "not-a-list", a simple string intentionally invalid for list ops)

  • fixtest:marker (initially absent)

We will also record the Lua script content and its SHA1 digest (for repeatability). For example:

cat << 'EOF' > script.lua
local n = redis.call('INCR', KEYS[1])
redis.call('LPUSH', KEYS[2], ARGV[1])
redis.call('SET', KEYS[3], 'done')
return n
EOF
echo -n "local n = redis.call('INCR', KEYS[1])
redis.call('LPUSH', KEYS[2], ARGV[1])
redis.call('SET', KEYS[3], 'done')
return n" | sha1sum

The above prints the SHA1 digest of the script bytes. We keep that in our records. We also note that this script, in RESP2 context, returns an integer or throws an error (not a status). The ACL and persistence settings are default; no restrictions or Redis Modules alter behavior. We will use redis-cli -n 0 to select database 0 for all commands. This strict inventory prevents missing details.

Exclude forked debugger state from acceptance evidence

A common pitfall is to accidentally run scripts under a debugger (SCRIPT DEBUG YES) and misinterpret results. The asynchronous debug mode forks the server and discards all changes at the end, which would hide the very state we want to inspect. We explicitly do not use any debug mode; we run the script normally.

In other words, our acceptance evidence comes only from a real EVAL execution, not a debug session. If any test log were produced under SCRIPT DEBUG, we would disregard it because it’s not representative of normal execution. We assume a plain execution path as a standard client would see.

Define the counter, ledger and marker contract

Our test fixture has these elements outside of the script:

  • Keys: We will always use exactly three keys:

- KEYS[1] = fixtest:counter,

- KEYS[2] = fixtest:ledger,

- KEYS[3] = fixtest:marker.

  • Arguments: The script takes one ARGV value: ARGV[1] = "e1" (an inert label for the event).

  • Initial State Ledger:

-  fixtest:counter (string "0").

-  fixtest:ledger (string "not-a-list").

-  fixtest:marker (absent).

We keep this ledger outside the script. When reading back state, we use appropriate commands: GET for strings, TYPE and LRANGE for lists. In particular, we will use TYPE fixtest:ledger rather than GET on it after it becomes a list, to distinguish between “no key” and “empty list”. The TYPE command returns, as a simple string, the type or none if the key doesn’t exist.

In Lua scripts under RESP2, calling redis.call('TYPE', key) yields a table with an ok field. So, for example, if ledger is absent, redis.call('TYPE', KEYS[2]).ok == "none"; if it's a list with elements, .ok == "list", and if it were still a string, .ok == "string".

We also define what “absent” means: a key either truly does not exist in DB or has been deleted. If fixtest:marker is absent, EXISTS will report 0. We will use EXISTS fixtest:marker (or check that GET fixtest:marker returns nil) to confirm absence. We treat an empty list differently: if fixtest:ledger were an empty list, TYPE would return “list” and LRANGE fixtest:ledger 0 -1 would return an empty array.

The contract of the unguarded script (the one we test first) is simply: run INCR, then LPUSH, then SET, returning the counter. The script has no guard and no explicit error handling.

We will compare its actual effects and return to the expected behavior for each scenario. For any test, we declare the initial ledger state, invoke the fixed script, then read back each key and the return or error. In that way we know exactly what happened before the script errored.

Build a resettable local fixture

We present here the exact shell commands to prepare and reset our fixture, run scripts, and probe final state. We always confine operations to the fixtest: prefix. No FLUSHALL, no cluster changes, no persistence toggles are used. We will show a successful control run first as baseline, then the failing cases.

# Fixture reset: set initial values for each case
redis-cli -n 0 SET fixtest:counter "0"  # initialize counter to "0"
redis-cli -n 0 SET fixtest:ledger "not-a-list"  # ledger as a plain string (wrong type)
redis-cli -n 0 DEL fixtest:marker  # ensure marker is absent

# Verify initial ledger (expected: counter=0, ledger type=string, marker absent)
redis-cli -n 0 GET fixtest:counter  # -> "0"
redis-cli -n 0 TYPE fixtest:ledger  # -> "string"
redis-cli -n 0 EXISTS fixtest:marker  # -> (integer) 0

We see the expected initial state: counter is "0", ledger is a string, and marker does not exist.

Next we load the unguarded script (the original body) into a file, although we could inline it. In these examples we’ll use redis-cli --eval with the script file for clarity:

# Save the original script to a file
cat << 'EOF' > script_unprotected.lua
local n = redis.call('INCR', KEYS[1])
redis.call('LPUSH', KEYS[2], ARGV[1])
redis.call('SET', KEYS[3], 'done')
return n
EOF

# Run the script (this is a healthy control case: assume ledger was NONE)
# First, simulate a case where ledger is absent (healthy scenario).
redis-cli -n 0 DEL fixtest:ledger  # remove fixtest:ledger (now nonexistent)
redis-cli -n 0 SET fixtest:counter "0"  # reset counter
redis-cli -n 0 EVAL "$(cat script_unprotected.lua)" 3 \
    fixtest:counter fixtest:ledger fixtest:marker e1
# Expected output (integer) 1 because counter was 0 -> 1

In this healthy control run, ledger was absent so LPUSH should create the list. The expected result is:

  • Return: 1 (the new counter)

  • fixtest:counter should now be "1"

  • fixtest:ledger should be a list with one element ["e1"]

  • fixtest:marker should be "done"

We verify:

redis-cli -n 0 GET fixtest:counter  # -> "1"
redis-cli -n 0 LRANGE fixtest:ledger 0 -1  # -> 1) "e1"
redis-cli -n 0 GET fixtest:marker  # -> "done"

All as expected. This is the normal success path.

However, our actual test is the wrong-type scenario: fixtest:ledger is a plain string. Before doing that, we should ensure the control leaves a clean state if needed. Now reset again for the failing case:

# Reset to initial wrong-type state for failing invocation
redis-cli -n 0 SET fixtest:counter "0"
redis-cli -n 0 SET fixtest:ledger "not-a-list"
redis-cli -n 0 DEL fixtest:marker

# Run the same script, expecting LPUSH to fail
redis-cli -n 0 EVAL "$(cat script_unprotected.lua)" 3 \
    fixtest:counter fixtest:ledger fixtest:marker e1

At this point, script_unprotected.lua tries to INCR fixtest:counter, then LPUSH fixtest:ledger e1. The INCR will succeed and set counter to 1. The subsequent LPUSH sees that fixtest:ledger is a string, not a list, so it throws a WRONGTYPE error. In RESP2, redis-cli will report an error message. For example, it might print:

(error) WRONGTYPE Operation against a key holding the wrong kind of value

At that point the script aborts before doing the SET. We label this transcript as Expected:

# Expected output, not an observed run
(error) WRONGTYPE Operation against a key holding the wrong kind of value

Now we record the post-error state by reading each key in a new connection:

redis-cli -n 0 GET fixtest:counter  # expected "1"
redis-cli -n 0 TYPE fixtest:ledger  # expected "string"
redis-cli -n 0 EXISTS fixtest:marker  # expected (integer) 0

Expected results: counter "1", ledger still type "string" with value "not-a-list", marker absent. The LPUSH did not run, and SET did not run. Critically, counter was incremented before the error. LPUSH’s error did not undo the INCR. This is our baseline evidence: Counter was changed (from 0 to 1), ledger and marker unchanged.

Locate the first completed mutation and first failing command

We can trace the order. The script made exactly two calls before failing:

  • redis.call('INCR', KEYS[1]) (completed, returning 1)

  • redis.call('LPUSH', KEYS[2], ARGV[1]) (this is where the runtime exception occurred).

To be precise, Redis compiled the script (no compile error happened), executed line 1, updated the counter, then hit line 2. The ledger being a string caused an immediate runtime error and control returned to the client as (error) WRONGTYPE. No further lines (SET or return) were executed. We distinguish this from, say, a script syntax error (which would fail before any execution) or an ACL violation (which would abort at the start); here we intentionally reached the middle.

At this point our independent reads confirm that INCR did happen. We must not conflate “script stopped” with “script rolled back”. The evidence table will clearly show counter=1, marker absent, which contradicts any idea that no changes happened.

Read every key after the failed invocation

It is crucial to always verify the database state independently of the error return. Simply seeing an (error) is not enough; we need to match it to what actually happened. Using a fresh client or CLI instance (to simulate an independent observer), we do:

redis-cli -n 0 GET fixtest:counter  # <- "1"
redis-cli -n 0 TYPE fixtest:ledger  # <- "string"
redis-cli -n 0 TYPE fixtest:marker  # <- "none"

(We could use EXISTS for marker, but TYPE is fine too.) The results should match our expected ledger: counter is "1", ledger is type "string" (still holding the scalar "not-a-list"), and marker is effectively non-existent. This triplet constitutes the post-error state.

Any test run where we cannot tie reads to that exact invocation (for instance, if another client had changed keys in between) should be invalid. But in our controlled lab, nothing else is touching fixtest: keys.

We log these values in a ledger for the scenario. If they had not matched expectation, we would flag that test as inconclusive. Fortunately, they match exactly: only counter changed.

Note: we have avoided using something like redis-cli --scan or key wildcards because we rely on exactly three known keys. Also, we do not inspect AOF or dump files. The only evidence is the key values and types.

Compare redis.call with an explicit pcall continuation

Next, we modify the script to use redis.pcall for the LPUSH, and to proceed even after the error. This shows how pcall changes control flow. We reset again first:

# Reset for pcall test
redis-cli -n 0 SET fixtest:counter "0"
redis-cli -n 0 SET fixtest:ledger "not-a-list"
redis-cli -n 0 DEL fixtest:marker

# Script with pcall
cat << 'EOF' > script_pcall.lua
local n = redis.call('INCR', KEYS[1])
local res = redis.pcall('LPUSH', KEYS[2], ARGV[1])
redis.call('SET', KEYS[3], 'continued')
return res
EOF

redis-cli -n 0 --eval script_pcall.lua \
    fixtest:counter fixtest:ledger fixtest:marker , e1

In this script, the pcall('LPUSH', ...) will catch the WRONGTYPE error and return an error object in res, rather than aborting. The script then continues to do SET marker 'continued' and finally returns the res value. When the script returns a Redis error reply object (from pcall), redis-cli will display it as an error, even though the script itself executed to completion.

We expect the client to still see an error (since we return the error result). However, state changes differ: here INCR runs (counter->1) and SET marker 'continued' runs (marker created), but LPUSH does nothing (it caught the error). The final return is an error table (same WRONGTYPE). So after this run, we check keys:

redis-cli -n 0 GET fixtest:counter  # -> "1"
redis-cli -n 0 TYPE fixtest:ledger  # -> "string"
redis-cli -n 0 GET fixtest:marker  # -> "continued"

Observed: counter = "1", ledger type = "string", marker = "continued".

(This is an unusual state: marker was set even though the LPUSH failed.) This tells us that using pcall changed which commands ran, not that it magically made LPUSH succeed. In particular, pcall turned the runtime error into a returned value, allowing later lines to execute.

The script log might show (error) WRONGTYPE ... as the return, but the state (counter incremented, marker set) is non-atomic from the caller’s perspective. This is why we will not treat this as a successful “all-or-nothing” operation. The returned error table is just a Lua value, not an implicit rollback.

An error table changes control flow, not history

It is important to explain that redis.pcall does not retroactively fail earlier commands. The pcall call causes only the LPUSH itself to produce an error object; it does not roll back the earlier INCR. And since we explicitly redis.call('SET', KEYS[3], 'continued') after that, the marker was set. One might be tempted to think “oh, pcall made everything succeed,” but it didn’t; it only prevented the script from aborting prematurely.

If we had returned immediately after the pcall, that would be a different program. Here we deliberately continue after catching the error. In short: the sequence of effects is what changed, not an atomic rollback.

Returning the error object at the end is also our choice. If instead we had returned n or some success message, the calling client would see a normal return. But we left it an error to emphasize that an error occurred (and to compare with the original variant).

Importantly, even though the marker is written, the client still sees an error. The caller must not assume that a returned Lua error implies nothing happened. Here, we have evidence the opposite: two writes happened (counter and marker). This shows that pcall does not align with any “undo” guarantee; it only alters control flow.

Expose the effect of an unguarded retry

Another scenario to examine: what if the caller, upon seeing the error, tries the same script again (maybe thinking it might succeed next time)? Since our fixture is disposable, we simulate a “retry” by simply invoking the unmodified script a second time without resetting the keys (but still in our test environment). Starting from counter=0, ledger=string, marker absent:

  • First call (we already did): counter went to 1, then error. Ledger still string.

  • Second call, immediately after, with counter="1", ledger="not-a-list": again INCR will make counter 2, then LPUSH errors.

For completeness:

# Already in wrong-type state (ledger is string) with counter=0
# First call (error) made counter=1, ledger string, marker absent.
# Now do a second call (retry) without resetting.
redis-cli -n 0 SET fixtest:counter "1"  # simulate post-first-call state
redis-cli -n 0 SET fixtest:ledger "not-a-list"
redis-cli -n 0 --eval script_unprotected.lua \
    fixtest:counter fixtest:ledger fixtest:marker , e1

After the second call, we expect:

redis-cli -n 0 GET fixtest:counter  # -> "2"
redis-cli -n 0 TYPE fixtest:ledger  # -> "string"
redis-cli -n 0 EXISTS fixtest:marker  # -> (integer) 0

So now counter=2, ledger still string, marker still absent. We would log this state too. The key point: if a retry is done blindly after the first failure, counter increments again (now 2) but we still have no new evidence in ledger or marker.

The error state is persistent. We only “know” that two calls occurred because counter is 2, but we have no separate log of the event except our counter.

In an application, a network timeout could look similar: maybe the client didn’t get the response but the script actually ran. We caution not to misinterpret: any retry of an unguarded script on the same keys will repeat the partial effect.

We cannot conclude “it didn’t run last time so try again” because it did run. This is why an error is not idempotent or equivalent to “no operation.” If a retry is attempted, it will create additional effect (as we see with counter).

Move the known validation before the writes

Given that we know the error happens only when fixtest:ledger is a string, the safest fix is to check the type first before incrementing. In other words, validate the input/state before doing any mutation. This is analogous to safe API input-validation boundaries: we refuse the bad case immediately. We write a new script:

local t = redis.call('TYPE', KEYS[2])
if t['ok'] == 'string' then
  return redis.error_reply("ERR ledger is string")
end
local n = redis.call('INCR', KEYS[1])
redis.call('LPUSH', KEYS[2], ARGV[1])
redis.call('SET', KEYS[3], 'done')
return n

In RESP2 mode, redis.call('TYPE', ...) returns a table like {ok="string"}, {ok="list"}, or {ok="none"}. We check if it’s "string", and if so we call redis.error_reply to send a Redis error (terminating the script with an error message of our choice). Otherwise (if ok is "none" or "list"), we proceed safely. We save this as script_guarded.lua.

Test cases:

  1. Guarded wrong-type state: Start with the original wrong-type initial state. Run the guarded script.

redis-cli -n 0 SET fixtest:counter "0"
redis-cli -n 0 SET fixtest:ledger "not-a-list"
redis-cli -n 0 DEL fixtest:marker
redis-cli -n 0 --eval script_guarded.lua \
    fixtest:counter fixtest:ledger fixtest:marker , e1

Since TYPE ledger will return "string", the script should immediately return an error. The expected reply is:

(error) ERR ledger is string

No writes should happen (counter remains 0, marker is still absent, ledger still string). We confirm:

redis-cli -n 0 GET fixtest:counter  # -> "0"
redis-cli -n 0 TYPE fixtest:ledger  # -> "string"
redis-cli -n 0 EXISTS fixtest:marker  # -> (integer) 0
  1. Guarded healthy state: Now test the guarded script when the ledger is not string (for example, absent). Reset:

redis-cli -n 0 SET fixtest:counter "0"
redis-cli -n 0 DEL fixtest:ledger  # ledger is now absent
redis-cli -n 0 DEL fixtest:marker
redis-cli -n 0 --eval script_guarded.lua \
    fixtest:counter fixtest:ledger fixtest:marker , e1

Because TYPE ledger returns "none", the guard passes. The script should run as normal: counter goes to 1, LPUSH creates a list, and marker="done". It should return 1. We check:

redis-cli -n 0 GET fixtest:counter  # -> "1"
redis-cli -n 0 LRANGE fixtest:ledger 0 -1  # -> 1) "e1"
redis-cli -n 0 GET fixtest:marker  # -> "done"

Everything is correct.

State exactly what the guard does not cover

It’s important to note: this guard only handles our specific known failure (ledger being a string). It does not magically prevent other errors. For instance, if counter were a string instead of an integer, INCR could still fail. Or if ARGV[1] were something unexpected, an LPUSH might fail on bad encoding (unlikely). Resource limit failures or permission errors are also possible.

Also, if we changed the script’s logic (adding more commands), new failure points could appear. This guard is a very narrow precondition check. We must not label it as a “rollback” mechanism; it’s simply an early validation. The script’s contract now says “if ledger is string, do nothing (error); otherwise do full writes.” Only that exact behavior is supported by our evidence.

Run the acceptance matrix from clean baselines

We now have a suite of cases and observations. For clarity, we summarize them in a table. Each row is a case we ran (with clean initial state where specified) and shows final key states and the result, plus our verdict.

Note the distinction: ACCEPT means “this script revision and state observations satisfy the contract,” REPAIR means “the validation or caller must change,” HOLD means “ambiguous or unsafe to assume,” and RECONCILE means “plan to fix the data effect before moving on.” We identify the owner and next action informally.

Case

Return/Error

counter

ledger

marker

Verdict

Action/Evidence

Healthy unguarded (ledger none)

1 (integer)

1

list ["e1"]

"done"

ACCEPT

Script as-is works; keys match expectation, with no error.

Wrong-type unguarded

(error) WRONGTYPE

1

string "not-a-list"

absent

HOLD

Script errored mid-way; counter=1 but ledger/marker as before (partial effect). Caller must handle.

Unprotected retried (2×)

(error) WRONGTYPE

2

string "not-a-list"

absent

HOLD

Double invocation left counter=2. Retry caused extra effect; no new ledger entry. Risk: lost event ordering.

Pcall continuation (unguarded)

(error) WRONGTYPE

1

string "not-a-list"

"continued"

HOLD

Script continued past LPUSH, set marker. State changed despite error. Caller needs to know marker set.

Guarded wrong-type (rejection)

(error) ERR ...

0

string "not-a-list"

absent

ACCEPT

Guard script did nothing, as intended. No writes changed. Good contract for bad input.

Guarded healthy (ledger none)

1 (integer)

1

list ["e1"]

"done"

ACCEPT

Guard script passes correctly for good input. State as expected.

Each observation column was confirmed by the independent reads above. The HOLD cases indicate that the script left partial effects or we have ambiguity about new events. We cannot blindly retry or assume failure means no-change in these cases.

For the ACCEPT cases, the script (original or guarded) produced exactly the known expected state. The guarded-script rejection case (ledger was string) gave zero changes, which matches its contract.

Notice that in all cases labeled ACCEPT, the script returned exactly the tested value and the keys match the declared outcomes. Thus in those cases we trust the script version for those inputs. In the HOLD cases, the returned error means we cannot trust a “no effects” guarantee: something happened (counter or marker changed). Those scenarios need special handling (see next sections).

Reconcile retained effects before making a correction

For cases where the script left partial effects (the HOLD rows), the caller needs to reconcile the actual effects with the intended event history. For example, in the unguarded case where LPUSH fails, we got counter=1 but no ledger entry. The business event e1 did occur (ledger should have it). Since we lost the list insert, one approach is to prevent blindly decrementing the counter. Instead, the caller should record that one event happened but wasn’t logged. Perhaps it should push the event into another durable log or retry once a fix is in place.

Critically, we say: do not try DECR fixtest:counter just because the script errored. That would risk another failure (if counter were 0 or not integer) and still wouldn’t fix the missing ledger entry. Similarly, in the pcall case the marker was set to "continued" in error; a fix might require appending or reconciling that too.

A safer strategy is: pause any further automatic processing of new events from this source. The caller (or a human operator) must decide how to fix the ledger. Possibly by separate corrections: e.g. LPUSH fixtest:ledger e1 after investigating, and mark the counter with an external log entry explaining the reconciliation.

The key point: treat this as a known exception path. The event e1 did happen, so the system’s understanding of what occurred must be adjusted. This might involve a “compensating write” to add missing info. For example:

redis-cli -n 0 LPUSH fixtest:ledger "e1-retry"

This is only an illustration; actual reconciliation would likely involve a different key or transaction. This is a correction step after a separate decision. Notice: this correction is not a rollback of the script; it is an independent write. As such, it can itself fail or race. We should log the old and new states if we do correct.

Do not rename compensation as rollback

It can be tempting to think “we’ll just decrement the counter to undo that write, then re-run the script.” That is dangerous. Even if you do DECR fixtest:counter, what if fixtest:counter was not an integer (though here it is)? Also, there is no transaction, so after DECR another client could intervene. Rolling back by hand can leave its own inconsistencies.

Instead, any fixing should be done with a clear ledger: record exactly what was undone, and keep evidence. For example, if we decremented, we should immediately log that we did so and why. But again, we prefer not to blindly drop the counter; it is better to treat it as persistent truth that two increments occurred in the retry case. The compensation is better applied to add missing ledger entries or similar, because the counter itself can be less critical if used for concurrency control or idempotency checking.

In short: compensation writes must be treated as new, separately approved changes. We mark the space with who authorized them and what values were before/after. For example, one could log to an audit key:

redis-cli -n 0 LPUSH event_corrections "Counter decremented from 2 to 1 after fix"

Again, just an example. The point is not to pretend a rollback happened invisibly, but to explicitly state “we are fixing data that was left inconsistent by the script.” We never call it a rollback, because it is not atomic and can fail on its own.

Choose ACCEPT, REPAIR, HOLD or RECONCILE

To summarize decisions and ownership, consider the following:

  • ACCEPT: We accept the script only under the exact tested contract. For example, the guarded script is accepted when the inputs match its guard conditions (ledger none or list). The original script was accepted in the healthy case (ledger was list/none). In each ACCEPT case, the evidence is complete: we saw the return value and the final state, and they matched the declared outcome. The owner responsible is the developer of the script or the backend team: their job is done if inputs always satisfy these preconditions.

  • REPAIR: We use REPAIR if we need to change the validation logic or have the caller avoid the scenario. The “guarded script” is effectively a repair of the original, changing validation. If this goes into production, the script owners should update it. Another REPAIR scenario would be improving the caller so that it never passes a string type in fixtest:ledger. For example, the backend developer who calls this script might add code to ensure ledger is a list or handle the error. In summary, REPAIR indicates changes either to the script or its callers’ assumptions. The evidence needed is failing test results (we saw WRONGTYPE). The owner is the application or script author.

  • HOLD: We HOLD when we see partial effects or uncertainty that we cannot automatically resolve. The unguarded wrong-type and pcall cases both landed on HOLD. It means “pause, do not retry blindly, escalate this.” The responsible party is usually the service or support team that monitors script failures. The next action might be to alert operators or prevent further automatic retries. Evidence here is the mismatch between expected “all-or-none” behavior and the observed state. In a mature incident process, this row would correspond to filing an incident or bug report.

  • RECONCILE: We require reconciling retained effects before declaring full success. For example, after an error, the developer might be asked to create a one-off remediation: adding the missing list item(s), or adjusting counter if agreed. That is a conscious RECONCILE action. We document that as separate tasks. The owner could be the database administrator or a programmer given permission to alter data. Evidence: the known partial state (e.g. counter=1, missing ledger entry) and a plan for how to correct it.

Below is a compact decision grid:

Verdict

Required evidence

Action / Owner

ACCEPT

All tested cases match expected state (counter, ledger, marker) under the script revision.

No change; script owner should freeze this version.

REPAIR

Inputs violated preconditions (e.g. wrong type) and script changed or caller must change.

Script developer: add guard or fix caller.

HOLD

Script errored after partial writes or state unclear; cannot retry safely.

DevOps/QA: halt retries, alert owners, assess issue.

RECONCILE

Partial effects observed (counter ≠ events logged); require a manual or system fix.

Data owner: apply compensating writes, log corrections.

Whenever the script or invocation conditions change (new Redis version, changed key types, ACL updates, etc.), the team must re-validate. For instance, if a Redis upgrade changed TYPE return format, we’d retest the guard code. If a new client version changes CLI quoting, we check the digest. Every fix or decision should be backed by our recorded evidence (script text, digest, key reads).

Assign owners and revalidation triggers

Ownership is critical for accountability. We assign:

  • Script logic & digest: Owned by the Backend Engineer or team that maintains this Lua script. They control revisions to the script and should update the guard if needed.

  • Key-type contract: Owned by the data layer or DB team. They define what types fixtest:ledger should have. If that changes (e.g. ledger becomes allowed to be a stream or sorted set), the guard should be re-evaluated.

  • Caller/application: Owned by the service that calls EVAL. They must know this script’s behavior. If their usage pattern changes (e.g. sending different ARGV or keys), they must re-run these tests.

  • Execution environment: Owned by the DevOps/SRE team. If Redis config (like lua-time-limit, stop-writes mode, AOF settings) or networking (timeouts) change, the test should be repeated.

On any change (script modified, Redis upgraded, ACL altered to block LPUSH, etc.), we should re-run the appropriate subset of tests. For example, an ACL change might cause INCR or LPUSH to error differently; we’d confirm that our detected effects still hold. We might store a manifest or test report (outside Redis) that lists the script digest, Redis build, and expected key states for reference. That way, resetting the fixture and re-running the tests becomes a quick smoke-check after any change.

Preserve evidence independently of mutable keys

All test results should be recorded outside of Redis itself. For instance, we might keep a Markdown or wiki page (this blog is one) containing the “Expected” key states, script SHA1, and transcripts. Do not rely on the keys we tested to hold the record, because resetting them would destroy the evidence.

For example, we might have a table like above (or even commit it to version control) so that audits later show exactly what was checked. If one ever logs into the actual Redis and finds an unexpected state, we can compare against these accepted profiles. Ideally, automated tests or CI jobs would run these commands and diff the outputs with stored expectations.

We should also redact any real credentials (though our fixture had none). If using a non-default user or password, record them in a secure variables store, not in this shared doc. All specifics here are synthetic, but the principle stands: keep a log of <case, initial state, final state> that is immutable.

Build backend skills around explicit failure boundaries

This exercise highlights a practical backend engineering skill: treat errors as opportunities to confirm assumptions, not as safe resets. In a real project, whenever a script or API call can partially succeed, one should explicitly check the resulting state. For example, after any critical EVAL, the client might always read back affected keys to verify consistency.

Training programs like the Refonte Learning Backend Developer Program, a three-month curriculum covering Node.js/Express, MongoDB/SQL, REST APIs, testing and deployment, emphasize writing robust code that gracefully handles such edge cases.

In summary, when Redis stops a Lua script with an error, do not assume “everything rolled back.” Instead, replay the scenario in a safe environment, confirm exactly which writes stuck, and make a conscious choice: ACCEPT the behavior (if verified), REPAIR code (add validation or adjust the caller), HOLD and investigate, or RECONCILE the data before proceeding. With that evidence-based approach, you turn a surprise partial failure into a known, documented boundary case.