The symptom was one row down, the bug was two rows down
A publisher workflow reads a Google Sheets queue, sorts by row, and picks the first item marked ready. For several weeks it kept failing on the same row — a stale duplicate of an already-published video whose source files were long gone — and never got far enough to look at the row underneath it. Skipping that duplicate row was a five-minute fix. What was underneath it was not: the real next video's row also said ready, but the task directory on disk had only the raw render. No captioned video, no thumbnail.
The Creator workflow that had produced that row had, according to n8n's execution log, succeeded. Both the captioning step and the thumbnail step showed green.
What the retry loop actually looked like
Both steps ran over an SSH node, wrapping the real command in a small retry loop to absorb the occasional transient failure — a locked file, a process that hadn't released a port yet, that sort of thing:
for i in 1 2 3; do
python3 burn_captions.py "$IN" "$OUT" && break
sleep 5
done
This reads naturally as "try up to three times, and stop as soon as it works." It does try up to three times. It does not report whether any of those three tries actually worked.
Why the loop's exit code is always 0
A shell script's (or a compound command's) exit status is the exit status of the last command it ran — not of "the loop" as a concept, and not of whichever command inside it seems like the important one. Walk through what happens when every attempt fails:
i=1: python3 burn_captions.py ... -> exit 1 (failed)
&& break -> short-circuits, does not run
sleep 5 -> exit 0
i=2: same as above -> exit 0 after sleep
i=3: same as above -> exit 0 after sleep
loop exits normally, last command run was `sleep 5`, exit code 0
The SSH node reads that final exit code. Three real, back-to-back failures, and n8n was told the step succeeded. There is no point at which the loop's own exit status reflects the python script's result — sleep never fails under normal conditions, so this shape cannot produce a nonzero exit code no matter how badly the wrapped command behaves.
This is the quiet sibling of a command that dies with a real, informative exit code: at least a timeout tells you something happened. A loop shaped like this tells you nothing ever happened, in either direction.
How it stayed invisible for weeks
Downstream, the workflow appended a row to the render queue with status: ready and a hardcoded path to the captioned video and thumbnail, without checking that either file existed. As long as the raw render itself worked — which it reliably did — the queue kept accepting rows that pointed at files nobody had produced. Nothing downstream ever looked at the disk to confirm; it trusted the exit code of the step that was supposed to have created those files, and that exit code was always green.
The fix: track a real result, exit on that
The loop needs its own variable that only a genuine success sets, checked once the loop is done:
ok=0
for i in 1 2 3; do
python3 burn_captions.py "$IN" "$OUT" && { ok=1; break; }
sleep 5
done
[ "$ok" = 1 ]
The last line is the important part: a bare test expression as the script's final statement, so the SSH node's exit code is 1 whenever every attempt genuinely failed, and 0 only when one of them actually succeeded. Same change applied to the thumbnail-generation step, which had the identical shape copy-pasted into it.
A validation gate as the second line of defense
Fixing the loop stops new false-positives from this exact bug, but it doesn't verify the files are actually usable — a script can exit 0 and still write a corrupt or truncated file. A gate was added right before the queue-append step: an ffprobe check that the captioned video is playable with an audio track over 60 seconds, and a dimension check that the thumbnail is exactly the expected size and non-empty. Only if both pass does the row get appended as ready; otherwise a Telegram alert fires and nothing reaches the queue. The two checks cover different failure layers on purpose — the exit-code fix catches "the step lied about succeeding," the ffprobe gate catches "the step succeeded but the output is still broken."
The pattern to search for
Any bash snippet inside an n8n SSH or Execute Command node with the shape for ...; do cmd && break; sleep N; done (or a while loop written the same way) has this bug, regardless of what cmd does. It is worth grepping existing workflows for && break inside a retry loop specifically — the failure is silent by construction, so nothing in n8n's UI will point at it until something downstream notices missing output, the same way a starved queue row eventually did here.
n8n Automation Hub