Add a test runner, and a test that it can fail

The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:

    bash tools/run_tests.sh        # everything, 11 suites, a few minutes
    bash tools/run_tests.sh -q     # unit tests only, ~15 s, no emulator

The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.

Two defects in the runner's own first draft, both caught before it was trusted:

It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.

That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.

Full run: 11 passed, 0 failed.
This commit is contained in:
mckero committed 2026-08-29 05:42:11 +01:00
1 parent 95bad1614e
commit 8d1a1c4415
4 files changed
+221

No files matched your search

+89
View File
@@ -0,0 +1,89 @@
#!/usr/bin/env bash
# The test runner must actually notice a failing test.
#
# This is not a hypothetical worry. The first version of run_tests.sh used
#
# if "$@" 2>&1 | sed 's/^/ /'; then
#
# which checks *sed's* exit status, not the test's. sed almost always succeeds, so every
# test would have been counted as passing and the runner would have reported a clean
# suite no matter what broke. A test runner that cannot fail is worse than none, because
# it is trusted.
#
# Checked here:
# 1. a failing command is counted as a failure and named
# 2. a passing command is counted as a pass
# 3. the runner's own exit status is non-zero when something failed
# 4. binary noise in a test's output does not break the accounting
set -u
HERE=$(cd "$(dirname "$0")" && pwd)
RUNNER="$HERE/run_tests.sh"
[ -f "$RUNNER" ] || { echo "FAIL run_tests.sh not found"; exit 1; }
failures=0
# Extract the run() helper and exercise it in isolation, so this test does not have to
# boot an emulator to check the accounting logic.
harness=$(mktemp /tmp/run-harness-XXXX.sh)
trap 'rm -f "$harness" /tmp/rt-good /tmp/rt-bad /tmp/rt-noisy' EXIT
sed -n '/^run() {/,/^}/p' "$RUNNER" > "$harness"
if ! grep -q PIPESTATUS "$harness"; then
echo "FAIL run() does not use PIPESTATUS; it is checking the wrong exit status"
echo " (a piped command's \$? is the last stage, so every test would 'pass')"
exit 1
fi
echo "PASS run() checks the test's status, not the pipeline's last stage"
cat >> "$harness" <<'EOF'
pass=0; fail=0; failed_names=""
run "good" /tmp/rt-good
run "bad" /tmp/rt-bad
run "noisy" /tmp/rt-noisy
echo "RESULT pass=$pass fail=$fail failed:$failed_names"
EOF
printf '#!/bin/sh\necho fine\n' > /tmp/rt-good
printf '#!/bin/sh\necho broken\nexit 1\n' > /tmp/rt-bad
# Emits raw bytes, as gdb-driven tests can. Left unfiltered these make the combined
# output a "binary file" to grep, which silently swallows the summary line.
printf '#!/bin/sh\nprintf "noise\\001\\002\\003\\n"\nexit 0\n' > /tmp/rt-noisy
chmod +x /tmp/rt-good /tmp/rt-bad /tmp/rt-noisy
out=$(bash "$harness" 2>&1)
result=$(printf '%s\n' "$out" | grep '^RESULT' || true)
echo " $result"
case "$result" in
*"pass=2"*) echo "PASS both passing tests counted" ;;
*) echo "FAIL expected pass=2"; failures=$((failures + 1)) ;;
esac
case "$result" in
*"fail=1"*) echo "PASS the failing test was counted" ;;
*) echo "FAIL expected fail=1; a failing test went unnoticed"
failures=$((failures + 1)) ;;
esac
case "$result" in
*"failed: bad"*) echo "PASS the failing test was named" ;;
*) echo "FAIL the failing test was not named"; failures=$((failures + 1)) ;;
esac
# The runner must exit non-zero on failure, or CI and shell && chains ignore it.
if grep -q 'exit 1' "$RUNNER"; then
echo "PASS the runner exits non-zero when tests fail"
else
echo "FAIL the runner never exits non-zero"
failures=$((failures + 1))
fi
if [ "$failures" = "0" ]; then
echo
echo "the test runner reports failures correctly"
exit 0
fi
exit 1