Add a test runner, and a test that it can fail

The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:

    bash tools/run_tests.sh        # everything, 11 suites, a few minutes
    bash tools/run_tests.sh -q     # unit tests only, ~15 s, no emulator

The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.

Two defects in the runner's own first draft, both caught before it was trusted:

It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.

That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.

Full run: 11 passed, 0 failed.
This commit is contained in:
mckero committed 2026-08-29 05:42:11 +01:00
1 parent 95bad1614e
commit 8d1a1c4415
4 files changed
+221

No files matched your search

+96
View File
@@ -0,0 +1,96 @@
#!/usr/bin/env bash
# Run every test, report what failed, and exit non-zero if anything did.
#
# This exists because the suite had grown to ten separate invocations that had to be
# remembered and pasted in the right order. That is how regressions slip through: it is
# too easy to run the two tests related to what you just changed and miss the one that
# broke.
#
# Emulator tests boot their own QEMU and take 20-30 s each, so a full run is a few
# minutes. Pass -q to run only the fast unit tests, which need no emulator at all and
# finish in about 15 s -- useful while iterating.
#
# The build is checked FIRST and a failure stops everything. ninja leaves the previous
# binary in place when it fails, so the tests would otherwise run happily against a
# stale build and report results for code that was never compiled. That has produced
# two rounds of meaningless output before now.
set -u
HERE=$(cd "$(dirname "$0")" && pwd)
SIM=$(dirname "$HERE")
QEMU_SRC=${QEMU_SRC:-/root/qemu-build/qemu-7.2+dfsg}
QUICK=0
[ "${1:-}" = "-q" ] && QUICK=1
pass=0
fail=0
failed_names=""
run() {
local name="$1"; shift
printf '\n=== %s\n' "$name"
# Strip control characters. Some tests shell out to gdb, whose output can carry
# escape sequences and stray bytes; left alone they make the combined log a
# "binary file" as far as grep is concerned, which silently swallows the summary.
#
# PIPESTATUS, not $?, because $? here is sed's status and would report success
# for every failing test.
"$@" 2>&1 | tr -cd '\11\12\15\40-\176' | sed 's/^/ /'
if [ "${PIPESTATUS[0]}" = "0" ]; then
pass=$((pass + 1))
else
fail=$((fail + 1))
failed_names="$failed_names $name"
fi
}
# --- build ---------------------------------------------------------------
# Only when the model source differs from what the QEMU tree holds, so a plain test
# run does not pay for a rebuild it does not need.
if [ -d "$QEMU_SRC/build" ] && \
! cmp -s "$SIM/qemu/py32f071.c" "$QEMU_SRC/hw/arm/py32f071.c"; then
echo "=== build (model source changed)"
cp "$SIM/qemu/py32f071.c" "$QEMU_SRC/hw/arm/py32f071.c"
if (cd "$QEMU_SRC/build" && ninja qemu-system-arm 2>&1 | tail -20) \
| grep -qE 'FAILED|error:'; then
echo " BUILD FAILED -- stopping before any test runs"
echo " (a stale binary would otherwise be tested silently)"
exit 1
fi
echo " ok"
fi
# --- unit tests, no emulator --------------------------------------------
cd "$HERE"
# First, that this script itself reports failures. A runner that silently counts every
# test as passing is worse than no runner, because it gets trusted.
run "runner self-check" bash "$HERE/test_run_tests.sh"
run "unit: model helpers" python3 -m unittest discover -p 'test_uvk5*.py' -q
run "unit: web UI" python3 -m unittest test_webui -q
if [ "$QUICK" = "1" ]; then
printf '\n%d passed, %d failed (unit tests only)\n' "$pass" "$fail"
[ "$fail" = "0" ] || { echo "failed:$failed_names"; exit 1; }
exit 0
fi
# --- emulator tests -----------------------------------------------------
# Ordered cheapest first, so an obvious breakage surfaces without waiting for the
# whole run.
cd "$SIM"
run "keypad" python3 tools/keypad_test.py
run "BK4819 registers" python3 tools/test_bk4819.py
run "register readback" bash tools/test_bk4819_readback.sh
run "S-meter" python3 tools/test_smeter.py
run "PTT" python3 tools/test_ptt.py
run "scan" python3 tools/test_scan.py
run "serial receive" python3 tools/test_serial_rx.py
run "flash persistence" python3 tools/test_flash_persist.py
run "frequency entry" python3 tools/test_freq_entry.py
printf '\n%d passed, %d failed\n' "$pass" "$fail"
if [ "$fail" != "0" ]; then
echo "failed:$failed_names"
exit 1
fi