Commit Graph
8 Commits
Author SHA1 Message Date
mckero ae8b48c85a Run anywhere: no author paths left, and CI that proves it
tools/run_tests.sh defaults QEMU_SRC to a sibling of the checkout, which is where setup_qemu.sh puts it; the two defaults disagreed, so a fresh clone rebuilt nothing and reported a build that was not there. The interpreter list was also reading an empty $PY.

tools/test_bk4819_readback.sh was the last test with the author's paths, and the only one that could not run elsewhere. It now takes QEMU, GDB and ELF from the environment or PATH like the python tests, skips with a reason when one is missing, and says so on a platform whose QEMU cannot make the unix socket it uses.

tools/check_docs.py points UVK5_FW_DIR at a sibling and skips the file:line checks, with a message, when there is no firmware tree -- a fresh clone used to see thirteen failures it could do nothing about.

Added .github/workflows/unit.yml (the fast half of run_tests.sh on every push and PR), requirements-dev.txt for the one pip dependency, and a Dockerfile. Flask is not always installed, so test_webui now skips through setUpModule rather than erroring.
2026-10-01 15:12:45 +08:00
mckero 2667e046e8 Emulator: multiboot slots from the page, flash controller, portable tests
flash controller: store ACR/OPTKEYR instead of swallowing them, which is what stopped the factory bootloader from starting

slots over the firmware's own serial protocol (0x0720 family); uvk5_socket/uvk5_testenv so a fresh checkout skips instead of failing; web UI slot table and Multiboot button; quick start, CONTRIBUTING, and stop tracking firmware images and radio dumps
2026-10-01 14:54:34 +08:00
mckero b32335d8c0 Check the docs' claims against the code mechanically
Translating everything into Chinese found four claims that had already drifted, and
none of them were caught by reading -- they were caught by comparing against source.
Proofreading does not find rot, so do the comparison mechanically and keep doing it.

tools/check_docs.py verifies that every tool a README names exists, that every test in
run_tests.sh is documented in both languages, that internal .md links resolve, that the
translation pairs have matching heading structure, that memory-map addresses match the
model's #defines, and that documented firmware file:line references still point at what
the prose claims. It runs in the quick tier of run_tests.sh, needing no emulator.

Confirmed it can actually fail, because a checker that cannot is worthless: renaming a
documented tool and deleting a heading from the Chinese side each produce one named
failure and exit 1, and reverting returns it to clean.

One thing it deliberately does not check. An early version compared firmware constants
with a regex that took the first number on a line, so `key_debounce_10ms = 20 / 10` read
as 20 and it declared the docs wrong for saying 2. The docs were right and the checker
was broken. A checker that cries wolf gets ignored, so claims it cannot verify
unambiguously are left out rather than guessed at.

Current state: 16 file:line references all accurate, 7 memory-map addresses all match,
zero broken links, all three translation pairs structurally aligned.
2026-08-29 16:54:25 +01:00
mckero 8b995aa610 Make RSSI depend on tuning instead of being a constant
The S-meter had a number to draw, but a fixed RSSI above squelch meant the band was
uniformly and permanently occupied. Scanning, squelch, and every "is this channel busy"
decision therefore faced a situation that never varied, so none of that logic was
really being tested -- the tests passed without testing much.

RSSI is now derived from where the firmware tuned. BK4819_SetFrequency splits the
frequency across REG_38 and REG_39 (driver/bk4819.c:743), which the model already
records; verified against a live guest that 0x0262/0x5A00 reads back as 400.00000 MHz,
matching the screen. A small table of virtual stations plus a noise floor and a fade
either side of centre gives a band with signals in some places and not others.

Measured through the firmware's own tuning path -- typing 410.000 on the keypad rather
than poking the registers, so the test does not check the model against itself:

    400.000 MHz (station)  RSSI 0x01E5
    410.000 MHz (empty)    RSSI 0x0091      a gap of 85 dB

What is honest and what is not, recorded in the code: the shape is real physics, power
falls off away from a carrier with a noise floor underneath. The station list is
invented. So this reproduces "the firmware copes with a band that is busy in places",
which is genuine coverage, and it reproduces no actual radio environment -- a dBm figure
from here is not a claim about the world.

Also records why backlight PWM is deliberately left stubbed. Intermediate brightness
runs TIM7 -> DMA rewriting GPIOA BSRR at 128 kHz, so modelling it costs 128,000 GPIO
writes per emulated second and changes nothing observable: backlight is LED brightness
and never touches the framebuffer. The two endpoints that are observable, off and full,
bypass the timer and already work.

Full run: 16 passed, 0 failed.
2026-08-29 08:52:52 +01:00
mckero fdcbe80056 Model TIM2, so millis() advances and timeouts can expire
Second finding from the audit. TIM2 was covered by the catch-all stub, which returns
the last value written, so

    uint32_t millis(void) { return LL_TIM_GetCounter(TIM2); }

returned 0 forever. All 17 call sites that measure elapsed milliseconds could never
see time pass -- a silent wrong answer rather than a hang, which is harder to notice
and was not noticed.

The counter is derived on read from QEMU_CLOCK_VIRTUAL rather than stored, with
CR1.CEN starting and freezing it and a CNT write rebasing it. Guest time here is not
proportional to wall time anyway, and code measuring elapsed milliseconds wants
something advancing at roughly the rate a human sees; this is explicitly not for
anything needing cycle accuracy.

Measured: 24358 ms, then 29527 ms five seconds later -- 5169 ms elapsed, so the rate
is right rather than merely non-zero. The test checks the rate for that reason: a
counter ticking at the wrong speed would satisfy "non-zero" and "increasing" and still
break every timeout.

AGENTS.md now carries the audit itself: a table of what the firmware actually drives
against what is modelled versus stubbed, and the point that answering reads is not the
same as being reproduced. The honest summary is that the digital side the firmware
depends on is reproduced, and the analogue side is not and cannot be.

Full run: 15 passed, 0 failed.
2026-08-29 08:00:27 +01:00
mckero e46cae2e48 Make the ADC settable, which reaches the battery behaviour
Prompted by a fair criticism: the reports said what runs, not what is actually
reproduced. An audit found the ADC was modelled but returned a hardcoded 2200 forever,
so gBatteryDisplayLevel, gLowBattery and the warning popup were all unreachable. A
peripheral that answers reads is not the same as a peripheral that is reproduced.

adc-result is now settable over QOM and clamped to 12 bits. Measured: 2200 gives
level 4 and no warning, 1200 gives level 0 and raises gLowBattery, and the level
recovers to 4 afterwards.

tools/test_battery.py covers it, and deliberately does NOT assert that gLowBattery
clears on recovery. helper/battery.c:190-204 only clears it when the level lands
exactly on 2; above that it clears gLowBatteryConfirmed and leaves gLowBattery set. So
4 -> 0 -> 4 really does leave the flag raised. The first version of this test called
that a failure -- the test was wrong, not the model. The emulator reproduces the
firmware, including behaviour that looks like a bug.
2026-08-29 07:49:25 +01:00
mckero 7ed9f61f71 Add the audio path test to the runner
Full run: 13 passed, 0 failed.
2026-08-29 06:13:13 +01:00
mckero 8d1a1c4415 Add a test runner, and a test that it can fail
The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:

    bash tools/run_tests.sh        # everything, 11 suites, a few minutes
    bash tools/run_tests.sh -q     # unit tests only, ~15 s, no emulator

The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.

Two defects in the runner's own first draft, both caught before it was trusted:

It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.

That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.

Full run: 11 passed, 0 failed.
2026-08-29 05:42:11 +01:00