The page edits a working copy of the flash image on purpose -- the file it was first pointed at may be a real calibration dump -- but a restart pointed at a different image switched to that file's copy instead, and the games installed through the page looked like they had vanished (measured: the app table came back holding an older slot). work/run-webui.ps1 now prefers the working copy the page has been editing, ahead of the original, and the page's own hint says which image it is editing and that the loaded one is never modified. A test asserts the order in the script.
Sixteen app rows (most of them empty or holding something that is not an app) plus five slot rows made the column long enough to be annoying, so each table is now a details pane: the firmware slots start open, the app list starts folded, and the state that used to sit in its own row moved into the summary line. Same ids, so loadSlots and loadApps are untouched. A front-end test asserts both panes exist and that the app list starts folded.
F then 7 opens the F4HWN APPS menu and MENU runs the selected app, which is upstream's own wording in UVStudio's locales/en.js and what App/apps/app_menu.c does (KEY_MENU on an installed row calls APP_LaunchOverlay, which runs until the app exits; EXIT leaves). Verified in the emulator with Tetris.app installed through the page's endpoint: F,7 reads the app region eight times and shows the boxed menu; MENU reads once inside slot 0's code and the game's screen replaces the radio's; DOWN moves the piece and the game is still there two seconds later. The page's own hint now says so.
Worth recording: twenty keys short and long, the whole 79-entry settings menu and the multiboot menu all left the app region untouched -- the multiboot menu reads firmware slots, not apps. The entry point was written down in the flasher's translation file, not in the firmware source: when a feature's way in is missing, the host tool that installs it is the document.
GET /api/apps/radio opens the firmware's serial port and sends 0x0730 for all sixteen slots, so the answer comes from the running firmware rather than from our reading of the file -- which is the check that matters, because the bytes can be right and the firmware still refuse a slot. Measured through the page after installing Beam.app into slot 0: slot 0 -> Beam 1.0, 1100 B, crc 0xd976058, shortcut beam, committed; slots 1..3 -> status 2 with unrelated data, the resource-block overlap the install guard refuses. A button beside the table asks it and shows the answer in its own column.
The server gives its own emulator a serial port (--serial-port, default 4445) and uvk5_slots_serial.Radio gained a public app_info(slot), so nothing reaches into a private helper. uvk5_apps.parse_radio_reply decodes the answer and is tested without a radio. Also recorded: QEMU needs the mingw64 DLLs on PATH, and started by hand without them it exits before opening QMP, which surfaces only as 'QMP socket never appeared'.
UVStudio's own js/flash.js names the protocol: MSG_APP_INFO 0x0730/0x0731, MSG_APP_ERASE 0x0732/0x0733, MSG_APP_WRITE 0x0734/0x0735, MSG_APP_VALIDATE 0x0736/0x0737, with APP_SLOT_COUNT 16, APP_IMG_OFFSET 0x1000, APP_HDR_SIZE 64 and APP_MAGIC 0x31504146 -- the constants this page already used. It uses slots 0..7 and labels them 1..8, and writes the header last so a partial write cannot validate; the page now labels slots the same way.
The Labs build answers 0x0730 over USART, so an installed Beam.app was queried on the radio: slot 0 came back status 0 with the header this page wrote (FAP1, code_size 1100, CRC 0x0d976058, flags 0x0801, name Beam), which confirms the region, the offset and the layout from the firmware's side rather than from a header I read. Slots 1 and 2 answered status 2 with unrelated data -- the overlap the install guard refuses to overwrite.
The endpoints existed; now the page shows them. An Overlay apps block beside the firmware slots lists all 16 app slots with their name, version, shortcut and size, a .app file picker and an Erase button per row, and says where the region is. It reads GET /api/apps and posts to POST /api/apps/<n> and /api/apps/<n>/erase, so installing a game is a file pick where the firmware slots already are -- no WebSerial, no browser permission.
An upload into a slot that holds something which is not an app is refused by the server; the page asks once and retries with ?force=1 rather than either failing quietly or destroying the factory resource data that overlaps this region on a localised image. Front-end tests assert the table and the endpoint are in the served page and that it never asks for a serial port or audio, and TestPageScriptParses keeps checking that the page's script parses.
The Labs edition's apps live in the external flash, in the region its own App/apps/app_overlay.h defines (16 slots of 8 KiB from 0x102000, a 64-byte FAP1 header at the slot base, code one 4 KiB sector later), and upstream installs them from UVStudio over WebSerial. The page owns the image, so this adds GET /api/apps, POST /api/apps/<n> and POST /api/apps/<n>/erase, which put the same bytes at the same offsets with no serial protocol and no browser permission.
Measured on a real image while wiring it up: every one of the 16 slots already held data that is neither empty nor an app -- the localised build's factory resource block overlaps 0x102000 -- so install now refuses to overwrite anything that is not an app unless asked (--force, ?force=1), naming what is there. And _edit_flash edits a copy, which is why the source image shows no changed bytes; the first test read that as 'the install did nothing' and now checks FlashSlot.path. Tests: test_uvk5_apps grew to 20, test_webui.TestAppEndpoints adds 7 over the endpoints.
The Labs edition runs overlay apps (Tetris, Breakout, Plasma, Cube3D, Beam, Beacon, FoxHunt, BroadcastFM) that upstream UVStudio installs over WebSerial. This page owns the flash image, so the same bytes go to the same offsets with no serial protocol and no browser permission: APP_REGION_BASE 0x102000, APP_SLOT_STRIDE 0x2000, APP_CODE_OFFSET 0x1000, 16 slots, taken from the firmware's own App/apps/app_overlay.h rather than inferred.
tools/uvk5_apps.py parses and validates the 64-byte FAP1 header (zlib CRC-32 over the code, vma 0x20000280, name, version, capabilities), lists, installs and erases slots, and refuses what the firmware would show as APP ERROR. test_uvk5_apps covers those refusals plus install/erase/list round trips, and parses a real upstream Beam.app when one has been downloaded. The header struct was 60 bytes at first -- a missing vma field -- which the real file's bytes showed at once.
The page only knew its own input. A build called f4hwn.fusion.bin reports EGZUMER+F4HWN v6.0.0.CN, and with the multi-system release a committed external slot makes the factory bootloader reflash the internal flash from that slot on every power-on -- so the uploaded image never runs and the page keeps naming it. The firmware prints its own banner on USART1; tools/uvk5_banner.py reads it back, /api/firmware returns running: {banner, matches_uploaded, note}, and the page shows what the device reports, flagging it only when the running version is not in the uploaded image at all.
That reader also exposed a regression of my own: _start_stderr_pump had been rewritten to read the pipe in 64 KB chunks, which kept QEMU from blocking but delivered nothing to the log until 64 KB had accumulated -- and the banner is forty bytes, so it never appeared. It reads lines again, still starting before anything waits on QEMU, and test_uvk5_supervisor passes either way.
The flags are optional now and the page needs no addresses at all, so the examples no longer paste one build's numbers: screenshot.py is shown asking the firmware through tools/uvk5_buffers.py, and the webui example omits them entirely. The paragraph that said they default to one known build is replaced with what actually happens.
The page was told --frame-addr 0x200012BE --status-addr 0x2000163E and used them as a fallback. The firmware the user actually flashed keeps its buffers at 0x2000129E/0x2000161E, 32 bytes earlier, so every line landed 32 bytes off: that is the "other firmware looks shifted" report. The images here are minimal ELFs with no symbol table, so there is nothing to read -- but the firmware's own buffers hold the same bytes the controller holds, and tools/uvk5_buffers.py finds them by matching (1024/1024 bytes for that file).
The two address flags are optional now, work/run-webui.ps1 passes no machine-specific values at all, and the page reports what it found in /api/status and /api/panel. tools/uvk5_testenv.qemu() also looks in the sibling qemu-7.2/build the rest of the repo assumes. Fixed /api/panel's emulator-off branch, which called jsonify with both a dict and kwargs and 500'd.
Tests: test_uvk5_buffers (the search must count matches, not pairs -- its first version scored every offset full marks and always answered the first one).
uvk5_stream.py used STATUS_BYTES without importing it, so the panel branch raised NameError on every frame and a bare except swallowed it: every screen the page drew came from guest RAM at one build's addresses. The pump now reports which source it used and why, webui exposes it (/api/panel and frame_source), and test_uvk5_stream asserts the panel wins when reachable and that a fallback is announced.
The ST7565 column counter wrapped at 128 instead of the controller's 132, so addresses 128..131 came back as 0..3, fell outside the col>=4 store, and were dropped: every row lost its last four pixels, which is where the battery icon lives. Pre-fix, filling a page with 0xFF left columns 124..127 blank; now they carry content and the page's frame matches the panel memory 8192/8192.
UVK5_BK4819_PROBE reassembles the sixteen bits a read clocks out and logs them against the register's own value: 1566 of 1566 agree on the working model. That agreement is not proof -- removing the skip_falling fix, which is exactly the historical left-shift regression, leaves the reassembled word unchanged, so this observation point is not the one the guest samples at. tools/test_bk4819_readback.sh therefore stays the guard.
tools/test_bk4819_readback.py is the working draft of a portable replacement (no source patch, no rebuild, no ARM gdb) and is deliberately NOT registered in run_tests.sh, so a proven guard is not swapped for an unproven one. Both the file and the two READMEs say so.
QMP takes a single client and the web UI holds it for its whole lifetime, so this is the normal outcome while the page is open rather than a broken emulator. Verified against a private instance: 128x64, 1819 pixels lit, 64 rows of ASCII.
tools/panel_dump.py prints the display controller's own memory as ASCII or PNG for any firmware, with the four mappings, so two builds can be compared instead of glanced at. The panel model gained a bounded UVK5_PANEL_PROBE diagnostic along the way.
Measured: the 5.9.0.CN panel agrees with its own framebuffer 8188 of 8192 pixels with the data untouched, and the fetched 6.0.0 build renders identically. Both program the same geometry registers, which is why the mapping is a driver convention and cannot be derived from the controller: it has to be measured.
Noted, not yet fixed: the display start line (0x40|n) is ignored, and the model stores pixels at col-4 with the column counter wrapping at 128 instead of the controller's 132 columns.
The probe scripts, run.sh, trace_run.sh, webui.py and two emulator tests each named the same hardcoded firmware from a source tree that is not in this repository. They now resolve QEMU and the firmware the way tools/uvk5_testenv.py does -- environment, then PATH, then whatever the checkout has -- and skip with a reason when there is nothing.
webui.py's --qemu and --elf lost their author defaults too: a bare qemu-system-arm through PATH, and no firmware until one is uploaded, which the page already reports.
tools/run_tests.sh defaults QEMU_SRC to a sibling of the checkout, which is where setup_qemu.sh puts it; the two defaults disagreed, so a fresh clone rebuilt nothing and reported a build that was not there. The interpreter list was also reading an empty $PY.
tools/test_bk4819_readback.sh was the last test with the author's paths, and the only one that could not run elsewhere. It now takes QEMU, GDB and ELF from the environment or PATH like the python tests, skips with a reason when one is missing, and says so on a platform whose QEMU cannot make the unix socket it uses.
tools/check_docs.py points UVK5_FW_DIR at a sibling and skips the file:line checks, with a message, when there is no firmware tree -- a fresh clone used to see thirteen failures it could do nothing about.
Added .github/workflows/unit.yml (the fast half of run_tests.sh on every push and PR), requirements-dev.txt for the one pip dependency, and a Dockerfile. Flask is not always installed, so test_webui now skips through setUpModule rather than erroring.
flash controller: store ACR/OPTKEYR instead of swallowing them, which is what stopped the factory bootloader from starting
slots over the firmware's own serial protocol (0x0720 family); uvk5_socket/uvk5_testenv so a fresh checkout skips instead of failing; web UI slot table and Multiboot button; quick start, CONTRIBUTING, and stop tracking firmware images and radio dumps
The remaining class of claim check_docs.py could not see: whether the commands in the
docs would actually run. A renamed or removed option is the classic form of command
rot, and the one that wastes a reader's time most directly -- they paste the line and
it fails. All 9 documented flags across screenshot.py, webui.py and restore_flash.sh
are real.
The check earned its own lesson, recorded in both languages. Its first version matched
only to the end of the line, so on a wrapped command it saw --frame-addr and nothing
after the backslash: 4 of 9 flags, and it reported a clean run. A check that silently
covers a quarter of what it claims is worse than no check, because the clean result is
believed. Continuations are joined before matching now.
Confirmed it fails when it should: renaming --frame-addr to something no tool accepts
produces two named failures and exit 1, and reverting returns it to clean.
check_docs.py now runs seven checks.
Translating everything into Chinese found four claims that had already drifted, and
none of them were caught by reading -- they were caught by comparing against source.
Proofreading does not find rot, so do the comparison mechanically and keep doing it.
tools/check_docs.py verifies that every tool a README names exists, that every test in
run_tests.sh is documented in both languages, that internal .md links resolve, that the
translation pairs have matching heading structure, that memory-map addresses match the
model's #defines, and that documented firmware file:line references still point at what
the prose claims. It runs in the quick tier of run_tests.sh, needing no emulator.
Confirmed it can actually fail, because a checker that cannot is worthless: renaming a
documented tool and deleting a heading from the Chinese side each produce one named
failure and exit 1, and reverting returns it to clean.
One thing it deliberately does not check. An early version compared firmware constants
with a regex that took the first number on a line, so `key_debounce_10ms = 20 / 10` read
as 20 and it declared the docs wrong for saying 2. The docs were right and the checker
was broken. A checker that cries wolf gets ignored, so claims it cannot verify
unambiguously are left out rather than guessed at.
Current state: 16 file:line references all accurate, 7 memory-map addresses all match,
zero broken links, all three translation pairs structurally aligned.
The S-meter had a number to draw, but a fixed RSSI above squelch meant the band was
uniformly and permanently occupied. Scanning, squelch, and every "is this channel busy"
decision therefore faced a situation that never varied, so none of that logic was
really being tested -- the tests passed without testing much.
RSSI is now derived from where the firmware tuned. BK4819_SetFrequency splits the
frequency across REG_38 and REG_39 (driver/bk4819.c:743), which the model already
records; verified against a live guest that 0x0262/0x5A00 reads back as 400.00000 MHz,
matching the screen. A small table of virtual stations plus a noise floor and a fade
either side of centre gives a band with signals in some places and not others.
Measured through the firmware's own tuning path -- typing 410.000 on the keypad rather
than poking the registers, so the test does not check the model against itself:
400.000 MHz (station) RSSI 0x01E5
410.000 MHz (empty) RSSI 0x0091 a gap of 85 dB
What is honest and what is not, recorded in the code: the shape is real physics, power
falls off away from a carrier with a noise floor underneath. The station list is
invented. So this reproduces "the firmware copes with a band that is busy in places",
which is genuine coverage, and it reproduces no actual radio environment -- a dBm figure
from here is not a claim about the world.
Also records why backlight PWM is deliberately left stubbed. Intermediate brightness
runs TIM7 -> DMA rewriting GPIOA BSRR at 128 kHz, so modelling it costs 128,000 GPIO
writes per emulated second and changes nothing observable: backlight is LED brightness
and never touches the framebuffer. The two endpoints that are observable, off and full,
bypass the timer and already work.
Full run: 16 passed, 0 failed.
Second finding from the audit. TIM2 was covered by the catch-all stub, which returns
the last value written, so
uint32_t millis(void) { return LL_TIM_GetCounter(TIM2); }
returned 0 forever. All 17 call sites that measure elapsed milliseconds could never
see time pass -- a silent wrong answer rather than a hang, which is harder to notice
and was not noticed.
The counter is derived on read from QEMU_CLOCK_VIRTUAL rather than stored, with
CR1.CEN starting and freezing it and a CNT write rebasing it. Guest time here is not
proportional to wall time anyway, and code measuring elapsed milliseconds wants
something advancing at roughly the rate a human sees; this is explicitly not for
anything needing cycle accuracy.
Measured: 24358 ms, then 29527 ms five seconds later -- 5169 ms elapsed, so the rate
is right rather than merely non-zero. The test checks the rate for that reason: a
counter ticking at the wrong speed would satisfy "non-zero" and "increasing" and still
break every timeout.
AGENTS.md now carries the audit itself: a table of what the firmware actually drives
against what is modelled versus stubbed, and the point that answering reads is not the
same as being reproduced. The honest summary is that the digital side the firmware
depends on is reproduced, and the analogue side is not and cannot be.
Full run: 15 passed, 0 failed.
Prompted by a fair criticism: the reports said what runs, not what is actually
reproduced. An audit found the ADC was modelled but returned a hardcoded 2200 forever,
so gBatteryDisplayLevel, gLowBattery and the warning popup were all unreachable. A
peripheral that answers reads is not the same as a peripheral that is reproduced.
adc-result is now settable over QOM and clamped to 12 bits. Measured: 2200 gives
level 4 and no warning, 1200 gives level 0 and raises gLowBattery, and the level
recovers to 4 afterwards.
tools/test_battery.py covers it, and deliberately does NOT assert that gLowBattery
clears on recovery. helper/battery.c:190-204 only clears it when the level lands
exactly on 2; above that it clears gLowBatteryConfirmed and leaves gLowBattery set. So
4 -> 0 -> 4 really does leave the flag raised. The first version of this test called
that a failure -- the test was wrong, not the model. The emulator reproduces the
firmware, including behaviour that looks like a bug.
Asked for a speaker and a microphone. The honest answer is that neither exists to
model: on the real radio neither passes through the MCU. Receive audio is demodulated
inside the BK4819 and leaves it as analogue on its AF pin; transmit audio goes from
the microphone into the chip's own ADC. The firmware's entire involvement is
- PA8, the amplifier enable (GPIO_EnableAudioPath, driver/gpio.h:34)
- REG_47, which AF source the chip routes
- REG_64, a level it displays
No audio samples exist anywhere in the MCU's address space, so a device model has
nothing to capture or play, and a browser has nothing to be granted permission for.
Synthesising sound would be inventing data the firmware never produced.
What is real is the firmware's intent, and PA8 states it exactly. TYPE_UVK5_AUDIO
watches that pin and exposes read-only speaker-on; the web UI shows it as a speaker
glyph beside the power state, and /api/status reports it. Read-only deliberately:
letting a test write it would only let the test lie to itself. The page asks for no
audio permission, and a test asserts it never will -- no getUserMedia, no
AudioContext, no <audio>.
Measured: amplifier off while idle in power save, on after SIDE1 engages monitor, and
still on afterwards rather than blipping.
One bug found the hard way, and the stub was the cause. QmpClient.command returns the
unwrapped value and raises on error, but the test stub returned {"return": ...}. So
webui.py was written to unwrap a second time, every test passed, and the live UI
returned 500 with "argument of type 'bool' is not iterable". The stub is now pinned to
the real contract by a test. A stub more forgiving than the real thing is worse than
no stub.
Also stopped swallowing the failure: a bare `except: return None` made the error
indistinguishable from a radio that simply was not making sound, and cost a detour
into looking for a stale process.
The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:
bash tools/run_tests.sh # everything, 11 suites, a few minutes
bash tools/run_tests.sh -q # unit tests only, ~15 s, no emulator
The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.
Two defects in the runner's own first draft, both caught before it was trusted:
It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.
That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.
Full run: 11 passed, 0 failed.
bk4819_eval_receiver reports RSSI above any sane squelch threshold on every poll, and
a scan halts when it finds a busy channel -- so the S-meter work could plausibly have
stopped scanning dead on its first step. Measured: it does not, 6 distinct tunings
over 6 samples.
The check compares the framebuffer rows holding the frequency digits rather than
counting distinct frames. Frame comparison would pass on a completely parked radio,
because RSSI varies every poll and the meter redraws constantly -- an initial run
scored 8/8 distinct frames and proved nothing. Narrowing to pages 1-2 answers the
actual question: did the radio retune.
Incidentally established that page 4 is not the meter row: it was byte-identical
across all six samples while the frequency changed.
Three places claimed there was no PTT line, and one of them named the wrong pin
(GPIOC rather than PB10). All now say the same thing: PTT works, but not through
the key table, because the firmware reads its own pin instead of scanning it as a
matrix key.
Documents the transmit bar's two gates -- FUNCTION_TRANSMIT and gSetting_mic_bar,
the latter already on because blank flash reads 0xFF -- and why the release path
gets more attention than the press: a stuck PTT leaves every later test running
against a transmitting radio.
Also records the '.key' vs '.key[data-key]' trap for whoever adds the next
non-key button.
PTT was the one input the model never had, so the radio could not be keyed and the
mic level bar was unreachable. It is not a matrix key -- GPIO_IsPttPressed reads
PB10 directly (driver/gpio.h:31, active low) -- so it gets its own GPIO line rather
than a column/row intersection.
With it the transmit screen is complete: TX annunciator, a running timer, and a
level bar at roughly 80% of scale, fed from REG_64 via BK4819_GetVoiceAmplitudeOut.
app/app.c:1700 draws that only while gCurrentFunction == FUNCTION_TRANSMIT with
gSetting_mic_bar set; the setting is Data[7] bit 4 at flash 0xA0A8, and blank flash
reads 0xFF, so it is already on.
Exposed as a boolean on the keypad device and as POST /api/ptt with an explicit
held flag, plus a button in the browser UI. Held rather than tapped, because
transmitting is a state the operator stays in and a fixed duration would be wrong
for it.
Releasing is treated as the important half:
- pointerleave and pointercancel release, so dragging off the button cannot leave
the radio keyed
- pagehide releases, so closing the tab cannot either
- /api/release-all clears PTT too, since it is outside the matrix and an empty
press does not touch it
- non-boolean bodies are rejected, so {"held": "false"} cannot key the transmitter
by truthiness
Two things the test suite caught, both real:
The browser wired '.key' handlers over every styled button, and the PTT button
carries no data-key, so it would have sent the key "undefined". Narrowed to
'.key[data-key]'.
StubClient.presses() collected every qom-set regardless of property, so PTT's
booleans landed among the key names and presses()[-1] reported False after a
release-all. Now filtered by property, with a matching ptts() accessor.
tools/test_ptt.py covers the path end to end and asserts the release as well as the
press: a PTT that stuck would leave every later test running against a transmitting
radio.
The firmware now draws a working meter: -53 dBm, +40 over S9, nine of thirteen
segments, next to a MONI label and a running receive timer. The two numbers agree
with each other -- S9 is -93 dBm on UHF, so -53 really is S9+40.
Three pieces had to line up, and the order they were found in was the hard part.
RSSI and audio amplitude are refreshed when the firmware polls REG_0C, not when it
configures the chip. Raising a flag at configuration time is a trap: REG_3F is
written 0 then 0x0C0C repeatedly during setup, so anything announced there is
disabled again before it can be collected.
The squelch flag is SQUELCH_LOST, bit 2 -- not SQUELCH_FOUND. Per app/app.c:1027
"squelch lost" is what sets g_SquelchLost = true, i.e. a signal is present.
SQUELCH_FOUND reads like "found a signal" and means the opposite.
Announcing is rate-limited to every 64th poll. Announcing once means the firmware
collects it during startup, before the flag leads anywhere. Announcing on every poll
re-arms the request bit inside the firmware's own collection loop, which uses REG_0C
as its condition and has no timeout, so it never exits. Periodic satisfies both.
What finally made the meter appear was not the interrupt at all. The radio idles in
power save and does not act on squelch there. ACTION_Monitor skips squelch entirely
-- app/app.c:482 picks FUNCTION_MONITOR over FUNCTION_RECEIVE when gMonitor is set --
and settings.c:263 defaults an out-of-range stored action to ACTION_OPT_MONITOR,
which blank flash (0xFF) is. So SIDE1 short-press is the way in. Measured: fn=5
idle=1 monitor=0 before, fn=2 idle=0 monitor=1 after.
Gating on RX_DSP (REG_30 bit 0, from App/driver/bk4819-regs.h:240) rather than the
whole register being zero: TX and tone paths leave other bits set with RX_DSP clear,
and would otherwise look like a live receiver.
tools/test_smeter.py covers the whole path -- boots pristine, confirms power save,
presses SIDE1, and checks the screen gained content. It compares lit-pixel counts
rather than matching pixels, so an unrelated UI change does not produce a mysterious
failure.
None of this is radio simulation. The levels are plausible numbers that move; they
are not the result of modelling a signal. What they buy is firmware control flow
running on live values instead of on zero.
Every read came back doubled: seed REG_0C with 0x1248 and the firmware received
0x2490. The command byte's own trailing falling edge was being treated as a data
clock, so bit 15 was shifted away before the guest sampled it and the whole word
landed one place too high.
Each firmware bit is read/raise/lower (BK4819_ReadU16), which means the eighth
command bit is followed by a falling edge before the data phase begins. Skip that
one edge.
Why it went unnoticed: writes were always fine -- 52 registers held exactly what
the firmware wrote -- and the register the firmware polls hardest, REG_0C, was
legitimately 0 in this model. Reading zero and getting zero looks like success.
The skew only surfaced when something tried to report a value through it.
It also explains four failed attempts at the squelch interrupt. The model raised
REG_0C bit 0; the firmware received bit 1. So
while (BK4819_ReadRegister(BK4819_REG_0C) & 1u)
was never true, the acknowledging write inside it never ran, and 1719 polls saw a
flag the guest could not act on. Every one of those attempts was diagnosed as a
timing or gating problem and was not.
tools/test_bk4819_readback.sh locks it down. It seeds REG_0C -- read ~1700 times
per 30s, so a sample is guaranteed -- with a value carrying bits in both halves,
so a shift either way is unmistakable, and names the direction on failure. Bit 0
is left clear on purpose: with it set the firmware enters an acknowledge loop that
has no timeout, and this test is about alignment only.
Confirmed by A/B: committed code 0x2490, patched 0x1248.
The transceiver was not modelled at all. Its bit-banged three-wire bus went nowhere,
so every register read returned whatever the floating GPIO happened to be, and PB9
had to be idled low as a workaround: with the line high, reads came back 0xFFFF and
RADIO_SetupRegisters spun forever on bit 0 of REG_0C.
Now a real device, wired to the pins the driver uses -- CS on PF9, SCL PB8, SDA PB9
with both directions connected -- decoding the protocol from App/driver/bk4819.c:
CS low, eight bits of register number MSB first with bit 7 set for a read, then
sixteen bits of data. Registers read back what the firmware wrote; the few it reads
without having written return plausible values.
Scope, deliberately narrow: this is the register interface, not the radio. The chip
has no public datasheet, so the driver is the only specification available and it can
only say which registers were written, never what left the antenna. Keying envelopes,
spurious emissions and sensitivity still need a real radio and a spectrum analyser.
The comments say so at the top of the device, so a passing test here is not mistaken
for evidence about RF.
What it buys is control flow that evaluates real values. RSSI was hard zero at 18
call sites -- -160 dBm -- so the S-meter read empty and squelch and scan decisions
saw a dead band. It now reports about -40 dBm. Measurably, the main screen comes up
on 400 MHz instead of the 18 MHz floor, because band setup no longer reads zeros.
Two things the untimed spins force:
- REG_0C bit 0 must stay clear. App/app/app.c:910 and :1417 loop on it with no
timeout whatsoever, so a stuck bit hangs the guest rather than degrading.
- A soft reset (REG_00 bit 15, which BK4819_Init issues first) has to re-seed the
measurement registers. Real hardware keeps measuring afterwards; this model would
be left holding zeros. That was not theoretical -- the first test run decoded 48
registers correctly and still reported RSSI as 0 for exactly this reason.
The register file is exposed over QOM as regNN so tests can inspect it without gdb.
That matters beyond convenience: attaching a debugger pauses the guest and changes
timing-sensitive behaviour, which has repeatedly produced conclusions that were
artefacts of the measurement rather than facts about the firmware.
tools/test_bk4819.py checks the guest still boots (i.e. the spin terminates), that
dozens of registers hold written values (52 currently, so the transfer really is
being decoded), that RSSI is not zero, and that REG_0C bit 0 is clear.
keypad_test.py, test_flash_persist.py, test_freq_entry.py, test_serial_rx.py and the
143 unit tests all still pass.
Nothing could be sent to the firmware. Three separate pieces were missing.
USART1 was a register stub with no chardev, so there was no source of incoming
bytes. It now takes a chardev property, defaulting to serial0, so -serial works.
The DMA model never serviced USART. App/driver/uart.c receives over a circular
peripheral-to-memory channel and never reads DR; it finds new data with
write_ptr = sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(DMA1, CHANNEL_2)
so leaving CNDTR at its programmed value made the buffer look permanently empty
however many bytes arrived. DMA now drains USART1's queue byte by byte, decrementing
CNDTR and reloading it in circular mode. Channels also remember the length they were
given, since CNDTR counts down and the offset into the buffer has to be derived from
the difference.
Servicing happens on a CNDTR read rather than from a timer: that read is precisely
how the driver looks for data, so no polling is needed and no byte can be delivered
before the guest asks.
And transmit was invisible to the far end. DR writes went to stderr only, so a host
tool would send a command, the firmware would answer, and the answer went nowhere the
tool could see -- indistinguishable from being ignored. This cost a debugging round:
the first test run reported "no reply at all" with 0 bytes of boot output, which
looked like receive failing when the boot banner was in fact being written to stderr
as always. DR now also writes the raw byte to the chardev when one is connected.
SR reports RXNE when bytes are queued and DR consumes one, so a polling firmware
would work too, even though this one uses DMA.
tools/test_serial_rx.py speaks the real wire protocol -- AB CD framing, the fixed XOR
obfuscation, CRC-16/XMODEM -- and checks two exchanges end to end:
0x0514 hello -> 0x0515 ack
0x051B EEPROM read -> 0x051C with the requested 8 bytes at 0x0E70
Both pass. CPS/CHIRP-style tools can now talk to the emulator. keypad_test.py,
test_flash_persist.py, test_freq_entry.py and the 143 unit tests still pass.
Two faults that combined to take the web UI down twice in one session, each time
surfacing to the user as a 502 through the reverse proxy.
`pkill -f 'M uv-k5-v3'` in run.sh and trace_run.sh matched far more than intended.
-f tests the whole command line, so it also matched the shell running the pkill
(the pattern sits in its own argv), any script mentioning the machine type, and the
QEMU child of a running webui.py. Cleanup now lives in tools/lib_kill_emulator.sh:
pgrep -x on the binary name, confirm uv-k5-v3 in /proc/PID/cmdline, and optionally
scope to one QMP socket so a caller only stops the instance it owns.
The supervisor then could not recover from it. power_on() began with
`if self._client is not None: return False`, but a client object is not proof of a
live guest -- after an external kill the stale client made power_on refuse forever,
so the Power button was dead until the whole service was restarted. It now checks
whether the process actually exited and relaunches, logging why.
Tests:
- tools/test_kill_emulator.sh checks a plain process, a process whose command line
merely mentions uv-k5-v3, and the calling script all survive; that a real emulator
on a named socket is stopped; that one on another socket is not; and that an
unscoped call still clears everything. Verified it leaves a live webui.py alone.
- Two supervisor unit tests cover relaunch-after-external-kill and the case that
must still refuse, so this cannot regress into starting two emulators at once.
Verified end to end against the running web UI: kill the emulator from outside,
status reports unreachable, and pressing Power brings it back to
{"status":"running"} where before it stayed dead.
While writing the first version of the test I modelled the failure as a client
raising BrokenPipeError, which is not what is_running() looks at -- it checks
poll(). The mock was wrong, not the code; the test now has the process report an
exit status, which is what really happens.
DMA moved bytes through address_space_memory, which cannot decode this SoC's memory
at all: the container region holding flash, SRAM and the peripherals is handed only
to the ARMv7M core and never registered with global system memory. Reads came back
MEMTX_DECODE_ERROR with all-zero data, and writes went nowhere. Proved directly --
an address_space_read of SRAM through it returns result=2 and 00000000, while the
same address read through the container returns the real contents.
This is what "the frequency will not change" and "flash behaves like RAM" had in
common. PY25Q16_WriteBuffer reads a 4 KB sector into SectorCache, patches the part
it wants, and programs the whole sector back. The read looked healthy from the flash
side -- the model handed over real 0xFF bytes -- but DMA dropped them, so the
write-back sourced 4096 zeros and cleared the sector, VFO frequencies at 0x9000
included. RADIO_ConfigureChannel only substitutes a band's lower limit for
0xFFFFFFFF, so a stored zero was used as-is and clamped to BX4819_band1_lower.
That is where the 18.000 MHz came from, every time.
DMA now runs over an AddressSpace built on the SoC container, and refuses to
transfer at all if none is configured rather than silently moving zeros.
tools/test_freq_entry.py covers the whole user-visible path: type 435000, confirm
435.00000 MHz lands in band 5, confirm no other band was zeroed, and confirm it is
still there after a power cycle. It drives QMP with no debugger attached, because
the input box times out in ~2.5 s and a gdb attach takes longer -- that alone
invalidated several earlier investigations.
Verified: 435 MHz now appears at flash 0x90A0 where before the entire sector read
zero. keypad_test.py, test_flash_persist.py and the 141 unit tests all pass.
Two related faults in the DMA model, both of which corrupted flash reads.
Transfers started when a channel was enabled. On hardware, enabling only arms a
channel; the transfer begins when the peripheral raises its DMA request. The flash
driver's SPI_ReadBuf arms RX, arms TX, then enables SPI and sets TXDMAEN -- so
firing at arm time clocked the bus before the read command had been sent. SPI now
kicks the armed channels from CR1/CR2 when SPE and a DMA request enable are both
set, which covers the read path (SPE last) and the write path alike.
Each channel also ran to completion independently. SPI is duplex: one clocked byte
is simultaneously sent and received, and the driver relies on that, pairing a
memory-to-peripheral channel feeding dummy bytes with a peripheral-to-memory
channel collecting the reply. Running them in sequence meant TX clocked the whole
transfer out before RX looked at the bus, so RX collected nothing. They are now
stepped together, one byte at a time.
Either fault alone made a 4 KB sector read return zeros. PY25Q16_WriteBuffer reads
a sector into SectorCache, patches it, and writes the whole thing back, so a zeroed
read turned into a zeroed sector -- including the per-band VFO frequencies at
0x9000. That is a second, independent cause of typed frequencies reverting to
18 MHz, on top of the missing page wrap fixed in da1ad7e.
Verified: the frequency area stays 0xFF across a boot where it was previously
zeroed every time, and an instrumented build shows the sector read now returning
0xFF rather than 0x00.
test_flash_persist.py now starts from the pristine image rather than
assets/flash.img. A dirty live image left nothing for the boot to write, which
surfaced as "the image is byte-identical" -- a failure that looks like broken
persistence but is really a dirty fixture. Runs twice in a row cleanly now.
keypad_test.py and the 141 unit tests still pass.
Real SPI NOR latches only the low address bits into its page buffer, so a program
burst that runs past the page boundary continues at the start of the same page. The
model incremented the address straight through instead.
Consequence: the firmware issues a 512-byte burst at 0x008F00 inside a single CS
assertion (measured -- the CS never drops mid-burst), which spilled into 0x009000.
That is the per-band VFO frequency area in eeprom_compat.c's map, so stored
frequencies were zeroed. RADIO_ConfigureChannel only substitutes the band's lower
limit when it reads 0xFFFFFFFF, so a stored 0 was taken literally and clamped to
BX4819_band1_lower -- which is why every typed frequency reverted to 18.000 MHz.
Verified: writes now align to sectors (0x8000-0x9000 and 0xA000-0xB000) and an
instrumented build records zero stores into 0x9000-0x90D6, where before it was
overwritten on every boot.
test_flash_persist.py had encoded the bug in its expectations: it watched 0x008100
and 0x00A100, which were only ever written *because* of the missing wrap. Those move
to 0x008000/0x00A000, and a MUST_NOT_CHANGE guard on the frequency area now fails if
a write spills there again.
keypad_test.py still passes.
Power on returned HTTP 500 with a ConnectionRefusedError traceback. A unix socket
file outlives the process that created it, so a killed QEMU left
/tmp/uvk5-qmp.sock behind; wait_for_socket only checked os.path.exists, returned
immediately, and the connect then failed. It now probes with a real connect, which
distinguishes "listening" from "leftover file".
Two related hardenings:
- power_on cleans up if connecting fails. Otherwise a half-started QEMU keeps
running untracked, holds the socket, and blocks the next power on -- which is
how one stale socket turned into a repeatable failure.
- The route reports a failed power action as 503 with the reason, instead of a 500
and a traceback the browser cannot display.
Three tests cover the stale socket, a real listener, and a path that never appears.
The SPI NOR model read its image at realize and never wrote back, so the "flash"
was a g_malloc buffer: everything the firmware saved -- settings, edited
frequencies, channel data -- vanished when the QEMU process exited. That is the
"it behaves like RAM" the user reported, and the image on disk still had its
original mtime and was byte-identical to what make_flash.py produces.
Page-program and sector-erase now mark the image dirty, and it is written out when
chip select is released. Flushing there rather than per byte means one file write
per settings save instead of thousands, because the firmware's driver holds CS for
a whole erase-and-program sequence.
Written via a temporary file and rename: an interrupted flush must not leave a
truncated image, since that file is the only copy of the radio's state. A short
write keeps the previous image rather than replacing it with a partial one.
Also flushes from an exit notifier. Deselect covers normal operation, but QMP quit
-- which is what the web UI's power off sends -- can arrive with the chip still
selected, and the last write would be dropped.
tools/test_flash_persist.py covers it end to end on a copy of the image, so it
cannot disturb a running session. It asserts specific regions rather than just a
changed hash: flash 0x008100 (MR/VFO attributes) and 0x00A100 (settings), which are
the two sectors a boot demonstrably writes. Both were 0xFF before and non-0xFF
after, and sha256 moved from 933d6974 to 7cdff6ce.
Three approaches were tried and abandoned first, all for the same reason: driving
the guest from gdb. `call EEPROM_WriteBuffer` and `call SETTINGS_SaveSettings`
both hang, because the main loop is running and the called function waits on
hardware the debugger has frozen. Observing the file is simpler and closer to what
the user actually sees. An earlier version of the test also watched EEPROM offsets
instead of flash offsets and reported "same" for every region while persistence was
in fact working -- the two address spaces are related by the table in
App/driver/eeprom_compat.c, not equal.
keypad_test.py still passes, which matters because this file is where deleting
three fprintfs once silently removed the keypad.
assets/flash.img was gitignored, so the only copy of the never-booted image lived
on one disk. The emulator writes to that image, so a session can leave edited
settings or a damaged EEPROM behind with nothing to restore from.
assets/pristine/ now holds the image as first generated, gzipped and checksummed,
and is tracked deliberately. Gzip takes it from 2 MiB to 2.3 KiB because the image
is nearly all 0xFF, which is what makes keeping it in git reasonable. The live
image and its .bak-* files stay ignored.
Two checksums are recorded, for the archive and for its contents, so a corrupted
archive is distinguishable from one that was replaced.
tools/restore_flash.sh verifies, diffs, or restores. Restore backs up the current
image first, then re-checks the result, since a restore that silently half-worked
would be worse than none.
Verified by deliberately corrupting the live image: --diff reported 32 differing
bytes, restore backed up and rewrote it, and --diff then reported no change. The
image is currently byte-identical to what make_flash.py produces, so this is the
genuine original rather than a copy of something already used.
Entries gain an "ip" field, rendered between the time and the source as asked.
The buffer is shared by every viewer, so without attribution a log of keypresses
from two people is unreadable.
Resolving the address matters more than it looks: behind the nginx reverse proxy
REMOTE_ADDR is always 127.0.0.1, so the first hop of X-Forwarded-For is what
identifies the real client. Only the first entry is trusted -- the rest of the
chain is set by the caller and a test covers that.
Power actions are logged at the route rather than in the supervisor, which has no
request context, so "who powered it off" is recorded.
Entries with no client behind them keep ip=None and render as "-": firmware serial
and QEMU stderr are not caused by a request.
The sharing and history the user asked for already worked and needed no change --
verified rather than assumed. The front end starts at logCursor=0, so a page
opened now receives the full buffer, including lines produced before it connected
and lines from other people. Confirmed live through the proxy: a new reader saw
entries attributed to 172.21.91.140, fd3c:3f9b:6424:2::5 and "-".
Reverts optimistic send. The browser times the press with performance.now() and
sends it once on release, so the firmware sees exactly the press that was made.
Optimistic send fired a speculative tap at pointerdown plus a held press if the
button was still down. It was 152 ms faster (505 vs 657 ms click-to-visible at
400 ms RTT) but it guessed, and a wrong guess sent both presses for the firmware
to act on. Raising the threshold to 900 ms hid the symptom without removing the
failure mode, and it also made hold-to-repeat unreachable, since the server
released the key after a fixed 900 ms no matter how long you held.
Measuring costs the click duration in latency and buys exactness plus real
hold-to-repeat. Verified against the firmware:
taps under 400 ms 120/250/390 ms -> cursor +1, submenu never opens
holds from 400 ms 500/900/1500 ms -> cursor +3/+8/+15
MENU tap 120/300/390 ms -> menu opens, submenu stays shut
One correction to my own expectations along the way: I first recorded the multi-step
moves at 500 and 800 ms as failures. They are not. App/misc.c has
key_repeat_10ms = 8, so past 400 ms the firmware auto-repeats every 80 ms, and the
counts match (duration - 400) / 80. That is what a real radio does when you hold a
button, so the note on FIRMWARE_HELD_MS now says not to filter it out.
MIN_HOLD_MS returns as the floor for a measured press, since a very fast click can
measure below the debounce window. LONG_PRESS_AFTER_MS and LONG_PRESS_MS are gone
with the scheme that needed them.
Three things reported from actual use, plus the bug the logging exposed.
1. Keys are logged (source "key"), including refusals and presses while powered
off. Without this there was no way to tell "the key never arrived" from "the
firmware did something else with it" -- which is exactly what was needed below.
2. /stream resends the current frame every IDLE_FRAME_INTERVAL_S even when nothing
changed. Change-detection alone made a static screen indistinguishable from a
dead connection, and a client joining mid-idle stayed blank. Measured: 6 frames
in 10 idle seconds, where before it was 0.
3. Off now actually blanks the screen. The dark panel moved to a .screenwrap
wrapper and the <img> is hidden; setting a background on the <img> alone did
nothing visible, because the image kept painting the last frame over it.
Then the reported bug: in the menu, UP/DOWN behaved like another MENU press. The
key log made it diagnosable and the cause was mine -- LONG_PRESS_AFTER_MS was set
to the firmware's own 400 ms boundary, but a deliberate click runs 100-500 ms, so
ordinary clicks sent tap AND held and the firmware acted on both:
held DOWN auto-repeated, gMenuCursor 3 -> 12 from one click
held MENU entered the submenu, gIsInSubMenu 0 -> 1
The UI threshold is now 900 ms, well clear of any click, and FIRMWARE_HELD_MS is a
separate constant so the two are not conflated again. Verified against the real
firmware: 120/300/500/800 ms clicks each move the cursor exactly +1 with
submenu=0, while a deliberate 1400 ms hold still auto-repeats (+9).
Three sources into one buffer: power events from the supervisor, QEMU's stderr
(which run.sh and the tests used to discard), and firmware serial, which the
machine model tags SERIAL. default_launcher now captures stderr rather than
sending it to DEVNULL, which is what made the last two reachable.
The pane is a fixed-height scroll box as asked: 180px with overflow-y:auto, so it
never grows with content -- older lines move up out of view and you scroll back to
read them.
Two details that make that usable rather than annoying:
- Autoscroll only sticks when you are already at the bottom. Otherwise a new line
arriving would yank the view away from whatever you had scrolled up to read.
- MAX_LOG_LINES caps the <pre> as well. The box is fixed-height either way, but an
unbounded DOM node would still grow memory across a long session.
Verified on the live server: power on produced power/qemu/serial lines including
"UV-K5 Firmware, EGZUMER-F4HWN+NR7Y c91cec95", each Reset logs the event and the
banner reappearing, and since= never resent a line. A capacity-500 buffer fed 2000
lines keeps exactly 500 and does not replay evicted entries to a stale cursor.
Collects power events, QEMU stderr, and firmware serial output. Bounded and in
memory on purpose: an unbounded buffer in a long-running server is a slow leak,
and anyone wanting a permanent record can redirect the server's stderr.
Entries carry a monotonic seq so a polling client can ask for "anything after N"
and receive each line exactly once, including after eviction has dropped older
entries -- a test covers that case specifically, since an index-based cursor would
silently repeat or skip lines there.
Stream decoding is lenient: serial bytes can be garbage before the firmware
configures the port, and losing the stream to one bad byte would be worse than a
replacement character.
Two changes, both aimed only at latency, since that is what the link makes
expensive.
1. Send on pointerdown instead of pointerup. Waiting for release left the network
idle for the entire click. The duration is unknown at that moment, so the
speculative request asks for a short press; holding past LONG_PRESS_AFTER_MS
sends a second, deliberately long press, which is how held events stay
reachable.
2. TAP_MS 200 -> 60 ms. The server blocks for hold_ms before replying, so this is
latency the user pays directly. 200 ms was a guess that gave back most of what
change 1 saved.
The 60 ms is measured, and the sample size mattered: at 4 trials per value 30 ms
looked reliable, but at 12 trials 20 ms registered only 5/12 while 30 ms was 12/12.
The nominal 20 ms debounce is not sufficient alone because KEYBOARD_Poll samples
each column 8 times wanting 2 matching reads. 60 ms is double the proven floor.
Click-to-visible at 400 ms RTT, where one round trip is an unavoidable 400 ms:
original (2 requests, on release) 2525 ms (+2125 over the floor)
one request, on release 657 ms (+257)
current (1 request, on press) 505 ms (+105)
Long press still works: a 60 ms press opens the menu (gScreenToDisplay 0 -> 1)
while a 900 ms press is treated as held and correctly does not, so the firmware
still distinguishes them.
Also drops MIN_HOLD_MS, now dead: the browser no longer measures press duration,
so there is no measurement to clamp. Its test asserted only that the string
appeared, which would have kept passing over dead code.
The firmware has been printing all along and nothing was listening. USART1 has no
real model here -- it is one of the logging catch-all stubs -- so every byte went
into qemu_log_mask(LOG_UNIMP) and vanished.
Two things were needed, and the second was not in the plan:
1. Print USART1 DR writes (+0x04, per the vendor CMSIS header) as SERIAL lines.
2. Report TXE|TC in USART1 SR. This is the part I had missed. UART_Send() in
App/driver/uart.c spins on LL_USART_IsActiveFlag_TXE() with a bounded timeout
and *skips the byte* when the flag never sets. A stub returning 0 for SR meant
the firmware discarded its own output before it ever reached DR -- the only
write arriving was UART_Init()'s priming zero. So step 1 alone produced
nothing, which is why the first attempt looked like "the build has no logging".
Then a bug of my own: the priming byte is 0x00, I buffered it, and fprintf("%s")
stopped at that NUL and printed an empty line while all 46 bytes sat behind it.
NULs are now dropped, and a line flushes on CR as well as LF.
Verified: SERIAL UV-K5 Firmware, EGZUMER-F4HWN+NR7Y c91cec95
keypad_test.py still passes, which matters because this file is where removing
three fprintfs once silently deleted the keypad.
The server now owns the QEMU process by default but does not launch it. You open
the page to a dark screen and press On, which is the behaviour asked for: like
walking up to a machine rather than finding it already booted.
This inverts the flag from the plan. Owning the process has to be the default,
since it is the only way On/Off can work at all; --attach is the opt-in for joining
a run.sh instance, where Off is refused.
Two tests guard the intent rather than the wiring: one asserts main() has --attach
and not --own-emulator, another asserts main() never calls power_on(), so a future
edit cannot quietly restore auto-boot.
Verified on the live server with no QEMU running beforehand:
startup 0 QEMU processes, powered=false, frame.png 503, page 200
On 1 QEMU process, powered=true, frame.png 200 (2920 bytes)
Off 0 QEMU processes, frame.png 503, page still 200, keys 409
On again 1 QEMU process, powered=true, frame.png 200
The web server stays up across Off, which is what you asked for: the screen goes
dark and waits for the next person to press On.
Buttons sit above the LCD as asked, with a state label that turns green when the
guest is up. Powered off dims the panel via a screen-off class, so a dark screen
is the signal rather than a frozen last frame.
Off asks for confirmation: it ends the guest, and a stray click should not do that
silently. Buttons disable while a power action is in flight, since On and Off take
a couple of seconds and a double click would race.
The stream is restarted after every power action. The old multipart response ends
when the emulator goes away, so without a fresh src the image would stay blank
after On.
Refusals are surfaced in the status line rather than swallowed -- that is how the
409 for an adopted emulator becomes visible instead of looking like a dead button.
49 unit tests, and the generated page passes node --check.
POST /api/power/{on,off,reset,pause,resume}, and every route now tolerates there
being no emulator: /api/status reports powered:false, /frame.png returns 503,
keypresses return 409 with "press On first". Previously create_app required a live
client and the whole page would 500.
The client is fetched through the supervisor per request rather than captured once,
because a power cycle replaces it and the captured one goes stale.
After any power action the pump is rebound, so Off actually goes dark instead of
freezing on the last frame.
Off is refused with 409 for an adopted emulator: we did not start that process.
Reset is allowed either way, since system_reset ends nothing.
Verified over HTTP with a real supervisor and real QEMU:
start powered=False qemu=0 frame=503
key 409 as expected
On powered=True qemu=1 frame=2920 bytes
Reset qemu=1 (process survived)
Off powered=False qemu=0 frame=503
On again powered=True qemu=1 frame=2920 bytes
The web server stayed up throughout, which is the requested behaviour.