pip install ziglang cross-compiles to thumb-freestanding-eabi, which is enough to build a .app with no arm-none-eabi-gcc and no Docker. Measured refusals: --defsym, -Ttext and --section-start come back as unsupported linker args, so the VMA is resolved into a copy of app.ld; -T is forwarded (a missing script errors); --image-base is accepted but page-aligns the segments into 0x200103C8 and 0x20020C94. tools/elf2bin.py extracts allocated sections rather than program headers, because lld maps the ELF header and phdr table as a 180-byte LOAD of its own -- following the headers starts the image at 0x20000000 and APP_ERR_VMA. One division pulled in __aeabi_uidiv, which a -nostdlib blob cannot have: Minesweeper now avoids division entirely. app_main carries the .text.entry attribute upstream's apps use, so the entry is first for the loader's jump to offset 0.
Minesweeper builds to 2408 bytes of code against a 4096-byte budget. Installed through the page, the firmware reads the slot header twice and then exactly code_size bytes from slot+0x1000 -- the only read of that size in the boot log -- so the blob shape, header, CRC, VMA and offset are all accepted. Whether control reaches the overlay is unproven: the 100 ms PC probe saw no overlay address, no APP ERROR screen appears, and the app does not draw. test_elf2bin pins the phantom-header-segment lesson; both AGENTS files record the rest.
apps/minesweeper/host_test.c includes the app source with a fake app_api_t, so the real state machine runs on a PC: every string it draws is recorded and the lit pixels are counted. Measured -- a reveal/flag/new-game/digit/quit script returns normally, draws the title, the mine count and lights 24 pixels; a script that blindly reveals 85 cells reaches a terminal state eleven times, draws BOOM, and MENU starts a new game after it (terminal at record 356, title again at 4180). The win path is the one branch blind play does not reach.
Both READMEs now say what is verified (compiles clean under gcc -Wall -Wextra -Werror against upstream's real app_api.h, which is what caught the API's true member names and the absence of left/right keys; the logic runs on the host) and what is not (never built for ARM, never run on the radio -- no toolchain and no Docker here).
The blank screen was self-inflicted and the earlier explanation was wrong. The FMP3 marker at 0x100000 is an ordinary state record -- generation, image_size, image_crc32, firmware_slot/slot_inv, config_bank/bank_inv and a matching state_crc32 (0x661286F1) -- not a pending flag. What actually happened: a firmware uploaded over a state that recorded a different identity made the firmware take the restore/adopt path, which draws nothing (panel 1024/1024 bytes zero) while the serial banner printed normally. Restoring the untouched dump fixed it at once (485/1024 bytes lit) and the three apps reinstalled and were confirmed by 0x0730.
Clearing the marker sectors was tried twice (a whole 8 KiB, then just the 24-byte headers) and is not a fix: it sends the firmware down MB_MARK_MISSING = fresh radio, which adopts the running firmware slowly and without drawing, and its own write-back restores the marker anyway. AGENTS.md and AGENTS.zh-CN.md now say so in place of the wrong claim.
apps/minesweeper/ adds our own 9x9 minesweeper for the 4 KiB overlay: no left/right keys exist on this radio (so the cursor walks with UP/DOWN and digits pick a row then a column), 81 cells need three 9-byte bit arrays rather than a uint16_t mask (the uint16_t version compiled fine and was wrong past cell 15), mines are placed after the first reveal so it cannot lose immediately, and the source compiles clean under gcc -Wall -Wextra -Werror against upstream's real app_api.h. It has not been built for ARM or run -- no toolchain here -- and the README says so.
A real image whose 0x100000 marker held FMP3 next to a committed slot 0 made the factory bootloader reflash the internal flash from that slot on every power-on: the serial banner reappeared once a cycle (7 -> 8 in 25 s) while the screen never changed. Clearing the two marker sectors (0x100000..0x101FFF, stopping just before the app region at 0x102000) ended it, and the radio booted once and stayed.
It also explained why page-side installs vanished: the emulator writes its in-memory image back on exit, and the looping guest's copy was older than the file, so powering it off overwrote the installs. With the loop gone the same installs survive a power cycle -- installed, powered off, still listed, powered on, still listed, and 0x0730 answered with all three. Both notes are in AGENTS.md and AGENTS.zh-CN.md; work/app-template/ got a starter app, its README and the upstream api/ld it needs (git-ignored).
The page edits a working copy of the flash image on purpose -- the file it was first pointed at may be a real calibration dump -- but a restart pointed at a different image switched to that file's copy instead, and the games installed through the page looked like they had vanished (measured: the app table came back holding an older slot). work/run-webui.ps1 now prefers the working copy the page has been editing, ahead of the original, and the page's own hint says which image it is editing and that the loaded one is never modified. A test asserts the order in the script.
Sixteen app rows (most of them empty or holding something that is not an app) plus five slot rows made the column long enough to be annoying, so each table is now a details pane: the firmware slots start open, the app list starts folded, and the state that used to sit in its own row moved into the summary line. Same ids, so loadSlots and loadApps are untouched. A front-end test asserts both panes exist and that the app list starts folded.
F then 7 opens the F4HWN APPS menu and MENU runs the selection -- upstream's own wording in UVStudio's locales/en.js, and what App/apps/app_menu.c does: KEY_MENU on an installed row calls APP_LaunchOverlay, which runs until the app exits; EXIT leaves the menu. The page's hint now says it too.
Recorded with it: twenty keys short and long, the whole 79-entry settings menu and the multiboot menu all left the app region untouched (the multiboot menu reads firmware slots, not apps), and the way in turned out to be written in the flasher's translation file rather than the firmware source. When a feature's entry point is missing, the host tool that installs it is the document.
F then 7 opens the F4HWN APPS menu and MENU runs the selected app, which is upstream's own wording in UVStudio's locales/en.js and what App/apps/app_menu.c does (KEY_MENU on an installed row calls APP_LaunchOverlay, which runs until the app exits; EXIT leaves). Verified in the emulator with Tetris.app installed through the page's endpoint: F,7 reads the app region eight times and shows the boxed menu; MENU reads once inside slot 0's code and the game's screen replaces the radio's; DOWN moves the piece and the game is still there two seconds later. The page's own hint now says so.
Worth recording: twenty keys short and long, the whole 79-entry settings menu and the multiboot menu all left the app region untouched -- the multiboot menu reads firmware slots, not apps. The entry point was written down in the flasher's translation file, not in the firmware source: when a feature's way in is missing, the host tool that installs it is the document.
GET /api/apps/radio opens the firmware's serial port and sends 0x0730 for all sixteen slots, so the answer comes from the running firmware rather than from our reading of the file -- which is the check that matters, because the bytes can be right and the firmware still refuse a slot. Measured through the page after installing Beam.app into slot 0: slot 0 -> Beam 1.0, 1100 B, crc 0xd976058, shortcut beam, committed; slots 1..3 -> status 2 with unrelated data, the resource-block overlap the install guard refuses. A button beside the table asks it and shows the answer in its own column.
The server gives its own emulator a serial port (--serial-port, default 4445) and uvk5_slots_serial.Radio gained a public app_info(slot), so nothing reaches into a private helper. uvk5_apps.parse_radio_reply decodes the answer and is tested without a radio. Also recorded: QEMU needs the mingw64 DLLs on PATH, and started by hand without them it exits before opening QMP, which surfaces only as 'QMP socket never appeared'.
UVStudio's own js/flash.js names the protocol: MSG_APP_INFO 0x0730/0x0731, MSG_APP_ERASE 0x0732/0x0733, MSG_APP_WRITE 0x0734/0x0735, MSG_APP_VALIDATE 0x0736/0x0737, with APP_SLOT_COUNT 16, APP_IMG_OFFSET 0x1000, APP_HDR_SIZE 64 and APP_MAGIC 0x31504146 -- the constants this page already used. It uses slots 0..7 and labels them 1..8, and writes the header last so a partial write cannot validate; the page now labels slots the same way.
The Labs build answers 0x0730 over USART, so an installed Beam.app was queried on the radio: slot 0 came back status 0 with the header this page wrote (FAP1, code_size 1100, CRC 0x0d976058, flags 0x0801, name Beam), which confirms the region, the offset and the layout from the firmware's side rather than from a header I read. Slots 1 and 2 answered status 2 with unrelated data -- the overlap the install guard refuses to overwrite.
The English line about the page's app table went in with the previous commit; the Chinese one did not, because the sentence it anchors to wraps mid-phrase and the marker I used did not match. Added, so the pair says the same thing again.
The endpoints existed; now the page shows them. An Overlay apps block beside the firmware slots lists all 16 app slots with their name, version, shortcut and size, a .app file picker and an Erase button per row, and says where the region is. It reads GET /api/apps and posts to POST /api/apps/<n> and /api/apps/<n>/erase, so installing a game is a file pick where the firmware slots already are -- no WebSerial, no browser permission.
An upload into a slot that holds something which is not an app is refused by the server; the page asks once and retries with ?force=1 rather than either failing quietly or destroying the factory resource data that overlaps this region on a localised image. Front-end tests assert the table and the endpoint are in the served page and that it never asks for a serial port or audio, and TestPageScriptParses keeps checking that the page's script parses.
The Labs edition's apps live in the external flash, in the region its own App/apps/app_overlay.h defines (16 slots of 8 KiB from 0x102000, a 64-byte FAP1 header at the slot base, code one 4 KiB sector later), and upstream installs them from UVStudio over WebSerial. The page owns the image, so this adds GET /api/apps, POST /api/apps/<n> and POST /api/apps/<n>/erase, which put the same bytes at the same offsets with no serial protocol and no browser permission.
Measured on a real image while wiring it up: every one of the 16 slots already held data that is neither empty nor an app -- the localised build's factory resource block overlaps 0x102000 -- so install now refuses to overwrite anything that is not an app unless asked (--force, ?force=1), naming what is there. And _edit_flash edits a copy, which is why the source image shows no changed bytes; the first test read that as 'the install did nothing' and now checks FlashSlot.path. Tests: test_uvk5_apps grew to 20, test_webui.TestAppEndpoints adds 7 over the endpoints.
UVK5_FLASH_PROBE is a file path, not a flag: passing 1 wrote the probe output to a file called 1 in the repository root, and the previous commit's git add -A picked it up. Deleted and untracked; the real probe logs live under work/, which is ignored.
The Labs edition runs overlay apps (Tetris, Breakout, Plasma, Cube3D, Beam, Beacon, FoxHunt, BroadcastFM) that upstream UVStudio installs over WebSerial. This page owns the flash image, so the same bytes go to the same offsets with no serial protocol and no browser permission: APP_REGION_BASE 0x102000, APP_SLOT_STRIDE 0x2000, APP_CODE_OFFSET 0x1000, 16 slots, taken from the firmware's own App/apps/app_overlay.h rather than inferred.
tools/uvk5_apps.py parses and validates the 64-byte FAP1 header (zlib CRC-32 over the code, vma 0x20000280, name, version, capabilities), lists, installs and erases slots, and refuses what the firmware would show as APP ERROR. test_uvk5_apps covers those refusals plus install/erase/list round trips, and parses a real upstream Beam.app when one has been downloaded. The header struct was 60 bytes at first -- a missing vma field -- which the real file's bytes showed at once.
The page only knew its own input. A build called f4hwn.fusion.bin reports EGZUMER+F4HWN v6.0.0.CN, and with the multi-system release a committed external slot makes the factory bootloader reflash the internal flash from that slot on every power-on -- so the uploaded image never runs and the page keeps naming it. The firmware prints its own banner on USART1; tools/uvk5_banner.py reads it back, /api/firmware returns running: {banner, matches_uploaded, note}, and the page shows what the device reports, flagging it only when the running version is not in the uploaded image at all.
That reader also exposed a regression of my own: _start_stderr_pump had been rewritten to read the pipe in 64 KB chunks, which kept QEMU from blocking but delivered nothing to the log until 64 KB had accumulated -- and the banner is forty bytes, so it never appeared. It reads lines again, still starting before anything waits on QEMU, and test_uvk5_supervisor passes either way.
The screenshot example pasted one build's numbers and told the reader to get them with nm -- which cannot work here, because the images are program-header-only ELFs with no symbol table. It now asks the firmware through tools/uvk5_buffers.py first. The webui example drops the flags entirely, since the page draws the controller's memory and needs none.
Also fixed the paragraph in both READMEs that had a tool name and a flag on one line, which the doc checker (correctly) read as passing that flag to that tool.
The flags are optional now and the page needs no addresses at all, so the examples no longer paste one build's numbers: screenshot.py is shown asking the firmware through tools/uvk5_buffers.py, and the webui example omits them entirely. The paragraph that said they default to one known build is replaced with what actually happens.
The page was told --frame-addr 0x200012BE --status-addr 0x2000163E and used them as a fallback. The firmware the user actually flashed keeps its buffers at 0x2000129E/0x2000161E, 32 bytes earlier, so every line landed 32 bytes off: that is the "other firmware looks shifted" report. The images here are minimal ELFs with no symbol table, so there is nothing to read -- but the firmware's own buffers hold the same bytes the controller holds, and tools/uvk5_buffers.py finds them by matching (1024/1024 bytes for that file).
The two address flags are optional now, work/run-webui.ps1 passes no machine-specific values at all, and the page reports what it found in /api/status and /api/panel. tools/uvk5_testenv.qemu() also looks in the sibling qemu-7.2/build the rest of the repo assumes. Fixed /api/panel's emulator-off branch, which called jsonify with both a dict and kwargs and 500'd.
Tests: test_uvk5_buffers (the search must count matches, not pairs -- its first version scored every offset full marks and always answered the first one).
uvk5_stream.py used STATUS_BYTES without importing it, so the panel branch raised NameError on every frame and a bare except swallowed it: every screen the page drew came from guest RAM at one build's addresses. The pump now reports which source it used and why, webui exposes it (/api/panel and frame_source), and test_uvk5_stream asserts the panel wins when reachable and that a fallback is announced.
The ST7565 column counter wrapped at 128 instead of the controller's 132, so addresses 128..131 came back as 0..3, fell outside the col>=4 store, and were dropped: every row lost its last four pixels, which is where the battery icon lives. Pre-fix, filling a page with 0xFF left columns 124..127 blank; now they carry content and the page's frame matches the panel memory 8192/8192.
UVK5_BK4819_PROBE reassembles the sixteen bits a read clocks out and logs them against the register's own value: 1566 of 1566 agree on the working model. That agreement is not proof -- removing the skip_falling fix, which is exactly the historical left-shift regression, leaves the reassembled word unchanged, so this observation point is not the one the guest samples at. tools/test_bk4819_readback.sh therefore stays the guard.
tools/test_bk4819_readback.py is the working draft of a portable replacement (no source patch, no rebuild, no ARM gdb) and is deliberately NOT registered in run_tests.sh, so a proven guard is not swapped for an unproven one. Both the file and the two READMEs say so.
QMP takes a single client and the web UI holds it for its whole lifetime, so this is the normal outcome while the page is open rather than a broken emulator. Verified against a private instance: 128x64, 1819 pixels lit, 64 rows of ASCII.
tools/panel_dump.py prints the display controller's own memory as ASCII or PNG for any firmware, with the four mappings, so two builds can be compared instead of glanced at. The panel model gained a bounded UVK5_PANEL_PROBE diagnostic along the way.
Measured: the 5.9.0.CN panel agrees with its own framebuffer 8188 of 8192 pixels with the data untouched, and the fetched 6.0.0 build renders identically. Both program the same geometry registers, which is why the mapping is a driver convention and cannot be derived from the controller: it has to be measured.
Noted, not yet fixed: the display start line (0x40|n) is ignored, and the model stores pixels at col-4 with the column counter wrapping at 128 instead of the controller's 132 columns.
The probe scripts, run.sh, trace_run.sh, webui.py and two emulator tests each named the same hardcoded firmware from a source tree that is not in this repository. They now resolve QEMU and the firmware the way tools/uvk5_testenv.py does -- environment, then PATH, then whatever the checkout has -- and skip with a reason when there is nothing.
webui.py's --qemu and --elf lost their author defaults too: a bare qemu-system-arm through PATH, and no firmware until one is uploaded, which the page already reports.
tools/run_tests.sh defaults QEMU_SRC to a sibling of the checkout, which is where setup_qemu.sh puts it; the two defaults disagreed, so a fresh clone rebuilt nothing and reported a build that was not there. The interpreter list was also reading an empty $PY.
tools/test_bk4819_readback.sh was the last test with the author's paths, and the only one that could not run elsewhere. It now takes QEMU, GDB and ELF from the environment or PATH like the python tests, skips with a reason when one is missing, and says so on a platform whose QEMU cannot make the unix socket it uses.
tools/check_docs.py points UVK5_FW_DIR at a sibling and skips the file:line checks, with a message, when there is no firmware tree -- a fresh clone used to see thirteen failures it could do nothing about.
Added .github/workflows/unit.yml (the fast half of run_tests.sh on every push and PR), requirements-dev.txt for the one pip dependency, and a Dockerfile. Flask is not always installed, so test_webui now skips through setUpModule rather than erroring.
The bootloader's DFU handler is reachable only when SRAM[0x20000020] is 3, which only a program that then resets can write. The application's 0x05DD takes that path only with ENABLE_OVERLAY; this build resets straight back into the application instead. Ruled out by measurement: PTT alone, PTT+SIDE1/SIDE2, MENU, a host byte in the boot window including 0x0530, and 0x05DD.
boot-key now reads a + separated list, because the firmware's own BOOT_GetMode() needs PTT and a matrix key together for every special boot mode. Tested against all four documented combinations.
flash controller: store ACR/OPTKEYR instead of swallowing them, which is what stopped the factory bootloader from starting
slots over the firmware's own serial protocol (0x0720 family); uvk5_socket/uvk5_testenv so a fresh checkout skips instead of failing; web UI slot table and Multiboot button; quick start, CONTRIBUTING, and stop tracking firmware images and radio dumps
The remaining class of claim check_docs.py could not see: whether the commands in the
docs would actually run. A renamed or removed option is the classic form of command
rot, and the one that wastes a reader's time most directly -- they paste the line and
it fails. All 9 documented flags across screenshot.py, webui.py and restore_flash.sh
are real.
The check earned its own lesson, recorded in both languages. Its first version matched
only to the end of the line, so on a wrapped command it saw --frame-addr and nothing
after the backslash: 4 of 9 flags, and it reported a clean run. A check that silently
covers a quarter of what it claims is worse than no check, because the clean result is
believed. Continuations are joined before matching now.
Confirmed it fails when it should: renaming --frame-addr to something no tool accepts
produces two named failures and exit 1, and reverting returns it to clean.
check_docs.py now runs seven checks.
Translating everything into Chinese found four claims that had already drifted, and
none of them were caught by reading -- they were caught by comparing against source.
Proofreading does not find rot, so do the comparison mechanically and keep doing it.
tools/check_docs.py verifies that every tool a README names exists, that every test in
run_tests.sh is documented in both languages, that internal .md links resolve, that the
translation pairs have matching heading structure, that memory-map addresses match the
model's #defines, and that documented firmware file:line references still point at what
the prose claims. It runs in the quick tier of run_tests.sh, needing no emulator.
Confirmed it can actually fail, because a checker that cannot is worthless: renaming a
documented tool and deleting a heading from the Chinese side each produce one named
failure and exit 1, and reverting returns it to clean.
One thing it deliberately does not check. An early version compared firmware constants
with a regex that took the first number on a line, so `key_debounce_10ms = 20 / 10` read
as 20 and it declared the docs wrong for saying 2. The docs were right and the checker
was broken. A checker that cries wolf gets ignored, so claims it cannot verify
unambiguously are left out rather than guessed at.
Current state: 16 file:line references all accurate, 7 memory-map addresses all match,
zero broken links, all three translation pairs structurally aligned.
Full translations rather than summaries, section-for-section with the English:
README (13 sections), AGENTS.md (21), and docs/reverse-proxy.md. Each pair
cross-links to the other and says the two are kept in step, since documentation
that has silently diverged is worse than documentation that does not exist.
Verified rather than eyeballed: heading counts and order match in both pairs,
every internal .md link resolves, every tool named in either README exists, and
every test in run_tests.sh appears in both.
Translating turned up four things that were already stale in the English, which
is the honest argument for having done it this way -- a summary would not have
touched them:
- the endpoint table was missing /api/ptt, /api/power/<action> and
/api/logs, and did not mention the speaker field on /api/status
- the modelled-peripheral list omitted TIM2
- the audit table still called TIM a stub, unchanged since fdcbe80 modelled
TIM2
- neither README listed uvk5_logs.py, uvk5_stream.py, uvk5_supervisor.py or
test_kill_emulator.sh, which are part of the repo rather than scratch
The ad-hoc probe scripts are now acknowledged in one line instead of being
silently absent, and described as what they are: quick to reach for, not
polished.
Unit tests: 89 passed.
The S-meter had a number to draw, but a fixed RSSI above squelch meant the band was
uniformly and permanently occupied. Scanning, squelch, and every "is this channel busy"
decision therefore faced a situation that never varied, so none of that logic was
really being tested -- the tests passed without testing much.
RSSI is now derived from where the firmware tuned. BK4819_SetFrequency splits the
frequency across REG_38 and REG_39 (driver/bk4819.c:743), which the model already
records; verified against a live guest that 0x0262/0x5A00 reads back as 400.00000 MHz,
matching the screen. A small table of virtual stations plus a noise floor and a fade
either side of centre gives a band with signals in some places and not others.
Measured through the firmware's own tuning path -- typing 410.000 on the keypad rather
than poking the registers, so the test does not check the model against itself:
400.000 MHz (station) RSSI 0x01E5
410.000 MHz (empty) RSSI 0x0091 a gap of 85 dB
What is honest and what is not, recorded in the code: the shape is real physics, power
falls off away from a carrier with a noise floor underneath. The station list is
invented. So this reproduces "the firmware copes with a band that is busy in places",
which is genuine coverage, and it reproduces no actual radio environment -- a dBm figure
from here is not a claim about the world.
Also records why backlight PWM is deliberately left stubbed. Intermediate brightness
runs TIM7 -> DMA rewriting GPIOA BSRR at 128 kHz, so modelling it costs 128,000 GPIO
writes per emulated second and changes nothing observable: backlight is LED brightness
and never touches the framebuffer. The two endpoints that are observable, off and full,
bypass the timer and already work.
Full run: 16 passed, 0 failed.
Second finding from the audit. TIM2 was covered by the catch-all stub, which returns
the last value written, so
uint32_t millis(void) { return LL_TIM_GetCounter(TIM2); }
returned 0 forever. All 17 call sites that measure elapsed milliseconds could never
see time pass -- a silent wrong answer rather than a hang, which is harder to notice
and was not noticed.
The counter is derived on read from QEMU_CLOCK_VIRTUAL rather than stored, with
CR1.CEN starting and freezing it and a CNT write rebasing it. Guest time here is not
proportional to wall time anyway, and code measuring elapsed milliseconds wants
something advancing at roughly the rate a human sees; this is explicitly not for
anything needing cycle accuracy.
Measured: 24358 ms, then 29527 ms five seconds later -- 5169 ms elapsed, so the rate
is right rather than merely non-zero. The test checks the rate for that reason: a
counter ticking at the wrong speed would satisfy "non-zero" and "increasing" and still
break every timeout.
AGENTS.md now carries the audit itself: a table of what the firmware actually drives
against what is modelled versus stubbed, and the point that answering reads is not the
same as being reproduced. The honest summary is that the digital side the firmware
depends on is reproduced, and the analogue side is not and cannot be.
Full run: 15 passed, 0 failed.
Prompted by a fair criticism: the reports said what runs, not what is actually
reproduced. An audit found the ADC was modelled but returned a hardcoded 2200 forever,
so gBatteryDisplayLevel, gLowBattery and the warning popup were all unreachable. A
peripheral that answers reads is not the same as a peripheral that is reproduced.
adc-result is now settable over QOM and clamped to 12 bits. Measured: 2200 gives
level 4 and no warning, 1200 gives level 0 and raises gLowBattery, and the level
recovers to 4 afterwards.
tools/test_battery.py covers it, and deliberately does NOT assert that gLowBattery
clears on recovery. helper/battery.c:190-204 only clears it when the level lands
exactly on 2; above that it clears gLowBatteryConfirmed and leaves gLowBattery set. So
4 -> 0 -> 4 really does leave the flag raised. The first version of this test called
that a failure -- the test was wrong, not the model. The emulator reproduces the
firmware, including behaviour that looks like a bug.
Records the scope plainly, because "add a speaker and a microphone" is the obvious
request and the answer is that neither is on the MCU: no audio samples exist in its
address space, so there is nothing to capture, nothing to play, and nothing for a
browser permission to carry. Same line as the analogue RF limit.
Also records the stub lesson: QmpClient.command returns the unwrapped value, the test
stub returned an envelope, and the divergence let 88 tests pass while the live UI
returned 500. A stub more forgiving than the real client is worse than none.
Asked for a speaker and a microphone. The honest answer is that neither exists to
model: on the real radio neither passes through the MCU. Receive audio is demodulated
inside the BK4819 and leaves it as analogue on its AF pin; transmit audio goes from
the microphone into the chip's own ADC. The firmware's entire involvement is
- PA8, the amplifier enable (GPIO_EnableAudioPath, driver/gpio.h:34)
- REG_47, which AF source the chip routes
- REG_64, a level it displays
No audio samples exist anywhere in the MCU's address space, so a device model has
nothing to capture or play, and a browser has nothing to be granted permission for.
Synthesising sound would be inventing data the firmware never produced.
What is real is the firmware's intent, and PA8 states it exactly. TYPE_UVK5_AUDIO
watches that pin and exposes read-only speaker-on; the web UI shows it as a speaker
glyph beside the power state, and /api/status reports it. Read-only deliberately:
letting a test write it would only let the test lie to itself. The page asks for no
audio permission, and a test asserts it never will -- no getUserMedia, no
AudioContext, no <audio>.
Measured: amplifier off while idle in power save, on after SIDE1 engages monitor, and
still on afterwards rather than blipping.
One bug found the hard way, and the stub was the cause. QmpClient.command returns the
unwrapped value and raises on error, but the test stub returned {"return": ...}. So
webui.py was written to unwrap a second time, every test passed, and the live UI
returned 500 with "argument of type 'bool' is not iterable". The stub is now pinned to
the real contract by a test. A stub more forgiving than the real thing is worse than
no stub.
Also stopped swallowing the failure: a bare `except: return None` made the error
indistinguishable from a radio that simply was not making sound, and cost a detour
into looking for a stale process.
The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:
bash tools/run_tests.sh # everything, 11 suites, a few minutes
bash tools/run_tests.sh -q # unit tests only, ~15 s, no emulator
The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.
Two defects in the runner's own first draft, both caught before it was trusted:
It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.
That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.
Full run: 11 passed, 0 failed.
With RSSI varying every poll, the meter redraws constantly, so "consecutive frames
differ" is true on a completely parked radio. Records the framebuffer page layout so
a screen test can compare the rows that answer its actual question, and the fact that
page 4 is not the meter row.
bk4819_eval_receiver reports RSSI above any sane squelch threshold on every poll, and
a scan halts when it finds a busy channel -- so the S-meter work could plausibly have
stopped scanning dead on its first step. Measured: it does not, 6 distinct tunings
over 6 samples.
The check compares the framebuffer rows holding the frequency digits rather than
counting distinct frames. Frame comparison would pass on a completely parked radio,
because RSSI varies every poll and the meter redraws constantly -- an initial run
scored 8/8 distinct frames and proved nothing. Narrowing to pages 1-2 answers the
actual question: did the radio retune.
Incidentally established that page 4 is not the meter row: it was byte-identical
across all six samples while the frequency changed.
Three places claimed there was no PTT line, and one of them named the wrong pin
(GPIOC rather than PB10). All now say the same thing: PTT works, but not through
the key table, because the firmware reads its own pin instead of scanning it as a
matrix key.
Documents the transmit bar's two gates -- FUNCTION_TRANSMIT and gSetting_mic_bar,
the latter already on because blank flash reads 0xFF -- and why the release path
gets more attention than the press: a stuck PTT leaves every later test running
against a transmitting radio.
Also records the '.key' vs '.key[data-key]' trap for whoever adds the next
non-key button.
PTT was the one input the model never had, so the radio could not be keyed and the
mic level bar was unreachable. It is not a matrix key -- GPIO_IsPttPressed reads
PB10 directly (driver/gpio.h:31, active low) -- so it gets its own GPIO line rather
than a column/row intersection.
With it the transmit screen is complete: TX annunciator, a running timer, and a
level bar at roughly 80% of scale, fed from REG_64 via BK4819_GetVoiceAmplitudeOut.
app/app.c:1700 draws that only while gCurrentFunction == FUNCTION_TRANSMIT with
gSetting_mic_bar set; the setting is Data[7] bit 4 at flash 0xA0A8, and blank flash
reads 0xFF, so it is already on.
Exposed as a boolean on the keypad device and as POST /api/ptt with an explicit
held flag, plus a button in the browser UI. Held rather than tapped, because
transmitting is a state the operator stays in and a fixed duration would be wrong
for it.
Releasing is treated as the important half:
- pointerleave and pointercancel release, so dragging off the button cannot leave
the radio keyed
- pagehide releases, so closing the tab cannot either
- /api/release-all clears PTT too, since it is outside the matrix and an empty
press does not touch it
- non-boolean bodies are rejected, so {"held": "false"} cannot key the transmitter
by truthiness
Two things the test suite caught, both real:
The browser wired '.key' handlers over every styled button, and the PTT button
carries no data-key, so it would have sent the key "undefined". Narrowed to
'.key[data-key]'.
StubClient.presses() collected every qom-set regardless of property, so PTT's
booleans landed among the key names and presses()[-1] reported False after a
release-all. Now filtered by property, with a matching ptts() accessor.
tools/test_ptt.py covers the path end to end and asserts the release as well as the
press: a PTT that stuck would leave every later test running against a transmitting
radio.
The squelch section said the blocker sat above the device model and told the reader
not to resume without new information. Both are now wrong, so it leads with the
outcome instead: SIDE1 engages monitor, which skips squelch entirely, and the meter
reads -53 dBm / S9+40.
The four failed attempts are kept rather than deleted. Each ended in a confident
wrong diagnosis -- power-save gating, trigger timing, bit semantics -- and the shared
cause was a one-bit read skew underneath all of them. That pattern is worth more to
the next reader than a clean account of the version that worked.
README gains the S-meter row and test_smeter.py.
The firmware now draws a working meter: -53 dBm, +40 over S9, nine of thirteen
segments, next to a MONI label and a running receive timer. The two numbers agree
with each other -- S9 is -93 dBm on UHF, so -53 really is S9+40.
Three pieces had to line up, and the order they were found in was the hard part.
RSSI and audio amplitude are refreshed when the firmware polls REG_0C, not when it
configures the chip. Raising a flag at configuration time is a trap: REG_3F is
written 0 then 0x0C0C repeatedly during setup, so anything announced there is
disabled again before it can be collected.
The squelch flag is SQUELCH_LOST, bit 2 -- not SQUELCH_FOUND. Per app/app.c:1027
"squelch lost" is what sets g_SquelchLost = true, i.e. a signal is present.
SQUELCH_FOUND reads like "found a signal" and means the opposite.
Announcing is rate-limited to every 64th poll. Announcing once means the firmware
collects it during startup, before the flag leads anywhere. Announcing on every poll
re-arms the request bit inside the firmware's own collection loop, which uses REG_0C
as its condition and has no timeout, so it never exits. Periodic satisfies both.
What finally made the meter appear was not the interrupt at all. The radio idles in
power save and does not act on squelch there. ACTION_Monitor skips squelch entirely
-- app/app.c:482 picks FUNCTION_MONITOR over FUNCTION_RECEIVE when gMonitor is set --
and settings.c:263 defaults an out-of-range stored action to ACTION_OPT_MONITOR,
which blank flash (0xFF) is. So SIDE1 short-press is the way in. Measured: fn=5
idle=1 monitor=0 before, fn=2 idle=0 monitor=1 after.
Gating on RX_DSP (REG_30 bit 0, from App/driver/bk4819-regs.h:240) rather than the
whole register being zero: TX and tone paths leave other bits set with RX_DSP clear,
and would otherwise look like a live receiver.
tools/test_smeter.py covers the whole path -- boots pristine, confirms power save,
presses SIDE1, and checks the screen gained content. It compares lit-pixel counts
rather than matching pixels, so an unrelated UI change does not produce a mysterious
failure.
None of this is radio simulation. The levels are plausible numbers that move; they
are not the result of modelling a signal. What they buy is firmware control flow
running on live values instead of on zero.
The shifted-read bug invalidated four earlier diagnoses, so the notes claiming a
power-save gate or a timing problem were wrong and are corrected.
With reads fixed the interrupt handshake demonstrably works: RAISE pending=0004,
ACK delivering flags=0004, correct bit, collected by the firmware, REG_0C back to
0x0000 so nothing hangs.
g_SquelchLost is still 0 and the S-meter still absent, but the shape of the problem
is now clear and recorded: announcing once lands during startup before the flag
leads anywhere; announcing every poll re-arms the bit inside the firmware's own
untimed collection loop and spins forever; announcing periodically avoids both and
still changes nothing. The blocker is getting the radio into a receiving state at
all, which sits above the device model.
Also records the general lesson, which cost the most time here: when several
independent attempts fail in the same way, suspect the shared transport rather than
the logic layered on top of it.
Every read came back doubled: seed REG_0C with 0x1248 and the firmware received
0x2490. The command byte's own trailing falling edge was being treated as a data
clock, so bit 15 was shifted away before the guest sampled it and the whole word
landed one place too high.
Each firmware bit is read/raise/lower (BK4819_ReadU16), which means the eighth
command bit is followed by a falling edge before the data phase begins. Skip that
one edge.
Why it went unnoticed: writes were always fine -- 52 registers held exactly what
the firmware wrote -- and the register the firmware polls hardest, REG_0C, was
legitimately 0 in this model. Reading zero and getting zero looks like success.
The skew only surfaced when something tried to report a value through it.
It also explains four failed attempts at the squelch interrupt. The model raised
REG_0C bit 0; the firmware received bit 1. So
while (BK4819_ReadRegister(BK4819_REG_0C) & 1u)
was never true, the acknowledging write inside it never ran, and 1719 polls saw a
flag the guest could not act on. Every one of those attempts was diagnosed as a
timing or gating problem and was not.
tools/test_bk4819_readback.sh locks it down. It seeds REG_0C -- read ~1700 times
per 30s, so a sample is guaranteed -- with a value carrying bits in both halves,
so a shift either way is unmistakable, and names the direction on failure. Bit 0
is left clear on purpose: with it set the firmware enters an acknowledge loop that
has no timeout, and this test is about alignment only.
Confirmed by A/B: committed code 0x2490, patched 0x1248.
The previous commit blamed the power-save gate at app/app.c:1697 for the interrupt
never being collected. That was wrong, and the test which disproves it is cheap:
BATTERY_SAVE lives at flash 0xA00B, app/app.c:1374 refuses power save when it is 0,
and patching the byte gives
BATTERY_SAVE=4: fn=5 idle=1 polls=2161 acks=0
BATTERY_SAVE=0: fn=0 idle=0 polls=2161 acks=0
The gate passes and nothing changes. A gdb backtrace confirms the loop runs --
CheckRadioInterrupts is inlined into APP_TimeSlice10ms, which is the caller of every
REG_0C read.
Where it really stands: a breakpoint on BK4819_GetRSSI never fires at all. The
firmware does not read RSSI in this idle state, so the missing S-meter is not the
model withholding a value; the receive state machine has to be entered first, and the
squelch interrupt is an input to that rather than the switch. Three rounds of
increasingly precise instrumentation all ended at "the firmware is not asking".
Four measurement mistakes are recorded because each produced a confident wrong
conclusion, and one of them made me change working code:
- Sampling PC at the REG_0C read lands in BK4819_WriteU8; sampling LR lands inside
BK4819_ReadRegister, since it calls BK4819_ReadU16. Use a backtrace.
- A probe printing shift_out before the assignment showed 0000 for a value about to
be sent as 0001.
- BK4819_ReadRegister returning 0x0 for REG_0C looked like a broken read path, and I
altered the bit timing over it. REG_0C legitimately holds 0 in the committed build.
A read returning the register's real contents proves nothing -- test against one the
firmware wrote, like REG_3F (0x0C0C) or REG_78 (0x2F5B).
- nexti after a breakpoint reported r0=0 from an unrelated location. finish gives the
actual return value.
Also noted: gdb cannot call guest functions on this target, and there is no
gCurrentRSSI global -- RSSI is read and discarded, so a breakpoint plus finish is the
only way to see what the firmware received.
Code is unchanged and at the committed state; test_bk4819.py passes.
Second attempt, and this one produced a definite answer rather than another guess.
Two real mistakes in the first version, both fixed along the way and worth recording:
squelch was evaluated when the firmware configured the chip, but the startup sequence
writes REG_3F as 0x0000 then 0x0C0C three times over, so a flag raised on the enabling
write was disabled again before anyone read it. And the threshold came from REG_4E,
whose low bits are the glitch threshold -- the RSSI open level is REG_78 bits 15:8 at
0.5 dB/step against REG_67's 0.25 dB/step, so squelch could never open at all.
With both corrected, every chip-side condition lines up: en=0x0C0C, rssi=0x01E0,
threshold 94, REG_0C correctly returning 1. The firmware still never acknowledged,
and the reason is outside the chip:
gCurrentFunction=5 (FUNCTION_POWER_SAVE), gRxIdleMode=1
against the gate at app/app.c:1697,
if (gCurrentFunction != FUNCTION_POWER_SAVE || !gRxIdleMode)
CheckRadioInterrupts();
Both halves false, so the interrupt loop is never entered and nothing can collect the
flag. The emulator idles in power save, so that is the steady state.
Gating on REG_30 -- zeroed by BK4819_Sleep on each power-save cycle -- does not help:
the chip is awake when the model is asked while the firmware still has gRxIdleMode=1.
Chip state and firmware state are not in step, so no condition available inside a
register model can decide this correctly.
So it is not a matter of a better trigger. A working version needs to keep the guest
out of power save, or drive the interrupt from something that knows the firmware's
receive state -- neither of which belongs in this device. Both attempts reverted; the
model is back at the committed state and test_bk4819.py passes, including the REG_0C
assertion that catches the armed-hang failure mode.