Measured by polling the PC over QMP every 30 ms, finer than the model's own 100 ms probe. After launching Tetris the overlay at 0x20000280 holds Tetris's code byte for byte, so the loader's copy is correct and aligned; the earlier claim in this session that it was shifted by one byte was my own parsing dropping the first value. But 172 PC samples over 4.5 s and 64 more over 2 s were all inside firmware flash at 0x08013260, a wait loop, and none landed in the overlay. A 16-byte app that calls nothing behaves identically, and Breakout's header is field-for-field the shape of ours, so neither the app's code nor its header is the reason. The failure is upstream of the app, which makes the earlier 'Tetris runs' note stale: per this file's own rule it is treated as unverified until it reproduces.
Also carries the CI-fix branch merge and a regression test for _edit_flash creating its working-copy directory (its fixture image has to be a real size: slot 0 sits at 0x102000).
pip install ziglang cross-compiles to thumb-freestanding-eabi, which is enough to build a .app with no arm-none-eabi-gcc and no Docker. Measured refusals: --defsym, -Ttext and --section-start come back as unsupported linker args, so the VMA is resolved into a copy of app.ld; -T is forwarded (a missing script errors); --image-base is accepted but page-aligns the segments into 0x200103C8 and 0x20020C94. tools/elf2bin.py extracts allocated sections rather than program headers, because lld maps the ELF header and phdr table as a 180-byte LOAD of its own -- following the headers starts the image at 0x20000000 and APP_ERR_VMA. One division pulled in __aeabi_uidiv, which a -nostdlib blob cannot have: Minesweeper now avoids division entirely. app_main carries the .text.entry attribute upstream's apps use, so the entry is first for the loader's jump to offset 0.
Minesweeper builds to 2408 bytes of code against a 4096-byte budget. Installed through the page, the firmware reads the slot header twice and then exactly code_size bytes from slot+0x1000 -- the only read of that size in the boot log -- so the blob shape, header, CRC, VMA and offset are all accepted. Whether control reaches the overlay is unproven: the 100 ms PC probe saw no overlay address, no APP ERROR screen appears, and the app does not draw. test_elf2bin pins the phantom-header-segment lesson; both AGENTS files record the rest.
The blank screen was self-inflicted and the earlier explanation was wrong. The FMP3 marker at 0x100000 is an ordinary state record -- generation, image_size, image_crc32, firmware_slot/slot_inv, config_bank/bank_inv and a matching state_crc32 (0x661286F1) -- not a pending flag. What actually happened: a firmware uploaded over a state that recorded a different identity made the firmware take the restore/adopt path, which draws nothing (panel 1024/1024 bytes zero) while the serial banner printed normally. Restoring the untouched dump fixed it at once (485/1024 bytes lit) and the three apps reinstalled and were confirmed by 0x0730.
Clearing the marker sectors was tried twice (a whole 8 KiB, then just the 24-byte headers) and is not a fix: it sends the firmware down MB_MARK_MISSING = fresh radio, which adopts the running firmware slowly and without drawing, and its own write-back restores the marker anyway. AGENTS.md and AGENTS.zh-CN.md now say so in place of the wrong claim.
apps/minesweeper/ adds our own 9x9 minesweeper for the 4 KiB overlay: no left/right keys exist on this radio (so the cursor walks with UP/DOWN and digits pick a row then a column), 81 cells need three 9-byte bit arrays rather than a uint16_t mask (the uint16_t version compiled fine and was wrong past cell 15), mines are placed after the first reveal so it cannot lose immediately, and the source compiles clean under gcc -Wall -Wextra -Werror against upstream's real app_api.h. It has not been built for ARM or run -- no toolchain here -- and the README says so.
A real image whose 0x100000 marker held FMP3 next to a committed slot 0 made the factory bootloader reflash the internal flash from that slot on every power-on: the serial banner reappeared once a cycle (7 -> 8 in 25 s) while the screen never changed. Clearing the two marker sectors (0x100000..0x101FFF, stopping just before the app region at 0x102000) ended it, and the radio booted once and stayed.
It also explained why page-side installs vanished: the emulator writes its in-memory image back on exit, and the looping guest's copy was older than the file, so powering it off overwrote the installs. With the loop gone the same installs survive a power cycle -- installed, powered off, still listed, powered on, still listed, and 0x0730 answered with all three. Both notes are in AGENTS.md and AGENTS.zh-CN.md; work/app-template/ got a starter app, its README and the upstream api/ld it needs (git-ignored).
F then 7 opens the F4HWN APPS menu and MENU runs the selection -- upstream's own wording in UVStudio's locales/en.js, and what App/apps/app_menu.c does: KEY_MENU on an installed row calls APP_LaunchOverlay, which runs until the app exits; EXIT leaves the menu. The page's hint now says it too.
Recorded with it: twenty keys short and long, the whole 79-entry settings menu and the multiboot menu all left the app region untouched (the multiboot menu reads firmware slots, not apps), and the way in turned out to be written in the flasher's translation file rather than the firmware source. When a feature's entry point is missing, the host tool that installs it is the document.
GET /api/apps/radio opens the firmware's serial port and sends 0x0730 for all sixteen slots, so the answer comes from the running firmware rather than from our reading of the file -- which is the check that matters, because the bytes can be right and the firmware still refuse a slot. Measured through the page after installing Beam.app into slot 0: slot 0 -> Beam 1.0, 1100 B, crc 0xd976058, shortcut beam, committed; slots 1..3 -> status 2 with unrelated data, the resource-block overlap the install guard refuses. A button beside the table asks it and shows the answer in its own column.
The server gives its own emulator a serial port (--serial-port, default 4445) and uvk5_slots_serial.Radio gained a public app_info(slot), so nothing reaches into a private helper. uvk5_apps.parse_radio_reply decodes the answer and is tested without a radio. Also recorded: QEMU needs the mingw64 DLLs on PATH, and started by hand without them it exits before opening QMP, which surfaces only as 'QMP socket never appeared'.
UVStudio's own js/flash.js names the protocol: MSG_APP_INFO 0x0730/0x0731, MSG_APP_ERASE 0x0732/0x0733, MSG_APP_WRITE 0x0734/0x0735, MSG_APP_VALIDATE 0x0736/0x0737, with APP_SLOT_COUNT 16, APP_IMG_OFFSET 0x1000, APP_HDR_SIZE 64 and APP_MAGIC 0x31504146 -- the constants this page already used. It uses slots 0..7 and labels them 1..8, and writes the header last so a partial write cannot validate; the page now labels slots the same way.
The Labs build answers 0x0730 over USART, so an installed Beam.app was queried on the radio: slot 0 came back status 0 with the header this page wrote (FAP1, code_size 1100, CRC 0x0d976058, flags 0x0801, name Beam), which confirms the region, the offset and the layout from the firmware's side rather than from a header I read. Slots 1 and 2 answered status 2 with unrelated data -- the overlap the install guard refuses to overwrite.
The endpoints existed; now the page shows them. An Overlay apps block beside the firmware slots lists all 16 app slots with their name, version, shortcut and size, a .app file picker and an Erase button per row, and says where the region is. It reads GET /api/apps and posts to POST /api/apps/<n> and /api/apps/<n>/erase, so installing a game is a file pick where the firmware slots already are -- no WebSerial, no browser permission.
An upload into a slot that holds something which is not an app is refused by the server; the page asks once and retries with ?force=1 rather than either failing quietly or destroying the factory resource data that overlaps this region on a localised image. Front-end tests assert the table and the endpoint are in the served page and that it never asks for a serial port or audio, and TestPageScriptParses keeps checking that the page's script parses.
The Labs edition's apps live in the external flash, in the region its own App/apps/app_overlay.h defines (16 slots of 8 KiB from 0x102000, a 64-byte FAP1 header at the slot base, code one 4 KiB sector later), and upstream installs them from UVStudio over WebSerial. The page owns the image, so this adds GET /api/apps, POST /api/apps/<n> and POST /api/apps/<n>/erase, which put the same bytes at the same offsets with no serial protocol and no browser permission.
Measured on a real image while wiring it up: every one of the 16 slots already held data that is neither empty nor an app -- the localised build's factory resource block overlaps 0x102000 -- so install now refuses to overwrite anything that is not an app unless asked (--force, ?force=1), naming what is there. And _edit_flash edits a copy, which is why the source image shows no changed bytes; the first test read that as 'the install did nothing' and now checks FlashSlot.path. Tests: test_uvk5_apps grew to 20, test_webui.TestAppEndpoints adds 7 over the endpoints.
The Labs edition runs overlay apps (Tetris, Breakout, Plasma, Cube3D, Beam, Beacon, FoxHunt, BroadcastFM) that upstream UVStudio installs over WebSerial. This page owns the flash image, so the same bytes go to the same offsets with no serial protocol and no browser permission: APP_REGION_BASE 0x102000, APP_SLOT_STRIDE 0x2000, APP_CODE_OFFSET 0x1000, 16 slots, taken from the firmware's own App/apps/app_overlay.h rather than inferred.
tools/uvk5_apps.py parses and validates the 64-byte FAP1 header (zlib CRC-32 over the code, vma 0x20000280, name, version, capabilities), lists, installs and erases slots, and refuses what the firmware would show as APP ERROR. test_uvk5_apps covers those refusals plus install/erase/list round trips, and parses a real upstream Beam.app when one has been downloaded. The header struct was 60 bytes at first -- a missing vma field -- which the real file's bytes showed at once.
The page only knew its own input. A build called f4hwn.fusion.bin reports EGZUMER+F4HWN v6.0.0.CN, and with the multi-system release a committed external slot makes the factory bootloader reflash the internal flash from that slot on every power-on -- so the uploaded image never runs and the page keeps naming it. The firmware prints its own banner on USART1; tools/uvk5_banner.py reads it back, /api/firmware returns running: {banner, matches_uploaded, note}, and the page shows what the device reports, flagging it only when the running version is not in the uploaded image at all.
That reader also exposed a regression of my own: _start_stderr_pump had been rewritten to read the pipe in 64 KB chunks, which kept QEMU from blocking but delivered nothing to the log until 64 KB had accumulated -- and the banner is forty bytes, so it never appeared. It reads lines again, still starting before anything waits on QEMU, and test_uvk5_supervisor passes either way.
The screenshot example pasted one build's numbers and told the reader to get them with nm -- which cannot work here, because the images are program-header-only ELFs with no symbol table. It now asks the firmware through tools/uvk5_buffers.py first. The webui example drops the flags entirely, since the page draws the controller's memory and needs none.
Also fixed the paragraph in both READMEs that had a tool name and a flag on one line, which the doc checker (correctly) read as passing that flag to that tool.
The page was told --frame-addr 0x200012BE --status-addr 0x2000163E and used them as a fallback. The firmware the user actually flashed keeps its buffers at 0x2000129E/0x2000161E, 32 bytes earlier, so every line landed 32 bytes off: that is the "other firmware looks shifted" report. The images here are minimal ELFs with no symbol table, so there is nothing to read -- but the firmware's own buffers hold the same bytes the controller holds, and tools/uvk5_buffers.py finds them by matching (1024/1024 bytes for that file).
The two address flags are optional now, work/run-webui.ps1 passes no machine-specific values at all, and the page reports what it found in /api/status and /api/panel. tools/uvk5_testenv.qemu() also looks in the sibling qemu-7.2/build the rest of the repo assumes. Fixed /api/panel's emulator-off branch, which called jsonify with both a dict and kwargs and 500'd.
Tests: test_uvk5_buffers (the search must count matches, not pairs -- its first version scored every offset full marks and always answered the first one).
uvk5_stream.py used STATUS_BYTES without importing it, so the panel branch raised NameError on every frame and a bare except swallowed it: every screen the page drew came from guest RAM at one build's addresses. The pump now reports which source it used and why, webui exposes it (/api/panel and frame_source), and test_uvk5_stream asserts the panel wins when reachable and that a fallback is announced.
The ST7565 column counter wrapped at 128 instead of the controller's 132, so addresses 128..131 came back as 0..3, fell outside the col>=4 store, and were dropped: every row lost its last four pixels, which is where the battery icon lives. Pre-fix, filling a page with 0xFF left columns 124..127 blank; now they carry content and the page's frame matches the panel memory 8192/8192.
The bootloader's DFU handler is reachable only when SRAM[0x20000020] is 3, which only a program that then resets can write. The application's 0x05DD takes that path only with ENABLE_OVERLAY; this build resets straight back into the application instead. Ruled out by measurement: PTT alone, PTT+SIDE1/SIDE2, MENU, a host byte in the boot window including 0x0530, and 0x05DD.
boot-key now reads a + separated list, because the firmware's own BOOT_GetMode() needs PTT and a matrix key together for every special boot mode. Tested against all four documented combinations.
flash controller: store ACR/OPTKEYR instead of swallowing them, which is what stopped the factory bootloader from starting
slots over the firmware's own serial protocol (0x0720 family); uvk5_socket/uvk5_testenv so a fresh checkout skips instead of failing; web UI slot table and Multiboot button; quick start, CONTRIBUTING, and stop tracking firmware images and radio dumps
The remaining class of claim check_docs.py could not see: whether the commands in the
docs would actually run. A renamed or removed option is the classic form of command
rot, and the one that wastes a reader's time most directly -- they paste the line and
it fails. All 9 documented flags across screenshot.py, webui.py and restore_flash.sh
are real.
The check earned its own lesson, recorded in both languages. Its first version matched
only to the end of the line, so on a wrapped command it saw --frame-addr and nothing
after the backslash: 4 of 9 flags, and it reported a clean run. A check that silently
covers a quarter of what it claims is worse than no check, because the clean result is
believed. Continuations are joined before matching now.
Confirmed it fails when it should: renaming --frame-addr to something no tool accepts
produces two named failures and exit 1, and reverting returns it to clean.
check_docs.py now runs seven checks.
Translating everything into Chinese found four claims that had already drifted, and
none of them were caught by reading -- they were caught by comparing against source.
Proofreading does not find rot, so do the comparison mechanically and keep doing it.
tools/check_docs.py verifies that every tool a README names exists, that every test in
run_tests.sh is documented in both languages, that internal .md links resolve, that the
translation pairs have matching heading structure, that memory-map addresses match the
model's #defines, and that documented firmware file:line references still point at what
the prose claims. It runs in the quick tier of run_tests.sh, needing no emulator.
Confirmed it can actually fail, because a checker that cannot is worthless: renaming a
documented tool and deleting a heading from the Chinese side each produce one named
failure and exit 1, and reverting returns it to clean.
One thing it deliberately does not check. An early version compared firmware constants
with a regex that took the first number on a line, so `key_debounce_10ms = 20 / 10` read
as 20 and it declared the docs wrong for saying 2. The docs were right and the checker
was broken. A checker that cries wolf gets ignored, so claims it cannot verify
unambiguously are left out rather than guessed at.
Current state: 16 file:line references all accurate, 7 memory-map addresses all match,
zero broken links, all three translation pairs structurally aligned.
Full translations rather than summaries, section-for-section with the English:
README (13 sections), AGENTS.md (21), and docs/reverse-proxy.md. Each pair
cross-links to the other and says the two are kept in step, since documentation
that has silently diverged is worse than documentation that does not exist.
Verified rather than eyeballed: heading counts and order match in both pairs,
every internal .md link resolves, every tool named in either README exists, and
every test in run_tests.sh appears in both.
Translating turned up four things that were already stale in the English, which
is the honest argument for having done it this way -- a summary would not have
touched them:
- the endpoint table was missing /api/ptt, /api/power/<action> and
/api/logs, and did not mention the speaker field on /api/status
- the modelled-peripheral list omitted TIM2
- the audit table still called TIM a stub, unchanged since fdcbe80 modelled
TIM2
- neither README listed uvk5_logs.py, uvk5_stream.py, uvk5_supervisor.py or
test_kill_emulator.sh, which are part of the repo rather than scratch
The ad-hoc probe scripts are now acknowledged in one line instead of being
silently absent, and described as what they are: quick to reach for, not
polished.
Unit tests: 89 passed.
The S-meter had a number to draw, but a fixed RSSI above squelch meant the band was
uniformly and permanently occupied. Scanning, squelch, and every "is this channel busy"
decision therefore faced a situation that never varied, so none of that logic was
really being tested -- the tests passed without testing much.
RSSI is now derived from where the firmware tuned. BK4819_SetFrequency splits the
frequency across REG_38 and REG_39 (driver/bk4819.c:743), which the model already
records; verified against a live guest that 0x0262/0x5A00 reads back as 400.00000 MHz,
matching the screen. A small table of virtual stations plus a noise floor and a fade
either side of centre gives a band with signals in some places and not others.
Measured through the firmware's own tuning path -- typing 410.000 on the keypad rather
than poking the registers, so the test does not check the model against itself:
400.000 MHz (station) RSSI 0x01E5
410.000 MHz (empty) RSSI 0x0091 a gap of 85 dB
What is honest and what is not, recorded in the code: the shape is real physics, power
falls off away from a carrier with a noise floor underneath. The station list is
invented. So this reproduces "the firmware copes with a band that is busy in places",
which is genuine coverage, and it reproduces no actual radio environment -- a dBm figure
from here is not a claim about the world.
Also records why backlight PWM is deliberately left stubbed. Intermediate brightness
runs TIM7 -> DMA rewriting GPIOA BSRR at 128 kHz, so modelling it costs 128,000 GPIO
writes per emulated second and changes nothing observable: backlight is LED brightness
and never touches the framebuffer. The two endpoints that are observable, off and full,
bypass the timer and already work.
Full run: 16 passed, 0 failed.
Second finding from the audit. TIM2 was covered by the catch-all stub, which returns
the last value written, so
uint32_t millis(void) { return LL_TIM_GetCounter(TIM2); }
returned 0 forever. All 17 call sites that measure elapsed milliseconds could never
see time pass -- a silent wrong answer rather than a hang, which is harder to notice
and was not noticed.
The counter is derived on read from QEMU_CLOCK_VIRTUAL rather than stored, with
CR1.CEN starting and freezing it and a CNT write rebasing it. Guest time here is not
proportional to wall time anyway, and code measuring elapsed milliseconds wants
something advancing at roughly the rate a human sees; this is explicitly not for
anything needing cycle accuracy.
Measured: 24358 ms, then 29527 ms five seconds later -- 5169 ms elapsed, so the rate
is right rather than merely non-zero. The test checks the rate for that reason: a
counter ticking at the wrong speed would satisfy "non-zero" and "increasing" and still
break every timeout.
AGENTS.md now carries the audit itself: a table of what the firmware actually drives
against what is modelled versus stubbed, and the point that answering reads is not the
same as being reproduced. The honest summary is that the digital side the firmware
depends on is reproduced, and the analogue side is not and cannot be.
Full run: 15 passed, 0 failed.
Records the scope plainly, because "add a speaker and a microphone" is the obvious
request and the answer is that neither is on the MCU: no audio samples exist in its
address space, so there is nothing to capture, nothing to play, and nothing for a
browser permission to carry. Same line as the analogue RF limit.
Also records the stub lesson: QmpClient.command returns the unwrapped value, the test
stub returned an envelope, and the divergence let 88 tests pass while the live UI
returned 500. A stub more forgiving than the real client is worse than none.
The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:
bash tools/run_tests.sh # everything, 11 suites, a few minutes
bash tools/run_tests.sh -q # unit tests only, ~15 s, no emulator
The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.
Two defects in the runner's own first draft, both caught before it was trusted:
It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.
That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.
Full run: 11 passed, 0 failed.
With RSSI varying every poll, the meter redraws constantly, so "consecutive frames
differ" is true on a completely parked radio. Records the framebuffer page layout so
a screen test can compare the rows that answer its actual question, and the fact that
page 4 is not the meter row.
Three places claimed there was no PTT line, and one of them named the wrong pin
(GPIOC rather than PB10). All now say the same thing: PTT works, but not through
the key table, because the firmware reads its own pin instead of scanning it as a
matrix key.
Documents the transmit bar's two gates -- FUNCTION_TRANSMIT and gSetting_mic_bar,
the latter already on because blank flash reads 0xFF -- and why the release path
gets more attention than the press: a stuck PTT leaves every later test running
against a transmitting radio.
Also records the '.key' vs '.key[data-key]' trap for whoever adds the next
non-key button.
The squelch section said the blocker sat above the device model and told the reader
not to resume without new information. Both are now wrong, so it leads with the
outcome instead: SIDE1 engages monitor, which skips squelch entirely, and the meter
reads -53 dBm / S9+40.
The four failed attempts are kept rather than deleted. Each ended in a confident
wrong diagnosis -- power-save gating, trigger timing, bit semantics -- and the shared
cause was a one-bit read skew underneath all of them. That pattern is worth more to
the next reader than a clean account of the version that worked.
README gains the S-meter row and test_smeter.py.
The shifted-read bug invalidated four earlier diagnoses, so the notes claiming a
power-save gate or a timing problem were wrong and are corrected.
With reads fixed the interrupt handshake demonstrably works: RAISE pending=0004,
ACK delivering flags=0004, correct bit, collected by the firmware, REG_0C back to
0x0000 so nothing hangs.
g_SquelchLost is still 0 and the S-meter still absent, but the shape of the problem
is now clear and recorded: announcing once lands during startup before the flag
leads anywhere; announcing every poll re-arms the bit inside the firmware's own
untimed collection loop and spins forever; announcing periodically avoids both and
still changes nothing. The blocker is getting the radio into a receiving state at
all, which sits above the device model.
Also records the general lesson, which cost the most time here: when several
independent attempts fail in the same way, suspect the shared transport rather than
the logic layered on top of it.
The previous commit blamed the power-save gate at app/app.c:1697 for the interrupt
never being collected. That was wrong, and the test which disproves it is cheap:
BATTERY_SAVE lives at flash 0xA00B, app/app.c:1374 refuses power save when it is 0,
and patching the byte gives
BATTERY_SAVE=4: fn=5 idle=1 polls=2161 acks=0
BATTERY_SAVE=0: fn=0 idle=0 polls=2161 acks=0
The gate passes and nothing changes. A gdb backtrace confirms the loop runs --
CheckRadioInterrupts is inlined into APP_TimeSlice10ms, which is the caller of every
REG_0C read.
Where it really stands: a breakpoint on BK4819_GetRSSI never fires at all. The
firmware does not read RSSI in this idle state, so the missing S-meter is not the
model withholding a value; the receive state machine has to be entered first, and the
squelch interrupt is an input to that rather than the switch. Three rounds of
increasingly precise instrumentation all ended at "the firmware is not asking".
Four measurement mistakes are recorded because each produced a confident wrong
conclusion, and one of them made me change working code:
- Sampling PC at the REG_0C read lands in BK4819_WriteU8; sampling LR lands inside
BK4819_ReadRegister, since it calls BK4819_ReadU16. Use a backtrace.
- A probe printing shift_out before the assignment showed 0000 for a value about to
be sent as 0001.
- BK4819_ReadRegister returning 0x0 for REG_0C looked like a broken read path, and I
altered the bit timing over it. REG_0C legitimately holds 0 in the committed build.
A read returning the register's real contents proves nothing -- test against one the
firmware wrote, like REG_3F (0x0C0C) or REG_78 (0x2F5B).
- nexti after a breakpoint reported r0=0 from an unrelated location. finish gives the
actual return value.
Also noted: gdb cannot call guest functions on this target, and there is no
gCurrentRSSI global -- RSSI is read and discarded, so a breakpoint plus finish is the
only way to see what the firmware received.
Code is unchanged and at the committed state; test_bk4819.py passes.
Second attempt, and this one produced a definite answer rather than another guess.
Two real mistakes in the first version, both fixed along the way and worth recording:
squelch was evaluated when the firmware configured the chip, but the startup sequence
writes REG_3F as 0x0000 then 0x0C0C three times over, so a flag raised on the enabling
write was disabled again before anyone read it. And the threshold came from REG_4E,
whose low bits are the glitch threshold -- the RSSI open level is REG_78 bits 15:8 at
0.5 dB/step against REG_67's 0.25 dB/step, so squelch could never open at all.
With both corrected, every chip-side condition lines up: en=0x0C0C, rssi=0x01E0,
threshold 94, REG_0C correctly returning 1. The firmware still never acknowledged,
and the reason is outside the chip:
gCurrentFunction=5 (FUNCTION_POWER_SAVE), gRxIdleMode=1
against the gate at app/app.c:1697,
if (gCurrentFunction != FUNCTION_POWER_SAVE || !gRxIdleMode)
CheckRadioInterrupts();
Both halves false, so the interrupt loop is never entered and nothing can collect the
flag. The emulator idles in power save, so that is the steady state.
Gating on REG_30 -- zeroed by BK4819_Sleep on each power-save cycle -- does not help:
the chip is awake when the model is asked while the firmware still has gRxIdleMode=1.
Chip state and firmware state are not in step, so no condition available inside a
register model can decide this correctly.
So it is not a matter of a better trigger. A working version needs to keep the guest
out of power save, or drive the interrupt from something that knows the firmware's
receive state -- neither of which belongs in this device. Both attempts reverted; the
model is back at the committed state and test_bk4819.py passes, including the REG_0C
assertion that catches the armed-hang failure mode.
Scanning does work now that RSSI reports a real level -- long-press * and the
frequency steps, 6 distinct frames over 7 seconds. The S-meter still does not appear,
because ui/main.c:2370 draws it only when FUNCTION_IsRx(), which needs the chip to
report a squelch opening rather than merely a healthy RSSI.
I implemented that interrupt and backed it out. The guest kept running, but REG_0C
bit 0 remained set afterwards, meaning the firmware never collected the interrupt.
That leaves a hang armed: app/app.c:910 and :1417 spin on that bit with no timeout, so
any path reaching them with it stuck never returns. A missing S-meter is a cosmetic
gap; a latent hang is not, and shipping the second to fix the first is a bad trade.
Documented with the mechanism (REG_0C pending bit, REG_02 acknowledge-then-read,
sqlFound at bit 3), the reason it failed, and the check that says a retry is correct:
REG_0C reading 0 afterwards, which tools/test_bk4819.py already asserts. The next
person should find out why the flag was not collected rather than raise it at a
different moment and hope.
The "what this cannot do" section said the transceiver was not modelled and treated
that as permanent. Half of it is now wrong: the register interface works. The other
half is still true and worth keeping sharp -- analogue behaviour is out of reach
because the chip has no public datasheet, so the driver is the only specification and
it can only say which registers were written.
Rewritten to separate the two: what is modelled and what it fixed (RSSI was hard zero
at 18 call sites), the two constraints the untimed spin loops impose, and then the
line that does not move. Explicitly warns against reading the new test as evidence
about RF.
Also corrects the GPIO comment about idling PB9 low. It described the pin as a
workaround pending a device model; that model now exists, and the idle level only
covers the window between reset and the bus being wired.
The previous commit documented serial receive as a permanent limitation, listing
what would be needed to build it. It has since been built, so that section was
actively misleading -- it told the reader not to try something that already works.
Replaced with what it takes to keep working, since each of the three pieces fails
silently on its own: USART1 needs a chardev, DMA must decrement CNDTR for USART
(serviced on the CNDTR read, which is where the driver looks), and DR writes must
reach the chardev and not just stderr. The last one is the trap -- transmit that
goes only to stderr is indistinguishable from the firmware ignoring the command.
README status table and build-verification steps updated to match.
Found while explaining a DMA channel that stayed permanently armed with cndtr=256
during the flash investigation. It is USART1's receive channel, running in
LL_DMA_MODE_CIRCULAR, so never completing is correct behaviour and not a bug.
But it exposed a real gap. driver/uart.c derives its write pointer from
sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(...), and the DMA model only
services SPI, so CNDTR never decrements for USART and that expression is always
zero. Combined with USART1 being a py32-stub with no chardev backend, nothing can
be sent *to* the firmware.
The cost is specific: UART_IsCommandAvailable never fires, so the UV-K5 programming
protocol in app/uart.c is unreachable -- 0x0514 handshake, 0x051B EEPROM read,
0x051D EEPROM write, 0x05DD reset. CPS/CHIRP-style tools cannot talk to this
emulator. Transmit is unaffected, which is why the firmware banner shows up fine
and this went unnoticed.
Documented rather than fixed: it needs a chardev on USART1 plus circular-mode DMA
driven by receive, which is a new feature rather than a repair. The notes say what
would be involved so the next person does not have to rediscover the mechanism.
Status table also updated to reflect what the flash and DMA fixes settled --
persistence and frequency entry now work.
Four faults in the SPI/DMA/flash models each produced the same symptom -- stored
frequencies zeroed, a typed frequency reverting to 18 MHz -- and each was invisible
from the layer above. AGENTS.md now lists all four with the reasoning that connects
them to what the user saw, so the next person does not rediscover them one at a
time.
Also records the methodology mistakes, because they cost more than the bugs did:
- Attaching gdb between digits clears the frequency input box: the timeout is
~2.5 s and an attach takes ~3 s. Three wrong conclusions came from this, so
instrument the model and read stderr instead of stopping the guest.
- A diagnostic log capped at six entries showed only 0xFF payloads and supported
precisely the wrong conclusion. Do not cap before the shape of the data is known.
- A failed ninja leaves the old binary in place and the test still runs. Two rounds
of results were meaningless. Check for FAILED and error: before trusting a run.
- assets/flash.img is written by every session, so a test starting from it can find
its work already done -- which looks identical to broken persistence. Start from
assets/pristine/, and power off before restoring, since shutdown flushes the old
image back over the file.
- Hand-computed struct offsets gave KEY_LOCK=4 and TX_VFO=11, impossible values,
because the ELF has no DWARF and the structs contain enums. Use an unambiguous
nm symbol, or find the field by toggling it and diffing.
- One probe printed phase before incrementing it, making a correct address decoder
look off by one. A working implementation was nearly "fixed" as a result.
README gains the two new tests in the build-verification step and the tools list.
There is no bootloader, kernel, partition table or filesystem here, and that is not
obvious from the code -- someone arriving with hosted-OS assumptions will look for
layers that do not exist.
Covers the parts that are easy to get wrong rather than restating the source:
- The reset handoff is two words. SP from 0x08002800, PC from +0x04. Verified
against the image: 00400020 492d0008 is SP 0x20004000 (top of the 16 KB SRAM,
matching PY32_SRAM_BASE + PY32_SRAM_SIZE) and PC 0x08002d49, which is also the
ELF entry and the Reset_Handler symbol. The odd address is Thumb-state bit 0.
- PY32_APP_OFFSET 0x2800 is load-bearing. The first 10 KB of flash is the factory
bootloader region, so armv7m_load_kernel gets that offset; without it the vector
table lands in the wrong place and the first fetch faults.
- Startup is 31 lines of assembly, and the .data copy plus .bss zero-fill are the
part worth understanding: on a hosted OS the kernel and loader do that, here
nobody does, so a fault in either loop shows up as globals that are silently
garbage rather than as a crash.
- SETTINGS_InitEEPROM reading fixed SPI offsets is the closest thing to mounting a
partition. No metadata and no checksum, just an address both sides must agree on,
so a setting that reads back wrong points at the offset before the transport.
Also notes that the ~15 s to the main loop is emulation overhead; a real radio is
up in about a second.
README gets a section with the endpoint table, the two constraints that will
otherwise surprise someone (single QMP client, no authentication), and the reason
frames go through memsave rather than pmemsave or gdb.
AGENTS.md gets the run instructions plus a new entry under 'Things that already
went wrong' for the pmemsave trap: it takes a physical address, returns zeros for
gFrameBuffer, and reports success. The web UI was built on it initially because a
benchmark showed it was fast -- the benchmark never checked the contents. Worth
recording as the general lesson, not just the specific fix.
AGENTS.md still carried a stale entry telling the reader to *lengthen* key
holds when a press seems ignored, which is the opposite of the fix and is
what broke the tooling in the first place. Replaced with the correction and
a pointer to the right section.
The keypad heading also claimed the hold time was the only cause. There were
two: the 2500 ms hold in key.py, and row_out missing volatile. Both are now
listed up front with a link to the detail.
Adds the regression test to the places someone would actually look: the
"How to run it" section in AGENTS.md, the layout listing, and a build step
in the README noting that a clean build is not evidence the keypad works,
since the -O2 dead-code elimination produces no warning.
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.
Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.
From the object code:
callers reaching keypad_update_rows
plain none -- the calls are gone
volatile keypad_key_changed, keypad_col_changed, keypad_set_press,
keypad_reset, uvk5_machine_init
keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.
Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.
Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.
Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).
README screenshots regenerated from the current build.
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.
The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.
2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.
Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.
Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
keypad scan and the model does not wake from it. Includes the two dead
ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
Notes for whoever works on this next, weighted toward what the code does not
say: that the firmware is the reference and must never be edited to suit the
emulator, that register layouts come from the vendor CMSIS header rather than
inference, and that the way to find the next peripheral worth modelling is to
watch where the firmware stops.
Records the mistakes that already cost time here, each with the symptom that
made it look like something else: GDB breakpoints halting the guest (which reads
as 'the keypress does nothing'), writing the SysTick counter back while
accelerating it (which hangs the delay loop outright), lowering the clock to
speed up busy-waits (measured, 32x, nowhere near enough), unnamed qdev GPIO lines
sharing one namespace, and a probe script whose own regex silently matched
nothing.
Also states plainly what the emulator cannot answer, so a passing test is not
mistaken for evidence about radio behaviour.