# Working on this repo Notes for whoever picks this up next. Focused on what is not obvious from the code, and on mistakes that already cost time here. *中文:[AGENTS.zh-CN.md](AGENTS.zh-CN.md) · the two are kept in step; change both.* ## What this is A QEMU machine for the Puya PY32F071 (Cortex-M0+), so Quansheng UV-K5 V3 firmware runs on a PC. Boots to the main loop in ~5 s; the LCD is readable. The machine and every device model live in one file, `qemu/py32f071.c`. That is deliberate: the models are small and tightly coupled to each other's wiring, and splitting them would spread the board layout out without making any of it clearer. ## How it boots Worth reading before debugging anything that looks like a startup problem. There is no bootloader, no kernel, no partition table and no filesystem -- the firmware is the only code on the machine and it owns the CPU outright. **The hardware knows two numbers.** A Cortex-M0+ coming out of reset does not run any boot logic. It loads SP from the first word of the vector table and PC from the second, and starts executing. That is the whole handoff. .isr_vector 0x08002800 (readelf -SW, size 0xc0) +0x00 0x20004000 initial SP, i.e. the top of the 16 KB SRAM +0x04 0x08002d49 Reset_Handler, and the ELF entry point Read it straight off the image when in doubt -- the bytes are little-endian, so `00400020 492d0008` is SP 0x20004000 followed by PC 0x08002d49: objdump -s -j .isr_vector firmware.elf | head -5 The odd address is not a typo: bit 0 flags Thumb state and the hardware masks it off when fetching. **`PY32_APP_OFFSET` 0x2800 is load-bearing.** Flash starts at `0x08000000` but the first 10 KB is the factory bootloader region, so the application sits after it. `armv7m_load_kernel()` is passed that offset for exactly this reason -- load at `0x08000000` instead and the vector table lands in the wrong place, so the very first fetch faults. **Startup is 31 lines of assembly**, in the firmware's `Core/startup_py32f071xx.s`: set SP from _estack bl SystemInit copy .data from flash (_sidata) into RAM (_sdata .. _edata) zero .bss (_sbss .. _ebss) bl __libc_init_array bl main LoopForever: b LoopForever @ main never returns The copy and the zero-fill are the interesting part. Initialised globals live in flash but have to be writable, so they are copied word by word into RAM; uninitialised globals must read as zero per the C standard, so `.bss` is cleared. On a hosted OS the kernel and the loader do this for you. Here nobody does, so if either loop is wrong you get globals that are silently garbage. **Then the application:** main() Core/Src/main.c -- clock config only, then Main() Main() App/main.c -- the actual firmware SYSTICK_Init() the 10 ms tick everything is timed against BOARD_Init() GPIO, SPI, LCD, keypad matrix UART_Init() where the SERIAL banner in the log comes from SETTINGS_InitEEPROM() reads settings over SPI from the flash image while (1) { ... } main loop, never exits **There is no filesystem.** The nearest thing to "mounting a partition" is `SETTINGS_InitEEPROM()` reading fixed byte offsets over SPI: `0xA008` for the power save byte, `0x0E70` for the VFO indices, and so on. No metadata, no directory, no checksum -- just an address that the code and the data both have to agree on. When a setting reads back wrong, suspect the offset before suspecting the transport. Boot time is emulation overhead. Measured on this machine: first pixels at ~1.6 s and a drawn main screen at ~3.6 s after QEMU starts, which is the "~5 s" the README quotes. A real radio is up in about a second. ## Ground rules **Never edit the firmware to make the emulator work.** The firmware is the reference. If something does not run, the model is wrong. A fix that changes firmware source makes every later test meaningless, because you are no longer testing what the radio runs. **Register layouts come from the vendor CMSIS header**, not from a datasheet search and not from inference: /Drivers/CMSIS/Device/PY32F071/Include/py32f071xB.h When you need a bit position, read it from there. Several details are unintuitive — `LL_ADC_FLAG_EOS` is really `ADC_SR_EOC` on this part — and guessing produces models that look right and hang. **Find the next thing to model by watching where the firmware stops**, not by reading the datasheet front to back. Every peripheral here was added because the firmware demonstrably waited on it: tools/where.sh 4 # sample the call stack a few times A stack that repeats in the same function across samples is a spin loop. Look at what it reads. ## How to run it python3 tools/make_flash.py # once; builds assets/flash.img tools/run.sh # GDB stub on :1234, QMP on /tmp/uvk5-qmp.sock tools/where.sh # where execution is tools/gpiob_dump.sh # GPIOB registers python3 tools/key.py MENU # inject a keypress python3 tools/screenshot.py --frame-addr 0x200013DC \ --status-addr 0x2000175C --port 1234 --out screen.png Screenshot addresses move between firmware builds. Get the current ones with: arm-none-eabi-nm firmware.elf | grep -E 'gFrameBuffer|gStatusLine' Rebuild after editing the machine: cd $QEMU/build && ninja qemu-system-arm # ~10 s incremental After any change near the keypad or the GPIO wiring, run the regression test. It boots its own instance on private ports, so it does not disturb a `run.sh` session: python3 tools/keypad_test.py There is also a browser UI, which is usually the quickest way to poke at the firmware by hand: python3 tools/webui.py --frame-addr 0x200013DC \ --status-addr 0x2000175C # then open http://127.0.0.1:8080/ Two things about it that matter when working on this repo: - **It holds the QMP socket for its lifetime**, so `key.py` cannot run at the same time. The socket accepts a single client. - **It reads frames with QMP `memsave`, deliberately.** Not `pmemsave`, which takes a *physical* address and silently returns zeros for `gFrameBuffer` -- a blank screen with no error. And not gdb, which halts the guest on every attach: that stutters the stream and perturbs key debounce timing. Its tests: `tools/test_uvk5_*.py` and `tools/test_webui.py` need no emulator, `tools/test_webui_e2e.py` boots its own. A firmware can also be loaded from the page rather than from the command line: `POST /api/firmware` takes the image as its request body, stores it in `work/firmware/`, and boots it -- restarting the emulator if it was running. The image's **shape** is read out of the image (`tools/uvk5_image.py` on the host, `uvk5_sniff_app_offset()` in the machine): an *application* image is linked for `0x08002800`, a *full-flash* image starts at `0x08000000`, and address 0 has to alias the matching base. Getting that wrong is silent -- the image lands 0x2800 bytes off and the first fetch reads whatever data is there -- which is why it is not a flag and not a file-name convention. A file that is not an image is refused without disturbing the running radio. Two things about that path are worth knowing, both found the hard way: - **The flash image travels in the environment, not in `-M`.** Through the launcher, QEMU rejected `-M uv-k5-v3,flash-image=...` with "unsupported machine type": the identical argv started fine when run by hand, `-M help` in the *same* context listed the machine, the argv `repr` was clean, and the environment diffed down to nothing conclusive. The property still works when it is set, so both are supported; the launcher now passes the bare machine name plus `UVK5_FLASH_IMAGE`, which the model reads as a fallback. The root cause is unexplained -- do not "clean this up" without re-testing a power-on from the page. - **The screen is read from the display controller, not from guest RAM.** The panel model keeps the controller's own display RAM (8 pages of 128 columns), and the web page renders that, so the picture is right for *any* firmware -- builds sharing an ancestor still differ in their display logic, and the multi-system release keeps its image somewhere else entirely. Do not apply the driver's `0xA1` segment reverse on top of the data: measured at one instant against the guest's own framebuffer, 8153 of 8192 pixels agree with no mirroring and 6557 with it. `memsave` of `gFrameBuffer` remains the fallback for an emulator built without the panel model. ## The flash bugs: four faults, one symptom "The frequency will not change" and "flash forgets everything after power off" looked like two complaints. They were one root cause plus three real bugs found on the way, all in this file. Worth reading before touching SPI, DMA or the flash model, because each was invisible from the layer above. 1. **DMA used the wrong address space** — the actual cause. It moved bytes through `address_space_memory`, which cannot decode this SoC's memory at all: the container region is handed only to the ARMv7M core and never registered with global system memory. Reads returned `MEMTX_DECODE_ERROR` and zeros; writes went nowhere. DMA now runs over an `AddressSpace` built on the container. 2. **Page program did not wrap.** Real SPI NOR latches only the low address bits, so a burst past the 256-byte page boundary continues at the start of the same page. The model walked straight through, and a 512-byte burst at 0x008F00 (which the firmware really does send in one CS assertion) spilled into 0x009000. 3. **DMA started too early.** Transfers ran when a channel was enabled, but on hardware they start when the peripheral raises its request. The driver arms both channels, then enables SPI, then sets TXDMAEN — so firing at arm time clocked the bus before the read command had been sent. 4. **DMA channels ran one after another.** SPI is duplex and the driver pairs a dummy-feeding TX channel with a data-collecting RX channel over one transfer. Running them in sequence let TX finish before RX ever sampled the bus. Any one of them zeroed the sector holding per-band VFO frequencies. `RADIO_ConfigureChannel` substitutes a band's lower limit only for `0xFFFFFFFF`, so a stored zero was taken literally and clamped to `BX4819_band1_lower` — 18 MHz. That is the whole explanation for a typed frequency always reverting. `tools/test_freq_entry.py` and the `MUST_NOT_CHANGE` guard in `tools/test_flash_persist.py` exist to catch a regression in any of the four. ### What made this hard to find, and what to do instead **Instrument the model, not the guest.** The frequency input box times out after `key_input_timeout_500ms / 3`, about 2.5 s, and a gdb attach takes roughly 3 s. So probing between digits clears the box, and the run reports a failure that the measurement caused. This produced at least three confident wrong conclusions, including "the firmware saved band 0" when the box had simply emptied. Add an `fprintf` to `qemu/py32f071.c` and read stderr instead — the guest never stops. **Never cap a diagnostic log before you know the shape of the data.** A probe limited to the first six transactions showed only `0xFF` payloads, which supported exactly the wrong conclusion. Without the cap, the writes that mattered were obvious. **Check that the build succeeded before believing a test.** A failed `ninja` leaves the previous binary in place and the test still runs, so a stale build silently answers the question. Two rounds of results were meaningless this way. Grep the build output for `FAILED` and `error:` and stop if either appears. **Reset the flash image between runs.** `assets/flash.img` is written by every session. A test that starts from it may find its work already done — which shows up as "the image is byte-identical", indistinguishable from broken persistence. Start from `assets/pristine/`, and power the emulator off *before* restoring, since shutdown flushes the old in-memory image back over the file. **Do not hand-compute struct offsets.** The ELF has no DWARF and the structs contain enums whose size cannot be assumed. Offsets computed by hand produced `KEY_LOCK=4` and `TX_VFO=11`, neither of which is a possible value. Either use a symbol that `nm` reports and whose type is unambiguous (`gInputBoxIndex` is a plain `uint8_t`), or locate a field by behaviour — toggling the keypad lock with a long `F` press and diffing the region found `KEY_LOCK` at `gEeprom+0x12` in one step. **Read your own probe output carefully.** One probe printed `phase` before it was incremented, which made a correct address decoder look off by one byte. Replaying the logic in Python cleared it up; without that, a working implementation would have been "fixed". ## Things that already went wrong **GDB breakpoints halt the guest.** A key held across a breakpoint session is never processed, because the main loop is not running. This produced a whole round of "the keypress does nothing" that was really "the machine is stopped". Use `tools/press_and_shot.sh` — it presses, lets the machine run, then reads the framebuffer, with no breakpoints anywhere. **Do not write the SysTick counter back when accelerating it.** Two attempts did that. Each read re-anchored the count, so the value the firmware saw stopped changing, its `if (cur != prev)` guard never fired, and the delay loop hung outright — worse than the slowness being fixed. The working approach reports a value that runs ahead of the real counter and leaves the timer alone. **Lowering the clock does not speed up delay loops.** The bottleneck is loop iterations per second, not counter speed. 48 MHz to 200 Hz bought 32x and was nowhere near enough. Measured, not assumed. **Unnamed qdev in and out lines share one namespace.** A device with both unnamed `qdev_init_gpio_in` and `qdev_init_gpio_out` makes `qdev_get_gpio_in()` ambiguous, and board wiring silently attaches to the wrong line. The GPIO model uses `"pin-in"` and `"pin-out"` for this reason. Keep it that way. **Key hold times must be SHORT, not generous.** This entry used to say the opposite -- that guest time runs fast so a press needs a long hold, and that `key.py` should hold for 2500 ms. That was wrong and it broke the keypad tooling for a long time. 2500 ms is ~250 firmware ticks, six times past the long-press threshold, so every press was dispatched as a *hold* and handlers that act on a short release did nothing. See the keypad section below; `key.py` now holds 200 ms. **Verify a tool's own parsing before trusting its output.** `gpio_watch.py` reported `IDR=0x0000` for several rounds because its regex did not match gdb's output format at all. The register was fine; the reader was broken. Cross-check with `tools/gpiob_dump.sh`, which uses a different path. The same trap one layer further out: **a redirect can change the encoding.** Three probe runs under `qemu ... 2> probe.log` reported zero SPI transfers, zero flash reads and zero chip-select changes, and "the firmware never touches SPI" was written down as a finding. PowerShell 5.1 writes `2>` as UTF-16LE, so every ASCII line a probe printed had a NUL between each character and a `startswith("LCDW")` filter could never match it. Decoding the same file as UTF-16 showed a complete ST7565 init sequence and 48 distinct settings reads. Before believing an empty probe, check that the probe *can* be seen: read the file, count its bytes, or write it from `cmd /c`, which does not re-encode. **QMP `pmemsave` is physical, `memsave` is virtual.** The framebuffer symbols are CPU virtual addresses, so `pmemsave` on `gFrameBuffer` returns a block of zeros and reports success -- a blank screen with nothing logged anywhere. The web UI was built on `pmemsave` first because a timing benchmark said it was fast; the benchmark never checked the *contents*. Measure the thing you actually care about: the bug surfaced only when a rendered frame came back with 0 lit pixels where the gdb path reported 1693. ### The page is generated by an f-string, so check the script it serves The web UI is one f-string. A stray backslash in a JavaScript string literal therefore produces a page whose **whole** `