The previous commit documented serial receive as a permanent limitation, listing
what would be needed to build it. It has since been built, so that section was
actively misleading -- it told the reader not to try something that already works.
Replaced with what it takes to keep working, since each of the three pieces fails
silently on its own: USART1 needs a chardev, DMA must decrement CNDTR for USART
(serviced on the CNDTR read, which is where the driver looks), and DR writes must
reach the chardev and not just stderr. The last one is the trap -- transmit that
goes only to stderr is indistinguishable from the firmware ignoring the command.
README status table and build-verification steps updated to match.
Found while explaining a DMA channel that stayed permanently armed with cndtr=256
during the flash investigation. It is USART1's receive channel, running in
LL_DMA_MODE_CIRCULAR, so never completing is correct behaviour and not a bug.
But it exposed a real gap. driver/uart.c derives its write pointer from
sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(...), and the DMA model only
services SPI, so CNDTR never decrements for USART and that expression is always
zero. Combined with USART1 being a py32-stub with no chardev backend, nothing can
be sent *to* the firmware.
The cost is specific: UART_IsCommandAvailable never fires, so the UV-K5 programming
protocol in app/uart.c is unreachable -- 0x0514 handshake, 0x051B EEPROM read,
0x051D EEPROM write, 0x05DD reset. CPS/CHIRP-style tools cannot talk to this
emulator. Transmit is unaffected, which is why the firmware banner shows up fine
and this went unnoticed.
Documented rather than fixed: it needs a chardev on USART1 plus circular-mode DMA
driven by receive, which is a new feature rather than a repair. The notes say what
would be involved so the next person does not have to rediscover the mechanism.
Status table also updated to reflect what the flash and DMA fixes settled --
persistence and frequency entry now work.
Four faults in the SPI/DMA/flash models each produced the same symptom -- stored
frequencies zeroed, a typed frequency reverting to 18 MHz -- and each was invisible
from the layer above. AGENTS.md now lists all four with the reasoning that connects
them to what the user saw, so the next person does not rediscover them one at a
time.
Also records the methodology mistakes, because they cost more than the bugs did:
- Attaching gdb between digits clears the frequency input box: the timeout is
~2.5 s and an attach takes ~3 s. Three wrong conclusions came from this, so
instrument the model and read stderr instead of stopping the guest.
- A diagnostic log capped at six entries showed only 0xFF payloads and supported
precisely the wrong conclusion. Do not cap before the shape of the data is known.
- A failed ninja leaves the old binary in place and the test still runs. Two rounds
of results were meaningless. Check for FAILED and error: before trusting a run.
- assets/flash.img is written by every session, so a test starting from it can find
its work already done -- which looks identical to broken persistence. Start from
assets/pristine/, and power off before restoring, since shutdown flushes the old
image back over the file.
- Hand-computed struct offsets gave KEY_LOCK=4 and TX_VFO=11, impossible values,
because the ELF has no DWARF and the structs contain enums. Use an unambiguous
nm symbol, or find the field by toggling it and diffing.
- One probe printed phase before incrementing it, making a correct address decoder
look off by one. A working implementation was nearly "fixed" as a result.
README gains the two new tests in the build-verification step and the tools list.
There is no bootloader, kernel, partition table or filesystem here, and that is not
obvious from the code -- someone arriving with hosted-OS assumptions will look for
layers that do not exist.
Covers the parts that are easy to get wrong rather than restating the source:
- The reset handoff is two words. SP from 0x08002800, PC from +0x04. Verified
against the image: 00400020 492d0008 is SP 0x20004000 (top of the 16 KB SRAM,
matching PY32_SRAM_BASE + PY32_SRAM_SIZE) and PC 0x08002d49, which is also the
ELF entry and the Reset_Handler symbol. The odd address is Thumb-state bit 0.
- PY32_APP_OFFSET 0x2800 is load-bearing. The first 10 KB of flash is the factory
bootloader region, so armv7m_load_kernel gets that offset; without it the vector
table lands in the wrong place and the first fetch faults.
- Startup is 31 lines of assembly, and the .data copy plus .bss zero-fill are the
part worth understanding: on a hosted OS the kernel and loader do that, here
nobody does, so a fault in either loop shows up as globals that are silently
garbage rather than as a crash.
- SETTINGS_InitEEPROM reading fixed SPI offsets is the closest thing to mounting a
partition. No metadata and no checksum, just an address both sides must agree on,
so a setting that reads back wrong points at the offset before the transport.
Also notes that the ~15 s to the main loop is emulation overhead; a real radio is
up in about a second.
README gets a section with the endpoint table, the two constraints that will
otherwise surprise someone (single QMP client, no authentication), and the reason
frames go through memsave rather than pmemsave or gdb.
AGENTS.md gets the run instructions plus a new entry under 'Things that already
went wrong' for the pmemsave trap: it takes a physical address, returns zeros for
gFrameBuffer, and reports success. The web UI was built on it initially because a
benchmark showed it was fast -- the benchmark never checked the contents. Worth
recording as the general lesson, not just the specific fix.
AGENTS.md still carried a stale entry telling the reader to *lengthen* key
holds when a press seems ignored, which is the opposite of the fix and is
what broke the tooling in the first place. Replaced with the correction and
a pointer to the right section.
The keypad heading also claimed the hold time was the only cause. There were
two: the 2500 ms hold in key.py, and row_out missing volatile. Both are now
listed up front with a link to the detail.
Adds the regression test to the places someone would actually look: the
"How to run it" section in AGENTS.md, the layout listing, and a build step
in the README noting that a clean build is not evidence the keypad works,
since the -O2 dead-code elimination produces no warning.
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.
Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.
From the object code:
callers reaching keypad_update_rows
plain none -- the calls are gone
volatile keypad_key_changed, keypad_col_changed, keypad_set_press,
keypad_reset, uvk5_machine_init
keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.
Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.
Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.
Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).
README screenshots regenerated from the current build.
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.
The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.
2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.
Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.
Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
keypad scan and the model does not wake from it. Includes the two dead
ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
Notes for whoever works on this next, weighted toward what the code does not
say: that the firmware is the reference and must never be edited to suit the
emulator, that register layouts come from the vendor CMSIS header rather than
inference, and that the way to find the next peripheral worth modelling is to
watch where the firmware stops.
Records the mistakes that already cost time here, each with the symptom that
made it look like something else: GDB breakpoints halting the guest (which reads
as 'the keypress does nothing'), writing the SysTick counter back while
accelerating it (which hangs the delay loop outright), lowering the clock to
speed up busy-waits (measured, 32x, nowhere near enough), unnamed qdev GPIO lines
sharing one namespace, and a probe script whose own regex silently matched
nothing.
Also states plainly what the emulator cannot answer, so a passing test is not
mistaken for evidence about radio behaviour.