The squelch section said the blocker sat above the device model and told the reader
not to resume without new information. Both are now wrong, so it leads with the
outcome instead: SIDE1 engages monitor, which skips squelch entirely, and the meter
reads -53 dBm / S9+40.
The four failed attempts are kept rather than deleted. Each ended in a confident
wrong diagnosis -- power-save gating, trigger timing, bit semantics -- and the shared
cause was a one-bit read skew underneath all of them. That pattern is worth more to
the next reader than a clean account of the version that worked.
README gains the S-meter row and test_smeter.py.
The shifted-read bug invalidated four earlier diagnoses, so the notes claiming a
power-save gate or a timing problem were wrong and are corrected.
With reads fixed the interrupt handshake demonstrably works: RAISE pending=0004,
ACK delivering flags=0004, correct bit, collected by the firmware, REG_0C back to
0x0000 so nothing hangs.
g_SquelchLost is still 0 and the S-meter still absent, but the shape of the problem
is now clear and recorded: announcing once lands during startup before the flag
leads anywhere; announcing every poll re-arms the bit inside the firmware's own
untimed collection loop and spins forever; announcing periodically avoids both and
still changes nothing. The blocker is getting the radio into a receiving state at
all, which sits above the device model.
Also records the general lesson, which cost the most time here: when several
independent attempts fail in the same way, suspect the shared transport rather than
the logic layered on top of it.
The previous commit blamed the power-save gate at app/app.c:1697 for the interrupt
never being collected. That was wrong, and the test which disproves it is cheap:
BATTERY_SAVE lives at flash 0xA00B, app/app.c:1374 refuses power save when it is 0,
and patching the byte gives
BATTERY_SAVE=4: fn=5 idle=1 polls=2161 acks=0
BATTERY_SAVE=0: fn=0 idle=0 polls=2161 acks=0
The gate passes and nothing changes. A gdb backtrace confirms the loop runs --
CheckRadioInterrupts is inlined into APP_TimeSlice10ms, which is the caller of every
REG_0C read.
Where it really stands: a breakpoint on BK4819_GetRSSI never fires at all. The
firmware does not read RSSI in this idle state, so the missing S-meter is not the
model withholding a value; the receive state machine has to be entered first, and the
squelch interrupt is an input to that rather than the switch. Three rounds of
increasingly precise instrumentation all ended at "the firmware is not asking".
Four measurement mistakes are recorded because each produced a confident wrong
conclusion, and one of them made me change working code:
- Sampling PC at the REG_0C read lands in BK4819_WriteU8; sampling LR lands inside
BK4819_ReadRegister, since it calls BK4819_ReadU16. Use a backtrace.
- A probe printing shift_out before the assignment showed 0000 for a value about to
be sent as 0001.
- BK4819_ReadRegister returning 0x0 for REG_0C looked like a broken read path, and I
altered the bit timing over it. REG_0C legitimately holds 0 in the committed build.
A read returning the register's real contents proves nothing -- test against one the
firmware wrote, like REG_3F (0x0C0C) or REG_78 (0x2F5B).
- nexti after a breakpoint reported r0=0 from an unrelated location. finish gives the
actual return value.
Also noted: gdb cannot call guest functions on this target, and there is no
gCurrentRSSI global -- RSSI is read and discarded, so a breakpoint plus finish is the
only way to see what the firmware received.
Code is unchanged and at the committed state; test_bk4819.py passes.
Second attempt, and this one produced a definite answer rather than another guess.
Two real mistakes in the first version, both fixed along the way and worth recording:
squelch was evaluated when the firmware configured the chip, but the startup sequence
writes REG_3F as 0x0000 then 0x0C0C three times over, so a flag raised on the enabling
write was disabled again before anyone read it. And the threshold came from REG_4E,
whose low bits are the glitch threshold -- the RSSI open level is REG_78 bits 15:8 at
0.5 dB/step against REG_67's 0.25 dB/step, so squelch could never open at all.
With both corrected, every chip-side condition lines up: en=0x0C0C, rssi=0x01E0,
threshold 94, REG_0C correctly returning 1. The firmware still never acknowledged,
and the reason is outside the chip:
gCurrentFunction=5 (FUNCTION_POWER_SAVE), gRxIdleMode=1
against the gate at app/app.c:1697,
if (gCurrentFunction != FUNCTION_POWER_SAVE || !gRxIdleMode)
CheckRadioInterrupts();
Both halves false, so the interrupt loop is never entered and nothing can collect the
flag. The emulator idles in power save, so that is the steady state.
Gating on REG_30 -- zeroed by BK4819_Sleep on each power-save cycle -- does not help:
the chip is awake when the model is asked while the firmware still has gRxIdleMode=1.
Chip state and firmware state are not in step, so no condition available inside a
register model can decide this correctly.
So it is not a matter of a better trigger. A working version needs to keep the guest
out of power save, or drive the interrupt from something that knows the firmware's
receive state -- neither of which belongs in this device. Both attempts reverted; the
model is back at the committed state and test_bk4819.py passes, including the REG_0C
assertion that catches the armed-hang failure mode.
Scanning does work now that RSSI reports a real level -- long-press * and the
frequency steps, 6 distinct frames over 7 seconds. The S-meter still does not appear,
because ui/main.c:2370 draws it only when FUNCTION_IsRx(), which needs the chip to
report a squelch opening rather than merely a healthy RSSI.
I implemented that interrupt and backed it out. The guest kept running, but REG_0C
bit 0 remained set afterwards, meaning the firmware never collected the interrupt.
That leaves a hang armed: app/app.c:910 and :1417 spin on that bit with no timeout, so
any path reaching them with it stuck never returns. A missing S-meter is a cosmetic
gap; a latent hang is not, and shipping the second to fix the first is a bad trade.
Documented with the mechanism (REG_0C pending bit, REG_02 acknowledge-then-read,
sqlFound at bit 3), the reason it failed, and the check that says a retry is correct:
REG_0C reading 0 afterwards, which tools/test_bk4819.py already asserts. The next
person should find out why the flag was not collected rather than raise it at a
different moment and hope.
The "what this cannot do" section said the transceiver was not modelled and treated
that as permanent. Half of it is now wrong: the register interface works. The other
half is still true and worth keeping sharp -- analogue behaviour is out of reach
because the chip has no public datasheet, so the driver is the only specification and
it can only say which registers were written.
Rewritten to separate the two: what is modelled and what it fixed (RSSI was hard zero
at 18 call sites), the two constraints the untimed spin loops impose, and then the
line that does not move. Explicitly warns against reading the new test as evidence
about RF.
Also corrects the GPIO comment about idling PB9 low. It described the pin as a
workaround pending a device model; that model now exists, and the idle level only
covers the window between reset and the bus being wired.
The previous commit documented serial receive as a permanent limitation, listing
what would be needed to build it. It has since been built, so that section was
actively misleading -- it told the reader not to try something that already works.
Replaced with what it takes to keep working, since each of the three pieces fails
silently on its own: USART1 needs a chardev, DMA must decrement CNDTR for USART
(serviced on the CNDTR read, which is where the driver looks), and DR writes must
reach the chardev and not just stderr. The last one is the trap -- transmit that
goes only to stderr is indistinguishable from the firmware ignoring the command.
README status table and build-verification steps updated to match.
Found while explaining a DMA channel that stayed permanently armed with cndtr=256
during the flash investigation. It is USART1's receive channel, running in
LL_DMA_MODE_CIRCULAR, so never completing is correct behaviour and not a bug.
But it exposed a real gap. driver/uart.c derives its write pointer from
sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(...), and the DMA model only
services SPI, so CNDTR never decrements for USART and that expression is always
zero. Combined with USART1 being a py32-stub with no chardev backend, nothing can
be sent *to* the firmware.
The cost is specific: UART_IsCommandAvailable never fires, so the UV-K5 programming
protocol in app/uart.c is unreachable -- 0x0514 handshake, 0x051B EEPROM read,
0x051D EEPROM write, 0x05DD reset. CPS/CHIRP-style tools cannot talk to this
emulator. Transmit is unaffected, which is why the firmware banner shows up fine
and this went unnoticed.
Documented rather than fixed: it needs a chardev on USART1 plus circular-mode DMA
driven by receive, which is a new feature rather than a repair. The notes say what
would be involved so the next person does not have to rediscover the mechanism.
Status table also updated to reflect what the flash and DMA fixes settled --
persistence and frequency entry now work.
Four faults in the SPI/DMA/flash models each produced the same symptom -- stored
frequencies zeroed, a typed frequency reverting to 18 MHz -- and each was invisible
from the layer above. AGENTS.md now lists all four with the reasoning that connects
them to what the user saw, so the next person does not rediscover them one at a
time.
Also records the methodology mistakes, because they cost more than the bugs did:
- Attaching gdb between digits clears the frequency input box: the timeout is
~2.5 s and an attach takes ~3 s. Three wrong conclusions came from this, so
instrument the model and read stderr instead of stopping the guest.
- A diagnostic log capped at six entries showed only 0xFF payloads and supported
precisely the wrong conclusion. Do not cap before the shape of the data is known.
- A failed ninja leaves the old binary in place and the test still runs. Two rounds
of results were meaningless. Check for FAILED and error: before trusting a run.
- assets/flash.img is written by every session, so a test starting from it can find
its work already done -- which looks identical to broken persistence. Start from
assets/pristine/, and power off before restoring, since shutdown flushes the old
image back over the file.
- Hand-computed struct offsets gave KEY_LOCK=4 and TX_VFO=11, impossible values,
because the ELF has no DWARF and the structs contain enums. Use an unambiguous
nm symbol, or find the field by toggling it and diffing.
- One probe printed phase before incrementing it, making a correct address decoder
look off by one. A working implementation was nearly "fixed" as a result.
README gains the two new tests in the build-verification step and the tools list.
There is no bootloader, kernel, partition table or filesystem here, and that is not
obvious from the code -- someone arriving with hosted-OS assumptions will look for
layers that do not exist.
Covers the parts that are easy to get wrong rather than restating the source:
- The reset handoff is two words. SP from 0x08002800, PC from +0x04. Verified
against the image: 00400020 492d0008 is SP 0x20004000 (top of the 16 KB SRAM,
matching PY32_SRAM_BASE + PY32_SRAM_SIZE) and PC 0x08002d49, which is also the
ELF entry and the Reset_Handler symbol. The odd address is Thumb-state bit 0.
- PY32_APP_OFFSET 0x2800 is load-bearing. The first 10 KB of flash is the factory
bootloader region, so armv7m_load_kernel gets that offset; without it the vector
table lands in the wrong place and the first fetch faults.
- Startup is 31 lines of assembly, and the .data copy plus .bss zero-fill are the
part worth understanding: on a hosted OS the kernel and loader do that, here
nobody does, so a fault in either loop shows up as globals that are silently
garbage rather than as a crash.
- SETTINGS_InitEEPROM reading fixed SPI offsets is the closest thing to mounting a
partition. No metadata and no checksum, just an address both sides must agree on,
so a setting that reads back wrong points at the offset before the transport.
Also notes that the ~15 s to the main loop is emulation overhead; a real radio is
up in about a second.
README gets a section with the endpoint table, the two constraints that will
otherwise surprise someone (single QMP client, no authentication), and the reason
frames go through memsave rather than pmemsave or gdb.
AGENTS.md gets the run instructions plus a new entry under 'Things that already
went wrong' for the pmemsave trap: it takes a physical address, returns zeros for
gFrameBuffer, and reports success. The web UI was built on it initially because a
benchmark showed it was fast -- the benchmark never checked the contents. Worth
recording as the general lesson, not just the specific fix.
AGENTS.md still carried a stale entry telling the reader to *lengthen* key
holds when a press seems ignored, which is the opposite of the fix and is
what broke the tooling in the first place. Replaced with the correction and
a pointer to the right section.
The keypad heading also claimed the hold time was the only cause. There were
two: the 2500 ms hold in key.py, and row_out missing volatile. Both are now
listed up front with a link to the detail.
Adds the regression test to the places someone would actually look: the
"How to run it" section in AGENTS.md, the layout listing, and a build step
in the README noting that a clean build is not evidence the keypad works,
since the -O2 dead-code elimination produces no warning.
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.
Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.
From the object code:
callers reaching keypad_update_rows
plain none -- the calls are gone
volatile keypad_key_changed, keypad_col_changed, keypad_set_press,
keypad_reset, uvk5_machine_init
keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.
Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.
Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.
Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).
README screenshots regenerated from the current build.
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.
The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.
2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.
Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.
Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
keypad scan and the model does not wake from it. Includes the two dead
ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
Notes for whoever works on this next, weighted toward what the code does not
say: that the firmware is the reference and must never be edited to suit the
emulator, that register layouts come from the vendor CMSIS header rather than
inference, and that the way to find the next peripheral worth modelling is to
watch where the firmware stops.
Records the mistakes that already cost time here, each with the symptom that
made it look like something else: GDB breakpoints halting the guest (which reads
as 'the keypress does nothing'), writing the SysTick counter back while
accelerating it (which hangs the delay loop outright), lowering the clock to
speed up busy-waits (measured, 32x, nowhere near enough), unnamed qdev GPIO lines
sharing one namespace, and a probe script whose own regex silently matched
nothing.
Also states plainly what the emulator cannot answer, so a passing test is not
mistaken for evidence about radio behaviour.