Commit Graph
96 Commits
Author SHA1 Message Date
mckero 6dd5e34445 Record the squelch interrupt attempt, and why it was backed out
Scanning does work now that RSSI reports a real level -- long-press * and the
frequency steps, 6 distinct frames over 7 seconds. The S-meter still does not appear,
because ui/main.c:2370 draws it only when FUNCTION_IsRx(), which needs the chip to
report a squelch opening rather than merely a healthy RSSI.

I implemented that interrupt and backed it out. The guest kept running, but REG_0C
bit 0 remained set afterwards, meaning the firmware never collected the interrupt.
That leaves a hang armed: app/app.c:910 and :1417 spin on that bit with no timeout, so
any path reaching them with it stuck never returns. A missing S-meter is a cosmetic
gap; a latent hang is not, and shipping the second to fix the first is a bad trade.

Documented with the mechanism (REG_0C pending bit, REG_02 acknowledge-then-read,
sqlFound at bit 3), the reason it failed, and the check that says a retry is correct:
REG_0C reading 0 afterwards, which tools/test_bk4819.py already asserts. The next
person should find out why the flag was not collected rather than raise it at a
different moment and hope.
2026-08-28 16:56:08 +01:00
mckero 4e29d11810 Update the docs now that the BK4819 is modelled
The "what this cannot do" section said the transceiver was not modelled and treated
that as permanent. Half of it is now wrong: the register interface works. The other
half is still true and worth keeping sharp -- analogue behaviour is out of reach
because the chip has no public datasheet, so the driver is the only specification and
it can only say which registers were written.

Rewritten to separate the two: what is modelled and what it fixed (RSSI was hard zero
at 18 call sites), the two constraints the untimed spin loops impose, and then the
line that does not move. Explicitly warns against reading the new test as evidence
about RF.

Also corrects the GPIO comment about idling PB9 low. It described the pin as a
workaround pending a device model; that model now exists, and the idle level only
covers the window between reset and the bus being wired.
2026-08-28 16:46:55 +01:00
mckero 76fc72d5ed Model the BK4819 register interface
The transceiver was not modelled at all. Its bit-banged three-wire bus went nowhere,
so every register read returned whatever the floating GPIO happened to be, and PB9
had to be idled low as a workaround: with the line high, reads came back 0xFFFF and
RADIO_SetupRegisters spun forever on bit 0 of REG_0C.

Now a real device, wired to the pins the driver uses -- CS on PF9, SCL PB8, SDA PB9
with both directions connected -- decoding the protocol from App/driver/bk4819.c:
CS low, eight bits of register number MSB first with bit 7 set for a read, then
sixteen bits of data. Registers read back what the firmware wrote; the few it reads
without having written return plausible values.

Scope, deliberately narrow: this is the register interface, not the radio. The chip
has no public datasheet, so the driver is the only specification available and it can
only say which registers were written, never what left the antenna. Keying envelopes,
spurious emissions and sensitivity still need a real radio and a spectrum analyser.
The comments say so at the top of the device, so a passing test here is not mistaken
for evidence about RF.

What it buys is control flow that evaluates real values. RSSI was hard zero at 18
call sites -- -160 dBm -- so the S-meter read empty and squelch and scan decisions
saw a dead band. It now reports about -40 dBm. Measurably, the main screen comes up
on 400 MHz instead of the 18 MHz floor, because band setup no longer reads zeros.

Two things the untimed spins force:

- REG_0C bit 0 must stay clear. App/app/app.c:910 and :1417 loop on it with no
  timeout whatsoever, so a stuck bit hangs the guest rather than degrading.
- A soft reset (REG_00 bit 15, which BK4819_Init issues first) has to re-seed the
  measurement registers. Real hardware keeps measuring afterwards; this model would
  be left holding zeros. That was not theoretical -- the first test run decoded 48
  registers correctly and still reported RSSI as 0 for exactly this reason.

The register file is exposed over QOM as regNN so tests can inspect it without gdb.
That matters beyond convenience: attaching a debugger pauses the guest and changes
timing-sensitive behaviour, which has repeatedly produced conclusions that were
artefacts of the measurement rather than facts about the firmware.

tools/test_bk4819.py checks the guest still boots (i.e. the spin terminates), that
dozens of registers hold written values (52 currently, so the transfer really is
being decoded), that RSSI is not zero, and that REG_0C bit 0 is clear.

keypad_test.py, test_flash_persist.py, test_freq_entry.py, test_serial_rx.py and the
143 unit tests all still pass.
2026-08-28 16:45:29 +01:00
mckero f3412f3316 Correct the docs: serial receive is implemented now
The previous commit documented serial receive as a permanent limitation, listing
what would be needed to build it. It has since been built, so that section was
actively misleading -- it told the reader not to try something that already works.

Replaced with what it takes to keep working, since each of the three pieces fails
silently on its own: USART1 needs a chardev, DMA must decrement CNDTR for USART
(serviced on the CNDTR read, which is where the driver looks), and DR writes must
reach the chardev and not just stderr. The last one is the trap -- transmit that
goes only to stderr is indistinguishable from the firmware ignoring the command.

README status table and build-verification steps updated to match.
2026-08-28 16:26:55 +01:00
mckero b84d229326 Implement serial receive, unlocking the UV-K5 programming protocol
Nothing could be sent to the firmware. Three separate pieces were missing.

USART1 was a register stub with no chardev, so there was no source of incoming
bytes. It now takes a chardev property, defaulting to serial0, so -serial works.

The DMA model never serviced USART. App/driver/uart.c receives over a circular
peripheral-to-memory channel and never reads DR; it finds new data with

    write_ptr = sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(DMA1, CHANNEL_2)

so leaving CNDTR at its programmed value made the buffer look permanently empty
however many bytes arrived. DMA now drains USART1's queue byte by byte, decrementing
CNDTR and reloading it in circular mode. Channels also remember the length they were
given, since CNDTR counts down and the offset into the buffer has to be derived from
the difference.

Servicing happens on a CNDTR read rather than from a timer: that read is precisely
how the driver looks for data, so no polling is needed and no byte can be delivered
before the guest asks.

And transmit was invisible to the far end. DR writes went to stderr only, so a host
tool would send a command, the firmware would answer, and the answer went nowhere the
tool could see -- indistinguishable from being ignored. This cost a debugging round:
the first test run reported "no reply at all" with 0 bytes of boot output, which
looked like receive failing when the boot banner was in fact being written to stderr
as always. DR now also writes the raw byte to the chardev when one is connected.

SR reports RXNE when bytes are queued and DR consumes one, so a polling firmware
would work too, even though this one uses DMA.

tools/test_serial_rx.py speaks the real wire protocol -- AB CD framing, the fixed XOR
obfuscation, CRC-16/XMODEM -- and checks two exchanges end to end:

    0x0514 hello        -> 0x0515 ack
    0x051B EEPROM read  -> 0x051C with the requested 8 bytes at 0x0E70

Both pass. CPS/CHIRP-style tools can now talk to the emulator. keypad_test.py,
test_flash_persist.py, test_freq_entry.py and the 143 unit tests still pass.
2026-08-28 16:25:37 +01:00
mckero cdb73f13a0 Record that serial receive is not implemented
Found while explaining a DMA channel that stayed permanently armed with cndtr=256
during the flash investigation. It is USART1's receive channel, running in
LL_DMA_MODE_CIRCULAR, so never completing is correct behaviour and not a bug.

But it exposed a real gap. driver/uart.c derives its write pointer from
sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(...), and the DMA model only
services SPI, so CNDTR never decrements for USART and that expression is always
zero. Combined with USART1 being a py32-stub with no chardev backend, nothing can
be sent *to* the firmware.

The cost is specific: UART_IsCommandAvailable never fires, so the UV-K5 programming
protocol in app/uart.c is unreachable -- 0x0514 handshake, 0x051B EEPROM read,
0x051D EEPROM write, 0x05DD reset. CPS/CHIRP-style tools cannot talk to this
emulator. Transmit is unaffected, which is why the firmware banner shows up fine
and this went unnoticed.

Documented rather than fixed: it needs a chardev on USART1 plus circular-mode DMA
driven by receive, which is a new feature rather than a repair. The notes say what
would be involved so the next person does not have to rediscover the mechanism.

Status table also updated to reflect what the flash and DMA fixes settled --
persistence and frequency entry now work.
2026-08-28 15:52:05 +01:00
mckero 65e50c0065 Stop cleanup killing unrelated processes, and recover when it happens anyway
Two faults that combined to take the web UI down twice in one session, each time
surfacing to the user as a 502 through the reverse proxy.

`pkill -f 'M uv-k5-v3'` in run.sh and trace_run.sh matched far more than intended.
-f tests the whole command line, so it also matched the shell running the pkill
(the pattern sits in its own argv), any script mentioning the machine type, and the
QEMU child of a running webui.py. Cleanup now lives in tools/lib_kill_emulator.sh:
pgrep -x on the binary name, confirm uv-k5-v3 in /proc/PID/cmdline, and optionally
scope to one QMP socket so a caller only stops the instance it owns.

The supervisor then could not recover from it. power_on() began with
`if self._client is not None: return False`, but a client object is not proof of a
live guest -- after an external kill the stale client made power_on refuse forever,
so the Power button was dead until the whole service was restarted. It now checks
whether the process actually exited and relaunches, logging why.

Tests:

- tools/test_kill_emulator.sh checks a plain process, a process whose command line
  merely mentions uv-k5-v3, and the calling script all survive; that a real emulator
  on a named socket is stopped; that one on another socket is not; and that an
  unscoped call still clears everything. Verified it leaves a live webui.py alone.
- Two supervisor unit tests cover relaunch-after-external-kill and the case that
  must still refuse, so this cannot regress into starting two emulators at once.

Verified end to end against the running web UI: kill the emulator from outside,
status reports unreachable, and pressing Power brings it back to
{"status":"running"} where before it stayed dead.

While writing the first version of the test I modelled the failure as a client
raising BrokenPipeError, which is not what is_running() looks at -- it checks
poll(). The mock was wrong, not the code; the test now has the process report an
exit status, which is what really happens.
2026-08-28 15:50:13 +01:00
mckero d09bf5f47e Write down the flash investigation and what made it slow
Four faults in the SPI/DMA/flash models each produced the same symptom -- stored
frequencies zeroed, a typed frequency reverting to 18 MHz -- and each was invisible
from the layer above. AGENTS.md now lists all four with the reasoning that connects
them to what the user saw, so the next person does not rediscover them one at a
time.

Also records the methodology mistakes, because they cost more than the bugs did:

- Attaching gdb between digits clears the frequency input box: the timeout is
  ~2.5 s and an attach takes ~3 s. Three wrong conclusions came from this, so
  instrument the model and read stderr instead of stopping the guest.
- A diagnostic log capped at six entries showed only 0xFF payloads and supported
  precisely the wrong conclusion. Do not cap before the shape of the data is known.
- A failed ninja leaves the old binary in place and the test still runs. Two rounds
  of results were meaningless. Check for FAILED and error: before trusting a run.
- assets/flash.img is written by every session, so a test starting from it can find
  its work already done -- which looks identical to broken persistence. Start from
  assets/pristine/, and power off before restoring, since shutdown flushes the old
  image back over the file.
- Hand-computed struct offsets gave KEY_LOCK=4 and TX_VFO=11, impossible values,
  because the ELF has no DWARF and the structs contain enums. Use an unambiguous
  nm symbol, or find the field by toggling it and diffing.
- One probe printed phase before incrementing it, making a correct address decoder
  look off by one. A working implementation was nearly "fixed" as a result.

README gains the two new tests in the build-verification step and the tools list.
2026-08-28 15:32:45 +01:00
mckero 798905f154 Give DMA the CPU's address space, so a typed frequency sticks
DMA moved bytes through address_space_memory, which cannot decode this SoC's memory
at all: the container region holding flash, SRAM and the peripherals is handed only
to the ARMv7M core and never registered with global system memory. Reads came back
MEMTX_DECODE_ERROR with all-zero data, and writes went nowhere. Proved directly --
an address_space_read of SRAM through it returns result=2 and 00000000, while the
same address read through the container returns the real contents.

This is what "the frequency will not change" and "flash behaves like RAM" had in
common. PY25Q16_WriteBuffer reads a 4 KB sector into SectorCache, patches the part
it wants, and programs the whole sector back. The read looked healthy from the flash
side -- the model handed over real 0xFF bytes -- but DMA dropped them, so the
write-back sourced 4096 zeros and cleared the sector, VFO frequencies at 0x9000
included. RADIO_ConfigureChannel only substitutes a band's lower limit for
0xFFFFFFFF, so a stored zero was used as-is and clamped to BX4819_band1_lower.
That is where the 18.000 MHz came from, every time.

DMA now runs over an AddressSpace built on the SoC container, and refuses to
transfer at all if none is configured rather than silently moving zeros.

tools/test_freq_entry.py covers the whole user-visible path: type 435000, confirm
435.00000 MHz lands in band 5, confirm no other band was zeroed, and confirm it is
still there after a power cycle. It drives QMP with no debugger attached, because
the input box times out in ~2.5 s and a gdb attach takes longer -- that alone
invalidated several earlier investigations.

Verified: 435 MHz now appears at flash 0x90A0 where before the entire sector read
zero. keypad_test.py, test_flash_persist.py and the 141 unit tests all pass.
2026-08-28 15:30:26 +01:00
mckero f114666b42 Start DMA on the peripheral's request, and clock both directions together
Two related faults in the DMA model, both of which corrupted flash reads.

Transfers started when a channel was enabled. On hardware, enabling only arms a
channel; the transfer begins when the peripheral raises its DMA request. The flash
driver's SPI_ReadBuf arms RX, arms TX, then enables SPI and sets TXDMAEN -- so
firing at arm time clocked the bus before the read command had been sent. SPI now
kicks the armed channels from CR1/CR2 when SPE and a DMA request enable are both
set, which covers the read path (SPE last) and the write path alike.

Each channel also ran to completion independently. SPI is duplex: one clocked byte
is simultaneously sent and received, and the driver relies on that, pairing a
memory-to-peripheral channel feeding dummy bytes with a peripheral-to-memory
channel collecting the reply. Running them in sequence meant TX clocked the whole
transfer out before RX looked at the bus, so RX collected nothing. They are now
stepped together, one byte at a time.

Either fault alone made a 4 KB sector read return zeros. PY25Q16_WriteBuffer reads
a sector into SectorCache, patches it, and writes the whole thing back, so a zeroed
read turned into a zeroed sector -- including the per-band VFO frequencies at
0x9000. That is a second, independent cause of typed frequencies reverting to
18 MHz, on top of the missing page wrap fixed in da1ad7e.

Verified: the frequency area stays 0xFF across a boot where it was previously
zeroed every time, and an instrumented build shows the sector read now returning
0xFF rather than 0x00.

test_flash_persist.py now starts from the pristine image rather than
assets/flash.img. A dirty live image left nothing for the boot to write, which
surfaced as "the image is byte-identical" -- a failure that looks like broken
persistence but is really a dirty fixture. Runs twice in a row cleanly now.

keypad_test.py and the 141 unit tests still pass.
2026-08-28 15:05:15 +01:00
mckero da1ad7e1e6 Wrap page-program writes within their 256-byte page
Real SPI NOR latches only the low address bits into its page buffer, so a program
burst that runs past the page boundary continues at the start of the same page. The
model incremented the address straight through instead.

Consequence: the firmware issues a 512-byte burst at 0x008F00 inside a single CS
assertion (measured -- the CS never drops mid-burst), which spilled into 0x009000.
That is the per-band VFO frequency area in eeprom_compat.c's map, so stored
frequencies were zeroed. RADIO_ConfigureChannel only substitutes the band's lower
limit when it reads 0xFFFFFFFF, so a stored 0 was taken literally and clamped to
BX4819_band1_lower -- which is why every typed frequency reverted to 18.000 MHz.

Verified: writes now align to sectors (0x8000-0x9000 and 0xA000-0xB000) and an
instrumented build records zero stores into 0x9000-0x90D6, where before it was
overwritten on every boot.

test_flash_persist.py had encoded the bug in its expectations: it watched 0x008100
and 0x00A100, which were only ever written *because* of the missing wrap. Those move
to 0x008000/0x00A000, and a MUST_NOT_CHANGE guard on the frequency area now fails if
a write spills there again.

keypad_test.py still passes.
2026-08-28 14:41:20 +01:00
mckero 0b879227e5 Do not mistake a leftover socket file for a running emulator
Power on returned HTTP 500 with a ConnectionRefusedError traceback. A unix socket
file outlives the process that created it, so a killed QEMU left
/tmp/uvk5-qmp.sock behind; wait_for_socket only checked os.path.exists, returned
immediately, and the connect then failed. It now probes with a real connect, which
distinguishes "listening" from "leftover file".

Two related hardenings:

- power_on cleans up if connecting fails. Otherwise a half-started QEMU keeps
  running untracked, holds the socket, and blocks the next power on -- which is
  how one stale socket turned into a repeatable failure.

- The route reports a failed power action as 503 with the reason, instead of a 500
  and a traceback the browser cannot display.

Three tests cover the stale socket, a real listener, and a path that never appears.
2026-08-28 14:17:50 +01:00
mckero 23385d2eba Persist flash writes to the backing file
The SPI NOR model read its image at realize and never wrote back, so the "flash"
was a g_malloc buffer: everything the firmware saved -- settings, edited
frequencies, channel data -- vanished when the QEMU process exited. That is the
"it behaves like RAM" the user reported, and the image on disk still had its
original mtime and was byte-identical to what make_flash.py produces.

Page-program and sector-erase now mark the image dirty, and it is written out when
chip select is released. Flushing there rather than per byte means one file write
per settings save instead of thousands, because the firmware's driver holds CS for
a whole erase-and-program sequence.

Written via a temporary file and rename: an interrupted flush must not leave a
truncated image, since that file is the only copy of the radio's state. A short
write keeps the previous image rather than replacing it with a partial one.

Also flushes from an exit notifier. Deselect covers normal operation, but QMP quit
-- which is what the web UI's power off sends -- can arrive with the chip still
selected, and the last write would be dropped.

tools/test_flash_persist.py covers it end to end on a copy of the image, so it
cannot disturb a running session. It asserts specific regions rather than just a
changed hash: flash 0x008100 (MR/VFO attributes) and 0x00A100 (settings), which are
the two sectors a boot demonstrably writes. Both were 0xFF before and non-0xFF
after, and sha256 moved from 933d6974 to 7cdff6ce.

Three approaches were tried and abandoned first, all for the same reason: driving
the guest from gdb. `call EEPROM_WriteBuffer` and `call SETTINGS_SaveSettings`
both hang, because the main loop is running and the called function waits on
hardware the debugger has frozen. Observing the file is simpler and closer to what
the user actually sees. An earlier version of the test also watched EEPROM offsets
instead of flash offsets and reported "same" for every region while persistence was
in fact working -- the two address spaces are related by the table in
App/driver/eeprom_compat.c, not equal.

keypad_test.py still passes, which matters because this file is where deleting
three fprintfs once silently removed the keypad.
2026-08-28 12:52:06 +01:00
mckero 3e743152c1 Document how the machine boots
There is no bootloader, kernel, partition table or filesystem here, and that is not
obvious from the code -- someone arriving with hosted-OS assumptions will look for
layers that do not exist.

Covers the parts that are easy to get wrong rather than restating the source:

- The reset handoff is two words. SP from 0x08002800, PC from +0x04. Verified
  against the image: 00400020 492d0008 is SP 0x20004000 (top of the 16 KB SRAM,
  matching PY32_SRAM_BASE + PY32_SRAM_SIZE) and PC 0x08002d49, which is also the
  ELF entry and the Reset_Handler symbol. The odd address is Thumb-state bit 0.

- PY32_APP_OFFSET 0x2800 is load-bearing. The first 10 KB of flash is the factory
  bootloader region, so armv7m_load_kernel gets that offset; without it the vector
  table lands in the wrong place and the first fetch faults.

- Startup is 31 lines of assembly, and the .data copy plus .bss zero-fill are the
  part worth understanding: on a hosted OS the kernel and loader do that, here
  nobody does, so a fault in either loop shows up as globals that are silently
  garbage rather than as a crash.

- SETTINGS_InitEEPROM reading fixed SPI offsets is the closest thing to mounting a
  partition. No metadata and no checksum, just an address both sides must agree on,
  so a setting that reads back wrong points at the offset before the transport.

Also notes that the ~15 s to the main loop is emulation overhead; a real radio is
up in about a second.
2026-08-28 12:38:49 +01:00
mckero 1364d46e97 Keep a pristine copy of the flash image, and a way back to it
assets/flash.img was gitignored, so the only copy of the never-booted image lived
on one disk. The emulator writes to that image, so a session can leave edited
settings or a damaged EEPROM behind with nothing to restore from.

assets/pristine/ now holds the image as first generated, gzipped and checksummed,
and is tracked deliberately. Gzip takes it from 2 MiB to 2.3 KiB because the image
is nearly all 0xFF, which is what makes keeping it in git reasonable. The live
image and its .bak-* files stay ignored.

Two checksums are recorded, for the archive and for its contents, so a corrupted
archive is distinguishable from one that was replaced.

tools/restore_flash.sh verifies, diffs, or restores. Restore backs up the current
image first, then re-checks the result, since a restore that silently half-worked
would be worse than none.

Verified by deliberately corrupting the live image: --diff reported 32 differing
bytes, restore backed up and rewrote it, and --diff then reported no change. The
image is currently byte-identical to what make_flash.py produces, so this is the
genuine original rather than a copy of something already used.
2026-08-28 12:20:13 +01:00
mckero 19997e85e1 Rename the vhost to k6v3.mckero.dn42
The name is what the user is putting in DNS. server_name has to match or SNI falls
through to another vhost on the same socket.

Addresses are unchanged: 172.21.91.140 and fd3c:3f9b:6424:2::5, still sharing 443
under the existing *.mckero.dn42 wildcard.

Verified after reload: 200 on both families with the certificate validating, 80
redirecting, and dns./mail. still 200. This time nginx -t ran after the symlink was
in place, which is the ordering that caught me out last time.
2026-08-28 11:49:51 +01:00
mckero a4a8f21d50 Document and version-control the HTTPS front end
The UI is served at https://k6v6.mckero.dn42/ with nginx terminating TLS and the
server itself now bound to loopback, so it is not directly reachable.

No new address and no new certificate: 443 is shared with the other vhosts on
these DN42 addresses and separated by SNI, and the existing *.mckero.dn42 wildcard
already covers the name. Only DN42 addresses are bound, so the public 443
listeners on this host are untouched.

docs/reverse-proxy.md records the settings that are not optional, because each has
a failure mode that is easy to misread:
  proxy_buffering off  -- otherwise the frame stream arrives in bursts
  X-Forwarded-For      -- otherwise every log line is attributed to 127.0.0.1
  long read timeout    -- a paused guest emits nothing at all
  tcp_nodelay          -- Nagle would delay exactly the latency-critical requests

Two pitfalls hit while setting it up are written down. "http2 on;" needs nginx
1.25.1+ and this host runs 1.22.1, and because nginx -t was run before the symlink
existed it passed, then reload failed and left nginx stopped, briefly taking the
other sites down. Separately, a newly added listen address needs a reload to be
bound: after the failed reload, v6 requests failed with nothing in the error log
until a second reload created the socket.

deploy/nginx-k6v6.conf keeps a copy in the repo, since nothing here
version-controls /etc.

Verified: HTTP 200 on both families with the certificate validating (no -k), 7
frames in a 20 KB stream sample, log entries attributed to real client addresses.
2026-08-28 11:39:17 +01:00
mckero f2c5c6b31b Attribute log entries to the client IP
Entries gain an "ip" field, rendered between the time and the source as asked.
The buffer is shared by every viewer, so without attribution a log of keypresses
from two people is unreadable.

Resolving the address matters more than it looks: behind the nginx reverse proxy
REMOTE_ADDR is always 127.0.0.1, so the first hop of X-Forwarded-For is what
identifies the real client. Only the first entry is trusted -- the rest of the
chain is set by the caller and a test covers that.

Power actions are logged at the route rather than in the supervisor, which has no
request context, so "who powered it off" is recorded.

Entries with no client behind them keep ip=None and render as "-": firmware serial
and QEMU stderr are not caused by a request.

The sharing and history the user asked for already worked and needed no change --
verified rather than assumed. The front end starts at logCursor=0, so a page
opened now receives the full buffer, including lines produced before it connected
and lines from other people. Confirmed live through the proxy: a new reader saw
entries attributed to 172.21.91.140, fd3c:3f9b:6424:2::5 and "-".
2026-08-28 11:37:06 +01:00
mckero dbe7607720 Go back to measuring the press instead of guessing it
Reverts optimistic send. The browser times the press with performance.now() and
sends it once on release, so the firmware sees exactly the press that was made.

Optimistic send fired a speculative tap at pointerdown plus a held press if the
button was still down. It was 152 ms faster (505 vs 657 ms click-to-visible at
400 ms RTT) but it guessed, and a wrong guess sent both presses for the firmware
to act on. Raising the threshold to 900 ms hid the symptom without removing the
failure mode, and it also made hold-to-repeat unreachable, since the server
released the key after a fixed 900 ms no matter how long you held.

Measuring costs the click duration in latency and buys exactness plus real
hold-to-repeat. Verified against the firmware:

  taps under 400 ms   120/250/390 ms  -> cursor +1, submenu never opens
  holds from 400 ms   500/900/1500 ms -> cursor +3/+8/+15
  MENU tap            120/300/390 ms  -> menu opens, submenu stays shut

One correction to my own expectations along the way: I first recorded the multi-step
moves at 500 and 800 ms as failures. They are not. App/misc.c has
key_repeat_10ms = 8, so past 400 ms the firmware auto-repeats every 80 ms, and the
counts match (duration - 400) / 80. That is what a real radio does when you hold a
button, so the note on FIRMWARE_HELD_MS now says not to filter it out.

MIN_HOLD_MS returns as the floor for a measured press, since a very fast click can
measure below the debounce window. LONG_PRESS_AFTER_MS and LONG_PRESS_MS are gone
with the scheme that needed them.
2026-08-28 10:56:29 +01:00
mckero 46f0f37d58 Log keypresses, keep frames flowing when idle, blank the screen when off
Three things reported from actual use, plus the bug the logging exposed.

1. Keys are logged (source "key"), including refusals and presses while powered
   off. Without this there was no way to tell "the key never arrived" from "the
   firmware did something else with it" -- which is exactly what was needed below.

2. /stream resends the current frame every IDLE_FRAME_INTERVAL_S even when nothing
   changed. Change-detection alone made a static screen indistinguishable from a
   dead connection, and a client joining mid-idle stayed blank. Measured: 6 frames
   in 10 idle seconds, where before it was 0.

3. Off now actually blanks the screen. The dark panel moved to a .screenwrap
   wrapper and the <img> is hidden; setting a background on the <img> alone did
   nothing visible, because the image kept painting the last frame over it.

Then the reported bug: in the menu, UP/DOWN behaved like another MENU press. The
key log made it diagnosable and the cause was mine -- LONG_PRESS_AFTER_MS was set
to the firmware's own 400 ms boundary, but a deliberate click runs 100-500 ms, so
ordinary clicks sent tap AND held and the firmware acted on both:

  held DOWN  auto-repeated, gMenuCursor 3 -> 12 from one click
  held MENU  entered the submenu, gIsInSubMenu 0 -> 1

The UI threshold is now 900 ms, well clear of any click, and FIRMWARE_HELD_MS is a
separate constant so the two are not conflated again. Verified against the real
firmware: 120/300/500/800 ms clicks each move the cursor exactly +1 with
submenu=0, while a deliberate 1400 ms hold still auto-repeats (+9).
2026-08-28 10:38:58 +01:00
mckero c164917632 Log power events, QEMU stderr and firmware serial; show them in a fixed pane
Three sources into one buffer: power events from the supervisor, QEMU's stderr
(which run.sh and the tests used to discard), and firmware serial, which the
machine model tags SERIAL. default_launcher now captures stderr rather than
sending it to DEVNULL, which is what made the last two reachable.

The pane is a fixed-height scroll box as asked: 180px with overflow-y:auto, so it
never grows with content -- older lines move up out of view and you scroll back to
read them.

Two details that make that usable rather than annoying:

- Autoscroll only sticks when you are already at the bottom. Otherwise a new line
  arriving would yank the view away from whatever you had scrolled up to read.
- MAX_LOG_LINES caps the <pre> as well. The box is fixed-height either way, but an
  unbounded DOM node would still grow memory across a long session.

Verified on the live server: power on produced power/qemu/serial lines including
"UV-K5 Firmware, EGZUMER-F4HWN+NR7Y c91cec95", each Reset logs the event and the
banner reappearing, and since= never resent a line. A capacity-500 buffer fed 2000
lines keeps exactly 500 and does not replay evicted entries to a stale cursor.
2026-08-28 10:06:18 +01:00
mckero 7db89231b6 Add a bounded log buffer
Collects power events, QEMU stderr, and firmware serial output. Bounded and in
memory on purpose: an unbounded buffer in a long-running server is a slow leak,
and anyone wanting a permanent record can redirect the server's stderr.

Entries carry a monotonic seq so a polling client can ask for "anything after N"
and receive each line exactly once, including after eviction has dropped older
entries -- a test covers that case specifically, since an index-based cursor would
silently repeat or skip lines there.

Stream decoding is lenient: serial bytes can be garbage before the firmware
configures the port, and losing the stream to one bad byte would be worse than a
replacement character.
2026-08-28 09:59:34 +01:00
mckero 20e7577257 Cut key latency: fire at pointerdown and shorten the tap
Two changes, both aimed only at latency, since that is what the link makes
expensive.

1. Send on pointerdown instead of pointerup. Waiting for release left the network
   idle for the entire click. The duration is unknown at that moment, so the
   speculative request asks for a short press; holding past LONG_PRESS_AFTER_MS
   sends a second, deliberately long press, which is how held events stay
   reachable.

2. TAP_MS 200 -> 60 ms. The server blocks for hold_ms before replying, so this is
   latency the user pays directly. 200 ms was a guess that gave back most of what
   change 1 saved.

The 60 ms is measured, and the sample size mattered: at 4 trials per value 30 ms
looked reliable, but at 12 trials 20 ms registered only 5/12 while 30 ms was 12/12.
The nominal 20 ms debounce is not sufficient alone because KEYBOARD_Poll samples
each column 8 times wanting 2 matching reads. 60 ms is double the proven floor.

Click-to-visible at 400 ms RTT, where one round trip is an unavoidable 400 ms:

  original (2 requests, on release)   2525 ms   (+2125 over the floor)
  one request, on release              657 ms   (+257)
  current (1 request, on press)        505 ms   (+105)

Long press still works: a 60 ms press opens the menu (gScreenToDisplay 0 -> 1)
while a 900 ms press is treated as held and correctly does not, so the firmware
still distinguishes them.

Also drops MIN_HOLD_MS, now dead: the browser no longer measures press duration,
so there is no measurement to clamp. Its test asserted only that the string
appeared, which would have kept passing over dead code.
2026-08-28 09:56:04 +01:00
mckero c6d58f5601 Surface firmware serial output instead of dropping it
The firmware has been printing all along and nothing was listening. USART1 has no
real model here -- it is one of the logging catch-all stubs -- so every byte went
into qemu_log_mask(LOG_UNIMP) and vanished.

Two things were needed, and the second was not in the plan:

1. Print USART1 DR writes (+0x04, per the vendor CMSIS header) as SERIAL lines.

2. Report TXE|TC in USART1 SR. This is the part I had missed. UART_Send() in
   App/driver/uart.c spins on LL_USART_IsActiveFlag_TXE() with a bounded timeout
   and *skips the byte* when the flag never sets. A stub returning 0 for SR meant
   the firmware discarded its own output before it ever reached DR -- the only
   write arriving was UART_Init()'s priming zero. So step 1 alone produced
   nothing, which is why the first attempt looked like "the build has no logging".

Then a bug of my own: the priming byte is 0x00, I buffered it, and fprintf("%s")
stopped at that NUL and printed an empty line while all 46 bytes sat behind it.
NULs are now dropped, and a line flushes on CR as well as LF.

Verified: SERIAL UV-K5 Firmware, EGZUMER-F4HWN+NR7Y c91cec95
keypad_test.py still passes, which matters because this file is where removing
three fprintfs once silently deleted the keypad.
2026-08-28 06:27:46 +01:00
mckero 5b53cbbc82 Start powered off; the user presses On
The server now owns the QEMU process by default but does not launch it. You open
the page to a dark screen and press On, which is the behaviour asked for: like
walking up to a machine rather than finding it already booted.

This inverts the flag from the plan. Owning the process has to be the default,
since it is the only way On/Off can work at all; --attach is the opt-in for joining
a run.sh instance, where Off is refused.

Two tests guard the intent rather than the wiring: one asserts main() has --attach
and not --own-emulator, another asserts main() never calls power_on(), so a future
edit cannot quietly restore auto-boot.

Verified on the live server with no QEMU running beforehand:
  startup   0 QEMU processes, powered=false, frame.png 503, page 200
  On        1 QEMU process, powered=true, frame.png 200 (2920 bytes)
  Off       0 QEMU processes, frame.png 503, page still 200, keys 409
  On again  1 QEMU process, powered=true, frame.png 200

The web server stays up across Off, which is what you asked for: the screen goes
dark and waits for the next person to press On.
2026-08-28 06:09:56 +01:00
mckero 6a9cf9f28a Add an On/Off/Reset bar above the screen
Buttons sit above the LCD as asked, with a state label that turns green when the
guest is up. Powered off dims the panel via a screen-off class, so a dark screen
is the signal rather than a frozen last frame.

Off asks for confirmation: it ends the guest, and a stray click should not do that
silently. Buttons disable while a power action is in flight, since On and Off take
a couple of seconds and a double click would race.

The stream is restarted after every power action. The old multipart response ends
when the emulator goes away, so without a fresh src the image would stay blank
after On.

Refusals are surfaced in the status line rather than swallowed -- that is how the
409 for an adopted emulator becomes visible instead of looking like a dead button.

49 unit tests, and the generated page passes node --check.
2026-08-28 06:02:30 +01:00
mckero 699561794a Add power endpoints and survive the emulator being off
POST /api/power/{on,off,reset,pause,resume}, and every route now tolerates there
being no emulator: /api/status reports powered:false, /frame.png returns 503,
keypresses return 409 with "press On first". Previously create_app required a live
client and the whole page would 500.

The client is fetched through the supervisor per request rather than captured once,
because a power cycle replaces it and the captured one goes stale.

After any power action the pump is rebound, so Off actually goes dark instead of
freezing on the last frame.

Off is refused with 409 for an adopted emulator: we did not start that process.
Reset is allowed either way, since system_reset ends nothing.

Verified over HTTP with a real supervisor and real QEMU:
  start     powered=False  qemu=0  frame=503
  key       409 as expected
  On        powered=True   qemu=1  frame=2920 bytes
  Reset     qemu=1 (process survived)
  Off       powered=False  qemu=0  frame=503
  On again  powered=True   qemu=1  frame=2920 bytes
The web server stayed up throughout, which is the requested behaviour.
2026-08-28 06:00:50 +01:00
mckero 71d72c2393 Add a supervisor that owns the QEMU process
Power on/off cannot live inside the QMP connection: QMP quit destroys the socket
a later power on would have to arrive through. So something outside it has to be
able to spawn the process again.

Off then On is a cold boot -- process replaced, guest from reset -- which is the
behaviour asked for: like cutting mains power and restoring it. system_reset is the
warm alternative and keeps the process.

adopt() is for attaching to a run.sh instance. power_off then refuses, because we
did not start that process. is_running() also reports False for a process that
exited on its own, rather than trusting our own bookkeeping.

power_off tolerates quit raising: the socket usually drops before the reply
arrives, so that is success rather than an error. The launcher clears a stale
socket first, since QEMU failing to bind presents as On doing nothing.

Verified against real QEMU: starts with 0 processes, On gives 1, Reset keeps the
same one, Off returns to 0, On again cold boots. 14 unit tests with a fake
launcher, so they need no emulator.
2026-08-28 05:57:22 +01:00
mckero b500697224 Serve frames from the pump instead of reading QMP per request
/stream and /frame.png now read the pump's shared buffer. A slow client falls
behind in frames rather than in QMP reads, and reconnects no longer multiply the
load on the emulator.

/frame.png returns 503 when there is no frame rather than raising, because that is
a real state: the emulator can be powered off and the page still has to load. The
same reason create_app now tolerates client=None.

Measured on the live server with 4 concurrent streams, which is what a reconnecting
browser produces: keypress latency went from 6 ms avg to 5 ms, so -1 ms, i.e. noise.
All four clients received the same ~27 frames. /stream first byte in 4 ms.

Note this rules out my earlier guess: I had assumed concurrent streams were
starving keypresses, and the numbers said otherwise both before and after. The
pump is worth having for constant QMP load, not because contention was the
slowness.
2026-08-28 05:55:03 +01:00
mckero b75ef5fdcd Add a background frame pump so QMP load is constant
One thread grabs the LCD at a fixed rate into a shared buffer; HTTP clients serve
from that buffer. Before, every /stream iteration issued its own QMP reads, so
load scaled with client count and reconnects, and a slow reader could stall the
grab loop.

Encodes only when the framebuffer bytes change, and exposes generation so a client
can tell a new frame from the same one without comparing bytes.

rebind() is here for the power work in a later task: the emulator comes and goes
under the server, and rebind(None) blanks the screen instead of leaving a stale
frame that looks live. The run loop re-checks the grabber identity under the lock
after a read, so a rebind landing mid-read cannot be undone by the frame it was
already fetching.

A client that raises is swallowed on purpose -- a dead emulator must not kill the
pump, since power may come back.

Verified against a real emulator: 5000 latest() calls in 1 ms with no QMP traffic,
generation flat at 1 over a second of static screen, guest still running, and
rebind(None) going dark. 9 unit tests.
2026-08-28 05:50:16 +01:00
mckero 7fe605eeaf Send one request per keypress carrying the measured duration
The browser now times the press with performance.now() and posts hold_ms once,
instead of sending down and up as two requests. Halves the round trips per key and
makes press duration independent of the link.

Deletes test_sends_down_and_up_not_just_tap: it asserted the behaviour being
replaced, so keeping it would have meant asserting the bug.

MIN_HOLD_MS is injected into the page from webui.py so the two agree on the floor,
which exists because anything under the firmware's 20 ms debounce does not register
at all -- a very fast click still has to ask for 60 ms.

Verified over a simulated 400 ms link: three keys in 1.58 s where the old design
needed ~2.45 s, and a 120 ms press opens the menu (gScreenToDisplay 0 -> 1) with
DOWN then moving the cursor. The generated page also passes node --check.
2026-08-28 05:41:59 +01:00
mckero e7f3f9da76 Hold keys for a server-side duration, not a network round trip
POST /api/key now accepts {"key": "MENU", "hold_ms": 120} and holds the key for
exactly that long, locally.

The old design sent down and up as two requests so the browser would own press
duration. That is correct on loopback and broken over a real link: the round trip
between the two requests *is* the press duration. Measured against this server at
400 ms RTT, an intended tap arrived as a 407 ms hold, and since the firmware reads
400 ms as held (key_repeat_delay_10ms = 40), every short press was dispatched to
the hold path where MAIN_Key_MENU does nothing. Jitter either side of that
threshold is why it felt intermittent rather than simply broken.

Verified at 400 ms simulated RTT:
  one request  531 ms total, firmware saw 120 ms -> short press
  two requests 816 ms total, firmware saw 408 ms -> held (the bug)

hold_ms is clamped to MAX_HOLD_MS, rejects negatives and non-numbers, and defaults
to TAP_MS. A deliberate 900 ms hold is preserved, so long-press events still work.
down and up stay for scripting on a fast link.
2026-08-28 05:35:37 +01:00
mckero 2c3a602d9f Capture the DN42 firewall rules as a script
The rules restricting port 8080 to DN42 were applied by hand and existed only in
the live kernel tables -- lost on reboot, with the port then wide open and nothing
in the repo to say it ever had been restricted.

The script is idempotent (checks before adding) and has remove/show. show prints
packet counters, which is the part that matters: this host's INPUT policy is
ACCEPT, so a rule that only allows DN42 does nothing at all. The final DROP is
what restricts anything, and a rising DROP count is the only proof it works
rather than the traffic simply not arriving.

Currently observed: 2203 accepted from DN42 v4, 103 from v6, 22 dropped.
2026-08-28 05:32:31 +01:00
mckero 2d0fc24b62 Document the web remote control
README gets a section with the endpoint table, the two constraints that will
otherwise surprise someone (single QMP client, no authentication), and the reason
frames go through memsave rather than pmemsave or gdb.

AGENTS.md gets the run instructions plus a new entry under 'Things that already
went wrong' for the pmemsave trap: it takes a physical address, returns zeros for
gFrameBuffer, and reports success. The web UI was built on it initially because a
benchmark showed it was fast -- the benchmark never checked the contents. Worth
recording as the general lesson, not just the specific fix.
2026-08-28 04:49:32 +01:00
mckero db1117d61f Add an end-to-end test for the web remote control
Boots its own emulator on a private socket and port, so it neither disturbs a
run.sh session nor fights it for the single-client QMP socket.

The unit tests stub QMP and therefore cannot catch a wrong QMP command. That is
not hypothetical: pmemsave and memsave differ only in physical versus virtual
addressing, and pmemsave fails by silently returning zeros. Check 2 asserts the
frame has lit pixels for exactly that reason.

Six checks, all passing: PNG renders, screen has content (1883 lit pixels), a
MENU tap changes the screen, the guest is still running after streaming, PTT is
rejected with 400, and /stream is multipart with a PNG part.

Waits 16 s before the first press, past power-save entry, so it also covers keys
working once the radio is asleep.
2026-08-28 04:47:16 +01:00
mckero b4f3b28a39 Add the web remote control: live LCD plus clickable keypad
Flask app on the existing QMP socket. GET / serves the page, GET /stream is a
multipart PNG sequence at up to 15 fps, POST /api/key drives the keypad model.

The keypad takes discrete down/up rather than a fixed-duration tap, because the
firmware reads anything past 400 ms as a held key and dispatches it differently.
Letting the browser own the timing is what makes both short and long presses
reachable; tap stays available for scripting and is pinned between the 20 ms
debounce and the 400 ms hold, with a test asserting that range.

The stream only re-encodes when the framebuffer bytes change. The LCD is static
most of the time, so idle CPU stays near zero.

Unknown key names are rejected with 400 before reaching QMP, so a PTT button
cannot be added by accident. The front end releases on pointerleave,
pointercancel and blur, and /api/release-all is the safety valve for a key that
somehow stays down.

Verified against a live emulator over HTTP: /frame.png renders the dual-VFO main
screen, a MENU tap opens the menu at 01/79, and two DOWN presses reach 03/79.
19 unit tests, 41 across the whole suite.
2026-08-28 04:45:51 +01:00
mckero 6d19a33f45 Add key-name validation mirroring the keypad model
Rejects unknown names in the server rather than letting them reach QMP as an
error. A test parses keypad_key_names out of qemu/py32f071.c and fails if the
lists drift, so adding a key to the model cannot silently leave the web UI
behind.

PTT is deliberately absent: the model wires it separately on GPIOC, so the press
property rejects the name.
2026-08-28 04:37:32 +01:00
mckero 4c820bdf79 Add a frame grabber that reads the LCD over QMP memsave
memsave, not pmemsave. The plan specified pmemsave on the strength of a timing
measurement that never checked the contents; it turns out pmemsave takes a
*physical* address and silently returns zeros for gFrameBuffer's virtual
address. No error, no warning -- just a permanently blank screen.

Caught by rendering a real frame and finding 0 lit pixels where the gdb path
reported 1693. With memsave the count matches exactly, and the image reads
correctly: both VFOs at 18.00000 MHz, PS/DWR/CL status bar.

Two tests guard the decisions rather than the current text: the stub client
raises if pmemsave is ever used, and a source check rejects subprocess/popen so
frame reads cannot regress onto gdb, which would halt the guest.

Measured 1.35 ms per frame with the guest still reporting status running.
2026-08-28 04:36:30 +01:00
mckero 1906e110f1 Extract LCD unpacking and PNG encoding into a module
Lifted from tools/screenshot.py so the web UI cannot drift from the CLI
screenshotter. Verified unpack() is identical to the original across 200
randomised trials, and the PNG IHDR matches byte for byte.

encode_png() returns bytes instead of writing a file, and drops to zlib level 6:
at streaming rates the CPU saving beats the last few bytes on loopback.
2026-08-28 04:30:45 +01:00
mckero 59ca82403f Add a QMP client module for the web UI
Wraps the connect/negotiate/command dance with a lock, since only one client
can hold the QMP socket at a time. Events interleave with replies, so command()
reads until it sees the matching return rather than trusting the first message.

Connection failure carries the reason and the likely cause: key.py already
holding the socket.
2026-08-28 04:29:11 +01:00
mckero b246132a35 Bring the docs in line with both keypad fixes
AGENTS.md still carried a stale entry telling the reader to *lengthen* key
holds when a press seems ignored, which is the opposite of the fix and is
what broke the tooling in the first place. Replaced with the correction and
a pointer to the right section.

The keypad heading also claimed the hold time was the only cause. There were
two: the 2500 ms hold in key.py, and row_out missing volatile. Both are now
listed up front with a link to the detail.

Adds the regression test to the places someone would actually look: the
"How to run it" section in AGENTS.md, the layout listing, and a build step
in the README noting that a clean build is not evidence the keypad works,
since the -O2 dead-code elimination produces no warning.
2026-08-28 03:58:19 +01:00
mckero 2154f80414 Fix the keypad: row_out must be volatile
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.

Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.

From the object code:

  callers reaching keypad_update_rows
    plain     none -- the calls are gone
    volatile  keypad_key_changed, keypad_col_changed, keypad_set_press,
              keypad_reset, uvk5_machine_init

keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.

Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.

Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.

Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).

README screenshots regenerated from the current build.
2026-08-27 18:34:41 +01:00
mckero 1c9a2afe55 Fix key.py hold times; the keypad was never broken
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.

The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.

2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.

Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.

Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
  buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
  the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
  keypad scan and the model does not wake from it. Includes the two dead
  ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
2026-08-27 16:45:54 +01:00
mckero c012a2a6f0 Add Apache 2.0 licence
Downloaded verbatim from apache.org, not transcribed.

qemu/py32f071.c stays GPL-2.0-or-later as its header says: it builds into
QEMU and derives from QEMU's device models, so it cannot be relicensed.
README notes the split.
2026-08-27 16:44:11 +01:00
mckero 9667fef147 Add AGENTS.md
Notes for whoever works on this next, weighted toward what the code does not
say: that the firmware is the reference and must never be edited to suit the
emulator, that register layouts come from the vendor CMSIS header rather than
inference, and that the way to find the next peripheral worth modelling is to
watch where the firmware stops.

Records the mistakes that already cost time here, each with the symptom that
made it look like something else: GDB breakpoints halting the guest (which reads
as 'the keypress does nothing'), writing the SysTick counter back while
accelerating it (which hangs the delay loop outright), lowering the clock to
speed up busy-waits (measured, 32x, nowhere near enough), unnamed qdev GPIO lines
sharing one namespace, and a probe script whose own regex silently matched
nothing.

Also states plainly what the emulator cannot answer, so a passing test is not
mistaken for evidence about radio behaviour.
2026-08-27 15:18:08 +01:00
mckero c0a09827ed UV-K5 V3 emulator: QEMU machine for the PY32F071
Adds a QEMU machine for the Puya PY32F071 (Cortex-M0+) so Quansheng UV-K5 V3
firmware can run on a PC. The firmware boots to its main loop in about five
seconds and the LCD contents are readable.

Register layouts come from the vendor CMSIS header shipped with the firmware
rather than guesswork. Modelled: RCC, GPIO, ADC, both SPI controllers, DMA1 and
the PY25Q16 flash; everything else answers through a logging catch-all, which is
how the next thing worth modelling gets identified.

Seven things had to be right before it would boot, each found by watching where
the firmware stopped: flash aliased at the application offset, clock ready bits,
self-clearing ADC calibration, SPI transfer flags, DMA-driven flash reads,
SysTick poll acceleration, and the bit-banged transceiver bus idling low.

SysTick needs explanation. SYSTICK_DelayUs polls the counter and accumulates
differences; under emulation a register read costs far more relative to guest
time, so a measured 120 ms delay would have taken about 7.7 hours. Lowering the
clock does not help because the bottleneck is loop iterations, not counter speed.
Reporting a value that runs ahead of the real counter does, via a new poll-boost
property on SysTick. Guest time therefore runs fast during delays: fine for
exercising menus and control flow, wrong for judging signal timing.

Also includes the host build of the CW timing chain (harness, stubs, shim,
tests), which compiles app/cwkeyer.c and app/cwmacro.c unmodified against stub
drivers with a virtual clock and scripted paddle input.

Known gap: keypad rows reach the firmware's scan and KEYBOARD_Poll returns the
right key code, but the UI does not react yet.

Not modelled, and not intended to be: radio behaviour. The transceiver chip has
no public datasheet, so keying envelopes and emissions need real hardware.
2026-08-27 14:59:21 +01:00