POST /api/power/{on,off,reset,pause,resume}, and every route now tolerates there
being no emulator: /api/status reports powered:false, /frame.png returns 503,
keypresses return 409 with "press On first". Previously create_app required a live
client and the whole page would 500.
The client is fetched through the supervisor per request rather than captured once,
because a power cycle replaces it and the captured one goes stale.
After any power action the pump is rebound, so Off actually goes dark instead of
freezing on the last frame.
Off is refused with 409 for an adopted emulator: we did not start that process.
Reset is allowed either way, since system_reset ends nothing.
Verified over HTTP with a real supervisor and real QEMU:
start powered=False qemu=0 frame=503
key 409 as expected
On powered=True qemu=1 frame=2920 bytes
Reset qemu=1 (process survived)
Off powered=False qemu=0 frame=503
On again powered=True qemu=1 frame=2920 bytes
The web server stayed up throughout, which is the requested behaviour.
Power on/off cannot live inside the QMP connection: QMP quit destroys the socket
a later power on would have to arrive through. So something outside it has to be
able to spawn the process again.
Off then On is a cold boot -- process replaced, guest from reset -- which is the
behaviour asked for: like cutting mains power and restoring it. system_reset is the
warm alternative and keeps the process.
adopt() is for attaching to a run.sh instance. power_off then refuses, because we
did not start that process. is_running() also reports False for a process that
exited on its own, rather than trusting our own bookkeeping.
power_off tolerates quit raising: the socket usually drops before the reply
arrives, so that is success rather than an error. The launcher clears a stale
socket first, since QEMU failing to bind presents as On doing nothing.
Verified against real QEMU: starts with 0 processes, On gives 1, Reset keeps the
same one, Off returns to 0, On again cold boots. 14 unit tests with a fake
launcher, so they need no emulator.
/stream and /frame.png now read the pump's shared buffer. A slow client falls
behind in frames rather than in QMP reads, and reconnects no longer multiply the
load on the emulator.
/frame.png returns 503 when there is no frame rather than raising, because that is
a real state: the emulator can be powered off and the page still has to load. The
same reason create_app now tolerates client=None.
Measured on the live server with 4 concurrent streams, which is what a reconnecting
browser produces: keypress latency went from 6 ms avg to 5 ms, so -1 ms, i.e. noise.
All four clients received the same ~27 frames. /stream first byte in 4 ms.
Note this rules out my earlier guess: I had assumed concurrent streams were
starving keypresses, and the numbers said otherwise both before and after. The
pump is worth having for constant QMP load, not because contention was the
slowness.
One thread grabs the LCD at a fixed rate into a shared buffer; HTTP clients serve
from that buffer. Before, every /stream iteration issued its own QMP reads, so
load scaled with client count and reconnects, and a slow reader could stall the
grab loop.
Encodes only when the framebuffer bytes change, and exposes generation so a client
can tell a new frame from the same one without comparing bytes.
rebind() is here for the power work in a later task: the emulator comes and goes
under the server, and rebind(None) blanks the screen instead of leaving a stale
frame that looks live. The run loop re-checks the grabber identity under the lock
after a read, so a rebind landing mid-read cannot be undone by the frame it was
already fetching.
A client that raises is swallowed on purpose -- a dead emulator must not kill the
pump, since power may come back.
Verified against a real emulator: 5000 latest() calls in 1 ms with no QMP traffic,
generation flat at 1 over a second of static screen, guest still running, and
rebind(None) going dark. 9 unit tests.
The browser now times the press with performance.now() and posts hold_ms once,
instead of sending down and up as two requests. Halves the round trips per key and
makes press duration independent of the link.
Deletes test_sends_down_and_up_not_just_tap: it asserted the behaviour being
replaced, so keeping it would have meant asserting the bug.
MIN_HOLD_MS is injected into the page from webui.py so the two agree on the floor,
which exists because anything under the firmware's 20 ms debounce does not register
at all -- a very fast click still has to ask for 60 ms.
Verified over a simulated 400 ms link: three keys in 1.58 s where the old design
needed ~2.45 s, and a 120 ms press opens the menu (gScreenToDisplay 0 -> 1) with
DOWN then moving the cursor. The generated page also passes node --check.
POST /api/key now accepts {"key": "MENU", "hold_ms": 120} and holds the key for
exactly that long, locally.
The old design sent down and up as two requests so the browser would own press
duration. That is correct on loopback and broken over a real link: the round trip
between the two requests *is* the press duration. Measured against this server at
400 ms RTT, an intended tap arrived as a 407 ms hold, and since the firmware reads
400 ms as held (key_repeat_delay_10ms = 40), every short press was dispatched to
the hold path where MAIN_Key_MENU does nothing. Jitter either side of that
threshold is why it felt intermittent rather than simply broken.
Verified at 400 ms simulated RTT:
one request 531 ms total, firmware saw 120 ms -> short press
two requests 816 ms total, firmware saw 408 ms -> held (the bug)
hold_ms is clamped to MAX_HOLD_MS, rejects negatives and non-numbers, and defaults
to TAP_MS. A deliberate 900 ms hold is preserved, so long-press events still work.
down and up stay for scripting on a fast link.
The rules restricting port 8080 to DN42 were applied by hand and existed only in
the live kernel tables -- lost on reboot, with the port then wide open and nothing
in the repo to say it ever had been restricted.
The script is idempotent (checks before adding) and has remove/show. show prints
packet counters, which is the part that matters: this host's INPUT policy is
ACCEPT, so a rule that only allows DN42 does nothing at all. The final DROP is
what restricts anything, and a rising DROP count is the only proof it works
rather than the traffic simply not arriving.
Currently observed: 2203 accepted from DN42 v4, 103 from v6, 22 dropped.
Boots its own emulator on a private socket and port, so it neither disturbs a
run.sh session nor fights it for the single-client QMP socket.
The unit tests stub QMP and therefore cannot catch a wrong QMP command. That is
not hypothetical: pmemsave and memsave differ only in physical versus virtual
addressing, and pmemsave fails by silently returning zeros. Check 2 asserts the
frame has lit pixels for exactly that reason.
Six checks, all passing: PNG renders, screen has content (1883 lit pixels), a
MENU tap changes the screen, the guest is still running after streaming, PTT is
rejected with 400, and /stream is multipart with a PNG part.
Waits 16 s before the first press, past power-save entry, so it also covers keys
working once the radio is asleep.
Flask app on the existing QMP socket. GET / serves the page, GET /stream is a
multipart PNG sequence at up to 15 fps, POST /api/key drives the keypad model.
The keypad takes discrete down/up rather than a fixed-duration tap, because the
firmware reads anything past 400 ms as a held key and dispatches it differently.
Letting the browser own the timing is what makes both short and long presses
reachable; tap stays available for scripting and is pinned between the 20 ms
debounce and the 400 ms hold, with a test asserting that range.
The stream only re-encodes when the framebuffer bytes change. The LCD is static
most of the time, so idle CPU stays near zero.
Unknown key names are rejected with 400 before reaching QMP, so a PTT button
cannot be added by accident. The front end releases on pointerleave,
pointercancel and blur, and /api/release-all is the safety valve for a key that
somehow stays down.
Verified against a live emulator over HTTP: /frame.png renders the dual-VFO main
screen, a MENU tap opens the menu at 01/79, and two DOWN presses reach 03/79.
19 unit tests, 41 across the whole suite.
Rejects unknown names in the server rather than letting them reach QMP as an
error. A test parses keypad_key_names out of qemu/py32f071.c and fails if the
lists drift, so adding a key to the model cannot silently leave the web UI
behind.
PTT is deliberately absent: the model wires it separately on GPIOC, so the press
property rejects the name.
memsave, not pmemsave. The plan specified pmemsave on the strength of a timing
measurement that never checked the contents; it turns out pmemsave takes a
*physical* address and silently returns zeros for gFrameBuffer's virtual
address. No error, no warning -- just a permanently blank screen.
Caught by rendering a real frame and finding 0 lit pixels where the gdb path
reported 1693. With memsave the count matches exactly, and the image reads
correctly: both VFOs at 18.00000 MHz, PS/DWR/CL status bar.
Two tests guard the decisions rather than the current text: the stub client
raises if pmemsave is ever used, and a source check rejects subprocess/popen so
frame reads cannot regress onto gdb, which would halt the guest.
Measured 1.35 ms per frame with the guest still reporting status running.
Lifted from tools/screenshot.py so the web UI cannot drift from the CLI
screenshotter. Verified unpack() is identical to the original across 200
randomised trials, and the PNG IHDR matches byte for byte.
encode_png() returns bytes instead of writing a file, and drops to zlib level 6:
at streaming rates the CPU saving beats the last few bytes on loopback.
Wraps the connect/negotiate/command dance with a lock, since only one client
can hold the QMP socket at a time. Events interleave with replies, so command()
reads until it sees the matching return rather than trusting the first message.
Connection failure carries the reason and the likely cause: key.py already
holding the socket.
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.
Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.
From the object code:
callers reaching keypad_update_rows
plain none -- the calls are gone
volatile keypad_key_changed, keypad_col_changed, keypad_set_press,
keypad_reset, uvk5_machine_init
keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.
Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.
Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.
Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).
README screenshots regenerated from the current build.
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.
The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.
2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.
Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.
Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
keypad scan and the model does not wake from it. Includes the two dead
ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
Adds a QEMU machine for the Puya PY32F071 (Cortex-M0+) so Quansheng UV-K5 V3
firmware can run on a PC. The firmware boots to its main loop in about five
seconds and the LCD contents are readable.
Register layouts come from the vendor CMSIS header shipped with the firmware
rather than guesswork. Modelled: RCC, GPIO, ADC, both SPI controllers, DMA1 and
the PY25Q16 flash; everything else answers through a logging catch-all, which is
how the next thing worth modelling gets identified.
Seven things had to be right before it would boot, each found by watching where
the firmware stopped: flash aliased at the application offset, clock ready bits,
self-clearing ADC calibration, SPI transfer flags, DMA-driven flash reads,
SysTick poll acceleration, and the bit-banged transceiver bus idling low.
SysTick needs explanation. SYSTICK_DelayUs polls the counter and accumulates
differences; under emulation a register read costs far more relative to guest
time, so a measured 120 ms delay would have taken about 7.7 hours. Lowering the
clock does not help because the bottleneck is loop iterations, not counter speed.
Reporting a value that runs ahead of the real counter does, via a new poll-boost
property on SysTick. Guest time therefore runs fast during delays: fine for
exercising menus and control flow, wrong for judging signal timing.
Also includes the host build of the CW timing chain (harness, stubs, shim,
tests), which compiles app/cwkeyer.c and app/cwmacro.c unmodified against stub
drivers with a virtual clock and scripted paddle input.
Known gap: keypad rows reach the firmware's scan and KEYBOARD_Poll returns the
right key code, but the UI does not react yet.
Not modelled, and not intended to be: radio behaviour. The transceiver chip has
no public datasheet, so keying envelopes and emissions need real hardware.