assets/flash.img was gitignored, so the only copy of the never-booted image lived
on one disk. The emulator writes to that image, so a session can leave edited
settings or a damaged EEPROM behind with nothing to restore from.
assets/pristine/ now holds the image as first generated, gzipped and checksummed,
and is tracked deliberately. Gzip takes it from 2 MiB to 2.3 KiB because the image
is nearly all 0xFF, which is what makes keeping it in git reasonable. The live
image and its .bak-* files stay ignored.
Two checksums are recorded, for the archive and for its contents, so a corrupted
archive is distinguishable from one that was replaced.
tools/restore_flash.sh verifies, diffs, or restores. Restore backs up the current
image first, then re-checks the result, since a restore that silently half-worked
would be worse than none.
Verified by deliberately corrupting the live image: --diff reported 32 differing
bytes, restore backed up and rewrote it, and --diff then reported no change. The
image is currently byte-identical to what make_flash.py produces, so this is the
genuine original rather than a copy of something already used.
The name is what the user is putting in DNS. server_name has to match or SNI falls
through to another vhost on the same socket.
Addresses are unchanged: 172.21.91.140 and fd3c:3f9b:6424:2::5, still sharing 443
under the existing *.mckero.dn42 wildcard.
Verified after reload: 200 on both families with the certificate validating, 80
redirecting, and dns./mail. still 200. This time nginx -t ran after the symlink was
in place, which is the ordering that caught me out last time.
The UI is served at https://k6v6.mckero.dn42/ with nginx terminating TLS and the
server itself now bound to loopback, so it is not directly reachable.
No new address and no new certificate: 443 is shared with the other vhosts on
these DN42 addresses and separated by SNI, and the existing *.mckero.dn42 wildcard
already covers the name. Only DN42 addresses are bound, so the public 443
listeners on this host are untouched.
docs/reverse-proxy.md records the settings that are not optional, because each has
a failure mode that is easy to misread:
proxy_buffering off -- otherwise the frame stream arrives in bursts
X-Forwarded-For -- otherwise every log line is attributed to 127.0.0.1
long read timeout -- a paused guest emits nothing at all
tcp_nodelay -- Nagle would delay exactly the latency-critical requests
Two pitfalls hit while setting it up are written down. "http2 on;" needs nginx
1.25.1+ and this host runs 1.22.1, and because nginx -t was run before the symlink
existed it passed, then reload failed and left nginx stopped, briefly taking the
other sites down. Separately, a newly added listen address needs a reload to be
bound: after the failed reload, v6 requests failed with nothing in the error log
until a second reload created the socket.
deploy/nginx-k6v6.conf keeps a copy in the repo, since nothing here
version-controls /etc.
Verified: HTTP 200 on both families with the certificate validating (no -k), 7
frames in a 20 KB stream sample, log entries attributed to real client addresses.
The rules restricting port 8080 to DN42 were applied by hand and existed only in
the live kernel tables -- lost on reboot, with the port then wide open and nothing
in the repo to say it ever had been restricted.
The script is idempotent (checks before adding) and has remove/show. show prints
packet counters, which is the part that matters: this host's INPUT policy is
ACCEPT, so a rule that only allows DN42 does nothing at all. The final DROP is
what restricts anything, and a rising DROP count is the only proof it works
rather than the traffic simply not arriving.
Currently observed: 2203 accepted from DN42 v4, 103 from v6, 22 dropped.
README gets a section with the endpoint table, the two constraints that will
otherwise surprise someone (single QMP client, no authentication), and the reason
frames go through memsave rather than pmemsave or gdb.
AGENTS.md gets the run instructions plus a new entry under 'Things that already
went wrong' for the pmemsave trap: it takes a physical address, returns zeros for
gFrameBuffer, and reports success. The web UI was built on it initially because a
benchmark showed it was fast -- the benchmark never checked the contents. Worth
recording as the general lesson, not just the specific fix.
AGENTS.md still carried a stale entry telling the reader to *lengthen* key
holds when a press seems ignored, which is the opposite of the fix and is
what broke the tooling in the first place. Replaced with the correction and
a pointer to the right section.
The keypad heading also claimed the hold time was the only cause. There were
two: the 2500 ms hold in key.py, and row_out missing volatile. Both are now
listed up front with a link to the detail.
Adds the regression test to the places someone would actually look: the
"How to run it" section in AGENTS.md, the layout listing, and a build step
in the README noting that a clean build is not evidence the keypad works,
since the -O2 dead-code elimination produces no warning.
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.
Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.
From the object code:
callers reaching keypad_update_rows
plain none -- the calls are gone
volatile keypad_key_changed, keypad_col_changed, keypad_set_press,
keypad_reset, uvk5_machine_init
keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.
Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.
Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.
Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).
README screenshots regenerated from the current build.
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.
The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.
2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.
Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.
Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
keypad scan and the model does not wake from it. Includes the two dead
ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
Adds a QEMU machine for the Puya PY32F071 (Cortex-M0+) so Quansheng UV-K5 V3
firmware can run on a PC. The firmware boots to its main loop in about five
seconds and the LCD contents are readable.
Register layouts come from the vendor CMSIS header shipped with the firmware
rather than guesswork. Modelled: RCC, GPIO, ADC, both SPI controllers, DMA1 and
the PY25Q16 flash; everything else answers through a logging catch-all, which is
how the next thing worth modelling gets identified.
Seven things had to be right before it would boot, each found by watching where
the firmware stopped: flash aliased at the application offset, clock ready bits,
self-clearing ADC calibration, SPI transfer flags, DMA-driven flash reads,
SysTick poll acceleration, and the bit-banged transceiver bus idling low.
SysTick needs explanation. SYSTICK_DelayUs polls the counter and accumulates
differences; under emulation a register read costs far more relative to guest
time, so a measured 120 ms delay would have taken about 7.7 hours. Lowering the
clock does not help because the bottleneck is loop iterations, not counter speed.
Reporting a value that runs ahead of the real counter does, via a new poll-boost
property on SysTick. Guest time therefore runs fast during delays: fine for
exercising menus and control flow, wrong for judging signal timing.
Also includes the host build of the CW timing chain (harness, stubs, shim,
tests), which compiles app/cwkeyer.c and app/cwmacro.c unmodified against stub
drivers with a virtual clock and scripted paddle input.
Known gap: keypad rows reach the firmware's scan and KEYBOARD_Poll returns the
right key code, but the UI does not react yet.
Not modelled, and not intended to be: radio behaviour. The transceiver chip has
no public datasheet, so keying envelopes and emissions need real hardware.