Commit Graph
21 Commits
Author SHA1 Message Date
mckero fdcbe80056 Model TIM2, so millis() advances and timeouts can expire
Second finding from the audit. TIM2 was covered by the catch-all stub, which returns
the last value written, so

    uint32_t millis(void) { return LL_TIM_GetCounter(TIM2); }

returned 0 forever. All 17 call sites that measure elapsed milliseconds could never
see time pass -- a silent wrong answer rather than a hang, which is harder to notice
and was not noticed.

The counter is derived on read from QEMU_CLOCK_VIRTUAL rather than stored, with
CR1.CEN starting and freezing it and a CNT write rebasing it. Guest time here is not
proportional to wall time anyway, and code measuring elapsed milliseconds wants
something advancing at roughly the rate a human sees; this is explicitly not for
anything needing cycle accuracy.

Measured: 24358 ms, then 29527 ms five seconds later -- 5169 ms elapsed, so the rate
is right rather than merely non-zero. The test checks the rate for that reason: a
counter ticking at the wrong speed would satisfy "non-zero" and "increasing" and still
break every timeout.

AGENTS.md now carries the audit itself: a table of what the firmware actually drives
against what is modelled versus stubbed, and the point that answering reads is not the
same as being reproduced. The honest summary is that the digital side the firmware
depends on is reproduced, and the analogue side is not and cannot be.

Full run: 15 passed, 0 failed.
2026-08-29 08:00:27 +01:00
mckero e46cae2e48 Make the ADC settable, which reaches the battery behaviour
Prompted by a fair criticism: the reports said what runs, not what is actually
reproduced. An audit found the ADC was modelled but returned a hardcoded 2200 forever,
so gBatteryDisplayLevel, gLowBattery and the warning popup were all unreachable. A
peripheral that answers reads is not the same as a peripheral that is reproduced.

adc-result is now settable over QOM and clamped to 12 bits. Measured: 2200 gives
level 4 and no warning, 1200 gives level 0 and raises gLowBattery, and the level
recovers to 4 afterwards.

tools/test_battery.py covers it, and deliberately does NOT assert that gLowBattery
clears on recovery. helper/battery.c:190-204 only clears it when the level lands
exactly on 2; above that it clears gLowBatteryConfirmed and leaves gLowBattery set. So
4 -> 0 -> 4 really does leave the flag raised. The first version of this test called
that a failure -- the test was wrong, not the model. The emulator reproduces the
firmware, including behaviour that looks like a bug.
2026-08-29 07:49:25 +01:00
mckero a1aca9395c Document why there is no audio to model
Records the scope plainly, because "add a speaker and a microphone" is the obvious
request and the answer is that neither is on the MCU: no audio samples exist in its
address space, so there is nothing to capture, nothing to play, and nothing for a
browser permission to carry. Same line as the analogue RF limit.

Also records the stub lesson: QmpClient.command returns the unwrapped value, the test
stub returned an envelope, and the divergence let 88 tests pass while the live UI
returned 500. A stub more forgiving than the real client is worse than none.
2026-08-29 06:06:55 +01:00
mckero 8d1a1c4415 Add a test runner, and a test that it can fail
The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:

    bash tools/run_tests.sh        # everything, 11 suites, a few minutes
    bash tools/run_tests.sh -q     # unit tests only, ~15 s, no emulator

The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.

Two defects in the runner's own first draft, both caught before it was trusted:

It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.

That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.

Full run: 11 passed, 0 failed.
2026-08-29 05:42:11 +01:00
mckero 95bad1614e Note that counting distinct frames proves less than it looks
With RSSI varying every poll, the meter redraws constantly, so "consecutive frames
differ" is true on a completely parked radio. Records the framebuffer page layout so
a screen test can compare the rows that answer its actual question, and the fact that
page 4 is not the meter row.
2026-08-29 05:29:10 +01:00
mckero a310429eff Correct the docs that said PTT did not exist
Three places claimed there was no PTT line, and one of them named the wrong pin
(GPIOC rather than PB10). All now say the same thing: PTT works, but not through
the key table, because the firmware reads its own pin instead of scanning it as a
matrix key.

Documents the transmit bar's two gates -- FUNCTION_TRANSMIT and gSetting_mic_bar,
the latter already on because blank flash reads 0xFF -- and why the release path
gets more attention than the press: a stuck PTT leaves every later test running
against a transmitting radio.

Also records the '.key' vs '.key[data-key]' trap for whoever adds the next
non-key button.
2026-08-29 05:18:57 +01:00
mckero 39195b8028 Update the notes now that the S-meter works
The squelch section said the blocker sat above the device model and told the reader
not to resume without new information. Both are now wrong, so it leads with the
outcome instead: SIDE1 engages monitor, which skips squelch entirely, and the meter
reads -53 dBm / S9+40.

The four failed attempts are kept rather than deleted. Each ended in a confident
wrong diagnosis -- power-save gating, trigger timing, bit semantics -- and the shared
cause was a one-bit read skew underneath all of them. That pattern is worth more to
the next reader than a clean account of the version that worked.

README gains the S-meter row and test_smeter.py.
2026-08-29 04:57:00 +01:00
mckero 7faacfd433 Record what the read-path fix revealed about the squelch attempts
The shifted-read bug invalidated four earlier diagnoses, so the notes claiming a
power-save gate or a timing problem were wrong and are corrected.

With reads fixed the interrupt handshake demonstrably works: RAISE pending=0004,
ACK delivering flags=0004, correct bit, collected by the firmware, REG_0C back to
0x0000 so nothing hangs.

g_SquelchLost is still 0 and the S-meter still absent, but the shape of the problem
is now clear and recorded: announcing once lands during startup before the flag
leads anywhere; announcing every poll re-arms the bit inside the firmware's own
untimed collection loop and spins forever; announcing periodically avoids both and
still changes nothing. The blocker is getting the radio into a receiving state at
all, which sits above the device model.

Also records the general lesson, which cost the most time here: when several
independent attempts fail in the same way, suspect the shared transport rather than
the logic layered on top of it.
2026-08-29 04:40:07 +01:00
mckero 4e29d11810 Update the docs now that the BK4819 is modelled
The "what this cannot do" section said the transceiver was not modelled and treated
that as permanent. Half of it is now wrong: the register interface works. The other
half is still true and worth keeping sharp -- analogue behaviour is out of reach
because the chip has no public datasheet, so the driver is the only specification and
it can only say which registers were written.

Rewritten to separate the two: what is modelled and what it fixed (RSSI was hard zero
at 18 call sites), the two constraints the untimed spin loops impose, and then the
line that does not move. Explicitly warns against reading the new test as evidence
about RF.

Also corrects the GPIO comment about idling PB9 low. It described the pin as a
workaround pending a device model; that model now exists, and the idle level only
covers the window between reset and the bus being wired.
2026-08-28 16:46:55 +01:00
mckero f3412f3316 Correct the docs: serial receive is implemented now
The previous commit documented serial receive as a permanent limitation, listing
what would be needed to build it. It has since been built, so that section was
actively misleading -- it told the reader not to try something that already works.

Replaced with what it takes to keep working, since each of the three pieces fails
silently on its own: USART1 needs a chardev, DMA must decrement CNDTR for USART
(serviced on the CNDTR read, which is where the driver looks), and DR writes must
reach the chardev and not just stderr. The last one is the trap -- transmit that
goes only to stderr is indistinguishable from the firmware ignoring the command.

README status table and build-verification steps updated to match.
2026-08-28 16:26:55 +01:00
mckero cdb73f13a0 Record that serial receive is not implemented
Found while explaining a DMA channel that stayed permanently armed with cndtr=256
during the flash investigation. It is USART1's receive channel, running in
LL_DMA_MODE_CIRCULAR, so never completing is correct behaviour and not a bug.

But it exposed a real gap. driver/uart.c derives its write pointer from
sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(...), and the DMA model only
services SPI, so CNDTR never decrements for USART and that expression is always
zero. Combined with USART1 being a py32-stub with no chardev backend, nothing can
be sent *to* the firmware.

The cost is specific: UART_IsCommandAvailable never fires, so the UV-K5 programming
protocol in app/uart.c is unreachable -- 0x0514 handshake, 0x051B EEPROM read,
0x051D EEPROM write, 0x05DD reset. CPS/CHIRP-style tools cannot talk to this
emulator. Transmit is unaffected, which is why the firmware banner shows up fine
and this went unnoticed.

Documented rather than fixed: it needs a chardev on USART1 plus circular-mode DMA
driven by receive, which is a new feature rather than a repair. The notes say what
would be involved so the next person does not have to rediscover the mechanism.

Status table also updated to reflect what the flash and DMA fixes settled --
persistence and frequency entry now work.
2026-08-28 15:52:05 +01:00
mckero d09bf5f47e Write down the flash investigation and what made it slow
Four faults in the SPI/DMA/flash models each produced the same symptom -- stored
frequencies zeroed, a typed frequency reverting to 18 MHz -- and each was invisible
from the layer above. AGENTS.md now lists all four with the reasoning that connects
them to what the user saw, so the next person does not rediscover them one at a
time.

Also records the methodology mistakes, because they cost more than the bugs did:

- Attaching gdb between digits clears the frequency input box: the timeout is
  ~2.5 s and an attach takes ~3 s. Three wrong conclusions came from this, so
  instrument the model and read stderr instead of stopping the guest.
- A diagnostic log capped at six entries showed only 0xFF payloads and supported
  precisely the wrong conclusion. Do not cap before the shape of the data is known.
- A failed ninja leaves the old binary in place and the test still runs. Two rounds
  of results were meaningless. Check for FAILED and error: before trusting a run.
- assets/flash.img is written by every session, so a test starting from it can find
  its work already done -- which looks identical to broken persistence. Start from
  assets/pristine/, and power off before restoring, since shutdown flushes the old
  image back over the file.
- Hand-computed struct offsets gave KEY_LOCK=4 and TX_VFO=11, impossible values,
  because the ELF has no DWARF and the structs contain enums. Use an unambiguous
  nm symbol, or find the field by toggling it and diffing.
- One probe printed phase before incrementing it, making a correct address decoder
  look off by one. A working implementation was nearly "fixed" as a result.

README gains the two new tests in the build-verification step and the tools list.
2026-08-28 15:32:45 +01:00
mckero 1364d46e97 Keep a pristine copy of the flash image, and a way back to it
assets/flash.img was gitignored, so the only copy of the never-booted image lived
on one disk. The emulator writes to that image, so a session can leave edited
settings or a damaged EEPROM behind with nothing to restore from.

assets/pristine/ now holds the image as first generated, gzipped and checksummed,
and is tracked deliberately. Gzip takes it from 2 MiB to 2.3 KiB because the image
is nearly all 0xFF, which is what makes keeping it in git reasonable. The live
image and its .bak-* files stay ignored.

Two checksums are recorded, for the archive and for its contents, so a corrupted
archive is distinguishable from one that was replaced.

tools/restore_flash.sh verifies, diffs, or restores. Restore backs up the current
image first, then re-checks the result, since a restore that silently half-worked
would be worse than none.

Verified by deliberately corrupting the live image: --diff reported 32 differing
bytes, restore backed up and rewrote it, and --diff then reported no change. The
image is currently byte-identical to what make_flash.py produces, so this is the
genuine original rather than a copy of something already used.
2026-08-28 12:20:13 +01:00
mckero 19997e85e1 Rename the vhost to k6v3.mckero.dn42
The name is what the user is putting in DNS. server_name has to match or SNI falls
through to another vhost on the same socket.

Addresses are unchanged: 172.21.91.140 and fd3c:3f9b:6424:2::5, still sharing 443
under the existing *.mckero.dn42 wildcard.

Verified after reload: 200 on both families with the certificate validating, 80
redirecting, and dns./mail. still 200. This time nginx -t ran after the symlink was
in place, which is the ordering that caught me out last time.
2026-08-28 11:49:51 +01:00
mckero a4a8f21d50 Document and version-control the HTTPS front end
The UI is served at https://k6v6.mckero.dn42/ with nginx terminating TLS and the
server itself now bound to loopback, so it is not directly reachable.

No new address and no new certificate: 443 is shared with the other vhosts on
these DN42 addresses and separated by SNI, and the existing *.mckero.dn42 wildcard
already covers the name. Only DN42 addresses are bound, so the public 443
listeners on this host are untouched.

docs/reverse-proxy.md records the settings that are not optional, because each has
a failure mode that is easy to misread:
  proxy_buffering off  -- otherwise the frame stream arrives in bursts
  X-Forwarded-For      -- otherwise every log line is attributed to 127.0.0.1
  long read timeout    -- a paused guest emits nothing at all
  tcp_nodelay          -- Nagle would delay exactly the latency-critical requests

Two pitfalls hit while setting it up are written down. "http2 on;" needs nginx
1.25.1+ and this host runs 1.22.1, and because nginx -t was run before the symlink
existed it passed, then reload failed and left nginx stopped, briefly taking the
other sites down. Separately, a newly added listen address needs a reload to be
bound: after the failed reload, v6 requests failed with nothing in the error log
until a second reload created the socket.

deploy/nginx-k6v6.conf keeps a copy in the repo, since nothing here
version-controls /etc.

Verified: HTTP 200 on both families with the certificate validating (no -k), 7
frames in a 20 KB stream sample, log entries attributed to real client addresses.
2026-08-28 11:39:17 +01:00
mckero 2c3a602d9f Capture the DN42 firewall rules as a script
The rules restricting port 8080 to DN42 were applied by hand and existed only in
the live kernel tables -- lost on reboot, with the port then wide open and nothing
in the repo to say it ever had been restricted.

The script is idempotent (checks before adding) and has remove/show. show prints
packet counters, which is the part that matters: this host's INPUT policy is
ACCEPT, so a rule that only allows DN42 does nothing at all. The final DROP is
what restricts anything, and a rising DROP count is the only proof it works
rather than the traffic simply not arriving.

Currently observed: 2203 accepted from DN42 v4, 103 from v6, 22 dropped.
2026-08-28 05:32:31 +01:00
mckero 2d0fc24b62 Document the web remote control
README gets a section with the endpoint table, the two constraints that will
otherwise surprise someone (single QMP client, no authentication), and the reason
frames go through memsave rather than pmemsave or gdb.

AGENTS.md gets the run instructions plus a new entry under 'Things that already
went wrong' for the pmemsave trap: it takes a physical address, returns zeros for
gFrameBuffer, and reports success. The web UI was built on it initially because a
benchmark showed it was fast -- the benchmark never checked the contents. Worth
recording as the general lesson, not just the specific fix.
2026-08-28 04:49:32 +01:00
mckero b246132a35 Bring the docs in line with both keypad fixes
AGENTS.md still carried a stale entry telling the reader to *lengthen* key
holds when a press seems ignored, which is the opposite of the fix and is
what broke the tooling in the first place. Replaced with the correction and
a pointer to the right section.

The keypad heading also claimed the hold time was the only cause. There were
two: the 2500 ms hold in key.py, and row_out missing volatile. Both are now
listed up front with a link to the detail.

Adds the regression test to the places someone would actually look: the
"How to run it" section in AGENTS.md, the layout listing, and a build step
in the README noting that a clean build is not evidence the keypad works,
since the -O2 dead-code elimination produces no warning.
2026-08-28 03:58:19 +01:00
mckero 2154f80414 Fix the keypad: row_out must be volatile
The previous commit removed three TRACE fprintfs from py32f071.c as
cleanup. That silently broke the keypad completely -- no press reached the
UI, and nothing warned about it.

Root cause is dead-code elimination, not the printing.
qdev_init_gpio_out_named() is inlinable and only records the row_out array;
the lines are filled in later by qdev_connect_gpio_out_named() from board
code, which GCC cannot see. At -O2 GCC therefore proves every element is
still NULL, sees that qemu_set_irq() returns immediately on a NULL irq, and
deletes the body of keypad_update_rows() along with all five calls to it. No
row line is ever driven and the firmware's scan reads all-high.

From the object code:

  callers reaching keypad_update_rows
    plain     none -- the calls are gone
    volatile  keypad_key_changed, keypad_col_changed, keypad_set_press,
              keypad_reset, uvk5_machine_init

keypad_col_changed compiles to a store and a ret with no call at all; with
volatile it ends in jmp keypad_update_rows. Declaring row_out volatile fixes
it at the cause. 10/10 on the press test, 3/3 on keypad_test.py, no build
warnings.

Scoped rather than assumed: PY32GpioState::out is not affected. Marking it
volatile too gives a byte-identical object file, because py32_gpio_write()
is only reachable through a MemoryRegionOps function-pointer table so GCC
cannot enumerate its callers. It stays plain.

Adds tools/keypad_test.py: boots its own instance on private ports and
checks that a short MENU press opens the menu, DOWN moves the cursor, and a
held key is visible to the scan. This is what should have caught the
breakage before it was pushed.

Docs corrected. The breakage had been written up as "power save stops the
keypad scan" and called a gap in the model; it was neither. AGENTS.md now
records the mechanism, the measurements, the objdump check, and the two
measurement traps that made this hard: reading gKeyReading0 after releasing
the key (always KEY_INVALID), and trusting a gdb breakpoint on
KEYBOARD_Poll (with the guest stopped the scan's delays cost no guest time,
so Poll returns KEY_MENU on a build where it fails when running free).

README screenshots regenerated from the current build.
2026-08-27 18:34:41 +01:00
mckero 1c9a2afe55 Fix key.py hold times; the keypad was never broken
key.py held every key for 2500 ms, on the assumption that guest time runs
fast during delays so a press needs a long wall-clock hold. That is wrong
for this path, and it is why the keypad looked dead.

The two SysTick mechanisms are separate. poll-boost accelerates counter
*reads* so SYSTICK_DelayUs converges; it does not speed up interrupt
delivery. Interrupts drive SysTick_Handler -> gNextTimeslice ->
APP_TimeSlice10ms -> CheckKeys at close to real time, so the firmware's
thresholds hold in wall clock as written: 20 ms to register a press,
400 ms to count as held.

2500 ms is ~250 ticks, six times past the long-press threshold, so every
press was dispatched as a hold. MAIN_Key_MENU acts only on a short release
and returns early when bKeyHeld is set, so nothing happened. Confirmed by
reading gDebounceCounter mid-hold: 317 after a 3 s hold, which also proves
the timeslice was running all along.

Now HOLD_MS=200 and LONG_HOLD_MS=900. Verified with screenshots: the menu
opens and UP/DOWN move through it.

Also here:
- Drop the three TRACE fprintfs. They fired on every keypad poll and
  buried the console; the matrix is confirmed working.
- Drop a redundant forward declaration of py32_spi_xfer_byte, silencing
  the only build warning.
- Document the real remaining gap: power save (~6 s after boot) stops the
  keypad scan and the model does not wake from it. Includes the two dead
  ends already ruled out by experiment, so nobody repeats them.
- Add README screenshots captured from guest memory.
2026-08-27 16:45:54 +01:00
mckero c0a09827ed UV-K5 V3 emulator: QEMU machine for the PY32F071
Adds a QEMU machine for the Puya PY32F071 (Cortex-M0+) so Quansheng UV-K5 V3
firmware can run on a PC. The firmware boots to its main loop in about five
seconds and the LCD contents are readable.

Register layouts come from the vendor CMSIS header shipped with the firmware
rather than guesswork. Modelled: RCC, GPIO, ADC, both SPI controllers, DMA1 and
the PY25Q16 flash; everything else answers through a logging catch-all, which is
how the next thing worth modelling gets identified.

Seven things had to be right before it would boot, each found by watching where
the firmware stopped: flash aliased at the application offset, clock ready bits,
self-clearing ADC calibration, SPI transfer flags, DMA-driven flash reads,
SysTick poll acceleration, and the bit-banged transceiver bus idling low.

SysTick needs explanation. SYSTICK_DelayUs polls the counter and accumulates
differences; under emulation a register read costs far more relative to guest
time, so a measured 120 ms delay would have taken about 7.7 hours. Lowering the
clock does not help because the bottleneck is loop iterations, not counter speed.
Reporting a value that runs ahead of the real counter does, via a new poll-boost
property on SysTick. Guest time therefore runs fast during delays: fine for
exercising menus and control flow, wrong for judging signal timing.

Also includes the host build of the CW timing chain (harness, stubs, shim,
tests), which compiles app/cwkeyer.c and app/cwmacro.c unmodified against stub
drivers with a virtual clock and scripted paddle input.

Known gap: keypad rows reach the firmware's scan and KEYBOARD_Poll returns the
right key code, but the UI does not react yet.

Not modelled, and not intended to be: radio behaviour. The transceiver chip has
no public datasheet, so keying envelopes and emissions need real hardware.
2026-08-27 14:59:21 +01:00