Commit Graph
78 Commits
Author SHA1 Message Date
mckero fc432d2055 Report what the device says it is running, not the file we handed it
The page only knew its own input. A build called f4hwn.fusion.bin reports EGZUMER+F4HWN v6.0.0.CN, and with the multi-system release a committed external slot makes the factory bootloader reflash the internal flash from that slot on every power-on -- so the uploaded image never runs and the page keeps naming it. The firmware prints its own banner on USART1; tools/uvk5_banner.py reads it back, /api/firmware returns running: {banner, matches_uploaded, note}, and the page shows what the device reports, flagging it only when the running version is not in the uploaded image at all.

That reader also exposed a regression of my own: _start_stderr_pump had been rewritten to read the pipe in 64 KB chunks, which kept QEMU from blocking but delivered nothing to the log until 64 KB had accumulated -- and the banner is forty bytes, so it never appeared. It reads lines again, still starting before anything waits on QEMU, and test_uvk5_supervisor passes either way.
2026-10-01 16:22:44 +08:00
mckero 11e3678140 Take the stale addresses out of the AGENTS examples too
The screenshot example pasted one build's numbers and told the reader to get them with nm -- which cannot work here, because the images are program-header-only ELFs with no symbol table. It now asks the firmware through tools/uvk5_buffers.py first. The webui example drops the flags entirely, since the page draws the controller's memory and needs none.

Also fixed the paragraph in both READMEs that had a tool name and a flag on one line, which the doc checker (correctly) read as passing that flag to that tool.
2026-10-01 16:14:01 +08:00
mckero a237bd92c6 Stop teaching the old address flags in the docs
The flags are optional now and the page needs no addresses at all, so the examples no longer paste one build's numbers: screenshot.py is shown asking the firmware through tools/uvk5_buffers.py, and the webui example omits them entirely. The paragraph that said they default to one known build is replaced with what actually happens.
2026-10-01 16:12:36 +08:00
mckero 1e9fdf685c Find the screen buffers in the firmware instead of hardcoding one build's
The page was told --frame-addr 0x200012BE --status-addr 0x2000163E and used them as a fallback. The firmware the user actually flashed keeps its buffers at 0x2000129E/0x2000161E, 32 bytes earlier, so every line landed 32 bytes off: that is the "other firmware looks shifted" report. The images here are minimal ELFs with no symbol table, so there is nothing to read -- but the firmware's own buffers hold the same bytes the controller holds, and tools/uvk5_buffers.py finds them by matching (1024/1024 bytes for that file).

The two address flags are optional now, work/run-webui.ps1 passes no machine-specific values at all, and the page reports what it found in /api/status and /api/panel. tools/uvk5_testenv.qemu() also looks in the sibling qemu-7.2/build the rest of the repo assumes. Fixed /api/panel's emulator-off branch, which called jsonify with both a dict and kwargs and 500'd.

Tests: test_uvk5_buffers (the search must count matches, not pairs -- its first version scored every offset full marks and always answered the first one).
2026-10-01 16:09:26 +08:00
mckero 57106c0a66 Fix the two pixel bugs behind "the other firmware looks shifted"
uvk5_stream.py used STATUS_BYTES without importing it, so the panel branch raised NameError on every frame and a bare except swallowed it: every screen the page drew came from guest RAM at one build's addresses. The pump now reports which source it used and why, webui exposes it (/api/panel and frame_source), and test_uvk5_stream asserts the panel wins when reachable and that a fallback is announced.

The ST7565 column counter wrapped at 128 instead of the controller's 132, so addresses 128..131 came back as 0..3, fell outside the col>=4 store, and were dropped: every row lost its last four pixels, which is where the battery icon lives. Pre-fix, filling a page with 0xFF left columns 124..127 blank; now they carry content and the page's frame matches the panel memory 8192/8192.
2026-10-01 15:59:39 +08:00
mckero f4c9d343fc Probe what a BK4819 read presents, and say plainly that it is not the guard yet
UVK5_BK4819_PROBE reassembles the sixteen bits a read clocks out and logs them against the register's own value: 1566 of 1566 agree on the working model. That agreement is not proof -- removing the skip_falling fix, which is exactly the historical left-shift regression, leaves the reassembled word unchanged, so this observation point is not the one the guest samples at. tools/test_bk4819_readback.sh therefore stays the guard.

tools/test_bk4819_readback.py is the working draft of a portable replacement (no source patch, no rebuild, no ARM gdb) and is deliberately NOT registered in run_tests.sh, so a proven guard is not swapped for an unproven one. Both the file and the two READMEs say so.
2026-10-01 15:28:47 +08:00
mckero e1ffdf8fdd panel_dump: say why QMP is unreachable, and exit 2 for it
QMP takes a single client and the web UI holds it for its whole lifetime, so this is the normal outcome while the page is open rather than a broken emulator. Verified against a private instance: 128x64, 1819 pixels lit, 64 rows of ASCII.
2026-10-01 15:23:27 +08:00
mckero 4bddccfe7b Measure how a build renders, and keep the tool that does it
tools/panel_dump.py prints the display controller's own memory as ASCII or PNG for any firmware, with the four mappings, so two builds can be compared instead of glanced at. The panel model gained a bounded UVK5_PANEL_PROBE diagnostic along the way.

Measured: the 5.9.0.CN panel agrees with its own framebuffer 8188 of 8192 pixels with the data untouched, and the fetched 6.0.0 build renders identically. Both program the same geometry registers, which is why the mapping is a driver convention and cannot be derived from the controller: it has to be measured.

Noted, not yet fixed: the display start line (0x40|n) is ignored, and the model stores pixels at col-4 with the column counter wrapping at 128 instead of the controller's 132 columns.
2026-10-01 15:22:02 +08:00
mckero 542515d5ad Finish the portability sweep: one firmware resolver, no author paths
The probe scripts, run.sh, trace_run.sh, webui.py and two emulator tests each named the same hardcoded firmware from a source tree that is not in this repository. They now resolve QEMU and the firmware the way tools/uvk5_testenv.py does -- environment, then PATH, then whatever the checkout has -- and skip with a reason when there is nothing.

webui.py's --qemu and --elf lost their author defaults too: a bare qemu-system-arm through PATH, and no firmware until one is uploaded, which the page already reports.
2026-10-01 15:15:31 +08:00
mckero ae8b48c85a Run anywhere: no author paths left, and CI that proves it
tools/run_tests.sh defaults QEMU_SRC to a sibling of the checkout, which is where setup_qemu.sh puts it; the two defaults disagreed, so a fresh clone rebuilt nothing and reported a build that was not there. The interpreter list was also reading an empty $PY.

tools/test_bk4819_readback.sh was the last test with the author's paths, and the only one that could not run elsewhere. It now takes QEMU, GDB and ELF from the environment or PATH like the python tests, skips with a reason when one is missing, and says so on a platform whose QEMU cannot make the unix socket it uses.

tools/check_docs.py points UVK5_FW_DIR at a sibling and skips the file:line checks, with a message, when there is no firmware tree -- a fresh clone used to see thirteen failures it could do nothing about.

Added .github/workflows/unit.yml (the fast half of run_tests.sh on every push and PR), requirements-dev.txt for the one pip dependency, and a Dockerfile. Flask is not always installed, so test_webui now skips through setUpModule rather than erroring.
2026-10-01 15:12:45 +08:00
mckero 1308c98769 Document where Moto/DFU entry is, and let a boot key hold PTT plus a key
The bootloader's DFU handler is reachable only when SRAM[0x20000020] is 3, which only a program that then resets can write. The application's 0x05DD takes that path only with ENABLE_OVERLAY; this build resets straight back into the application instead. Ruled out by measurement: PTT alone, PTT+SIDE1/SIDE2, MENU, a host byte in the boot window including 0x0530, and 0x05DD.

boot-key now reads a + separated list, because the firmware's own BOOT_GetMode() needs PTT and a matrix key together for every special boot mode. Tested against all four documented combinations.
2026-10-01 15:06:20 +08:00
mckero 2667e046e8 Emulator: multiboot slots from the page, flash controller, portable tests
flash controller: store ACR/OPTKEYR instead of swallowing them, which is what stopped the factory bootloader from starting

slots over the firmware's own serial protocol (0x0720 family); uvk5_socket/uvk5_testenv so a fresh checkout skips instead of failing; web UI slot table and Multiboot button; quick start, CONTRIBUTING, and stop tracking firmware images and radio dumps
2026-10-01 14:54:34 +08:00
mckero ee80939c78 Check that documented tool flags exist
The remaining class of claim check_docs.py could not see: whether the commands in the
docs would actually run. A renamed or removed option is the classic form of command
rot, and the one that wastes a reader's time most directly -- they paste the line and
it fails. All 9 documented flags across screenshot.py, webui.py and restore_flash.sh
are real.

The check earned its own lesson, recorded in both languages. Its first version matched
only to the end of the line, so on a wrapped command it saw --frame-addr and nothing
after the backslash: 4 of 9 flags, and it reported a clean run. A check that silently
covers a quarter of what it claims is worse than no check, because the clean result is
believed. Continuations are joined before matching now.

Confirmed it fails when it should: renaming --frame-addr to something no tool accepts
produces two named failures and exit 1, and reverting returns it to clean.

check_docs.py now runs seven checks.
2026-08-29 16:58:28 +01:00
mckero b32335d8c0 Check the docs' claims against the code mechanically
Translating everything into Chinese found four claims that had already drifted, and
none of them were caught by reading -- they were caught by comparing against source.
Proofreading does not find rot, so do the comparison mechanically and keep doing it.

tools/check_docs.py verifies that every tool a README names exists, that every test in
run_tests.sh is documented in both languages, that internal .md links resolve, that the
translation pairs have matching heading structure, that memory-map addresses match the
model's #defines, and that documented firmware file:line references still point at what
the prose claims. It runs in the quick tier of run_tests.sh, needing no emulator.

Confirmed it can actually fail, because a checker that cannot is worthless: renaming a
documented tool and deleting a heading from the Chinese side each produce one named
failure and exit 1, and reverting returns it to clean.

One thing it deliberately does not check. An early version compared firmware constants
with a regex that took the first number on a line, so `key_debounce_10ms = 20 / 10` read
as 20 and it declared the docs wrong for saying 2. The docs were right and the checker
was broken. A checker that cries wolf gets ignored, so claims it cannot verify
unambiguously are left out rather than guessed at.

Current state: 16 file:line references all accurate, 7 memory-map addresses all match,
zero broken links, all three translation pairs structurally aligned.
2026-08-29 16:54:25 +01:00
mckero 3df3c1b16d Add Chinese translations of all three documents
Full translations rather than summaries, section-for-section with the English:
README (13 sections), AGENTS.md (21), and docs/reverse-proxy.md. Each pair
cross-links to the other and says the two are kept in step, since documentation
that has silently diverged is worse than documentation that does not exist.

Verified rather than eyeballed: heading counts and order match in both pairs,
every internal .md link resolves, every tool named in either README exists, and
every test in run_tests.sh appears in both.

Translating turned up four things that were already stale in the English, which
is the honest argument for having done it this way -- a summary would not have
touched them:

  - the endpoint table was missing /api/ptt, /api/power/<action> and
    /api/logs, and did not mention the speaker field on /api/status
  - the modelled-peripheral list omitted TIM2
  - the audit table still called TIM a stub, unchanged since fdcbe80 modelled
    TIM2
  - neither README listed uvk5_logs.py, uvk5_stream.py, uvk5_supervisor.py or
    test_kill_emulator.sh, which are part of the repo rather than scratch

The ad-hoc probe scripts are now acknowledged in one line instead of being
silently absent, and described as what they are: quick to reach for, not
polished.

Unit tests: 89 passed.
2026-08-29 16:48:27 +01:00
mckero 8b995aa610 Make RSSI depend on tuning instead of being a constant
The S-meter had a number to draw, but a fixed RSSI above squelch meant the band was
uniformly and permanently occupied. Scanning, squelch, and every "is this channel busy"
decision therefore faced a situation that never varied, so none of that logic was
really being tested -- the tests passed without testing much.

RSSI is now derived from where the firmware tuned. BK4819_SetFrequency splits the
frequency across REG_38 and REG_39 (driver/bk4819.c:743), which the model already
records; verified against a live guest that 0x0262/0x5A00 reads back as 400.00000 MHz,
matching the screen. A small table of virtual stations plus a noise floor and a fade
either side of centre gives a band with signals in some places and not others.

Measured through the firmware's own tuning path -- typing 410.000 on the keypad rather
than poking the registers, so the test does not check the model against itself:

    400.000 MHz (station)  RSSI 0x01E5
    410.000 MHz (empty)    RSSI 0x0091      a gap of 85 dB

What is honest and what is not, recorded in the code: the shape is real physics, power
falls off away from a carrier with a noise floor underneath. The station list is
invented. So this reproduces "the firmware copes with a band that is busy in places",
which is genuine coverage, and it reproduces no actual radio environment -- a dBm figure
from here is not a claim about the world.

Also records why backlight PWM is deliberately left stubbed. Intermediate brightness
runs TIM7 -> DMA rewriting GPIOA BSRR at 128 kHz, so modelling it costs 128,000 GPIO
writes per emulated second and changes nothing observable: backlight is LED brightness
and never touches the framebuffer. The two endpoints that are observable, off and full,
bypass the timer and already work.

Full run: 16 passed, 0 failed.
2026-08-29 08:52:52 +01:00
mckero fdcbe80056 Model TIM2, so millis() advances and timeouts can expire
Second finding from the audit. TIM2 was covered by the catch-all stub, which returns
the last value written, so

    uint32_t millis(void) { return LL_TIM_GetCounter(TIM2); }

returned 0 forever. All 17 call sites that measure elapsed milliseconds could never
see time pass -- a silent wrong answer rather than a hang, which is harder to notice
and was not noticed.

The counter is derived on read from QEMU_CLOCK_VIRTUAL rather than stored, with
CR1.CEN starting and freezing it and a CNT write rebasing it. Guest time here is not
proportional to wall time anyway, and code measuring elapsed milliseconds wants
something advancing at roughly the rate a human sees; this is explicitly not for
anything needing cycle accuracy.

Measured: 24358 ms, then 29527 ms five seconds later -- 5169 ms elapsed, so the rate
is right rather than merely non-zero. The test checks the rate for that reason: a
counter ticking at the wrong speed would satisfy "non-zero" and "increasing" and still
break every timeout.

AGENTS.md now carries the audit itself: a table of what the firmware actually drives
against what is modelled versus stubbed, and the point that answering reads is not the
same as being reproduced. The honest summary is that the digital side the firmware
depends on is reproduced, and the analogue side is not and cannot be.

Full run: 15 passed, 0 failed.
2026-08-29 08:00:27 +01:00
mckero e46cae2e48 Make the ADC settable, which reaches the battery behaviour
Prompted by a fair criticism: the reports said what runs, not what is actually
reproduced. An audit found the ADC was modelled but returned a hardcoded 2200 forever,
so gBatteryDisplayLevel, gLowBattery and the warning popup were all unreachable. A
peripheral that answers reads is not the same as a peripheral that is reproduced.

adc-result is now settable over QOM and clamped to 12 bits. Measured: 2200 gives
level 4 and no warning, 1200 gives level 0 and raises gLowBattery, and the level
recovers to 4 afterwards.

tools/test_battery.py covers it, and deliberately does NOT assert that gLowBattery
clears on recovery. helper/battery.c:190-204 only clears it when the level lands
exactly on 2; above that it clears gLowBatteryConfirmed and leaves gLowBattery set. So
4 -> 0 -> 4 really does leave the flag raised. The first version of this test called
that a failure -- the test was wrong, not the model. The emulator reproduces the
firmware, including behaviour that looks like a bug.
2026-08-29 07:49:25 +01:00
mckero 7ed9f61f71 Add the audio path test to the runner
Full run: 13 passed, 0 failed.
2026-08-29 06:13:13 +01:00
mckero a1aca9395c Document why there is no audio to model
Records the scope plainly, because "add a speaker and a microphone" is the obvious
request and the answer is that neither is on the MCU: no audio samples exist in its
address space, so there is nothing to capture, nothing to play, and nothing for a
browser permission to carry. Same line as the analogue RF limit.

Also records the stub lesson: QmpClient.command returns the unwrapped value, the test
stub returned an envelope, and the divergence let 88 tests pass while the live UI
returned 500. A stub more forgiving than the real client is worse than none.
2026-08-29 06:06:55 +01:00
mckero 7d299dbad2 Model the audio path, which is a single enable line and no samples
Asked for a speaker and a microphone. The honest answer is that neither exists to
model: on the real radio neither passes through the MCU. Receive audio is demodulated
inside the BK4819 and leaves it as analogue on its AF pin; transmit audio goes from
the microphone into the chip's own ADC. The firmware's entire involvement is

  - PA8, the amplifier enable (GPIO_EnableAudioPath, driver/gpio.h:34)
  - REG_47, which AF source the chip routes
  - REG_64, a level it displays

No audio samples exist anywhere in the MCU's address space, so a device model has
nothing to capture or play, and a browser has nothing to be granted permission for.
Synthesising sound would be inventing data the firmware never produced.

What is real is the firmware's intent, and PA8 states it exactly. TYPE_UVK5_AUDIO
watches that pin and exposes read-only speaker-on; the web UI shows it as a speaker
glyph beside the power state, and /api/status reports it. Read-only deliberately:
letting a test write it would only let the test lie to itself. The page asks for no
audio permission, and a test asserts it never will -- no getUserMedia, no
AudioContext, no <audio>.

Measured: amplifier off while idle in power save, on after SIDE1 engages monitor, and
still on afterwards rather than blipping.

One bug found the hard way, and the stub was the cause. QmpClient.command returns the
unwrapped value and raises on error, but the test stub returned {"return": ...}. So
webui.py was written to unwrap a second time, every test passed, and the live UI
returned 500 with "argument of type 'bool' is not iterable". The stub is now pinned to
the real contract by a test. A stub more forgiving than the real thing is worse than
no stub.

Also stopped swallowing the failure: a bare `except: return None` made the error
indistinguishable from a radio that simply was not making sound, and cost a detour
into looking for a stale process.
2026-08-29 06:05:54 +01:00
mckero 8d1a1c4415 Add a test runner, and a test that it can fail
The suite had grown to ten separate invocations that had to be remembered and pasted
in the right order, which is how regressions slip through: it is too easy to run the
two tests near what you changed and miss the one that broke. Now:

    bash tools/run_tests.sh        # everything, 11 suites, a few minutes
    bash tools/run_tests.sh -q     # unit tests only, ~15 s, no emulator

The build is checked first and a failure stops everything, because ninja leaves the
previous binary in place and the tests would otherwise report results for code that
was never compiled.

Two defects in the runner's own first draft, both caught before it was trusted:

It used `if "$@" | sed ...; then`, which tests sed's exit status rather than the
test's. sed practically always succeeds, so every test would have been counted as
passing no matter what failed -- a runner that silently cannot fail is worse than no
runner. Fixed with PIPESTATUS[0], and tools/test_run_tests.sh now asserts that a
failing test is counted and named, that the runner exits non-zero, and that the
accounting survives binary noise in test output.

That noise was the second defect: gdb-driven tests emit stray bytes, which made the
combined log a "binary file" as far as grep was concerned and silently swallowed the
summary line. Output now passes through tr -cd first.

Full run: 11 passed, 0 failed.
2026-08-29 05:42:11 +01:00
mckero 95bad1614e Note that counting distinct frames proves less than it looks
With RSSI varying every poll, the meter redraws constantly, so "consecutive frames
differ" is true on a completely parked radio. Records the framebuffer page layout so
a screen test can compare the rows that answer its actual question, and the fact that
page 4 is not the meter row.
2026-08-29 05:29:10 +01:00
mckero 3d1dad0617 Guard scanning against the always-busy receiver
bk4819_eval_receiver reports RSSI above any sane squelch threshold on every poll, and
a scan halts when it finds a busy channel -- so the S-meter work could plausibly have
stopped scanning dead on its first step. Measured: it does not, 6 distinct tunings
over 6 samples.

The check compares the framebuffer rows holding the frequency digits rather than
counting distinct frames. Frame comparison would pass on a completely parked radio,
because RSSI varies every poll and the meter redraws constantly -- an initial run
scored 8/8 distinct frames and proved nothing. Narrowing to pages 1-2 answers the
actual question: did the radio retune.

Incidentally established that page 4 is not the meter row: it was byte-identical
across all six samples while the frequency changed.
2026-08-29 05:28:31 +01:00
mckero a310429eff Correct the docs that said PTT did not exist
Three places claimed there was no PTT line, and one of them named the wrong pin
(GPIOC rather than PB10). All now say the same thing: PTT works, but not through
the key table, because the firmware reads its own pin instead of scanning it as a
matrix key.

Documents the transmit bar's two gates -- FUNCTION_TRANSMIT and gSetting_mic_bar,
the latter already on because blank flash reads 0xFF -- and why the release path
gets more attention than the press: a stuck PTT leaves every later test running
against a transmitting radio.

Also records the '.key' vs '.key[data-key]' trap for whoever adds the next
non-key button.
2026-08-29 05:18:57 +01:00
mckero 88994b1bd0 Add PTT, which brings the transmit level bar with it
PTT was the one input the model never had, so the radio could not be keyed and the
mic level bar was unreachable. It is not a matrix key -- GPIO_IsPttPressed reads
PB10 directly (driver/gpio.h:31, active low) -- so it gets its own GPIO line rather
than a column/row intersection.

With it the transmit screen is complete: TX annunciator, a running timer, and a
level bar at roughly 80% of scale, fed from REG_64 via BK4819_GetVoiceAmplitudeOut.
app/app.c:1700 draws that only while gCurrentFunction == FUNCTION_TRANSMIT with
gSetting_mic_bar set; the setting is Data[7] bit 4 at flash 0xA0A8, and blank flash
reads 0xFF, so it is already on.

Exposed as a boolean on the keypad device and as POST /api/ptt with an explicit
held flag, plus a button in the browser UI. Held rather than tapped, because
transmitting is a state the operator stays in and a fixed duration would be wrong
for it.

Releasing is treated as the important half:
  - pointerleave and pointercancel release, so dragging off the button cannot leave
    the radio keyed
  - pagehide releases, so closing the tab cannot either
  - /api/release-all clears PTT too, since it is outside the matrix and an empty
    press does not touch it
  - non-boolean bodies are rejected, so {"held": "false"} cannot key the transmitter
    by truthiness

Two things the test suite caught, both real:

The browser wired '.key' handlers over every styled button, and the PTT button
carries no data-key, so it would have sent the key "undefined". Narrowed to
'.key[data-key]'.

StubClient.presses() collected every qom-set regardless of property, so PTT's
booleans landed among the key names and presses()[-1] reported False after a
release-all. Now filtered by property, with a matching ptts() accessor.

tools/test_ptt.py covers the path end to end and asserts the release as well as the
press: a PTT that stuck would leave every later test running against a transmitting
radio.
2026-08-29 05:16:42 +01:00
mckero 39195b8028 Update the notes now that the S-meter works
The squelch section said the blocker sat above the device model and told the reader
not to resume without new information. Both are now wrong, so it leads with the
outcome instead: SIDE1 engages monitor, which skips squelch entirely, and the meter
reads -53 dBm / S9+40.

The four failed attempts are kept rather than deleted. Each ended in a confident
wrong diagnosis -- power-save gating, trigger timing, bit semantics -- and the shared
cause was a one-bit read skew underneath all of them. That pattern is worth more to
the next reader than a clean account of the version that worked.

README gains the S-meter row and test_smeter.py.
2026-08-29 04:57:00 +01:00
mckero e6cebed84e Report a receiver with a signal, so the S-meter reads
The firmware now draws a working meter: -53 dBm, +40 over S9, nine of thirteen
segments, next to a MONI label and a running receive timer. The two numbers agree
with each other -- S9 is -93 dBm on UHF, so -53 really is S9+40.

Three pieces had to line up, and the order they were found in was the hard part.

RSSI and audio amplitude are refreshed when the firmware polls REG_0C, not when it
configures the chip. Raising a flag at configuration time is a trap: REG_3F is
written 0 then 0x0C0C repeatedly during setup, so anything announced there is
disabled again before it can be collected.

The squelch flag is SQUELCH_LOST, bit 2 -- not SQUELCH_FOUND. Per app/app.c:1027
"squelch lost" is what sets g_SquelchLost = true, i.e. a signal is present.
SQUELCH_FOUND reads like "found a signal" and means the opposite.

Announcing is rate-limited to every 64th poll. Announcing once means the firmware
collects it during startup, before the flag leads anywhere. Announcing on every poll
re-arms the request bit inside the firmware's own collection loop, which uses REG_0C
as its condition and has no timeout, so it never exits. Periodic satisfies both.

What finally made the meter appear was not the interrupt at all. The radio idles in
power save and does not act on squelch there. ACTION_Monitor skips squelch entirely
-- app/app.c:482 picks FUNCTION_MONITOR over FUNCTION_RECEIVE when gMonitor is set --
and settings.c:263 defaults an out-of-range stored action to ACTION_OPT_MONITOR,
which blank flash (0xFF) is. So SIDE1 short-press is the way in. Measured: fn=5
idle=1 monitor=0 before, fn=2 idle=0 monitor=1 after.

Gating on RX_DSP (REG_30 bit 0, from App/driver/bk4819-regs.h:240) rather than the
whole register being zero: TX and tone paths leave other bits set with RX_DSP clear,
and would otherwise look like a live receiver.

tools/test_smeter.py covers the whole path -- boots pristine, confirms power save,
presses SIDE1, and checks the screen gained content. It compares lit-pixel counts
rather than matching pixels, so an unrelated UI change does not produce a mysterious
failure.

None of this is radio simulation. The levels are plausible numbers that move; they
are not the result of modelling a signal. What they buy is firmware control flow
running on live values instead of on zero.
2026-08-29 04:55:33 +01:00
mckero 7faacfd433 Record what the read-path fix revealed about the squelch attempts
The shifted-read bug invalidated four earlier diagnoses, so the notes claiming a
power-save gate or a timing problem were wrong and are corrected.

With reads fixed the interrupt handshake demonstrably works: RAISE pending=0004,
ACK delivering flags=0004, correct bit, collected by the firmware, REG_0C back to
0x0000 so nothing hangs.

g_SquelchLost is still 0 and the S-meter still absent, but the shape of the problem
is now clear and recorded: announcing once lands during startup before the flag
leads anywhere; announcing every poll re-arms the bit inside the firmware's own
untimed collection loop and spins forever; announcing periodically avoids both and
still changes nothing. The blocker is getting the radio into a receiving state at
all, which sits above the device model.

Also records the general lesson, which cost the most time here: when several
independent attempts fail in the same way, suspect the shared transport rather than
the logic layered on top of it.
2026-08-29 04:40:07 +01:00
mckero ad88ee1519 Fix BK4819 register reads arriving shifted one bit left
Every read came back doubled: seed REG_0C with 0x1248 and the firmware received
0x2490. The command byte's own trailing falling edge was being treated as a data
clock, so bit 15 was shifted away before the guest sampled it and the whole word
landed one place too high.

Each firmware bit is read/raise/lower (BK4819_ReadU16), which means the eighth
command bit is followed by a falling edge before the data phase begins. Skip that
one edge.

Why it went unnoticed: writes were always fine -- 52 registers held exactly what
the firmware wrote -- and the register the firmware polls hardest, REG_0C, was
legitimately 0 in this model. Reading zero and getting zero looks like success.
The skew only surfaced when something tried to report a value through it.

It also explains four failed attempts at the squelch interrupt. The model raised
REG_0C bit 0; the firmware received bit 1. So

    while (BK4819_ReadRegister(BK4819_REG_0C) & 1u)

was never true, the acknowledging write inside it never ran, and 1719 polls saw a
flag the guest could not act on. Every one of those attempts was diagnosed as a
timing or gating problem and was not.

tools/test_bk4819_readback.sh locks it down. It seeds REG_0C -- read ~1700 times
per 30s, so a sample is guaranteed -- with a value carrying bits in both halves,
so a shift either way is unmistakable, and names the direction on failure. Bit 0
is left clear on purpose: with it set the firmware enters an acknowledge loop that
has no timeout, and this test is about alignment only.

Confirmed by A/B: committed code 0x2490, patched 0x1248.
2026-08-29 04:27:00 +01:00
mckero 655e7883ad Correct the squelch notes: my power-save explanation was wrong
The previous commit blamed the power-save gate at app/app.c:1697 for the interrupt
never being collected. That was wrong, and the test which disproves it is cheap:
BATTERY_SAVE lives at flash 0xA00B, app/app.c:1374 refuses power save when it is 0,
and patching the byte gives

    BATTERY_SAVE=4:  fn=5 idle=1   polls=2161  acks=0
    BATTERY_SAVE=0:  fn=0 idle=0   polls=2161  acks=0

The gate passes and nothing changes. A gdb backtrace confirms the loop runs --
CheckRadioInterrupts is inlined into APP_TimeSlice10ms, which is the caller of every
REG_0C read.

Where it really stands: a breakpoint on BK4819_GetRSSI never fires at all. The
firmware does not read RSSI in this idle state, so the missing S-meter is not the
model withholding a value; the receive state machine has to be entered first, and the
squelch interrupt is an input to that rather than the switch. Three rounds of
increasingly precise instrumentation all ended at "the firmware is not asking".

Four measurement mistakes are recorded because each produced a confident wrong
conclusion, and one of them made me change working code:

- Sampling PC at the REG_0C read lands in BK4819_WriteU8; sampling LR lands inside
  BK4819_ReadRegister, since it calls BK4819_ReadU16. Use a backtrace.
- A probe printing shift_out before the assignment showed 0000 for a value about to
  be sent as 0001.
- BK4819_ReadRegister returning 0x0 for REG_0C looked like a broken read path, and I
  altered the bit timing over it. REG_0C legitimately holds 0 in the committed build.
  A read returning the register's real contents proves nothing -- test against one the
  firmware wrote, like REG_3F (0x0C0C) or REG_78 (0x2F5B).
- nexti after a breakpoint reported r0=0 from an unrelated location. finish gives the
  actual return value.

Also noted: gdb cannot call guest functions on this target, and there is no
gCurrentRSSI global -- RSSI is read and discarded, so a breakpoint plus finish is the
only way to see what the firmware received.

Code is unchanged and at the committed state; test_bk4819.py passes.
2026-08-29 02:06:50 +01:00
mckero 4b1bf14140 Settle the squelch interrupt: the blocker is not in the device model
Second attempt, and this one produced a definite answer rather than another guess.

Two real mistakes in the first version, both fixed along the way and worth recording:
squelch was evaluated when the firmware configured the chip, but the startup sequence
writes REG_3F as 0x0000 then 0x0C0C three times over, so a flag raised on the enabling
write was disabled again before anyone read it. And the threshold came from REG_4E,
whose low bits are the glitch threshold -- the RSSI open level is REG_78 bits 15:8 at
0.5 dB/step against REG_67's 0.25 dB/step, so squelch could never open at all.

With both corrected, every chip-side condition lines up: en=0x0C0C, rssi=0x01E0,
threshold 94, REG_0C correctly returning 1. The firmware still never acknowledged,
and the reason is outside the chip:

    gCurrentFunction=5 (FUNCTION_POWER_SAVE), gRxIdleMode=1

against the gate at app/app.c:1697,

    if (gCurrentFunction != FUNCTION_POWER_SAVE || !gRxIdleMode)
        CheckRadioInterrupts();

Both halves false, so the interrupt loop is never entered and nothing can collect the
flag. The emulator idles in power save, so that is the steady state.

Gating on REG_30 -- zeroed by BK4819_Sleep on each power-save cycle -- does not help:
the chip is awake when the model is asked while the firmware still has gRxIdleMode=1.
Chip state and firmware state are not in step, so no condition available inside a
register model can decide this correctly.

So it is not a matter of a better trigger. A working version needs to keep the guest
out of power save, or drive the interrupt from something that knows the firmware's
receive state -- neither of which belongs in this device. Both attempts reverted; the
model is back at the committed state and test_bk4819.py passes, including the REG_0C
assertion that catches the armed-hang failure mode.
2026-08-28 17:21:33 +01:00
mckero 6dd5e34445 Record the squelch interrupt attempt, and why it was backed out
Scanning does work now that RSSI reports a real level -- long-press * and the
frequency steps, 6 distinct frames over 7 seconds. The S-meter still does not appear,
because ui/main.c:2370 draws it only when FUNCTION_IsRx(), which needs the chip to
report a squelch opening rather than merely a healthy RSSI.

I implemented that interrupt and backed it out. The guest kept running, but REG_0C
bit 0 remained set afterwards, meaning the firmware never collected the interrupt.
That leaves a hang armed: app/app.c:910 and :1417 spin on that bit with no timeout, so
any path reaching them with it stuck never returns. A missing S-meter is a cosmetic
gap; a latent hang is not, and shipping the second to fix the first is a bad trade.

Documented with the mechanism (REG_0C pending bit, REG_02 acknowledge-then-read,
sqlFound at bit 3), the reason it failed, and the check that says a retry is correct:
REG_0C reading 0 afterwards, which tools/test_bk4819.py already asserts. The next
person should find out why the flag was not collected rather than raise it at a
different moment and hope.
2026-08-28 16:56:08 +01:00
mckero 4e29d11810 Update the docs now that the BK4819 is modelled
The "what this cannot do" section said the transceiver was not modelled and treated
that as permanent. Half of it is now wrong: the register interface works. The other
half is still true and worth keeping sharp -- analogue behaviour is out of reach
because the chip has no public datasheet, so the driver is the only specification and
it can only say which registers were written.

Rewritten to separate the two: what is modelled and what it fixed (RSSI was hard zero
at 18 call sites), the two constraints the untimed spin loops impose, and then the
line that does not move. Explicitly warns against reading the new test as evidence
about RF.

Also corrects the GPIO comment about idling PB9 low. It described the pin as a
workaround pending a device model; that model now exists, and the idle level only
covers the window between reset and the bus being wired.
2026-08-28 16:46:55 +01:00
mckero 76fc72d5ed Model the BK4819 register interface
The transceiver was not modelled at all. Its bit-banged three-wire bus went nowhere,
so every register read returned whatever the floating GPIO happened to be, and PB9
had to be idled low as a workaround: with the line high, reads came back 0xFFFF and
RADIO_SetupRegisters spun forever on bit 0 of REG_0C.

Now a real device, wired to the pins the driver uses -- CS on PF9, SCL PB8, SDA PB9
with both directions connected -- decoding the protocol from App/driver/bk4819.c:
CS low, eight bits of register number MSB first with bit 7 set for a read, then
sixteen bits of data. Registers read back what the firmware wrote; the few it reads
without having written return plausible values.

Scope, deliberately narrow: this is the register interface, not the radio. The chip
has no public datasheet, so the driver is the only specification available and it can
only say which registers were written, never what left the antenna. Keying envelopes,
spurious emissions and sensitivity still need a real radio and a spectrum analyser.
The comments say so at the top of the device, so a passing test here is not mistaken
for evidence about RF.

What it buys is control flow that evaluates real values. RSSI was hard zero at 18
call sites -- -160 dBm -- so the S-meter read empty and squelch and scan decisions
saw a dead band. It now reports about -40 dBm. Measurably, the main screen comes up
on 400 MHz instead of the 18 MHz floor, because band setup no longer reads zeros.

Two things the untimed spins force:

- REG_0C bit 0 must stay clear. App/app/app.c:910 and :1417 loop on it with no
  timeout whatsoever, so a stuck bit hangs the guest rather than degrading.
- A soft reset (REG_00 bit 15, which BK4819_Init issues first) has to re-seed the
  measurement registers. Real hardware keeps measuring afterwards; this model would
  be left holding zeros. That was not theoretical -- the first test run decoded 48
  registers correctly and still reported RSSI as 0 for exactly this reason.

The register file is exposed over QOM as regNN so tests can inspect it without gdb.
That matters beyond convenience: attaching a debugger pauses the guest and changes
timing-sensitive behaviour, which has repeatedly produced conclusions that were
artefacts of the measurement rather than facts about the firmware.

tools/test_bk4819.py checks the guest still boots (i.e. the spin terminates), that
dozens of registers hold written values (52 currently, so the transfer really is
being decoded), that RSSI is not zero, and that REG_0C bit 0 is clear.

keypad_test.py, test_flash_persist.py, test_freq_entry.py, test_serial_rx.py and the
143 unit tests all still pass.
2026-08-28 16:45:29 +01:00
mckero f3412f3316 Correct the docs: serial receive is implemented now
The previous commit documented serial receive as a permanent limitation, listing
what would be needed to build it. It has since been built, so that section was
actively misleading -- it told the reader not to try something that already works.

Replaced with what it takes to keep working, since each of the three pieces fails
silently on its own: USART1 needs a chardev, DMA must decrement CNDTR for USART
(serviced on the CNDTR read, which is where the driver looks), and DR writes must
reach the chardev and not just stderr. The last one is the trap -- transmit that
goes only to stderr is indistinguishable from the firmware ignoring the command.

README status table and build-verification steps updated to match.
2026-08-28 16:26:55 +01:00
mckero b84d229326 Implement serial receive, unlocking the UV-K5 programming protocol
Nothing could be sent to the firmware. Three separate pieces were missing.

USART1 was a register stub with no chardev, so there was no source of incoming
bytes. It now takes a chardev property, defaulting to serial0, so -serial works.

The DMA model never serviced USART. App/driver/uart.c receives over a circular
peripheral-to-memory channel and never reads DR; it finds new data with

    write_ptr = sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(DMA1, CHANNEL_2)

so leaving CNDTR at its programmed value made the buffer look permanently empty
however many bytes arrived. DMA now drains USART1's queue byte by byte, decrementing
CNDTR and reloading it in circular mode. Channels also remember the length they were
given, since CNDTR counts down and the offset into the buffer has to be derived from
the difference.

Servicing happens on a CNDTR read rather than from a timer: that read is precisely
how the driver looks for data, so no polling is needed and no byte can be delivered
before the guest asks.

And transmit was invisible to the far end. DR writes went to stderr only, so a host
tool would send a command, the firmware would answer, and the answer went nowhere the
tool could see -- indistinguishable from being ignored. This cost a debugging round:
the first test run reported "no reply at all" with 0 bytes of boot output, which
looked like receive failing when the boot banner was in fact being written to stderr
as always. DR now also writes the raw byte to the chardev when one is connected.

SR reports RXNE when bytes are queued and DR consumes one, so a polling firmware
would work too, even though this one uses DMA.

tools/test_serial_rx.py speaks the real wire protocol -- AB CD framing, the fixed XOR
obfuscation, CRC-16/XMODEM -- and checks two exchanges end to end:

    0x0514 hello        -> 0x0515 ack
    0x051B EEPROM read  -> 0x051C with the requested 8 bytes at 0x0E70

Both pass. CPS/CHIRP-style tools can now talk to the emulator. keypad_test.py,
test_flash_persist.py, test_freq_entry.py and the 143 unit tests still pass.
2026-08-28 16:25:37 +01:00
mckero cdb73f13a0 Record that serial receive is not implemented
Found while explaining a DMA channel that stayed permanently armed with cndtr=256
during the flash investigation. It is USART1's receive channel, running in
LL_DMA_MODE_CIRCULAR, so never completing is correct behaviour and not a bug.

But it exposed a real gap. driver/uart.c derives its write pointer from
sizeof(UART_DMA_Buffer) - LL_DMA_GetDataLength(...), and the DMA model only
services SPI, so CNDTR never decrements for USART and that expression is always
zero. Combined with USART1 being a py32-stub with no chardev backend, nothing can
be sent *to* the firmware.

The cost is specific: UART_IsCommandAvailable never fires, so the UV-K5 programming
protocol in app/uart.c is unreachable -- 0x0514 handshake, 0x051B EEPROM read,
0x051D EEPROM write, 0x05DD reset. CPS/CHIRP-style tools cannot talk to this
emulator. Transmit is unaffected, which is why the firmware banner shows up fine
and this went unnoticed.

Documented rather than fixed: it needs a chardev on USART1 plus circular-mode DMA
driven by receive, which is a new feature rather than a repair. The notes say what
would be involved so the next person does not have to rediscover the mechanism.

Status table also updated to reflect what the flash and DMA fixes settled --
persistence and frequency entry now work.
2026-08-28 15:52:05 +01:00
mckero 65e50c0065 Stop cleanup killing unrelated processes, and recover when it happens anyway
Two faults that combined to take the web UI down twice in one session, each time
surfacing to the user as a 502 through the reverse proxy.

`pkill -f 'M uv-k5-v3'` in run.sh and trace_run.sh matched far more than intended.
-f tests the whole command line, so it also matched the shell running the pkill
(the pattern sits in its own argv), any script mentioning the machine type, and the
QEMU child of a running webui.py. Cleanup now lives in tools/lib_kill_emulator.sh:
pgrep -x on the binary name, confirm uv-k5-v3 in /proc/PID/cmdline, and optionally
scope to one QMP socket so a caller only stops the instance it owns.

The supervisor then could not recover from it. power_on() began with
`if self._client is not None: return False`, but a client object is not proof of a
live guest -- after an external kill the stale client made power_on refuse forever,
so the Power button was dead until the whole service was restarted. It now checks
whether the process actually exited and relaunches, logging why.

Tests:

- tools/test_kill_emulator.sh checks a plain process, a process whose command line
  merely mentions uv-k5-v3, and the calling script all survive; that a real emulator
  on a named socket is stopped; that one on another socket is not; and that an
  unscoped call still clears everything. Verified it leaves a live webui.py alone.
- Two supervisor unit tests cover relaunch-after-external-kill and the case that
  must still refuse, so this cannot regress into starting two emulators at once.

Verified end to end against the running web UI: kill the emulator from outside,
status reports unreachable, and pressing Power brings it back to
{"status":"running"} where before it stayed dead.

While writing the first version of the test I modelled the failure as a client
raising BrokenPipeError, which is not what is_running() looks at -- it checks
poll(). The mock was wrong, not the code; the test now has the process report an
exit status, which is what really happens.
2026-08-28 15:50:13 +01:00
mckero d09bf5f47e Write down the flash investigation and what made it slow
Four faults in the SPI/DMA/flash models each produced the same symptom -- stored
frequencies zeroed, a typed frequency reverting to 18 MHz -- and each was invisible
from the layer above. AGENTS.md now lists all four with the reasoning that connects
them to what the user saw, so the next person does not rediscover them one at a
time.

Also records the methodology mistakes, because they cost more than the bugs did:

- Attaching gdb between digits clears the frequency input box: the timeout is
  ~2.5 s and an attach takes ~3 s. Three wrong conclusions came from this, so
  instrument the model and read stderr instead of stopping the guest.
- A diagnostic log capped at six entries showed only 0xFF payloads and supported
  precisely the wrong conclusion. Do not cap before the shape of the data is known.
- A failed ninja leaves the old binary in place and the test still runs. Two rounds
  of results were meaningless. Check for FAILED and error: before trusting a run.
- assets/flash.img is written by every session, so a test starting from it can find
  its work already done -- which looks identical to broken persistence. Start from
  assets/pristine/, and power off before restoring, since shutdown flushes the old
  image back over the file.
- Hand-computed struct offsets gave KEY_LOCK=4 and TX_VFO=11, impossible values,
  because the ELF has no DWARF and the structs contain enums. Use an unambiguous
  nm symbol, or find the field by toggling it and diffing.
- One probe printed phase before incrementing it, making a correct address decoder
  look off by one. A working implementation was nearly "fixed" as a result.

README gains the two new tests in the build-verification step and the tools list.
2026-08-28 15:32:45 +01:00
mckero 798905f154 Give DMA the CPU's address space, so a typed frequency sticks
DMA moved bytes through address_space_memory, which cannot decode this SoC's memory
at all: the container region holding flash, SRAM and the peripherals is handed only
to the ARMv7M core and never registered with global system memory. Reads came back
MEMTX_DECODE_ERROR with all-zero data, and writes went nowhere. Proved directly --
an address_space_read of SRAM through it returns result=2 and 00000000, while the
same address read through the container returns the real contents.

This is what "the frequency will not change" and "flash behaves like RAM" had in
common. PY25Q16_WriteBuffer reads a 4 KB sector into SectorCache, patches the part
it wants, and programs the whole sector back. The read looked healthy from the flash
side -- the model handed over real 0xFF bytes -- but DMA dropped them, so the
write-back sourced 4096 zeros and cleared the sector, VFO frequencies at 0x9000
included. RADIO_ConfigureChannel only substitutes a band's lower limit for
0xFFFFFFFF, so a stored zero was used as-is and clamped to BX4819_band1_lower.
That is where the 18.000 MHz came from, every time.

DMA now runs over an AddressSpace built on the SoC container, and refuses to
transfer at all if none is configured rather than silently moving zeros.

tools/test_freq_entry.py covers the whole user-visible path: type 435000, confirm
435.00000 MHz lands in band 5, confirm no other band was zeroed, and confirm it is
still there after a power cycle. It drives QMP with no debugger attached, because
the input box times out in ~2.5 s and a gdb attach takes longer -- that alone
invalidated several earlier investigations.

Verified: 435 MHz now appears at flash 0x90A0 where before the entire sector read
zero. keypad_test.py, test_flash_persist.py and the 141 unit tests all pass.
2026-08-28 15:30:26 +01:00
mckero f114666b42 Start DMA on the peripheral's request, and clock both directions together
Two related faults in the DMA model, both of which corrupted flash reads.

Transfers started when a channel was enabled. On hardware, enabling only arms a
channel; the transfer begins when the peripheral raises its DMA request. The flash
driver's SPI_ReadBuf arms RX, arms TX, then enables SPI and sets TXDMAEN -- so
firing at arm time clocked the bus before the read command had been sent. SPI now
kicks the armed channels from CR1/CR2 when SPE and a DMA request enable are both
set, which covers the read path (SPE last) and the write path alike.

Each channel also ran to completion independently. SPI is duplex: one clocked byte
is simultaneously sent and received, and the driver relies on that, pairing a
memory-to-peripheral channel feeding dummy bytes with a peripheral-to-memory
channel collecting the reply. Running them in sequence meant TX clocked the whole
transfer out before RX looked at the bus, so RX collected nothing. They are now
stepped together, one byte at a time.

Either fault alone made a 4 KB sector read return zeros. PY25Q16_WriteBuffer reads
a sector into SectorCache, patches it, and writes the whole thing back, so a zeroed
read turned into a zeroed sector -- including the per-band VFO frequencies at
0x9000. That is a second, independent cause of typed frequencies reverting to
18 MHz, on top of the missing page wrap fixed in da1ad7e.

Verified: the frequency area stays 0xFF across a boot where it was previously
zeroed every time, and an instrumented build shows the sector read now returning
0xFF rather than 0x00.

test_flash_persist.py now starts from the pristine image rather than
assets/flash.img. A dirty live image left nothing for the boot to write, which
surfaced as "the image is byte-identical" -- a failure that looks like broken
persistence but is really a dirty fixture. Runs twice in a row cleanly now.

keypad_test.py and the 141 unit tests still pass.
2026-08-28 15:05:15 +01:00
mckero da1ad7e1e6 Wrap page-program writes within their 256-byte page
Real SPI NOR latches only the low address bits into its page buffer, so a program
burst that runs past the page boundary continues at the start of the same page. The
model incremented the address straight through instead.

Consequence: the firmware issues a 512-byte burst at 0x008F00 inside a single CS
assertion (measured -- the CS never drops mid-burst), which spilled into 0x009000.
That is the per-band VFO frequency area in eeprom_compat.c's map, so stored
frequencies were zeroed. RADIO_ConfigureChannel only substitutes the band's lower
limit when it reads 0xFFFFFFFF, so a stored 0 was taken literally and clamped to
BX4819_band1_lower -- which is why every typed frequency reverted to 18.000 MHz.

Verified: writes now align to sectors (0x8000-0x9000 and 0xA000-0xB000) and an
instrumented build records zero stores into 0x9000-0x90D6, where before it was
overwritten on every boot.

test_flash_persist.py had encoded the bug in its expectations: it watched 0x008100
and 0x00A100, which were only ever written *because* of the missing wrap. Those move
to 0x008000/0x00A000, and a MUST_NOT_CHANGE guard on the frequency area now fails if
a write spills there again.

keypad_test.py still passes.
2026-08-28 14:41:20 +01:00
mckero 0b879227e5 Do not mistake a leftover socket file for a running emulator
Power on returned HTTP 500 with a ConnectionRefusedError traceback. A unix socket
file outlives the process that created it, so a killed QEMU left
/tmp/uvk5-qmp.sock behind; wait_for_socket only checked os.path.exists, returned
immediately, and the connect then failed. It now probes with a real connect, which
distinguishes "listening" from "leftover file".

Two related hardenings:

- power_on cleans up if connecting fails. Otherwise a half-started QEMU keeps
  running untracked, holds the socket, and blocks the next power on -- which is
  how one stale socket turned into a repeatable failure.

- The route reports a failed power action as 503 with the reason, instead of a 500
  and a traceback the browser cannot display.

Three tests cover the stale socket, a real listener, and a path that never appears.
2026-08-28 14:17:50 +01:00
mckero 23385d2eba Persist flash writes to the backing file
The SPI NOR model read its image at realize and never wrote back, so the "flash"
was a g_malloc buffer: everything the firmware saved -- settings, edited
frequencies, channel data -- vanished when the QEMU process exited. That is the
"it behaves like RAM" the user reported, and the image on disk still had its
original mtime and was byte-identical to what make_flash.py produces.

Page-program and sector-erase now mark the image dirty, and it is written out when
chip select is released. Flushing there rather than per byte means one file write
per settings save instead of thousands, because the firmware's driver holds CS for
a whole erase-and-program sequence.

Written via a temporary file and rename: an interrupted flush must not leave a
truncated image, since that file is the only copy of the radio's state. A short
write keeps the previous image rather than replacing it with a partial one.

Also flushes from an exit notifier. Deselect covers normal operation, but QMP quit
-- which is what the web UI's power off sends -- can arrive with the chip still
selected, and the last write would be dropped.

tools/test_flash_persist.py covers it end to end on a copy of the image, so it
cannot disturb a running session. It asserts specific regions rather than just a
changed hash: flash 0x008100 (MR/VFO attributes) and 0x00A100 (settings), which are
the two sectors a boot demonstrably writes. Both were 0xFF before and non-0xFF
after, and sha256 moved from 933d6974 to 7cdff6ce.

Three approaches were tried and abandoned first, all for the same reason: driving
the guest from gdb. `call EEPROM_WriteBuffer` and `call SETTINGS_SaveSettings`
both hang, because the main loop is running and the called function waits on
hardware the debugger has frozen. Observing the file is simpler and closer to what
the user actually sees. An earlier version of the test also watched EEPROM offsets
instead of flash offsets and reported "same" for every region while persistence was
in fact working -- the two address spaces are related by the table in
App/driver/eeprom_compat.c, not equal.

keypad_test.py still passes, which matters because this file is where deleting
three fprintfs once silently removed the keypad.
2026-08-28 12:52:06 +01:00
mckero 3e743152c1 Document how the machine boots
There is no bootloader, kernel, partition table or filesystem here, and that is not
obvious from the code -- someone arriving with hosted-OS assumptions will look for
layers that do not exist.

Covers the parts that are easy to get wrong rather than restating the source:

- The reset handoff is two words. SP from 0x08002800, PC from +0x04. Verified
  against the image: 00400020 492d0008 is SP 0x20004000 (top of the 16 KB SRAM,
  matching PY32_SRAM_BASE + PY32_SRAM_SIZE) and PC 0x08002d49, which is also the
  ELF entry and the Reset_Handler symbol. The odd address is Thumb-state bit 0.

- PY32_APP_OFFSET 0x2800 is load-bearing. The first 10 KB of flash is the factory
  bootloader region, so armv7m_load_kernel gets that offset; without it the vector
  table lands in the wrong place and the first fetch faults.

- Startup is 31 lines of assembly, and the .data copy plus .bss zero-fill are the
  part worth understanding: on a hosted OS the kernel and loader do that, here
  nobody does, so a fault in either loop shows up as globals that are silently
  garbage rather than as a crash.

- SETTINGS_InitEEPROM reading fixed SPI offsets is the closest thing to mounting a
  partition. No metadata and no checksum, just an address both sides must agree on,
  so a setting that reads back wrong points at the offset before the transport.

Also notes that the ~15 s to the main loop is emulation overhead; a real radio is
up in about a second.
2026-08-28 12:38:49 +01:00
mckero 1364d46e97 Keep a pristine copy of the flash image, and a way back to it
assets/flash.img was gitignored, so the only copy of the never-booted image lived
on one disk. The emulator writes to that image, so a session can leave edited
settings or a damaged EEPROM behind with nothing to restore from.

assets/pristine/ now holds the image as first generated, gzipped and checksummed,
and is tracked deliberately. Gzip takes it from 2 MiB to 2.3 KiB because the image
is nearly all 0xFF, which is what makes keeping it in git reasonable. The live
image and its .bak-* files stay ignored.

Two checksums are recorded, for the archive and for its contents, so a corrupted
archive is distinguishable from one that was replaced.

tools/restore_flash.sh verifies, diffs, or restores. Restore backs up the current
image first, then re-checks the result, since a restore that silently half-worked
would be worse than none.

Verified by deliberately corrupting the live image: --diff reported 32 differing
bytes, restore backed up and rewrote it, and --diff then reported no change. The
image is currently byte-identical to what make_flash.py produces, so this is the
genuine original rather than a copy of something already used.
2026-08-28 12:20:13 +01:00
mckero 19997e85e1 Rename the vhost to k6v3.mckero.dn42
The name is what the user is putting in DNS. server_name has to match or SNI falls
through to another vhost on the same socket.

Addresses are unchanged: 172.21.91.140 and fd3c:3f9b:6424:2::5, still sharing 443
under the existing *.mckero.dn42 wildcard.

Verified after reload: 200 on both families with the certificate validating, 80
redirecting, and dns./mail. still 200. This time nginx -t ran after the symlink was
in place, which is the ordering that caught me out last time.
2026-08-28 11:49:51 +01:00
mckero a4a8f21d50 Document and version-control the HTTPS front end
The UI is served at https://k6v6.mckero.dn42/ with nginx terminating TLS and the
server itself now bound to loopback, so it is not directly reachable.

No new address and no new certificate: 443 is shared with the other vhosts on
these DN42 addresses and separated by SNI, and the existing *.mckero.dn42 wildcard
already covers the name. Only DN42 addresses are bound, so the public 443
listeners on this host are untouched.

docs/reverse-proxy.md records the settings that are not optional, because each has
a failure mode that is easy to misread:
  proxy_buffering off  -- otherwise the frame stream arrives in bursts
  X-Forwarded-For      -- otherwise every log line is attributed to 127.0.0.1
  long read timeout    -- a paused guest emits nothing at all
  tcp_nodelay          -- Nagle would delay exactly the latency-critical requests

Two pitfalls hit while setting it up are written down. "http2 on;" needs nginx
1.25.1+ and this host runs 1.22.1, and because nginx -t was run before the symlink
existed it passed, then reload failed and left nginx stopped, briefly taking the
other sites down. Separately, a newly added listen address needs a reload to be
bound: after the failed reload, v6 requests failed with nothing in the error log
until a second reload created the socket.

deploy/nginx-k6v6.conf keeps a copy in the repo, since nothing here
version-controls /etc.

Verified: HTTP 200 on both families with the certificate validating (no -k), 7
frames in a 20 KB stream sample, log entries attributed to real client addresses.
2026-08-28 11:39:17 +01:00
mckero f2c5c6b31b Attribute log entries to the client IP
Entries gain an "ip" field, rendered between the time and the source as asked.
The buffer is shared by every viewer, so without attribution a log of keypresses
from two people is unreadable.

Resolving the address matters more than it looks: behind the nginx reverse proxy
REMOTE_ADDR is always 127.0.0.1, so the first hop of X-Forwarded-For is what
identifies the real client. Only the first entry is trusted -- the rest of the
chain is set by the caller and a test covers that.

Power actions are logged at the route rather than in the supervisor, which has no
request context, so "who powered it off" is recorded.

Entries with no client behind them keep ip=None and render as "-": firmware serial
and QEMU stderr are not caused by a request.

The sharing and history the user asked for already worked and needed no change --
verified rather than assumed. The front end starts at logCursor=0, so a page
opened now receives the full buffer, including lines produced before it connected
and lines from other people. Confirmed live through the proxy: a new reader saw
entries attributed to 172.21.91.140, fd3c:3f9b:6424:2::5 and "-".
2026-08-28 11:37:06 +01:00