flash controller: store ACR/OPTKEYR instead of swallowing them, which is what stopped the factory bootloader from starting
slots over the firmware's own serial protocol (0x0720 family); uvk5_socket/uvk5_testenv so a fresh checkout skips instead of failing; web UI slot table and Multiboot button; quick start, CONTRIBUTING, and stop tracking firmware images and radio dumps
Two faults that combined to take the web UI down twice in one session, each time
surfacing to the user as a 502 through the reverse proxy.
`pkill -f 'M uv-k5-v3'` in run.sh and trace_run.sh matched far more than intended.
-f tests the whole command line, so it also matched the shell running the pkill
(the pattern sits in its own argv), any script mentioning the machine type, and the
QEMU child of a running webui.py. Cleanup now lives in tools/lib_kill_emulator.sh:
pgrep -x on the binary name, confirm uv-k5-v3 in /proc/PID/cmdline, and optionally
scope to one QMP socket so a caller only stops the instance it owns.
The supervisor then could not recover from it. power_on() began with
`if self._client is not None: return False`, but a client object is not proof of a
live guest -- after an external kill the stale client made power_on refuse forever,
so the Power button was dead until the whole service was restarted. It now checks
whether the process actually exited and relaunches, logging why.
Tests:
- tools/test_kill_emulator.sh checks a plain process, a process whose command line
merely mentions uv-k5-v3, and the calling script all survive; that a real emulator
on a named socket is stopped; that one on another socket is not; and that an
unscoped call still clears everything. Verified it leaves a live webui.py alone.
- Two supervisor unit tests cover relaunch-after-external-kill and the case that
must still refuse, so this cannot regress into starting two emulators at once.
Verified end to end against the running web UI: kill the emulator from outside,
status reports unreachable, and pressing Power brings it back to
{"status":"running"} where before it stayed dead.
While writing the first version of the test I modelled the failure as a client
raising BrokenPipeError, which is not what is_running() looks at -- it checks
poll(). The mock was wrong, not the code; the test now has the process report an
exit status, which is what really happens.
Power on returned HTTP 500 with a ConnectionRefusedError traceback. A unix socket
file outlives the process that created it, so a killed QEMU left
/tmp/uvk5-qmp.sock behind; wait_for_socket only checked os.path.exists, returned
immediately, and the connect then failed. It now probes with a real connect, which
distinguishes "listening" from "leftover file".
Two related hardenings:
- power_on cleans up if connecting fails. Otherwise a half-started QEMU keeps
running untracked, holds the socket, and blocks the next power on -- which is
how one stale socket turned into a repeatable failure.
- The route reports a failed power action as 503 with the reason, instead of a 500
and a traceback the browser cannot display.
Three tests cover the stale socket, a real listener, and a path that never appears.
Three sources into one buffer: power events from the supervisor, QEMU's stderr
(which run.sh and the tests used to discard), and firmware serial, which the
machine model tags SERIAL. default_launcher now captures stderr rather than
sending it to DEVNULL, which is what made the last two reachable.
The pane is a fixed-height scroll box as asked: 180px with overflow-y:auto, so it
never grows with content -- older lines move up out of view and you scroll back to
read them.
Two details that make that usable rather than annoying:
- Autoscroll only sticks when you are already at the bottom. Otherwise a new line
arriving would yank the view away from whatever you had scrolled up to read.
- MAX_LOG_LINES caps the <pre> as well. The box is fixed-height either way, but an
unbounded DOM node would still grow memory across a long session.
Verified on the live server: power on produced power/qemu/serial lines including
"UV-K5 Firmware, EGZUMER-F4HWN+NR7Y c91cec95", each Reset logs the event and the
banner reappearing, and since= never resent a line. A capacity-500 buffer fed 2000
lines keeps exactly 500 and does not replay evicted entries to a stale cursor.
Power on/off cannot live inside the QMP connection: QMP quit destroys the socket
a later power on would have to arrive through. So something outside it has to be
able to spawn the process again.
Off then On is a cold boot -- process replaced, guest from reset -- which is the
behaviour asked for: like cutting mains power and restoring it. system_reset is the
warm alternative and keeps the process.
adopt() is for attaching to a run.sh instance. power_off then refuses, because we
did not start that process. is_running() also reports False for a process that
exited on its own, rather than trusting our own bookkeeping.
power_off tolerates quit raising: the socket usually drops before the reply
arrives, so that is success rather than an error. The launcher clears a stale
socket first, since QEMU failing to bind presents as On doing nothing.
Verified against real QEMU: starts with 0 processes, On gives 1, Reset keeps the
same one, Off returns to 0, On again cold boots. 14 unit tests with a fake
launcher, so they need no emulator.