Decoding around the return address 0x0801701c puts the call inside a function whose prologue is at 0x08016ff0. The literal pool near it holds 0x20001b40..0x20002811 and 0x2000000d/0x0e -- settings and EEPROM-area RAM addresses -- and 0x20000280 appears nowhere, so the pointer handed to memset is computed, which is what PY25Q16_OverlayBuffer() looks like.
My own disassembler misaligned badly here, printing nonsense branch targets like 0x8741e36, which is the tell that the halfwords are not paired the way Thumb-2 requires. The two facts above survive because they rest on the prologue shape and on the literal values, but no instruction-level claim from this round should be trusted.
Round 59 still stands and is what matters: the overlay is zeroed about a hundred times a second while the code is loaded once, so nothing the loader writes can survive. Next measurement needs no disassembler: at those hits SP is 0x20003c50, so reading the words above it gives the return-address chain and names the retry loop.
Logging every DMA run over 512 bytes with destination and count for twelve seconds after MENU gives exactly one run into the overlay, count=2568, which is code_size exactly. Round 58's short-read conclusion is withdrawn: the copy is complete and correct, as rounds 53 and 56 showed from the byte side.
What matters is the comparison with round 58's watchpoint: that word is written 1200 times in the same sort of window, every time by memset(0x20000280, 0, 0x1000), while the DMA writes the overlay once. The code is loaded once and the 4 KiB overlay is zeroed roughly a hundred times a second.
That fits what was puzzling: the 24-byte marker app ran because its CRC window is 24 bytes and the check follows the load immediately; Minesweeper's window is 2568 bytes and is far more likely to be hit first. It also explains the size ladder -- short apps survive the race, long ones do not. Next: name what zeroes the overlay, one function up from the memset's return address 0x0801701c.
A hardware write watchpoint on 0x20000C0C (overlay + 2444, the first byte wrong for the 2568-byte app) yields exactly one distinct writer: PC=0x0801ae7a, LR=0x0801701c, fill byte 0x00, 1200 hits, with r0=0x20000280, r2=0x20001280, r3=0x20000c0c. That is memset(ws, 0, 0x1000) -- APP_LaunchOverlay's own zeroing of the 4 KiB overlay -- called twelve hundred times: the launch path is being retried.
That the only writer to that word is the zeroing is the important part, because the word ends up holding 40 d6 01 08. Something leaves it non-zero without any store the watchpoint saw, and the one way that happens is if the copy never covered it: the memset zeroes 4096 bytes, the read fills code_size, and a read short by about 124 bytes leaves the tail zero and fails the CRC.
That reconciles round 56, whose window probe looked at one transfer -- the one that landed -- and found it correct; there are many transfers and it did not measure whether each carries code_size bytes. Next: log the count of every DMA run whose destination is the overlay and see whether some are 124 bytes short.
Polling the overlay hash every 250 ms from MENU, with no probes: it holds the previous content before MENU (169b5c51), becomes 4f23f6f3 at +0.25 s, and every later sample is the same value. So the load completes, the overlay ends up wrong, and nothing touches it afterwards.
With round 56 showing the transfer that lands in the overlay carries the correct bytes at indices 0, 1, 2444 and 2445, and no other run in that log having an rx address inside the overlay, the writer is not the DMA: it is a CPU store inside the quarter second after MENU, at overlay + code_size - 124. The instrument to name it is a write watchpoint on that word, which the gdbstub supports as Z2 -- the halt is the measurement rather than a perturbation, the one case where the advice against attaching a debugger does not apply.
Also: the round 56 Chinese note is in, appended at the end of the file, because the anchor the append script picks keeps landing on a fenced code block -- the same failure recorded in round 54. Heading counts still match.
A window probe on the DMA loop logs the byte the transfer actually carried at fixed indices. For the run whose rx_addr is the overlay the last lines are idx 0 f0 (correct), idx 1 b5 (correct), idx 2444 00 (correct), idx 2445 00 (correct). Earlier lines belong to other, shorter transfers carrying 0xff or the FAP1 bytes, so the load is a sequence and the one that lands carries the right bytes everywhere, including the window that ends up wrong.
DMA and flash paths are exonerated for the last time. That also opens the possibility that the overlay is correct at transfer time, the CRC passes, the app actually runs, and the six bytes are a post-mortem symptom rather than the cause -- the next measurement reads the overlay back within a fraction of a second of MENU and hashes it.
Own mistake recorded: the window probe was meant to write four lines and wrote 2088, because a condition on the loop index alone also matches every short transfer that starts there. This file warns about capping a diagnostic before knowing the shape of the data; this was the same rule in the other direction.
tools/check_docs.py enforces that the two files have the same number of headings, and the restored section added one. Made it bold text instead, which keeps the content and the order intact. Rounds 19 and 22 are covered by the combined 18/23 notes and round 52 by the 51 note, so the content parity is complete even though three snippet files were never written.
A round-by-round comparison shows AGENTS.zh-CN.md missing rounds 18, 19, 22, 23, 44, 49, 51, 52 and 54 while AGENTS.md has them all. The append helper picks its anchor from the fifth non-blank line of a window after a marker and refuses a non-unique anchor; in the Chinese file that anchor is often a code fence or a repeated phrase, so the write was skipped each time and nothing complained. The bilingual pair is a stated requirement, so this drifted for many rounds unnoticed.
The missing notes are restored from the work/rNN-zh.md snippets they were generated from, appended under a heading that says so, rather than being silently re-run.
The read path returns s->data[(s->addr++) % PY25Q16_SIZE] and the model writes its 2 MB back on exit, so the file a run leaves behind is what the model believed the flash held: flash-r52.img and flash-r53.img both hash to the app's 60234c72 with zero differing bytes at slot 1's code offset. The flash model hands back the right bytes.
Round 53 showed the DMA run is right too (count=2568, rx_addr=0x20000280, rx_inc=1, counts equal), so the transfer wrote the correct bytes to the correct place. Yet the overlay ends with 40 d6 01 08 28 0a at code_size-124 -- a firmware flash address at a position only the loader and the DMA know about, which is what a staged write or an interrupt frame looks like.
Next: have the DMA probe dump the bytes it actually wrote in that window, which separates 'the transfer wrote them' from 'something wrote them afterwards' in one run. Three independent instruments now agree the source bytes and the transfer are correct, so every earlier reading of this as a flash or DMA fault is retired.
A model diagnostic (UVK5_DMA_PROBE) prints every DMA run over 512 bytes. The app load is one run and it is correct in every field: count=2568, rx_addr=0x20000280 (the overlay), rx_inc=1, tx_cndtr=rx_cndtr=2568. So the DMA delivers everything to the right place.
The overlay still ends up with six wrong bytes at offsets 2444..2474, exactly code_size-124 for both sizes measured (2744->2620, 2568->2444). With the addresses and counts right, the bytes themselves must be wrong: they are what the flash model returned through py25q16_xfer. The firmware's flash probe logs the command and length but not the bytes, which is why this read as a memory corruption for so long. Next: log the bytes the flash model hands back near the end of a long read and compare with the image.
At 2744 bytes the six corrupted bytes sit at offsets 2620..2650; after trimming the same app to 2568 they sit at 2444..2474. Both are exactly 124 bytes from the end, so it is not a fixed address being clobbered but the last 124 bytes of the read. That fits what already worked: the 24-byte marker app passed and ran because 24 is shorter than the damaged tail, while Minesweeper at 2568 and 2744 always fails with APP_ERR_CRC. Same signature as the four DMA/flash faults already documented, and it retires the size ladder.
Direct evidence that the model, not the image, is at fault: the flash image at slot 1's code offset hashes to the header's 0x3c12630d with zero differing bytes, while the guest's overlay comes out 0x1312d98c.
Also recorded: trimming the app (removing the three temporary marker probes, shortening the header string, dropping the cursor readout) took it from 2744 to 2568 bytes with no warnings -- worth keeping, but a tail bug cannot be dodged by shrinking. Next: the DMA-driven flash read in py32f071.c, looking for where the final partial burst is counted or addressed.
Installing the 24-byte marker app into slot 1 instead of slot 0 changed everything: after F, 7, DOWN, MENU the fixed address reads efbeadde (the app's first instruction executed), the overlay starts with the app's own bytes 81b0034803490160, and the panel holds 615 non-zero bytes instead of the launcher's 43. entry() is reached, the app runs and it paints -- the display path, loader, header and CRC were never the problem.
Two independent causes. The app menu's selected row is not slot 0, which the overlay contents gave away (it held Minesweeper's code when Mark was in slot 0). And the overlay's tail is clobbered before the CRC: six bytes at offsets 2620..2650 that the app holds as zero and the overlay holds as the firmware pointer 0x0801D640.
The second explains the size ladder that cost several rounds: the clobbered offset is fixed, so an app shorter than it passes and runs, and one that reaches past it fails with APP_ERR_CRC -- Minesweeper at 2744 bytes is only about 124 bytes past the first corrupted byte. Next: shrink the app under that point, or find what writes at SectorCache+2620 (any cached sector read does ReadBufferRaw(SecAddr, SectorCache, SECTOR_SIZE), so a firmware path caching a sector during the load is the likely writer).
The header says code_crc32 0x3c12630d and the app's own code hashes to exactly that, so the app is not at fault. But the 2744 bytes actually sitting in the overlay, read with memsave from the running emulator, hash to 0x1312d98c. APP_LaunchOverlay computes MB_Crc32Bytes over that buffer and compares it with the header, so it takes the APP_ERR_CRC branch at line 1117 and returns before entry(&app_api) -- which is why the app never runs and why the code is nevertheless in the overlay: the copy happened, the verification did not pass.
Six bytes differ, at overlay offsets 2620, 2621, 2622, 2623, 2627 and 2650 (0x20000CBC onwards). The app's code is zero at every one; the overlay holds 40 d6 01 08 -- a firmware address, 0x0801D640 -- and two single bytes, which is what a stack frame or a global written by other code looks like.
That fits the warning app_api.h itself carries: the app runs from the RAM buffer that is also the PY25Q16 sector cache. Next: find what writes at 0x20000CBC, most likely an interrupt handler, which would make this timing-dependent and would explain why the size ladder once looked like a copy that finished late.
The function is in App/radio.c, 179 lines with exactly one loop: while (1) { if ((ReadRegister(REG_0C) & 1) == 0) break; write(REG_02,0); delay(1); }. The model keeps REG_0C at 0x0000 -- its own comment records that this exact hang was fixed once -- and reading the running emulator's reg0c over QOM shows bit 0 clear at every sample before and after the launch, so the loop exits immediately and cannot be where the firmware stops.
The marker test survives review: the blob decodes as sub sp,#4; ldr r0,[pc,#12]; ldr r1,[pc,#12]; str r1,[r0] with 0x20003F00 and 0xDEADBEEF in the literal pool, so an app that runs writes the magic and it never appears -- entry() is not reached. Why is open: with the copy present in the overlay and the CRC matching on paper, the candidates are the two early returns before the call (VMA and CRC) and whether that copy came from this launch.
Method notes: my Thumb bit-field decode was wrong twice this round (Rn/Rt are the low six bits and halfwords are little-endian), which made a correct instruction look like another -- the bytes were right and my reading was not; and the BK4819 register file is readable over QOM at /machine/bk4819, a probe-free way to see what a polling program sees.
A 24-byte app whose first instruction stores 0xDEADBEEF at a fixed address settles it without inference: installed in slot 0 and launched with F, 7, DOWN, MENU, that address reads back all 0xFF at t+2, +5, +10 and +15 s and never as the magic word, so the app's entry point is never executed and APP_LaunchOverlay does not reach entry(&app_api) at line 1162.
Between the copy at line 1113 and the call at 1162 the only return is the CRC check, and the CRC is right (MB_Crc32Bytes is an ordinary CRC-32 and the host tool's value is exactly what it computes), so the function is stuck in between -- and the only hardware-touching statement there is RADIO_SetupRegisters(true) at line 1159. The PC in the PY32 SPI routine at 0x08005118 and the panel frozen at the launcher's box both agree.
Method note: 0x20003F00 is not a free address -- it sits under the top of SRAM and the firmware was using it. The marker test is unaffected, but the free-address assumption was mine and wrong. Also adds the round-40 FillFF variant, which now builds with the documented command line (the build script had been breaking its own compile command); both new apps compile with no warnings.
A memsave of 0x20000280 after MENU holds the app's own code, 2738 of 2744 bytes identical; the PC probe's 14505 samples contain none inside it (the CPU never enters the app) and sit at 0x08005118/0x08005122, which decodes as a PY32 SPI byte transfer polling TXE then RXNE on 0x40013000 -- so the firmware is looping, not hung, matching the 1.93 million panel transfers.
APP_LaunchOverlay, read from the firmware's own app_overlay.c, validates the slot, requires link_vma == the overlay buffer, invalidates the cache, memsets, reads exactly code_size bytes, checks the CRC, then calls entry(&app_api) at line 1162. For this app every gate passes (magic FAP1, hdr 1, abi 1, api_min 1, committed, caps 0, code_size 2744 <= APP_OVERLAY_MAX, entry_off 0, link_vma 0x20000280) and MB_Crc32Bytes is an ordinary CRC-32, so the host tool's 0x3c12630d is exactly what the firmware expects.
So the copy happens, the checks pass, and the call does not take. Between them the only hardware-touching statement is RADIO_SetupRegisters(true) at line 1159, and the PC is in a PY32 SPI routine: the next round is one function wide.
The panel probe now accumulates in memory and writes a buffer at a time instead of once per byte; nothing else changed. Transfer volume went from 28416 pixel-data lines to 1930061, sixty-eight times more, which measures how much the old probe suppressed the guest.
Replaying the model's own store rule over the whole log (a0=1, col 4..131, gram[page & 7][col - 4]) and counting non-zero bytes gives 43 -- exactly what the gram property reported and matching round 42's frozen frame md5. All 1.93 million pixel bytes are accounted for, across all eight pages and all 128 columns, so nothing is dropped and no page is stale.
So the firmware is not failing to draw: it redraws the same 43-lit-byte screen -- the launcher's title box -- about nineteen thousand times. That is also why zeros dominate, and why the glyphs round 40 read as Minesweeper's frame (41, 40, 7f) are the launcher's own characters sent repeatedly. Rounds 39 and 40 were reading the launcher and calling it the app; rounds 32-35 were right that the app does not run.
Eight reads of the controller's RAM over thirty-five seconds on one QMP connection with no probe on: 485 non-zero before the keys, 43 at t+2 s, then identical md5 (43) at t+5, 10, 15, 20, 25, 30 and 35 s. Nothing repaints and the app's screen never appears.
That cannot be reconciled with rounds 39 and 40, where the panel probe recorded 114028 pixel-data transfers after the keys. The probe per-byte opens, writes and closes its file on the vCPU thread, so it is slow enough to change the timing of what it watches -- a trap this file already records twice. The transfer-count conclusion is withdrawn until the count is taken without per-byte file I/O.
What is not in doubt comes from the controller's own memory and needs no probe: after the handover the screen is the launcher's box (43, this file's recorded signature for it) and stays that way, so the app does not paint -- which agrees with rounds 32-35. Next: make the probe accumulate in memory and write once at exit, then re-run the launch.
Read the controller's display RAM over QMP (qom-get /machine/panel gram) with the keypad's press property on one connection: before the keys 485 non-zero, the radio's main screen with F4HWN legible; after F, 7, DOWN, MENU only 43 non-zero, rendered as one boxed title about 42 columns wide and seven rows tall with text inside and fifty-seven empty rows below -- exactly the picture the user described.
With rounds 39 and 40 both things hold: the app runs and blits (105 consecutive full-screen writes from the toggling app; Minesweeper's glyph bytes 7f, 41, 08, 40 in the panel stream), and then dies, after which the launcher repaints its own title box. What remains on the glass is the launcher's frame, so the open question is the app's lifetime, not its drawing.
Method: the QMP socket takes a single client, so pressing keys on one connection and reading the panel on another times out -- do both on one; and the screens were captured with a minimal PNG writer and re-rendered as ASCII, which is what makes them readable given this session has no image input.
Hand-run on a copy of the working image with the panel probe on and F, 7, DOWN, MENU over QMP, now with the real Minesweeper: 103370 pixel bytes after the keys, dominated by 00 (100418) with a longest run of 97640 (about 95 full-screen clears), and underneath the app's own drawing bytes 7f 287, 41 234, 08 192, 40 174 -- its text and box characters. Since the app calls display_clear() once per frame, a clear plus a few hundred glyph bytes is the expected shape, so the app paints and the earlier blank-screen reading was not about the path.
The FillFF variant failed to build through my own mistake: the build script is a textual rewrite, so replacing bigapp with fillff also rewrote the compiler's -o argument and source names; substitute only the names intended, or write that variant's script by hand.
Next: the model exposes the panel's display RAM as the QOM property gram (st7565_get_gram), so a hand-run can read the actual picture over QMP instead of inferring it from transfers.
Raise the panel probe cap (40000 transfers is reached about forty seconds after boot) and rebuild qemu-system-arm, then run the toggling app on a hand-copy: after F, 7, DOWN, MENU the probe grew from 22568 to 139404 lines, 114028 of them pixel data, with the longest run of byte 00 at 108273 transfers -- about 105 full-screen writes in a row, histogram 111061 zeros against a longest 0xFF run of 68.
A hundred consecutive full-screen fills is a loop, not a launcher repainting once: the app runs and blits, which retires the conclusion carried since round 32 that an app's calls do nothing. What hid it was the cap -- the instrument went blind forty seconds after boot and the blindness was read as an app that never acted; the third time this file records capping a diagnostic before knowing the shape of the data, and the first where the cap was inside the model.
Next: the 0xFF phase never appears. An app that fills 0xFF only, with no zeros at all, separates 'it never gets there' from 'the firmware repaints over it' in one run.
With a hand-started emulator on a copy of the working image, the page stayed untouched and healthy. Two of my own call signatures were wrong: uvk5_apps.install takes (image: bytearray, slot, blob: bytes, force) and edits the image in memory, so a path string is read as a blob, and key.py's Qmp wants a bare host:port -- the tcp: prefix fails getaddrinfo, the very trap the socket section records.
The useful find: the panel probe stops after 40000 transfers (panel_probe_n < 40000 in st7565_xfer) and a hand-run reaches that within about forty seconds, so a launch measured after that looks like an app sending nothing -- the same mistake this file records twice about capping a diagnostic. Raise the cap (or make it a knob), rebuild qemu-system-arm with the page's job stopped first, and re-run the launch on a hand-copy.
A hand-run with the variable definitely present wrote 29299 probe lines in forty-five seconds, 28416 of them pixel data, so the panel path is driven as intended and the instrument is sound. uvk5_supervisor.py builds QEMU's env as dict(os.environ) yet the page's emulator never sees the variable, so it is lost between the shell that starts the page and the server process it becomes; pass it like the other launcher options instead of relying on inheritance.
Corrections: the probe prints s->selected, not a chip-select level, so cs=1 means the panel IS selected and the bytes ARE stored; and the old note that a hand-started emulator never draws is not true of a hand-run against the page's working image -- this one booted, drew and streamed pixels, so that note should be re-derived rather than believed.
Replacing the page to get UVK5_PANEL_PROBE into the emulator's environment cost the user their running page, because the port owner of 8080 was the managed launcher job. It was restored as a managed background job and verified healthy (485).
The instrument is present: the running qemu-system-arm.exe contains UVK5_PANEL_PROBE and 'PANEL a0=' and was built after the source was last modified, yet no panel.log appears anywhere, so the variable is not in that process's environment. run-webui.ps1 passes only --qemu/--elf/--flash and webui.py has no env= dict in the paths read, so the spawn site -- most likely tools/uvk5_supervisor.py -- still needs reading.
Recorded two process-inspection traps: Get-CimInstance and Get-Process return nothing in this sandbox while netstat and taskkill work, and a second server instance dies silently while the first holds the port.
The panel is on SPI1 and the flash on SPI2 (the model says so and wires st7565_xfer to soc.spi[0]), so a blit cannot touch the external flash and round 34's sector-cache-overwrite mechanism is withdrawn -- which also matches the flash probe seeing nothing after the app's code load.
The panel model reads correctly (stores only when cs low and a0 high, keeps its own gram, 132-column counter with the col-4 store including the four-column fix), so nothing there explains a dropped blit. What is missing is the measurement: UVK5_PANEL_PROBE logs every byte sent to the panel with a0/cs/page/col, which separates 'the app's blit never reaches the driver' from 'the driver's bytes never land'. It needs the page restarted with that variable; two attempts failed because port 8080 is occupied and the port-owner lookup came back empty. The page stayed healthy (485).
At key.py's timing the launch is coherent and reproducible (485 -> 484 F -> 580 7 -> 617 DOWN -> 43 MENU, the handover). A toggling app -- fill FF blit, fill 00 blit, forever, a signature no menu can produce -- gave fifty readings over twenty-five seconds all exactly 43: its blits never reach the glass.
Offsets checked properly: abi_major (u8) + api_level (u8) + api_size (u16) is four bytes, not twelve, so fb is at 4, display_clear at 8 and print_tiny at 28, exactly where the app calls them; the earlier draw_rect/api_size reading was my own arithmetic error. With the header identical, the entry at 0, the offsets right, the firmware answering status 0 and the handover confirmed, every part of the app side now has evidence.
Left is the path between the app's call and the glass, and the firmware's app_api.h states the mechanism: the app runs from the RAM buffer that is also the PY25Q16 sector cache, and blit_full drives the panel over the same SPI bus, so a blit can reload the sector cache and overwrite the running app. Next: instrument the model's sector-cache path around a blit from an app.
Replicating round 17 exactly (332-byte fill-and-blit in slot 0, page-default taps) gave 485 -> 580 after F,7 -> 627 after MENU, and 627 held: the screen never left the app menu, so this run says nothing about the app. It did expose that 537/549/552/580/596/617/627 are one family -- the app menu, whose ink moves with the highlighted row -- while a genuinely all-lit panel would be ~1024 and has never been seen. At least one earlier 'it painted' reading was therefore a menu row.
Next: drive the launch with key.py's timing and confirm the handover by the 43-byte launcher box, then judge the app only by a signature no menu can produce.
The working copy's app_api.h and the firmware's own (fetched from armel/uv-k1-k5v3-firmware-custom) are the same 14403 bytes with the same sha256, so the layout theory of round 31 is withdrawn; the assembly confirms the calls land on print_tiny at offset 28 and display_clear at 8 as expected.
The measurement those conclusions rested on is now in doubt: 1446 font-region reads early in the session look like the font being cached in RAM, in which case print_tiny never touches the external flash and 'no font read' says nothing about whether draw() ran. The next oracle must be panel-only, and the sharpest unexplained contrast is that the same fill-and-blit lights the panel from a 320-byte app and not from the first statement of this one.
Header exonerated against four upstream apps (entry_off 0, name at 20, version at 36, link_vma 0x20000280 identical), entry point exonerated by the listing and linker script, acceptance confirmed by the firmware answering status 0 over 0x0730. What remains: the app calls the api members at offsets 8 and 28 and neither call changes the panel nor produces a flash read, while the app itself loops in its own first 180 bytes.
The Chinese append in round 30 used a fragment that does not exist in the file, so only the English note landed. Appended after the round-29 Chinese note and verified both files contain the new note.
The firmware answers status 0 over 0x0730 for all eight installed slots, the game and the 256-byte toy alike, so round 29's claim that the launcher refuses some apps is withdrawn. Measured with the instruments: at the key.py timing MENU drops the panel to 43 within a second, the 2 ms PC probe sees 51 of 11439 samples inside the overlay spread over about ten addresses, and the flash probe shows the app's code load as the last transaction of the session with zero after it -- no font reads, so print_tiny never ran and draw() is never reached.
place_mines() is the one loop on that path that can spin without touching the firmware, and it is now capped at 4000 attempts with a deterministic fallback, so a hang there can no longer look like an app that never started.
The page's key endpoint defaults to TAP_MS = 60 with no gap when hold_ms is omitted (what every curl did), while tools/key.py uses 200 ms plus GAP_MS = 400 for the release to debounce. At key.py's timing the same sequence is coherent: 485 -> 484 (F) -> 580 (7) -> 617 (DOWN) -> 43 (MENU, the handover); at the page default it is not.
With the handover reachable the real Minesweeper reaches 43 while a 256-byte fill-and-blit toy never leaves the menu (582): the launcher distinguishes them before either runs, so the question is which apps the launcher accepts -- and the code-size framing of rounds 27-28 is retired.
The English note was appended by the round-28 script but the Chinese one was not, because the anchor fragment did not match and the script asserted. Appended it after the round-27 Chinese note and re-checked both files.
The ladder app calls a chain of no-op functions then fills every framebuffer byte with 0xFF and blits, so the panel either goes all lit or stays as it was. The chain costs about 5.2 bytes per function (20 -> 256 B, 70 -> 620 B).
Neither rung could be measured because the launch never reached the handover: with the 256-byte app in all four slots the readings were 485 (main), 460 (after F, 7), 537 (after DOWN, the menu is open) and then 582 for twelve consecutive readings across three MENU presses -- the first press moved something, the rest did nothing. That is the inert-launcher state already documented, so the entry point rather than the app is the subject for the next round.
The game's app_main was given a fill-every-byte-and-blit first act. Inserted before A = api; it dereferenced an unset pointer and faulted (my own bug, that run is void); moved after the assignment (source +199 B, app 2508 B) the panel still never leaves 43 over 48 s, so the blit never happens, while a 300-byte app doing the same fill-and-blit paints 596 non-zero bytes.
The variable that fits every measurement is the app's code size rather than its blob size: a 3028-byte blob with a few dozen bytes of code runs and can call print_tiny, display_clear and get_key; Phases at 372 bytes of code runs; Ladder at 544 dies; Minesweeper at 2400+ dies with its first instruction having no effect. Next: sweep a trivial app at 400, 512, 600, 800 and 1200 bytes of code through the handover to find where the panel stops changing.
After the pristine restore only slot 3 held a real app; slots 0-2 read as factory data (the page calls them unknown), and every failed launch had the cursor left on one of those. Installing the game into slots 0-2 as well, so every row holds it, makes the open sequence work: the menu comes up at 617, MENU drops the screen to 43 (the handover) and at t+40s it reaches 484, the radio's own main screen -- so the app is entered and then returns.
The radio is left with the game in slots 0 through 3 so the launch is reliable from the page whatever row the cursor lands on. Next: launching the game and getting out of it, and why the app returns rather than staying up.
A 20-byte app that calls nothing was installed in slot 3 (the slot round 24 saw loaded and entered) and driven with the same F, 7, DOWN, DOWN, MENU sequence: the menu came up at 511 non-zero bytes, MENU changed nothing, the screen stayed there for twenty seconds and the PC probe recorded zero overlay samples out of 7316. So the experiment answered nothing, and the honest reading is that the launch is not reliably reproducible through the page's synthetic key events -- the same sequence produced a handover in round 24 (ink 210). That is a statement about the entry point, not the app: rounds 24 and 17 measured the load, the entry and a painted frame from this same artifact.
Minesweeper was restored into slot 3 so the radio is left with the game installed rather than the test stub.
Reading the panel after each single press from a fresh power-on gives the sequence the launcher wants: F, 7, DOWN opens the app menu (489 -> 552), DOWN again moves the selection (552 -> 526, ink 1799 -> 1677), MENU hands the screen over (ink 210). The earlier note that F then 7 opens the menu is wrong, and that error is why several rounds of MENU presses looked inert.
At the handover the flash probe shows the load as the session's last two transactions: addr=108000 len=64 first=46415031 (the FAP1 header of slot 3) then addr=109000 len=2444 first=f0b599b0 (the app's code, exactly code_size). The 2 ms PC probe catches four samples inside the overlay in three seconds, so the loader copies and the CPU enters the app, which then disappears and leaves the launcher's box on the glass.
After a day of app installs the working copy was no longer the image that worked in round 17, so the pristine dump was restored (with the emulator powered off first, since exit writes the in-memory image back over the file). The page now draws its own main screen (485 non-zero bytes) and Minesweeper is installed in slot 3 at 2444 bytes, the same artifact that painted a complete frame in round 17.
The launch does not reproduce: F, 7, DOWN, MENU and a second MENU leave the radio on a fully drawn app menu (ink 1799) that stops responding to keys, identical readings before and after every press. Round 17 measured the game's frame with this same app, slot and a F, 7, DOWN x3, MENU sequence, so the app is not in question; what differs is the launcher's key handling. An A/B that removed this round's instrumentation and rebuilt the round-17 source did not bring the frame back, which is what moved the suspicion off the app.
The panel reading that looked like a paint was the app menu itself: rows 0-6 the title box, rows 18-22 the slot list, 26 lit pixels on row 2 where the probe's marker bar would be 120. So MENU on the initially opened menu selects a row rather than launching; round 17's Minesweeper run worked because it pressed DOWN three times first.
That accounts for every unexplained probe failure in the last three rounds and rehabilitates the probes: with a press first, an app that clears the framebuffer, draws through a helper and blits does run (499 on the panel, 10.9% of PC samples inside the overlay). The launch sequence is F, 7, DOWN to select, MENU to run.
Round 15 measured get_key() repeated in an app loop at 27.1%, the same as the control, so calling it every pass is not what kills an overlay app; the previous section's use of the two key probes as evidence for that is wrong. Four probe builds (with and without an opening spin, with a short and with a 2M-iteration per-pass spin) all left the panel byte-identical at the launcher's own frame, so they did not run, and why is not established.
What that is not: not the shape other apps survived in, not get_key, not the framebuffer writes -- an app in the same slot that fills every framebuffer byte through api->fb and blits does paint (43 -> 596 non-zero bytes). The probes differ by zeroing the framebuffer and drawing with their own helper before blitting, which is the sort of difference that has to be isolated one change at a time rather than reasoned about.
Two font-free probes meant to draw the raw key code left the panel byte-identical across seven presses (26 lit pixels on row 2 and a constant row-20 pattern = the launcher's own frame), so neither painted. The contrast matters: an app of the same shape that fills the framebuffer through api->fb and blits does paint (43 -> 596 non-zero bytes), and its only substantive difference from these probes is that they call api->get_key() every pass. With round 17's result -- the full Minesweeper painted a frame and vanished the moment a key was pressed, returning the panel to the launcher's title box, which only happens on the key-as-EXIT path -- input is now one call wide.
The sector-cache question returns with it, since the key path is a plausible place for the firmware to reach the external flash, but the font reads print_tiny provokes do not kill the app, so this is specific to the key path rather than to flash reads in general.
A clear-and-blit app left the panel byte-identical, but filling the whole framebuffer through api->fb and blitting turned it from 43 non-zero bytes to 596: an overlay app's framebuffer writes and blit_full do reach the panel. What blocked the frame was Minesweeper's own delay_ms(40) on the invalid-key path -- 40 ms of guest time is seconds of wall time here, on top of roughly twenty seconds of font reads per frame, so frames were minutes apart. With it removed the panel went from 43 to 390 non-zero bytes within 24 s, ink 1275, with the title, counters and field legible.
Input is the remaining gap: MENU changed nothing and DOWN sent the app out of its loop (the panel returned to the launcher's 43-byte title box, which only happens on the path that treats a key as EXIT). Next: have the app draw the raw key code it receives and read it off the panel instead of guessing at the mapping.
With the validated shape (for(;;) { long spin; one piece; }) the header block scores 26.1% and the 81-cell loop plus cursor, with locals only, also 26.1% -- both equal to the control, so both are alive. The full Minesweeper is running too: 116 samples inside it over fifteen seconds across eight distinct addresses, with the firmware still serving it (1024 font-region reads), yet the panel never leaves the launcher's title box over ninety seconds, and that box is painted before the app is entered. Removing the opening settle spin changed nothing, which retires the idea that it was still inside that spin.
Two candidates remain: the statics, which draw() reads and neither surviving variant touched, and whether blit_full from an overlay app reaches the panel model at all -- the gutted variant was only ever measured by its PC share, never by its screen.
Each variant is for(;;) { long spin; one call; } so a live app is in the overlay most of the time: control 27.2%, display_clear 26.9%, get_key 27.1%, delay_ms(40) 26.8%, print_tiny 27.0% -- all four calls survive repetition, and the control shows the metric separates the two cases.
That also means the previous round's variant B (four calls in a loop, no spin) was most likely not dead at all: it spends nearly all its time inside firmware functions, which look exactly like a dead app from the trace. The bisect has since closed on one function: Minesweeper with draw() gutted to display_clear() + blit_full() runs (97 samples inside the app, alternating with firmware PCs), while the full draw() draws nothing in twenty seconds and ink never leaves 210, so display_clear() never ran. The end of the app is inside draw()'s body.
Same file, same call order, one variable, measured on the page's own emulator with the 2 ms PC probe. Variant A calls display_clear, print_tiny, get_key and delay_ms(40) once and then spins: it runs (2059 samples inside the app, 1024 font reads). Variant B wraps exactly those four in for(;;): the app is never seen again (0 samples inside it, and the same 1024 font reads, so the first pass really did execute). So the calls are fine individually and fine in sequence once; repeating them ends the app.
Also records that the metric needs fixing first: samples inside the app under-count a live app, because time spent in long firmware functions lands on the firmware side of the trace, which is what a dead app looks like too. The display_clear-only loop scored 6 and the get_key-only and delay_ms-only loops scored 0, and those are one good measurement and two unreadable ones, not three results. The fix is to alternate a long self-contained spin with each call.
The ladder conflated two variables -- every small test app spun first, every big one called the firmware immediately. Separated, all measured with the 2 ms PC probe on the page's own emulator: a 3028-byte app that calls nothing loops in the overlay indefinitely (1704 samples, all inside its first 512 bytes), so blob size is not the problem; adding one print_tiny changes nothing (2039 samples, with 1024 font-region reads in the flash log, the first at 0x1e0000), so the font path and the sector-cache worry are both fine; adding get_key as well still runs (2029 samples); adding display_clear too still runs (1978 of 13195 samples inside the app).
So the three calls no surviving app had ever made are harmless, and the API surface Minesweeper uses is covered. Its failure is its own bug; the next step is an ordinary bisect of its draw() path through the page.
Measured with the 2 ms PC probe: Spin (16 B) loops in the overlay forever; Delay (248 B) is still running 13 s later, 54 of the last 60 samples at 0x20000348; Phases (372 B) survives bar, blit and led; Ladder (544 B) dies before drawing; Minesweeper (2408 B, and 2428 B with an opening spin) is dead within about 4 ms either way. An opening spin does not save the big apps, so the difference is how much of the app has been copied rather than when the first call is made -- and that also explains why Tetris's full 3300 bytes were present in the overlay when checked afterwards: the copy finishes after the app has been entered.
Same family as the four flash bugs already recorded here: the loader is told the transfer is over while the guest is still feeding it. Next: watch the SPI/DMA busy and complete flags during a large read instead of reasoning about them.
uvk5_pc_probe_interval_ms() reads UVK5_PC_PROBE_MS and falls back to 100, so existing probe runs are unchanged. At 2 ms a launch gives 468 samples in about 0.94 s and exactly one of them is inside the 4 KiB overlay: the app lives for roughly 2 ms, not 100.
The trace settles two things and adds one. The loader copies and jumps (the single overlay sample; a 16-byte app that calls nothing stays there forever). The sector-cache overwrite is dead: the flash log's last transaction for the run is the app's own code load (addr=103000 len=2408) and nothing follows it, no font read included. New: after the app dies the firmware loops inside its own flash (PCs near 0x08005118/0x08005122/0x08005248) with zero flash reads and no key response at all -- F then 7 no longer opens the menu and an explicit MENU down/up changes nothing -- so the launcher or its fault path is wedged.
The flash probe records every transaction in order and the last one for that launch is the app's own code load (addr=103000 len=544, first bytes f0b583b0); nothing follows before the app is gone, so no font or resource read lands on the running app in that window. The 7848 font-region reads in the log belong to the firmware's own repainting.
What stands measured: the loader copies and jumps (a PC sample lands in the overlay, and a 16-byte app that calls nothing loops there forever); the app dies well under 100 ms (panel reads at t+113 ms still show the menu, at t+193 ms only the launcher's title box); an app that spins first survives a bar, a blit and led and dies around delay_ms; one that calls blit immediately is gone before it can be seen. Next: drop the PC probe interval from 100 ms to a couple of milliseconds, which is a one-line model change and turns it into a real trace of the app's short life.
Measured on the page's emulator with UVK5_PC_PROBE and UVK5_FLASH_PROBE inherited through its own launcher, so the instance measured is the one that draws. Spin (16 bytes, calls nothing) sits at 0x20000286 in 42 of 383 PC samples after MENU: it loops in the overlay forever, so the loader works. Minesweeper (2408 bytes) is loaded -- the only big flash read is 0x109000 len 2408, first bytes its own -- and a single PC sample catches it inside the overlay before it is gone: it lives well under 100 ms. Phases (372 bytes) draws a progress bar straight into the framebuffer between calls, and the bars for framebuffer-only and led are on screen while the bar after delay_ms never appears.
The overlay is the PY25Q16 sector cache (app_overlay.h says so), so a service reading the external flash -- the font table is at 0x1E0000 -- lands on top of the running app. Real hardware cannot behave that way, since upstream's own apps call print_tiny, so the firmware must gate the sector cache while an app is loaded; that gate is what the model is missing. Next: find it in the flash driver and check what the model answers.