Rounds 95 and 96 left one mechanism standing: KEYBOARD_Poll calls SYSTICK_DelayUs(10) for every column on every poll, so a delay that does not return on the overlay path would hang the app inside its first poll (MsKeys) or inside get_key() after a completed draw() (the game). Reading the machine settles it, and the answer is no.
qemu/py32f071.c already documents and implements exactly this concern -- SYSTICK_DelayUs busy-reads SysTick->VAL, under emulation the counter barely moves between reads, so the machine advances it on each read with poll-boost 24000 -- and a second note says the loop depends on how many ticks a read spans, not on real time. The counter therefore advances whoever is reading it, and SYSTICK_DelayUs converges in the main loop and inside an overlay app alike.
So the best available explanation is gone, and the key path is verified correct at every layer that can be read: the loader's app_get_key, KEYBOARD_GetKey and KEYBOARD_Poll, the keypad model's column and row logic, and the delay the poll depends on. The key still does not arrive.
What is left is a model-side question: whether the columns and rows are actually driven and read while an overlay app owns the foreground. Next instrument, per this file's own rule: put the probe in the model and count keypad_update_rows() calls and what they compute, comparing an overlay app's polls with the resident loop's.
The check round 95 asked for, done on the model side. The keypad device models the matrix properly: keypad_col_changed records which column the firmware pulls low and recomputes the four rows; keypad_update_rows sets a row low when a held key sits on a low column; side keys read low when every real column is high; line 0 is ignored, matching a driver that never uses it; and the column mapping agrees with the firmware's PIN_COL(c - 1).
It also carries a warning that describes this exact symptom: without volatile on row_out, GCC at -O2 proves every element is still NULL, notices qemu_set_irq returns immediately on a NULL irq, and deletes keypad_update_rows and its callers, so no row is ever driven and the scan reads nothing -- silent, and looks like a broken keypad. The guard is present, verified from the object code.
So the model is not where the key is lost, which rules out the one alternative that would have been a model bug rather than a firmware-path behaviour.
That leaves round 95's hypothesis, which now accounts for every measurement in twenty rounds: KEYBOARD_Poll calls SYSTICK_DelayUs(10) for every column on every poll, so a delay that does not return on the overlay path sticks MsKeys inside its first poll (it painted nothing) and sticks the game inside get_key() after a completed draw() (frozen frame, keys doing nothing). Next: read SYSTICK_DelayUs and the machine's SysTick counter model.
The firmware sources are in the repo as work/dl_*.c, so the chain is read rather than inferred: api->get_key points at app_get_key (fw_app_overlay.c:85, table at 1025), which calls K5VIEWER_ParseInput under ENABLE_FEAT_F4HWN_K5VIEWER and then KEYBOARD_GetKey -> KEYBOARD_Poll (dl_App_driver_keyboard.c:283, 185).
That rules one thing in and one thing out. In: the overlay path polls the matrix on demand, so the app does not need the resident loop running to see a key. Out: the serial-injection branch only short-circuits when gKeyFromSerial is set, and keys here come from the matrix.
The serial hypothesis was tested because app_get_key calls K5VIEWER_ParseInput first and the hand-started emulators never had a serial port. With one present and drained the marker still never appears, so it is not that either.
What is left is the scan's stability rule: three consecutive identical reads of the rows, each separated by SYSTICK_DelayUs(10), or the column is treated as having no key. If that delay does not actually wait on the overlay path, every column is skipped and no key is ever detected -- which would explain both halves of twenty rounds of measurement. Next: watch the keypad's column and row GPIO traffic while an overlay app polls, from the model side, without touching the firmware.
Four MENU presses held five seconds each left the game's frame bit-identical (total ink 57, field ink 24), which is what a missed key looks like and also what a key path that never runs looks like.
So an app asks directly: 132 bytes that poll api->get_key() in a tight loop and, on the first value other than APP_KEY_INVALID, paint that code as eight pixels plus a marker pixel at (0,0). Seven keys -- MENU, DOWN, UP, STAR, 1, F, EXIT -- each held five seconds, and the marker never appears. get_key() does not hand a keypress to an overlay app. The marker exists so a genuine zero is not confused with nothing arriving, and it stays dark.
The goal's state, plainly: the blank screen is fixed and understood; the Labs firmware displays; the app's source, build and page upload work and it compiles with no diagnostics; it installs, launches with F, 7, DOWN, MENU and renders its header, counter and cursor. What is not demonstrated is playable input, and the reason is a named service rather than the app. The game's MENU branch is place_mines, reveal, check_win, and it has never been reached.
Minesweeper 1.3 launched with F, 7, DOWN, MENU and read row by row: the header print_tiny draws at y=1 (the M, the two-digit counter and MINES) and the cursor frame the app inverts around its cell. Both are the app's own drawing and both are legible.
The reveal did not happen: ink 57 before MENU, 57 after it, 57 after a DOWN and another MENU. Either the key does not reach the running app or the reveal paints into cells that were already lit. That is now the whole of what stands between this and playable, and the panel can answer it -- the field's numbers appear only when a cell is opened.
Noted for the next round rather than smoothed over: the cursor frame reads at rows 42 to 48 while the game's arithmetic puts it at 33, so there is about a nine-row offset between the app's coordinates and the reader's. It does not affect what the game draws, but a reader that is nine rows off would mislead whoever uses it.
Round 91 put the count on the glass instead of in a bar: eleven pixels of n on row 0 and a marker at (0,0). The result was ink 52, marker 0, all bits 0, n = 0 -- draw() leaves the framebuffer empty, and the app's own marker did not appear either. That ruled out the bar's width and pointed at api->fb itself.
Reading app_api.h:71 rather than reasoning: 'The framebuffer is the resident gFrameBuffer[FRAME_LINES][128]'. gFrameBuffer is 1024 bytes, eight pages of 128, so fb's first index is the PAGE and the pixel within it is a bit: A->fb[y >> 3][x] |= 1u << (y & 7). The app had been writing fb[y][x] with y up to 63, eight times past the end, and the round 89 change from bit-packed to one byte per pixel moved it from wrong to differently wrong.
The frame now arrives and the ASCII reader makes it readable: the game's header, drawn by print_tiny at y=1, in a box from columns 42 to 84. Rows 7 and below are empty and that is correct -- every cell of a fresh game is closed and unflagged, so the 81-cell loop has nothing to paint.
Two habits: when a contradictory measurement appears, make the next one unarguable, and when a type is in doubt read the header that defines it -- the answer was two lines above the typedef. Next: press MENU inside the game and watch the field appear.
The counter probe -- new_game, draw, count the framebuffer's non-zero bytes, paint that count as an n/8-wide bar plus a fixed marker in columns 126-127, one entry -- gives ink 834 after MENU, so draw() left roughly a thousand non-zero bytes and a blit right after it reached the glass.
The same draw() inside the real game, with a second blit added because a repeated blit looked like the difference, gives ink 43 and row 2 reading columns 42..84 with 26 lit: the launcher's title box.
The same code, two opposite answers, and this round does not establish which reading is wrong, so both are suspect until a measurement that does not depend on reasoning about the bar settles it. The second run does rule out 'blit twice' as the answer, and the page-major decode mistake in the first script is a reminder that a reader written minutes ago needs the same scepticism as one written rounds ago.
Next: have the app write the count itself into known bit positions of the framebuffer and read it straight out of the panel hex. No bar, no width, no decode.
The marker now goes inside the game's own draw() instead of renaming its entry, so there is exactly one .text.entry and one app_main, counted in the source before each build. Marker at the start of draw() paints; marker at the end, after blit_full(), also paints -- so draw() is entered and completes, and the rounds 85-86 'draw() never returned' reading was the duplicate-entry artifact. Retracted.
Then a real bug, found by reading app_api.h line 73: typedef uint8_t (*app_fb_t)[128], one byte per pixel, while put() and invert() wrote it as bit-packed (fb[y][x >> 3] |= 1u << (7 - (x & 7))). Every field and cursor pixel went to the wrong byte with the wrong value, which is why MsStripes, which writes fb[y][x], painted and the game did not. Both helpers now write one byte per pixel; build clean at 2388 bytes; no x >> 3 remains.
It still paints nothing, which is now well posed: draw() completes, so blit_full ran, so the frame it sent is one the driver skips, and after display_clear nothing wrote to api->fb. Next: have the app count the framebuffer's non-zero bytes right after draw() and blit that count as a bar.
new_game() is twelve trivial statements: zero three 11-byte bit arrays (BYTES = ((81+7)/8) = 11) and set six counters. An app that shares nothing with the game's source and does exactly those statements -- 60 bytes, crc32 0x85397B54 -- paints: boot 485, then 491 after MENU. So the statements are not the fault.
Then the harness. Every round 85-86 variant renamed the game's app_main to game_main and appended a probe, and both carry the entry attribute. The script keeps both with KEEP(*(.text.entry)) and nothing decides which lands at offset 0, where the loader jumps. Those variants may have run the game's real main instead of the probe, so their readings cannot be trusted -- the same class of mistake this file keeps recording.
The real Minesweeper.app has one entry, so its own failure stands. The bisect has to be rebuilt with the probe as the only entry, and the tooling limit is now written down: under ziglang -Wl,-Map is refused, objdump -h prints nothing and nm returns no symbols.
The working stripe app padded with untouched volatile data in steps paints at code sizes 80, 588, 1100, 1612, 2124, 2380, 2636, 2892 and 3148. So the overlay runs three-kilobyte apps happily and the size ladder suspected since round 40 is not the fault.
Then the game's own material added piece by piece to the same base: new_game() alone, new_game() plus a rowcol loop, and three new_game() calls, all under three kilobytes and all failing to paint. What they share is new_game() and nothing else in this round does. That is also where round 68's observation points -- the app wrote its own globals in the first tenth of a second, which is new_game(), and then nothing further.
Next is small because new_game() is short: a handful of counters, the three 81-cell arrays, and the mine placement. The same draw-a-marker-afterwards harness names which statement stops control.
MsDraw1 (2772 B) calls new_game, then the game's own draw, then paints stripes and blits: boot 485, menu 526, MENU -> 43 for 120 seconds and no stripes, so the marker never ran and draw() never returned. NoFont (2392 B), the same game with every print_tiny routed to a no-op, behaves the same and ignores keys, so the font is not why. Round 73 drew the opposite conclusion from that app on a launch row that does not launch it.
The culprit is inside draw() but not the text: display_clear, the framebuffer writes, or blit_full. All three work in the small probes, so what differs is scale -- the app runs from the 4096-byte overlay and MsDraw1 is 2772 bytes of code there.
Next cut is the one round 16 attempted with the wrong instrument: split draw() into its three parts and run each with the metric that now works.
MsClearStripes -- display_clear, a stripe fill standing in for the game's drawing, then blit_full -- is 80 bytes, crc32 0x0D75B311, and gives boot 485, menu 526, then 491 and it stays. 491 is exactly what plain MsStripes produces, so the sequence arrives on the glass unchanged.
That retires round 79's reading too: MsOnce and MsDraw blitted a black frame, the driver skipped its all-zero pages, and the panel kept the launcher's box. Their failure was correct behaviour.
An overlay app's display path now works in every combination measured -- fill then blit, and clear then fill then blit -- and the game alone produces no visible change over three minutes. Next cut separates its two possibilities: a variant that calls the game's own draw() once and then fills stripes and blits. Stripes appearing means draw() returned and its output is what the driver skips; stripes absent means draw() never came back.
Minesweeper in slot 1 launched with F, 7, DOWN, MENU and polled every three seconds for three minutes: boot 485, menu on its row 526, then 43 at every sample, and still 43 after a DOWN and twenty more seconds.
That is a real finding now that the metric is calibrated against three fills the app controls -- 0xFF moves the panel to 939, a stripe pattern to 491, and 0x00 does not move it, correctly, since the driver does not send all-zero pages. A frame that reaches the glass moves the panel; the game's does not.
So the display path is fine and the game's own frame never arrives, which is where round 73 left it but with an instrument that could not then tell a black frame from no frame. Next cut: display_clear, then the app's own stripe fill, then blit_full -- if that paints, the sequence is fine and the failure is in what the game draws.
MsStripes fills the framebuffer with (x & 1) ? 0xFF : 0x00 and blits: 72 bytes, crc32 0xD176B01A, boot 485, 525 after the row is selected, then 491 after MENU and it stays there. 491 is not 525, so the panel took the app's frame and the whole display path works from an overlay app.
That confirms round 80's suspicion and overturns a reading used for many rounds. A blit of 0xFF reaches the panel and a blit of 0x00 does not change it, because the driver does not send pages that are all zero. An app whose frame is black therefore looks exactly like an app that never drew, and 'the panel still shows the launcher's title box' was never evidence that the app failed. MsZero, MsOnce, MsDraw and the game have all been running and drawing.
Two habits, both already in the file: measure the thing you actually care about -- those rounds asked 'did the app run' and the number answered 'is its frame mostly non-black' -- and a device that omits work it considers unnecessary makes a correct result indistinguishable from no result, so a probe must use content it cannot skip. Next: the game itself on the row that launches, which the panel can now actually answer.
MsZero zeros the framebuffer with the app's own loop and blits once: 68 bytes, crc32 0x85057463, boot 485, 525 after the row is selected, then framebuffer 0 nonzero and panel 43. The framebuffer going to zero proves the app ran and wrote where it meant to; the panel not changing is the puzzle.
fillff answered half of it already: a blit of 0xFF takes the panel to 939, a blit of 0x00 does not change it. The likely reason is ordinary driver behaviour -- ST7565_BlitFullScreen need not send pages that are all zero, so a black frame sends nothing and the panel keeps what it had. That makes ink 43 a coincidence, since it is also the count of non-zero bytes in the launcher's title box, and it means the metric has been lying: an app whose frame is black looks exactly like an app that never drew.
Cheap to test and it changes the shape of the work: fill the framebuffer with a pattern rather than a constant and blit. If the panel shows the pattern, the whole display path works from an overlay app and MsZero, MsOnce, MsDraw and the game have all been running.
MsClear (44 B, display_clear once) runs: the framebuffer it owns goes from 525 non-zero bytes to 43 and the app survives, with the panel staying at 525 because nothing blitted, which is correct for a clear-only app. fillff (44 B, direct fill then blit) paints, reaching 1024/1024 of 0xFF and panel 939. MsOnce (48 B, both calls once each) and MsDraw (52 B, both in a loop) both fall to panel 43.
So each call is fine on its own and calling them one after the other is what fails, once or in a loop. The specific difference: after a direct fill blit_full repaints the panel, and after display_clear it does not. That is the plain behavioural difference everything earlier was standing on.
One cell missing, separating the call from the act: replace display_clear with the app's own memset of api->fb and then blit. If that paints, the call and its side effects are the problem; if not, clearing and blitting in that order is, and the game can be written to avoid it.
The smallest app using the game's display path -- display_clear, blit_full, spin, nothing else -- is 52 bytes (crc32 0x9D6B4FA3) and behaves exactly like the game: boot 485, 525 once the row is selected, then 43 after MENU and nothing further.
Beside round 72 that is narrow. fillff, 44 bytes, writes api->fb directly and calls blit_full once, and works: the framebuffer ends 1024 bytes of 0xFF and the panel reaches 939. MsDraw does the same blit and differs in one visible way -- it calls display_clear every pass, which fillff never does.
So the failure is one API call wide at this point. Two 52-byte apps settle it: one calling only display_clear, one calling only blit_full. Whichever reproduces 43 is the call that ends an app, and it is also the shape of the fix, since a game that draws without it is a game that runs.
Widening the probe to a kilobyte ruled out access width. The cause was alignment: a region's alignment is its size, so asking for eight bytes at 0x20000C0C placed it at 0x20000C08, watching the wrong four bytes silently; aimed at 0x20001342 it landed on 0x20001340 or 0x20001000 and covered what the display path needed, which is why the radio stopped drawing.
Aligning the request down and warning fixes it, and the round-74 acceptance test now passes: probe off gives 485, 526, 43 and the probe at 0x20001342 also boots to 485, logging 10378 stores and saying it watched 0x20001000 instead.
That log answers round 74's question: every writer of that kilobyte during a whole run is in firmware flash -- 0x08004874 and 0x08004d34 with 1041 each, the memset at 0x0801ae7a with 1311, 0x0801af42 with 543, 0x08003cfe with 764 -- and not one is inside the overlay, so the app never writes the framebuffer at all. Before its next use the probe needs a cap or a program-counter filter: ten thousand fprintf lines in half a minute ended the run early.
Giving the probe its own backing store did not stop it blanking the radio, so re-entrancy was not the cause. Three runs of one image settle it: probe off gives 485, 526, 43; probe at 0x20000C0C gives the same three numbers with 8 stores logged, all from 0x0801ae7a; probe at 0x20001342 gives 0 throughout with 1 store.
So rounds 67 and 68, which ran with the probe at the overlay address, are valid and their conclusions stand. At the framebuffer the display breaks, and the difference is what lives there: 0x20000C0C is not the framebuffer, and the firmware's display path must reach gFrameBuffer in a way an eight-byte io region cannot serve, most likely a wider access.
The instrument's limitation and its fix are now both precise: make the probe region large enough to cover what it watches, with a backing store the same size, and let the handlers split or assemble whatever width the guest uses. That would allow the framebuffer measurement round 74 wanted. Also noted: at 0x20000C0C the only writer across a launch is the memset eight times, and round 68's stores at 0x20000280 came from the game, which is not the app in this image, so the two runs agree.
Two runs of the same image, same keys, one variable apart: with UVK5_RAM_PROBE unset the panel reads 485 at boot, 526 after F, 7, DOWN and 43 after MENU; with UVK5_RAM_PROBE=0x20001342 it is 0 from boot onwards. The probe built in round 67 therefore changes the guest's behaviour, and only one store was ever logged through it, so it is breaking the firmware's access to those bytes rather than merely observing.
Rounds 67 and 68 both ran with the probe active and must be re-checked. Round 68's reading is probably still sound -- the probe was then at 0x20000C0C rather than the framebuffer, and the PC it logged, 0x20000280, is inside the overlay with a value matching the bytes seen -- but that is the strongest claim available until it is re-run with a control.
The same A/B gave a free result: with the probe off, the no-font variant reproduces the failure cleanly (485, 526 once the row is selected, 43 after MENU), so rounds 72 and 73's conclusions stand and only the instrument used to look at them is suspect. Likely fault: an io region overlapped over RAM wins for those bytes, and an access wider or differently aligned than its valid range allows becomes a guest error instead of being forwarded. Its acceptance test is the same A/B.
The one-change experiment round 72 asked for: the game with every print_tiny routed to a no-op, board still drawn into the framebuffer and still blitted. It compiles clean, code 2328 B against the game's 2568, and gives the identical outcome -- ink 526 after F, 7, DOWN with the row selected, then ink 43 after MENU and nothing further.
So the font is not what ends a large app. Beside round 72 the rule is narrow: a 44-byte app that fills the framebuffer and blits works, a 2328-byte app doing the same does not. The difference is size and it only shows for apps that draw, which is where rounds 30-38 were: the app executes from the same RAM the PY25Q16 driver uses as its sector cache.
Also recorded: the build recipe's -mcpu takes its value joined; passing it as a separate argument makes zig reject the command, and round 70 never surfaced that because the .app already existed. Next: aim the round-67 model-side write probe at gFrameBuffer (0x20001342) and see which PC stores there, and whether it is inside the overlay.
Two of my own mistakes had to be separated. The menu does not start on the row that launches an app: Mark in slot 1 with F, 7, DOWN, MENU runs and the panel goes 485 to 525, while the same app with F, 7, MENU never appears. So one DOWN selects the row, and my earlier second DOWN aimed at slot 2 landed on the empty row below it, which is why fillff appeared to do nothing in rounds 70-71. That test was invalid and its conclusions must be re-derived.
With the sequence right: Mark (24 B) writes its marker and the panel goes to 525; fillff (44 B) leaves all 1024 framebuffer bytes 0xFF and the panel goes 485 -> 525 -> 939; Minesweeper (2568 B) goes 526 -> 43 and no further. So the loader, entry point, api pointer, api->fb and blit_full all work, and an app that fills the framebuffer and blits really does repaint the glass. Ink 939 rather than 8192 is the panel storing a bit per pixel.
The shape is sharper than the old size ladder: a small app that draws works, a large app that does not draw works (the 3028-byte pad and spin in round 47), and a large app that draws through the firmware -- Minesweeper with print_tiny and display_clear -- does not. Next: strip the print_tiny calls out of draw() and render with framebuffer writes only.
fillff -- fill the framebuffer with 0xFF, one blit_full, then spin, with no text and no service call -- installed through the page into slot 2 reports code_size 44; the panel reads 485 after power on, 587 after the menu opens on its row, and 43 from 1.5 s onwards. Ink 43 is the launcher's title box, not a framebuffer full of 0xFF, so the app does not get as far as its blit.
That retires the font theory: forty-four bytes with no text fail exactly as the 2600-byte game does, and the wall sits where rounds 16 to 21 kept arriving -- between the launcher handing over the screen and the app painting the glass.
It also means round 17's positive result (the game painting a frame once its delay was removed) was measured on a different image and build, so by this file's own rule it is unverified and does not reproduce here. Next: watch gFrameBuffer while the 44-byte app runs, which separates the app not reaching its fill from its blit not reaching the controller.
Everything goes through the browser UI's own HTTP interface: POST /api/apps/1 with the app bytes reports code_size 2568 and crc32 0x60234C72, the panel reads 485 before the menu, 580 after F+7, 617 after DOWN, and 43 after MENU, and stays at 43 for ninety seconds and a further MENU tap. That is the user-visible complaint on the interface the user uses, with source: panel, so the page draws the controller's own memory.
The app is not the old problem any more: round 68 showed the six bytes read as corruption are the app writing its own variables, and the source has no delay_ms left, so this is the build round 17 measured painting a frame. Its loop is spin, draw, get_key and it draws every pass.
Ink 43 is the launcher's title box, drawn before the app is entered, and one display_clear() from the app would change it. So the app reaches its own initialisation -- round 68 saw it store its globals in the first tenth of a second -- and does not get through its first draw(). Next: time one draw() from inside the app, or watch the same page for ten minutes to separate much-slower-than-the-window from stuck.
Printing every store the probe logged, not the first sixteen, gives two non-zero ones: pc=0x20000280 writes 0x0801d640 and pc=0x2000029e writes 0x28. Both PCs are inside the overlay, so the app itself is writing them, and the values are exactly the six bytes read as corruption for twenty rounds.
So the CRC check passes, control reaches the app, and the app runs. Every measurement in the last twenty rounds that read the overlay back and hashed it was reading a live app's memory; the loader's copy is gone by then because the app is using that space for its own variables. The corruption that explained the size ladder, the CRC failures and the retries does not exist.
The open question is therefore the earlier one -- what the app does after it starts -- and the round-17 finding stands as the last measured answer: it paints a complete frame and a key press ends it. The next round should start from a running app and watch its behaviour, with no further attention to the overlay's contents.
The probe is an io region overlapped over eight bytes of SRAM (UVK5_RAM_PROBE=0xADDR) whose write handler logs the PC and value and forwards to RAM, so no gdb and no halting. Watching 0x20000C0C: the overlay holds the app's code exactly at +0.05 s, one byte differs at +0.10 s, and the launch happens.
Two probe bugs came first. An io region over RAM intercepts reads too, and with no read handler it answered zero for those eight bytes, so the firmware read zeros where it expected its own data and the guest died with a QMP connection reset every run; forwarding reads fixed it. And a subregion callback's address is relative to that subregion, so forwarding it unchanged wrote to 0x20000000..7 rather than 0x20000C0C..13.
With the probe working the shape is clear: byte-perfect at +0.05 s and one byte wrong 50 ms later. 21 stores were logged into those eight bytes -- eight from 0x0801ae7a, eight from 0x08004874, both writing zero, plus five the summary cut off. Those five are the next thing to read.
Arming the watchpoint before the keys and polling the overlay while draining stops gives matched=None, diverged=None, 60 memset hits and no other hits: the launch did not happen again. The watched word is inside the region the firmware memsets about a hundred times a second, and every hit halts the guest until the debugger resumes it. Sixty hits here, four hundred in round 64, four thousand in round 63.
That explains three consecutive null results and matches the runs that worked: round 62 had no watchpoint and the overlay held the code byte for byte. The correlation is not that the app sometimes fails to launch; attaching this watchpoint stops it.
It is this file's own first rule in another form: instrument the model, not the guest. The fix is a write callback in qemu/py32f071.c over those four bytes -- memory_region_init_io plus memory_region_add_subregion_overlap, logging the PC and forwarding to RAM -- so the guest never stops. The gdb watchpoint cannot be conditional (Z2 has no condition in this stub) and the address cannot leave the memset's range.
Running the watchpoint and the control in one run, and reading memory only on hits whose PC is not the memset's, gives 400 hits all memset and zero others, and the control shows the overlay is all zeros with crc 169b5c51 against the app's 60234c72 -- 262 matching bytes that are just the app's own zeros coinciding.
So round 63's zero result is fully explained: there was nothing to observe. The same key sequence launched the app in round 62, where the overlay briefly held the code byte for byte, and did not here; the sequence is timing-sensitive and every conclusion from a run that did not verify the launch is vacuous, including the claim that the memset is the only writer.
Fix: make the harness deterministic before watching anything -- press the keys, poll the overlay until it matches the app or fail loudly after a deadline, and only then arm the watchpoint. The control exists; it must run before the measurement rather than after it.
Watching the four bytes at 0x20000C0C and reading their value at every hit -- rather than the fill register, which a byte copy would fool -- gives 4000 hits in forty-odd seconds, all of them the memset at PC=0x0801ae7a / LR=0x0801701c, and the word never read back non-zero.
That is not a finding: the same run never checked whether the overlay ended up corrupted. If the launch did not happen on this run, there was nothing to observe, and a probe that cannot show the thing it is looking for returns zero -- the lesson this file already carries.
The fix is one line in the same run: after the watchpoint pass, read the overlay and report how many bytes match the app, so the corruption is proven present before any claim about its writer. Two smaller facts kept: the memset writes one byte per hit, so the watched word is transiently part-written; and reading memory per hit costs a gdb round trip, which is why 4000 hits took most of the window.
Sampling the overlay every 50 ms from MENU: before it, crc 169b5c51 matching 262/2568; at +0.00 s, crc 60234c72 matching 2568/2568, which is the app's own CRC; at +0.06 s, crc 4f23f6f3 matching 2562/2568; and every later sample identical. So the load is byte-perfect and the loader's own check would pass at that instant, and exactly six bytes -- the pointer 0x0801D640 plus 28 0a -- change sixty milliseconds later and never change again.
That retires round 59's reading: a memset zeroing the whole 4 KiB would change thousands of bytes, not six, so the hundred-a-second memset is not what stops the app. The damage is a six-byte structure write.
Together with the offset tracking code_size (code_size - 124 for both the 2744 and 2568 builds), something on the load path that knows code_size writes a six-byte structure at overlay + code_size - 124 just after the copy and before the app is entered. Next: watch those bytes again but ignore the memset, reporting only the first hit whose fill byte is not zero.
Reading the stack at the watchpoint hits gives the same chain every time: PC=0x0801ae7a (inside memset, entry 0x0801ae70), SP=0x20003c50, with return addresses 0x080047ca and 0x080175fe. So 0x080175fe calls the routine at 0x08016ff0, which calls memset(0x20000280, 0, 0x1000) from 0x0801701a, about a hundred times a second.
Caveat worth stating: the sources quoted in recent rounds -- app_overlay.c, py25q16.c, radio.c -- came from armel/uv-k1-k5v3-firmware-custom's main branch while the running image is f4hwn's build. They need not be the same lineage, so claims like 'this routine is APP_LaunchOverlay' are inferences from those files rather than facts about this image.
What is measured and source-independent: a routine at 0x08016ff0 is entered about a hundred times a second and each time zeroes the whole 4 KiB overlay, while the code is read into it exactly once -- so nothing the loader writes can survive, which is why a 24-byte app passes its CRC check and a 2568-byte one does not.
Decoding around the return address 0x0801701c puts the call inside a function whose prologue is at 0x08016ff0. The literal pool near it holds 0x20001b40..0x20002811 and 0x2000000d/0x0e -- settings and EEPROM-area RAM addresses -- and 0x20000280 appears nowhere, so the pointer handed to memset is computed, which is what PY25Q16_OverlayBuffer() looks like.
My own disassembler misaligned badly here, printing nonsense branch targets like 0x8741e36, which is the tell that the halfwords are not paired the way Thumb-2 requires. The two facts above survive because they rest on the prologue shape and on the literal values, but no instruction-level claim from this round should be trusted.
Round 59 still stands and is what matters: the overlay is zeroed about a hundred times a second while the code is loaded once, so nothing the loader writes can survive. Next measurement needs no disassembler: at those hits SP is 0x20003c50, so reading the words above it gives the return-address chain and names the retry loop.
Logging every DMA run over 512 bytes with destination and count for twelve seconds after MENU gives exactly one run into the overlay, count=2568, which is code_size exactly. Round 58's short-read conclusion is withdrawn: the copy is complete and correct, as rounds 53 and 56 showed from the byte side.
What matters is the comparison with round 58's watchpoint: that word is written 1200 times in the same sort of window, every time by memset(0x20000280, 0, 0x1000), while the DMA writes the overlay once. The code is loaded once and the 4 KiB overlay is zeroed roughly a hundred times a second.
That fits what was puzzling: the 24-byte marker app ran because its CRC window is 24 bytes and the check follows the load immediately; Minesweeper's window is 2568 bytes and is far more likely to be hit first. It also explains the size ladder -- short apps survive the race, long ones do not. Next: name what zeroes the overlay, one function up from the memset's return address 0x0801701c.
A hardware write watchpoint on 0x20000C0C (overlay + 2444, the first byte wrong for the 2568-byte app) yields exactly one distinct writer: PC=0x0801ae7a, LR=0x0801701c, fill byte 0x00, 1200 hits, with r0=0x20000280, r2=0x20001280, r3=0x20000c0c. That is memset(ws, 0, 0x1000) -- APP_LaunchOverlay's own zeroing of the 4 KiB overlay -- called twelve hundred times: the launch path is being retried.
That the only writer to that word is the zeroing is the important part, because the word ends up holding 40 d6 01 08. Something leaves it non-zero without any store the watchpoint saw, and the one way that happens is if the copy never covered it: the memset zeroes 4096 bytes, the read fills code_size, and a read short by about 124 bytes leaves the tail zero and fails the CRC.
That reconciles round 56, whose window probe looked at one transfer -- the one that landed -- and found it correct; there are many transfers and it did not measure whether each carries code_size bytes. Next: log the count of every DMA run whose destination is the overlay and see whether some are 124 bytes short.
Polling the overlay hash every 250 ms from MENU, with no probes: it holds the previous content before MENU (169b5c51), becomes 4f23f6f3 at +0.25 s, and every later sample is the same value. So the load completes, the overlay ends up wrong, and nothing touches it afterwards.
With round 56 showing the transfer that lands in the overlay carries the correct bytes at indices 0, 1, 2444 and 2445, and no other run in that log having an rx address inside the overlay, the writer is not the DMA: it is a CPU store inside the quarter second after MENU, at overlay + code_size - 124. The instrument to name it is a write watchpoint on that word, which the gdbstub supports as Z2 -- the halt is the measurement rather than a perturbation, the one case where the advice against attaching a debugger does not apply.
Also: the round 56 Chinese note is in, appended at the end of the file, because the anchor the append script picks keeps landing on a fenced code block -- the same failure recorded in round 54. Heading counts still match.
A window probe on the DMA loop logs the byte the transfer actually carried at fixed indices. For the run whose rx_addr is the overlay the last lines are idx 0 f0 (correct), idx 1 b5 (correct), idx 2444 00 (correct), idx 2445 00 (correct). Earlier lines belong to other, shorter transfers carrying 0xff or the FAP1 bytes, so the load is a sequence and the one that lands carries the right bytes everywhere, including the window that ends up wrong.
DMA and flash paths are exonerated for the last time. That also opens the possibility that the overlay is correct at transfer time, the CRC passes, the app actually runs, and the six bytes are a post-mortem symptom rather than the cause -- the next measurement reads the overlay back within a fraction of a second of MENU and hashes it.
Own mistake recorded: the window probe was meant to write four lines and wrote 2088, because a condition on the loop index alone also matches every short transfer that starts there. This file warns about capping a diagnostic before knowing the shape of the data; this was the same rule in the other direction.
tools/check_docs.py enforces that the two files have the same number of headings, and the restored section added one. Made it bold text instead, which keeps the content and the order intact. Rounds 19 and 22 are covered by the combined 18/23 notes and round 52 by the 51 note, so the content parity is complete even though three snippet files were never written.
A round-by-round comparison shows AGENTS.zh-CN.md missing rounds 18, 19, 22, 23, 44, 49, 51, 52 and 54 while AGENTS.md has them all. The append helper picks its anchor from the fifth non-blank line of a window after a marker and refuses a non-unique anchor; in the Chinese file that anchor is often a code fence or a repeated phrase, so the write was skipped each time and nothing complained. The bilingual pair is a stated requirement, so this drifted for many rounds unnoticed.
The missing notes are restored from the work/rNN-zh.md snippets they were generated from, appended under a heading that says so, rather than being silently re-run.
The read path returns s->data[(s->addr++) % PY25Q16_SIZE] and the model writes its 2 MB back on exit, so the file a run leaves behind is what the model believed the flash held: flash-r52.img and flash-r53.img both hash to the app's 60234c72 with zero differing bytes at slot 1's code offset. The flash model hands back the right bytes.
Round 53 showed the DMA run is right too (count=2568, rx_addr=0x20000280, rx_inc=1, counts equal), so the transfer wrote the correct bytes to the correct place. Yet the overlay ends with 40 d6 01 08 28 0a at code_size-124 -- a firmware flash address at a position only the loader and the DMA know about, which is what a staged write or an interrupt frame looks like.
Next: have the DMA probe dump the bytes it actually wrote in that window, which separates 'the transfer wrote them' from 'something wrote them afterwards' in one run. Three independent instruments now agree the source bytes and the transfer are correct, so every earlier reading of this as a flash or DMA fault is retired.
A model diagnostic (UVK5_DMA_PROBE) prints every DMA run over 512 bytes. The app load is one run and it is correct in every field: count=2568, rx_addr=0x20000280 (the overlay), rx_inc=1, tx_cndtr=rx_cndtr=2568. So the DMA delivers everything to the right place.
The overlay still ends up with six wrong bytes at offsets 2444..2474, exactly code_size-124 for both sizes measured (2744->2620, 2568->2444). With the addresses and counts right, the bytes themselves must be wrong: they are what the flash model returned through py25q16_xfer. The firmware's flash probe logs the command and length but not the bytes, which is why this read as a memory corruption for so long. Next: log the bytes the flash model hands back near the end of a long read and compare with the image.
At 2744 bytes the six corrupted bytes sit at offsets 2620..2650; after trimming the same app to 2568 they sit at 2444..2474. Both are exactly 124 bytes from the end, so it is not a fixed address being clobbered but the last 124 bytes of the read. That fits what already worked: the 24-byte marker app passed and ran because 24 is shorter than the damaged tail, while Minesweeper at 2568 and 2744 always fails with APP_ERR_CRC. Same signature as the four DMA/flash faults already documented, and it retires the size ladder.
Direct evidence that the model, not the image, is at fault: the flash image at slot 1's code offset hashes to the header's 0x3c12630d with zero differing bytes, while the guest's overlay comes out 0x1312d98c.
Also recorded: trimming the app (removing the three temporary marker probes, shortening the header string, dropping the cursor readout) took it from 2744 to 2568 bytes with no warnings -- worth keeping, but a tail bug cannot be dodged by shrinking. Next: the DMA-driven flash read in py32f071.c, looking for where the final partial burst is counted or addressed.
Installing the 24-byte marker app into slot 1 instead of slot 0 changed everything: after F, 7, DOWN, MENU the fixed address reads efbeadde (the app's first instruction executed), the overlay starts with the app's own bytes 81b0034803490160, and the panel holds 615 non-zero bytes instead of the launcher's 43. entry() is reached, the app runs and it paints -- the display path, loader, header and CRC were never the problem.
Two independent causes. The app menu's selected row is not slot 0, which the overlay contents gave away (it held Minesweeper's code when Mark was in slot 0). And the overlay's tail is clobbered before the CRC: six bytes at offsets 2620..2650 that the app holds as zero and the overlay holds as the firmware pointer 0x0801D640.
The second explains the size ladder that cost several rounds: the clobbered offset is fixed, so an app shorter than it passes and runs, and one that reaches past it fails with APP_ERR_CRC -- Minesweeper at 2744 bytes is only about 124 bytes past the first corrupted byte. Next: shrink the app under that point, or find what writes at SectorCache+2620 (any cached sector read does ReadBufferRaw(SecAddr, SectorCache, SECTOR_SIZE), so a firmware path caching a sector during the load is the likely writer).
The header says code_crc32 0x3c12630d and the app's own code hashes to exactly that, so the app is not at fault. But the 2744 bytes actually sitting in the overlay, read with memsave from the running emulator, hash to 0x1312d98c. APP_LaunchOverlay computes MB_Crc32Bytes over that buffer and compares it with the header, so it takes the APP_ERR_CRC branch at line 1117 and returns before entry(&app_api) -- which is why the app never runs and why the code is nevertheless in the overlay: the copy happened, the verification did not pass.
Six bytes differ, at overlay offsets 2620, 2621, 2622, 2623, 2627 and 2650 (0x20000CBC onwards). The app's code is zero at every one; the overlay holds 40 d6 01 08 -- a firmware address, 0x0801D640 -- and two single bytes, which is what a stack frame or a global written by other code looks like.
That fits the warning app_api.h itself carries: the app runs from the RAM buffer that is also the PY25Q16 sector cache. Next: find what writes at 0x20000CBC, most likely an interrupt handler, which would make this timing-dependent and would explain why the size ladder once looked like a copy that finished late.
The function is in App/radio.c, 179 lines with exactly one loop: while (1) { if ((ReadRegister(REG_0C) & 1) == 0) break; write(REG_02,0); delay(1); }. The model keeps REG_0C at 0x0000 -- its own comment records that this exact hang was fixed once -- and reading the running emulator's reg0c over QOM shows bit 0 clear at every sample before and after the launch, so the loop exits immediately and cannot be where the firmware stops.
The marker test survives review: the blob decodes as sub sp,#4; ldr r0,[pc,#12]; ldr r1,[pc,#12]; str r1,[r0] with 0x20003F00 and 0xDEADBEEF in the literal pool, so an app that runs writes the magic and it never appears -- entry() is not reached. Why is open: with the copy present in the overlay and the CRC matching on paper, the candidates are the two early returns before the call (VMA and CRC) and whether that copy came from this launch.
Method notes: my Thumb bit-field decode was wrong twice this round (Rn/Rt are the low six bits and halfwords are little-endian), which made a correct instruction look like another -- the bytes were right and my reading was not; and the BK4819 register file is readable over QOM at /machine/bk4819, a probe-free way to see what a polling program sees.
A 24-byte app whose first instruction stores 0xDEADBEEF at a fixed address settles it without inference: installed in slot 0 and launched with F, 7, DOWN, MENU, that address reads back all 0xFF at t+2, +5, +10 and +15 s and never as the magic word, so the app's entry point is never executed and APP_LaunchOverlay does not reach entry(&app_api) at line 1162.
Between the copy at line 1113 and the call at 1162 the only return is the CRC check, and the CRC is right (MB_Crc32Bytes is an ordinary CRC-32 and the host tool's value is exactly what it computes), so the function is stuck in between -- and the only hardware-touching statement there is RADIO_SetupRegisters(true) at line 1159. The PC in the PY32 SPI routine at 0x08005118 and the panel frozen at the launcher's box both agree.
Method note: 0x20003F00 is not a free address -- it sits under the top of SRAM and the firmware was using it. The marker test is unaffected, but the free-address assumption was mine and wrong. Also adds the round-40 FillFF variant, which now builds with the documented command line (the build script had been breaking its own compile command); both new apps compile with no warnings.
A memsave of 0x20000280 after MENU holds the app's own code, 2738 of 2744 bytes identical; the PC probe's 14505 samples contain none inside it (the CPU never enters the app) and sit at 0x08005118/0x08005122, which decodes as a PY32 SPI byte transfer polling TXE then RXNE on 0x40013000 -- so the firmware is looping, not hung, matching the 1.93 million panel transfers.
APP_LaunchOverlay, read from the firmware's own app_overlay.c, validates the slot, requires link_vma == the overlay buffer, invalidates the cache, memsets, reads exactly code_size bytes, checks the CRC, then calls entry(&app_api) at line 1162. For this app every gate passes (magic FAP1, hdr 1, abi 1, api_min 1, committed, caps 0, code_size 2744 <= APP_OVERLAY_MAX, entry_off 0, link_vma 0x20000280) and MB_Crc32Bytes is an ordinary CRC-32, so the host tool's 0x3c12630d is exactly what the firmware expects.
So the copy happens, the checks pass, and the call does not take. Between them the only hardware-touching statement is RADIO_SetupRegisters(true) at line 1159, and the PC is in a PY32 SPI routine: the next round is one function wide.
The panel probe now accumulates in memory and writes a buffer at a time instead of once per byte; nothing else changed. Transfer volume went from 28416 pixel-data lines to 1930061, sixty-eight times more, which measures how much the old probe suppressed the guest.
Replaying the model's own store rule over the whole log (a0=1, col 4..131, gram[page & 7][col - 4]) and counting non-zero bytes gives 43 -- exactly what the gram property reported and matching round 42's frozen frame md5. All 1.93 million pixel bytes are accounted for, across all eight pages and all 128 columns, so nothing is dropped and no page is stale.
So the firmware is not failing to draw: it redraws the same 43-lit-byte screen -- the launcher's title box -- about nineteen thousand times. That is also why zeros dominate, and why the glyphs round 40 read as Minesweeper's frame (41, 40, 7f) are the launcher's own characters sent repeatedly. Rounds 39 and 40 were reading the launcher and calling it the app; rounds 32-35 were right that the app does not run.
Eight reads of the controller's RAM over thirty-five seconds on one QMP connection with no probe on: 485 non-zero before the keys, 43 at t+2 s, then identical md5 (43) at t+5, 10, 15, 20, 25, 30 and 35 s. Nothing repaints and the app's screen never appears.
That cannot be reconciled with rounds 39 and 40, where the panel probe recorded 114028 pixel-data transfers after the keys. The probe per-byte opens, writes and closes its file on the vCPU thread, so it is slow enough to change the timing of what it watches -- a trap this file already records twice. The transfer-count conclusion is withdrawn until the count is taken without per-byte file I/O.
What is not in doubt comes from the controller's own memory and needs no probe: after the handover the screen is the launcher's box (43, this file's recorded signature for it) and stays that way, so the app does not paint -- which agrees with rounds 32-35. Next: make the probe accumulate in memory and write once at exit, then re-run the launch.
Read the controller's display RAM over QMP (qom-get /machine/panel gram) with the keypad's press property on one connection: before the keys 485 non-zero, the radio's main screen with F4HWN legible; after F, 7, DOWN, MENU only 43 non-zero, rendered as one boxed title about 42 columns wide and seven rows tall with text inside and fifty-seven empty rows below -- exactly the picture the user described.
With rounds 39 and 40 both things hold: the app runs and blits (105 consecutive full-screen writes from the toggling app; Minesweeper's glyph bytes 7f, 41, 08, 40 in the panel stream), and then dies, after which the launcher repaints its own title box. What remains on the glass is the launcher's frame, so the open question is the app's lifetime, not its drawing.
Method: the QMP socket takes a single client, so pressing keys on one connection and reading the panel on another times out -- do both on one; and the screens were captured with a minimal PNG writer and re-rendered as ASCII, which is what makes them readable given this session has no image input.
Hand-run on a copy of the working image with the panel probe on and F, 7, DOWN, MENU over QMP, now with the real Minesweeper: 103370 pixel bytes after the keys, dominated by 00 (100418) with a longest run of 97640 (about 95 full-screen clears), and underneath the app's own drawing bytes 7f 287, 41 234, 08 192, 40 174 -- its text and box characters. Since the app calls display_clear() once per frame, a clear plus a few hundred glyph bytes is the expected shape, so the app paints and the earlier blank-screen reading was not about the path.
The FillFF variant failed to build through my own mistake: the build script is a textual rewrite, so replacing bigapp with fillff also rewrote the compiler's -o argument and source names; substitute only the names intended, or write that variant's script by hand.
Next: the model exposes the panel's display RAM as the QOM property gram (st7565_get_gram), so a hand-run can read the actual picture over QMP instead of inferring it from transfers.