Commit Graph
137 Commits
Author SHA1 Message Date
mckero e30d3db681 Round 47: RADIO_SetupRegisters is exonerated and round 46's narrowing is withdrawn
The function is in App/radio.c, 179 lines with exactly one loop: while (1) { if ((ReadRegister(REG_0C) & 1) == 0) break; write(REG_02,0); delay(1); }. The model keeps REG_0C at 0x0000 -- its own comment records that this exact hang was fixed once -- and reading the running emulator's reg0c over QOM shows bit 0 clear at every sample before and after the launch, so the loop exits immediately and cannot be where the firmware stops.

The marker test survives review: the blob decodes as sub sp,#4; ldr r0,[pc,#12]; ldr r1,[pc,#12]; str r1,[r0] with 0x20003F00 and 0xDEADBEEF in the literal pool, so an app that runs writes the magic and it never appears -- entry() is not reached. Why is open: with the copy present in the overlay and the CRC matching on paper, the candidates are the two early returns before the call (VMA and CRC) and whether that copy came from this launch.

Method notes: my Thumb bit-field decode was wrong twice this round (Rn/Rt are the low six bits and halfwords are little-endian), which made a correct instruction look like another -- the bytes were right and my reading was not; and the BK4819 register file is readable over QOM at /machine/bk4819, a probe-free way to see what a polling program sees.
2026-10-02 13:07:44 +08:00
mckero 1aba41b284 Round 46: the app's first instruction never executes; the hang is in RADIO_SetupRegisters
A 24-byte app whose first instruction stores 0xDEADBEEF at a fixed address settles it without inference: installed in slot 0 and launched with F, 7, DOWN, MENU, that address reads back all 0xFF at t+2, +5, +10 and +15 s and never as the magic word, so the app's entry point is never executed and APP_LaunchOverlay does not reach entry(&app_api) at line 1162.

Between the copy at line 1113 and the call at 1162 the only return is the CRC check, and the CRC is right (MB_Crc32Bytes is an ordinary CRC-32 and the host tool's value is exactly what it computes), so the function is stuck in between -- and the only hardware-touching statement there is RADIO_SetupRegisters(true) at line 1159. The PC in the PY32 SPI routine at 0x08005118 and the panel frozen at the launcher's box both agree.

Method note: 0x20003F00 is not a free address -- it sits under the top of SRAM and the firmware was using it. The marker test is unaffected, but the free-address assumption was mine and wrong. Also adds the round-40 FillFF variant, which now builds with the documented command line (the build script had been breaking its own compile command); both new apps compile with no warnings.
2026-10-02 13:04:45 +08:00
mckero dd1c4136e8 Rounds 44-45: the loader read line by line; the overlay is loaded and every gate before the call passes
A memsave of 0x20000280 after MENU holds the app's own code, 2738 of 2744 bytes identical; the PC probe's 14505 samples contain none inside it (the CPU never enters the app) and sit at 0x08005118/0x08005122, which decodes as a PY32 SPI byte transfer polling TXE then RXNE on 0x40013000 -- so the firmware is looping, not hung, matching the 1.93 million panel transfers.

APP_LaunchOverlay, read from the firmware's own app_overlay.c, validates the slot, requires link_vma == the overlay buffer, invalidates the cache, memsets, reads exactly code_size bytes, checks the CRC, then calls entry(&app_api) at line 1162. For this app every gate passes (magic FAP1, hdr 1, abi 1, api_min 1, committed, caps 0, code_size 2744 <= APP_OVERLAY_MAX, entry_off 0, link_vma 0x20000280) and MB_Crc32Bytes is an ordinary CRC-32, so the host tool's 0x3c12630d is exactly what the firmware expects.

So the copy happens, the checks pass, and the call does not take. Between them the only hardware-touching statement is RADIO_SetupRegisters(true) at line 1159, and the PC is in a PY32 SPI routine: the next round is one function wide.
2026-10-02 13:02:11 +08:00
mckero 0690770f2c Round 43: the probe fixed, the two instruments agree, and the app never runs
The panel probe now accumulates in memory and writes a buffer at a time instead of once per byte; nothing else changed. Transfer volume went from 28416 pixel-data lines to 1930061, sixty-eight times more, which measures how much the old probe suppressed the guest.

Replaying the model's own store rule over the whole log (a0=1, col 4..131, gram[page & 7][col - 4]) and counting non-zero bytes gives 43 -- exactly what the gram property reported and matching round 42's frozen frame md5. All 1.93 million pixel bytes are accounted for, across all eight pages and all 128 columns, so nothing is dropped and no page is stale.

So the firmware is not failing to draw: it redraws the same 43-lit-byte screen -- the launcher's title box -- about nineteen thousand times. That is also why zeros dominate, and why the glyphs round 40 read as Minesweeper's frame (41, 40, 7f) are the launcher's own characters sent repeatedly. Rounds 39 and 40 were reading the launcher and calling it the app; rounds 32-35 were right that the app does not run.
2026-10-02 12:57:19 +08:00
mckero 40d24b74c5 Round 42: the panel memory is frozen at the launcher's box; the probe-based conclusion is withdrawn
Eight reads of the controller's RAM over thirty-five seconds on one QMP connection with no probe on: 485 non-zero before the keys, 43 at t+2 s, then identical md5 (43) at t+5, 10, 15, 20, 25, 30 and 35 s. Nothing repaints and the app's screen never appears.

That cannot be reconciled with rounds 39 and 40, where the panel probe recorded 114028 pixel-data transfers after the keys. The probe per-byte opens, writes and closes its file on the vCPU thread, so it is slow enough to change the timing of what it watches -- a trap this file already records twice. The transfer-count conclusion is withdrawn until the count is taken without per-byte file I/O.

What is not in doubt comes from the controller's own memory and needs no probe: after the handover the screen is the launcher's box (43, this file's recorded signature for it) and stays that way, so the app does not paint -- which agrees with rounds 32-35. Next: make the probe accumulate in memory and write once at exit, then re-run the launch.
2026-10-02 12:53:51 +08:00
mckero 58c030d213 Round 41: the panel's own memory reproduces the user's screen; the app runs and then dies
Read the controller's display RAM over QMP (qom-get /machine/panel gram) with the keypad's press property on one connection: before the keys 485 non-zero, the radio's main screen with F4HWN legible; after F, 7, DOWN, MENU only 43 non-zero, rendered as one boxed title about 42 columns wide and seven rows tall with text inside and fifty-seven empty rows below -- exactly the picture the user described.

With rounds 39 and 40 both things hold: the app runs and blits (105 consecutive full-screen writes from the toggling app; Minesweeper's glyph bytes 7f, 41, 08, 40 in the panel stream), and then dies, after which the launcher repaints its own title box. What remains on the glass is the launcher's frame, so the open question is the app's lifetime, not its drawing.

Method: the QMP socket takes a single client, so pressing keys on one connection and reading the panel on another times out -- do both on one; and the screens were captured with a minimal PNG writer and re-rendered as ASCII, which is what makes them readable given this session has no image input.
2026-10-02 12:50:06 +08:00
mckero 51a8c3935e Round 40: Minesweeper's frame content reaches the panel; the log is dominated by its per-frame clear
Hand-run on a copy of the working image with the panel probe on and F, 7, DOWN, MENU over QMP, now with the real Minesweeper: 103370 pixel bytes after the keys, dominated by 00 (100418) with a longest run of 97640 (about 95 full-screen clears), and underneath the app's own drawing bytes 7f 287, 41 234, 08 192, 40 174 -- its text and box characters. Since the app calls display_clear() once per frame, a clear plus a few hundred glyph bytes is the expected shape, so the app paints and the earlier blank-screen reading was not about the path.

The FillFF variant failed to build through my own mistake: the build script is a textual rewrite, so replacing bigapp with fillff also rewrote the compiler's -o argument and source names; substitute only the names intended, or write that variant's script by hand.

Next: the model exposes the panel's display RAM as the QOM property gram (st7565_get_gram), so a hand-run can read the actual picture over QMP instead of inferring it from transfers.
2026-10-02 12:43:36 +08:00
mckero 33e2d7f38e Round 39: an overlay app's blits DO reach the panel; the probe cap was hiding them
Raise the panel probe cap (40000 transfers is reached about forty seconds after boot) and rebuild qemu-system-arm, then run the toggling app on a hand-copy: after F, 7, DOWN, MENU the probe grew from 22568 to 139404 lines, 114028 of them pixel data, with the longest run of byte 00 at 108273 transfers -- about 105 full-screen writes in a row, histogram 111061 zeros against a longest 0xFF run of 68.

A hundred consecutive full-screen fills is a loop, not a launcher repainting once: the app runs and blits, which retires the conclusion carried since round 32 that an app's calls do nothing. What hid it was the cap -- the instrument went blind forty seconds after boot and the blindness was read as an app that never acted; the third time this file records capping a diagnostic before knowing the shape of the data, and the first where the cap was inside the model.

Next: the 0xFF phase never appears. An app that fills 0xFF only, with no zeros at all, separates 'it never gets there' from 'the firmware repaints over it' in one run.
2026-10-02 12:40:55 +08:00
mckero 6045de09b1 Round 38: a hand-run carries the measurement; the panel probe has a 40000-transfer cap
With a hand-started emulator on a copy of the working image, the page stayed untouched and healthy. Two of my own call signatures were wrong: uvk5_apps.install takes (image: bytearray, slot, blob: bytes, force) and edits the image in memory, so a path string is read as a blob, and key.py's Qmp wants a bare host:port -- the tcp: prefix fails getaddrinfo, the very trap the socket section records.

The useful find: the panel probe stops after 40000 transfers (panel_probe_n < 40000 in st7565_xfer) and a hand-run reaches that within about forty seconds, so a launch measured after that looks like an app sending nothing -- the same mistake this file records twice about capping a diagnostic. Raise the cap (or make it a knob), rebuild qemu-system-arm with the page's job stopped first, and re-run the launch on a hand-copy.
2026-10-02 12:36:53 +08:00
mckero 11286a3817 Round 37: the panel probe works; the page's environment drops the variable; two notes corrected
A hand-run with the variable definitely present wrote 29299 probe lines in forty-five seconds, 28416 of them pixel data, so the panel path is driven as intended and the instrument is sound. uvk5_supervisor.py builds QEMU's env as dict(os.environ) yet the page's emulator never sees the variable, so it is lost between the shell that starts the page and the server process it becomes; pass it like the other launcher options instead of relying on inheritance.

Corrections: the probe prints s->selected, not a chip-select level, so cs=1 means the panel IS selected and the bytes ARE stored; and the old note that a hand-started emulator never draws is not true of a hand-run against the page's working image -- this one booted, drew and streamed pixels, so that note should be re-derived rather than believed.
2026-10-02 12:34:05 +08:00
mckero fc18c157e1 Round 36: the panel probe is in the binary but not in the environment; the page was restored
Replacing the page to get UVK5_PANEL_PROBE into the emulator's environment cost the user their running page, because the port owner of 8080 was the managed launcher job. It was restored as a managed background job and verified healthy (485).

The instrument is present: the running qemu-system-arm.exe contains UVK5_PANEL_PROBE and 'PANEL a0=' and was built after the source was last modified, yet no panel.log appears anywhere, so the variable is not in that process's environment. run-webui.ps1 passes only --qemu/--elf/--flash and webui.py has no env= dict in the paths read, so the spawn site -- most likely tools/uvk5_supervisor.py -- still needs reading.

Recorded two process-inspection traps: Get-CimInstance and Get-Process return nothing in this sandbox while netstat and taskkill work, and a second server instance dies silently while the first holds the port.
2026-10-02 12:31:13 +08:00
mckero c425af4d31 Round 35: the model exonerates the sector cache; the panel instrument needs a page restart
The panel is on SPI1 and the flash on SPI2 (the model says so and wires st7565_xfer to soc.spi[0]), so a blit cannot touch the external flash and round 34's sector-cache-overwrite mechanism is withdrawn -- which also matches the flash probe seeing nothing after the app's code load.

The panel model reads correctly (stores only when cs low and a0 high, keeps its own gram, 132-column counter with the col-4 store including the four-column fix), so nothing there explains a dropped blit. What is missing is the measurement: UVK5_PANEL_PROBE logs every byte sent to the panel with a0/cs/page/col, which separates 'the app's blit never reaches the driver' from 'the driver's bytes never land'. It needs the page restarted with that variable; two attempts failed because port 8080 is occupied and the port-owner lookup came back empty. The page stayed healthy (485).
2026-10-02 12:25:44 +08:00
mckero 46b077233b Round 34: the app side is exonerated; the firmware header names the mechanism
At key.py's timing the launch is coherent and reproducible (485 -> 484 F -> 580 7 -> 617 DOWN -> 43 MENU, the handover). A toggling app -- fill FF blit, fill 00 blit, forever, a signature no menu can produce -- gave fifty readings over twenty-five seconds all exactly 43: its blits never reach the glass.

Offsets checked properly: abi_major (u8) + api_level (u8) + api_size (u16) is four bytes, not twelve, so fb is at 4, display_clear at 8 and print_tiny at 28, exactly where the app calls them; the earlier draw_rect/api_size reading was my own arithmetic error. With the header identical, the entry at 0, the offsets right, the firmware answering status 0 and the handover confirmed, every part of the app side now has evidence.

Left is the path between the app's call and the glass, and the firmware's app_api.h states the mechanism: the app runs from the RAM buffer that is also the PY25Q16 sector cache, and blit_full drives the panel over the same SPI bus, so a blit can reload the sector cache and overwrite the running app. Next: instrument the model's sector-cache path around a blit from an app.
2026-10-02 12:22:21 +08:00
mckero 73b4fb161b Round 33: the fill-and-blit probe does not reproduce; one of my readings was a menu row
Replicating round 17 exactly (332-byte fill-and-blit in slot 0, page-default taps) gave 485 -> 580 after F,7 -> 627 after MENU, and 627 held: the screen never left the app menu, so this run says nothing about the app. It did expose that 537/549/552/580/596/617/627 are one family -- the app menu, whose ink moves with the highlighted row -- while a genuinely all-lit panel would be ~1024 and has never been seen. At least one earlier 'it painted' reading was therefore a menu row.

Next: drive the launch with key.py's timing and confirm the handover by the 43-byte launcher box, then judge the app only by a signature no menu can produce.
2026-10-02 12:19:32 +08:00
mckero 4c842e7b32 Round 32: the api header is byte-identical to the firmware's, and the font-read inference is suspect
The working copy's app_api.h and the firmware's own (fetched from armel/uv-k1-k5v3-firmware-custom) are the same 14403 bytes with the same sha256, so the layout theory of round 31 is withdrawn; the assembly confirms the calls land on print_tiny at offset 28 and display_clear at 8 as expected.

The measurement those conclusions rested on is now in doubt: 1446 font-region reads early in the session look like the font being cached in RAM, in which case print_tiny never touches the external flash and 'no font read' says nothing about whether draw() ran. The next oracle must be panel-only, and the sharpest unexplained contrast is that the same fill-and-blit lights the panel from a 320-byte app and not from the first statement of this one.
2026-10-02 12:17:34 +08:00
mckero d3fcff7ac3 Round 31 notes: the app's first API call has no effect; the struct layout is the suspect
Header exonerated against four upstream apps (entry_off 0, name at 20, version at 36, link_vma 0x20000280 identical), entry point exonerated by the listing and linker script, acceptance confirmed by the firmware answering status 0 over 0x0730. What remains: the app calls the api members at offsets 8 and 28 and neither call changes the panel nor produces a flash read, while the app itself loops in its own first 180 bytes.
2026-10-02 12:15:44 +08:00
mckero a64ea430a5 Keep the two AGENTS files in step: add the Chinese round-30 note
The Chinese append in round 30 used a fragment that does not exist in the file, so only the English note landed. Appended after the round-29 Chinese note and verified both files contain the new note.
2026-10-02 12:11:46 +08:00
mckero 0c15b63c34 Round 30: the app runs in the overlay and never calls the firmware
The firmware answers status 0 over 0x0730 for all eight installed slots, the game and the 256-byte toy alike, so round 29's claim that the launcher refuses some apps is withdrawn. Measured with the instruments: at the key.py timing MENU drops the panel to 43 within a second, the 2 ms PC probe sees 51 of 11439 samples inside the overlay spread over about ten addresses, and the flash probe shows the app's code load as the last transaction of the session with zero after it -- no font reads, so print_tiny never ran and draw() is never reached.

place_mines() is the one loop on that path that can spin without touching the firmware, and it is now capped at 4000 attempts with a deterministic fallback, so a hang there can no longer look like an app that never started.
2026-10-02 12:11:14 +08:00
mckero 865c752173 Round 29 notes: key timing decides coherence; the launcher refuses some apps
The page's key endpoint defaults to TAP_MS = 60 with no gap when hold_ms is omitted (what every curl did), while tools/key.py uses 200 ms plus GAP_MS = 400 for the release to debounce. At key.py's timing the same sequence is coherent: 485 -> 484 (F) -> 580 (7) -> 617 (DOWN) -> 43 (MENU, the handover); at the page default it is not.

With the handover reachable the real Minesweeper reaches 43 while a 256-byte fill-and-blit toy never leaves the menu (582): the launcher distinguishes them before either runs, so the question is which apps the launcher accepts -- and the code-size framing of rounds 27-28 is retired.
2026-10-02 12:07:25 +08:00
mckero 10ec032d38 Keep the two AGENTS files in step: add the Chinese round-28 note
The English note was appended by the round-28 script but the Chinese one was not, because the anchor fragment did not match and the script asserted. Appended it after the round-27 Chinese note and re-checked both files.
2026-10-02 12:00:53 +08:00
mckero 46be242809 Round 28: the code-size ladder is built; the entry point is what blocks it
The ladder app calls a chain of no-op functions then fills every framebuffer byte with 0xFF and blits, so the panel either goes all lit or stays as it was. The chain costs about 5.2 bytes per function (20 -> 256 B, 70 -> 620 B).

Neither rung could be measured because the launch never reached the handover: with the 256-byte app in all four slots the readings were 485 (main), 460 (after F, 7), 537 (after DOWN, the menu is open) and then 582 for twelve consecutive readings across three MENU presses -- the first press moved something, the rest did nothing. That is the inert-launcher state already documented, so the entry point rather than the app is the subject for the next round.
2026-10-02 12:00:35 +08:00
mckero cad35e85dd Round 27: a correctly placed first action still does not run; the fitting variable is code size
The game's app_main was given a fill-every-byte-and-blit first act. Inserted before A = api; it dereferenced an unset pointer and faulted (my own bug, that run is void); moved after the assignment (source +199 B, app 2508 B) the panel still never leaves 43 over 48 s, so the blit never happens, while a 300-byte app doing the same fill-and-blit paints 596 non-zero bytes.

The variable that fits every measurement is the app's code size rather than its blob size: a 3028-byte blob with a few dozen bytes of code runs and can call print_tiny, display_clear and get_key; Phases at 372 bytes of code runs; Ladder at 544 dies; Minesweeper at 2400+ dies with its first instruction having no effect. Next: sweep a trivial app at 400, 512, 600, 800 and 1200 bytes of code through the handover to find where the panel stops changing.
2026-10-02 11:54:33 +08:00
mckero 82a34fe39a Round 26: the launch depends on which row is selected, and the handover works
After the pristine restore only slot 3 held a real app; slots 0-2 read as factory data (the page calls them unknown), and every failed launch had the cursor left on one of those. Installing the game into slots 0-2 as well, so every row holds it, makes the open sequence work: the menu comes up at 617, MENU drops the screen to 43 (the handover) and at t+40s it reaches 484, the radio's own main screen -- so the app is entered and then returns.

The radio is left with the game in slots 0 through 3 so the launch is reliable from the page whatever row the cursor lands on. Next: launching the game and getting out of it, and why the app returns rather than staying up.
2026-10-02 11:50:03 +08:00
mckero a027997f6c Round 25: the handover discriminator could not run, since the launch did not reproduce
A 20-byte app that calls nothing was installed in slot 3 (the slot round 24 saw loaded and entered) and driven with the same F, 7, DOWN, DOWN, MENU sequence: the menu came up at 511 non-zero bytes, MENU changed nothing, the screen stayed there for twenty seconds and the PC probe recorded zero overlay samples out of 7316. So the experiment answered nothing, and the honest reading is that the launch is not reliably reproducible through the page's synthetic key events -- the same sequence produced a handover in round 24 (ink 210). That is a statement about the entry point, not the app: rounds 24 and 17 measured the load, the entry and a painted frame from this same artifact.

Minesweeper was restored into slot 3 so the radio is left with the game installed rather than the test stub.
2026-10-02 11:46:40 +08:00
mckero 2547de9b2d Round 24: the real launch sequence, and proof the app is loaded and entered
Reading the panel after each single press from a fresh power-on gives the sequence the launcher wants: F, 7, DOWN opens the app menu (489 -> 552), DOWN again moves the selection (552 -> 526, ink 1799 -> 1677), MENU hands the screen over (ink 210). The earlier note that F then 7 opens the menu is wrong, and that error is why several rounds of MENU presses looked inert.

At the handover the flash probe shows the load as the session's last two transactions: addr=108000 len=64 first=46415031 (the FAP1 header of slot 3) then addr=109000 len=2444 first=f0b599b0 (the app's code, exactly code_size). The 2 ms PC probe catches four samples inside the overlay in three seconds, so the loader copies and the CPU enters the app, which then disappears and leaves the launcher's box on the glass.
2026-10-02 11:44:40 +08:00
mckero 62b7b5f7c3 Round 22-23: pristine image restored; the launch itself does not reproduce
After a day of app installs the working copy was no longer the image that worked in round 17, so the pristine dump was restored (with the emulator powered off first, since exit writes the in-memory image back over the file). The page now draws its own main screen (485 non-zero bytes) and Minesweeper is installed in slot 3 at 2444 bytes, the same artifact that painted a complete frame in round 17.

The launch does not reproduce: F, 7, DOWN, MENU and a second MENU leave the radio on a fully drawn app menu (ink 1799) that stops responding to keys, identical readings before and after every press. Round 17 measured the game's frame with this same app, slot and a F, 7, DOWN x3, MENU sequence, so the app is not in question; what differs is the launcher's key handling. An A/B that removed this round's instrumentation and rebuilt the round-17 source did not bring the frame back, which is what moved the suspicion off the app.
2026-10-02 11:37:01 +08:00
mckero db01005874 Round 21: the app menu needs a press before MENU, which explains the probe failures
The panel reading that looked like a paint was the app menu itself: rows 0-6 the title box, rows 18-22 the slot list, 26 lit pixels on row 2 where the probe's marker bar would be 120. So MENU on the initially opened menu selects a row rather than launching; round 17's Minesweeper run worked because it pressed DOWN three times first.

That accounts for every unexplained probe failure in the last three rounds and rehabilitates the probes: with a press first, an app that clears the framebuffer, draws through a helper and blits does run (499 on the panel, 10.9% of PC samples inside the overlay). The launch sequence is F, 7, DOWN to select, MENU to run.
2026-10-02 11:18:11 +08:00
mckero d099e2ecfc Round 20: the key-path conclusion is contradicted, and the probes' failure is unexplained
Round 15 measured get_key() repeated in an app loop at 27.1%, the same as the control, so calling it every pass is not what kills an overlay app; the previous section's use of the two key probes as evidence for that is wrong. Four probe builds (with and without an opening spin, with a short and with a 2M-iteration per-pass spin) all left the panel byte-identical at the launcher's own frame, so they did not run, and why is not established.

What that is not: not the shape other apps survived in, not get_key, not the framebuffer writes -- an app in the same slot that fills every framebuffer byte through api->fb and blits does paint (43 -> 596 non-zero bytes). The probes differ by zeroing the framebuffer and drawing with their own helper before blitting, which is the sort of difference that has to be isolated one change at a time rather than reasoned about.
2026-10-02 11:08:15 +08:00
mckero 3b1d4abbbf Round 18-19: the key probe never painted, and the contrast narrows input to get_key()
Two font-free probes meant to draw the raw key code left the panel byte-identical across seven presses (26 lit pixels on row 2 and a constant row-20 pattern = the launcher's own frame), so neither painted. The contrast matters: an app of the same shape that fills the framebuffer through api->fb and blits does paint (43 -> 596 non-zero bytes), and its only substantive difference from these probes is that they call api->get_key() every pass. With round 17's result -- the full Minesweeper painted a frame and vanished the moment a key was pressed, returning the panel to the launcher's title box, which only happens on the key-as-EXIT path -- input is now one call wide.

The sector-cache question returns with it, since the key path is a plausible place for the firmware to reach the external flash, but the font reads print_tiny provokes do not kill the app, so this is specific to the key path rather than to flash reads in general.
2026-10-02 11:05:32 +08:00
mckero 90d5a1f687 Round 17: the game paints -- the missing piece was a per-frame delay
A clear-and-blit app left the panel byte-identical, but filling the whole framebuffer through api->fb and blitting turned it from 43 non-zero bytes to 596: an overlay app's framebuffer writes and blit_full do reach the panel. What blocked the frame was Minesweeper's own delay_ms(40) on the invalid-key path -- 40 ms of guest time is seconds of wall time here, on top of roughly twenty seconds of font reads per frame, so frames were minutes apart. With it removed the panel went from 43 to 390 non-zero bytes within 24 s, ink 1275, with the title, counters and field legible.

Input is the remaining gap: MENU changed nothing and DOWN sent the app out of its loop (the panel returned to the launcher's 43-byte title box, which only happens on the path that treats a key as EXIT). Next: have the app draw the raw key code it receives and read it off the panel instead of guessing at the mapping.
2026-10-02 10:59:41 +08:00
mckero a630500e22 Round 16: draw()'s halves are each alive, so the blit is the next suspect
With the validated shape (for(;;) { long spin; one piece; }) the header block scores 26.1% and the 81-cell loop plus cursor, with locals only, also 26.1% -- both equal to the control, so both are alive. The full Minesweeper is running too: 116 samples inside it over fifteen seconds across eight distinct addresses, with the firmware still serving it (1024 font-region reads), yet the panel never leaves the launcher's title box over ninety seconds, and that box is painted before the app is entered. Removing the opening settle spin changed nothing, which retires the idea that it was still inside that spin.

Two candidates remain: the statics, which draw() reads and neither surviving variant touched, and whether blit_full from an overlay app reaches the panel model at all -- the gutted variant was only ever measured by its PC share, never by its screen.
2026-10-02 10:53:22 +08:00
mckero 5b861f82ab Validated metric: every individual call survives repetition, and the killer is inside draw()
Each variant is for(;;) { long spin; one call; } so a live app is in the overlay most of the time: control 27.2%, display_clear 26.9%, get_key 27.1%, delay_ms(40) 26.8%, print_tiny 27.0% -- all four calls survive repetition, and the control shows the metric separates the two cases.

That also means the previous round's variant B (four calls in a loop, no spin) was most likely not dead at all: it spends nearly all its time inside firmware functions, which look exactly like a dead app from the trace. The bisect has since closed on one function: Minesweeper with draw() gutted to display_clear() + blit_full() runs (97 samples inside the app, alternating with firmware PCs), while the full draw() draws nothing in twenty seconds and ink never leaves 210, so display_clear() never ran. The end of the app is inside draw()'s body.
2026-10-02 10:44:18 +08:00
mckero 057e408016 Clean A/B: the app dies when the same firmware calls are repeated, not when they are made
Same file, same call order, one variable, measured on the page's own emulator with the 2 ms PC probe. Variant A calls display_clear, print_tiny, get_key and delay_ms(40) once and then spins: it runs (2059 samples inside the app, 1024 font reads). Variant B wraps exactly those four in for(;;): the app is never seen again (0 samples inside it, and the same 1024 font reads, so the first pass really did execute). So the calls are fine individually and fine in sequence once; repeating them ends the app.

Also records that the metric needs fixing first: samples inside the app under-count a live app, because time spent in long firmware functions lands on the firmware side of the trace, which is what a dead app looks like too. The display_clear-only loop scored 6 and the get_key-only and delay_ms-only loops scored 0, and those are one good measurement and two unreadable ones, not three results. The fix is to alternate a long self-contained spin with each call.
2026-10-02 10:35:00 +08:00
mckero f6c0115aa5 The size ladder was a red herring: Minesweeper fails on its own code
The ladder conflated two variables -- every small test app spun first, every big one called the firmware immediately. Separated, all measured with the 2 ms PC probe on the page's own emulator: a 3028-byte app that calls nothing loops in the overlay indefinitely (1704 samples, all inside its first 512 bytes), so blob size is not the problem; adding one print_tiny changes nothing (2039 samples, with 1024 font-region reads in the flash log, the first at 0x1e0000), so the font path and the sector-cache worry are both fine; adding get_key as well still runs (2029 samples); adding display_clear too still runs (1978 of 13195 samples inside the app).

So the three calls no surviving app had ever made are harmless, and the API surface Minesweeper uses is covered. Its failure is its own bug; the next step is an ordinary bisect of its draw() path through the page.
2026-10-02 10:24:22 +08:00
mckero 36c8c66006 The app-size ladder: the launcher jumps before the app is fully copied
Measured with the 2 ms PC probe: Spin (16 B) loops in the overlay forever; Delay (248 B) is still running 13 s later, 54 of the last 60 samples at 0x20000348; Phases (372 B) survives bar, blit and led; Ladder (544 B) dies before drawing; Minesweeper (2408 B, and 2428 B with an opening spin) is dead within about 4 ms either way. An opening spin does not save the big apps, so the difference is how much of the app has been copied rather than when the first call is made -- and that also explains why Tetris's full 3300 bytes were present in the overlay when checked afterwards: the copy finishes after the app has been entered.

Same family as the four flash bugs already recorded here: the loader is told the transfer is over while the guest is still feeding it. Next: watch the SPI/DMA busy and complete flags during a large read instead of reasoning about them.
2026-10-02 10:19:15 +08:00
mckero f3a369d38f Make the PC probe interval configurable, and trace the app's life as one tick
uvk5_pc_probe_interval_ms() reads UVK5_PC_PROBE_MS and falls back to 100, so existing probe runs are unchanged. At 2 ms a launch gives 468 samples in about 0.94 s and exactly one of them is inside the 4 KiB overlay: the app lives for roughly 2 ms, not 100.

The trace settles two things and adds one. The loader copies and jumps (the single overlay sample; a 16-byte app that calls nothing stays there forever). The sector-cache overwrite is dead: the flash log's last transaction for the run is the app's own code load (addr=103000 len=2408) and nothing follows it, no font read included. New: after the app dies the firmware loops inside its own flash (PCs near 0x08005118/0x08005122/0x08005248) with zero flash reads and no key response at all -- F then 7 no longer opens the menu and an explicit MENU down/up changes nothing -- so the launcher or its fault path is wedged.
2026-10-02 10:14:50 +08:00
mckero 394f473df0 Correct the overlay cause: the log order rules the sector-cache overwrite out
The flash probe records every transaction in order and the last one for that launch is the app's own code load (addr=103000 len=544, first bytes f0b583b0); nothing follows before the app is gone, so no font or resource read lands on the running app in that window. The 7848 font-region reads in the log belong to the firmware's own repainting.

What stands measured: the loader copies and jumps (a PC sample lands in the overlay, and a 16-byte app that calls nothing loops there forever); the app dies well under 100 ms (panel reads at t+113 ms still show the menu, at t+193 ms only the launcher's title box); an app that spins first survives a bar, a blit and led and dies around delay_ms; one that calls blit immediately is gone before it can be seen. Next: drop the PC probe interval from 100 ms to a couple of milliseconds, which is a one-line model change and turns it into a real trace of the app's short life.
2026-10-02 10:08:40 +08:00
mckero e9c344b52a Name the overlay failure: an app runs until it calls a service that reads the external flash
Measured on the page's emulator with UVK5_PC_PROBE and UVK5_FLASH_PROBE inherited through its own launcher, so the instance measured is the one that draws. Spin (16 bytes, calls nothing) sits at 0x20000286 in 42 of 383 PC samples after MENU: it loops in the overlay forever, so the loader works. Minesweeper (2408 bytes) is loaded -- the only big flash read is 0x109000 len 2408, first bytes its own -- and a single PC sample catches it inside the overlay before it is gone: it lives well under 100 ms. Phases (372 bytes) draws a progress bar straight into the framebuffer between calls, and the bars for framebuffer-only and led are on screen while the bar after delay_ms never appears.

The overlay is the PY25Q16 sector cache (app_overlay.h says so), so a service reading the external flash -- the font table is at 0x1E0000 -- lands on top of the running app. Real hardware cannot behave that way, since upstream's own apps call print_tiny, so the firmware must gate the sector cache while an app is loaded; that gate is what the model is missing. Next: find it in the flash driver and check what the model answers.
2026-10-02 10:05:24 +08:00
mckero 53a04ee869 Measure the launch on the page's own healthy instance: the menu works, the app never runs
With the pristine dump restored the page draws (485 lit bytes) and answers 0x0730 with four committed apps. Driving its /api/key and reading its /api/panel: F then 7 opens the menu (ink 1810 -> 1976), DOWN x3 moves the selection (1976 -> 1832), MENU on the app row drops the screen to the title box alone (1832 -> 210) after which MENU, DOWN, UP and F change nothing, and EXIT brings the menu back (1832). So the launcher is entered and left, and the app neither draws nor reads keys.

This supersedes the inconclusive PC measurements taken on hand-started QEMU instances, which came up without a picture while the page's instance draws. Recorded in both AGENTS files: restore work/user-flash.img before testing (the copy carrying the adopted FMP3 marker, image_crc32 0x9d27c3db rather than 0x4d87ce48, comes up blank and makes keys dead, which earlier rounds mistook for the app failing), and measure through the page rather than a hand-started emulator.
2026-10-02 09:59:14 +08:00
mckero bfd1a9ed2c Correct the overlay measurements: they were taken on blank instances, and this build has no F+7 menu
Hand-started QEMU instances all came up with a blank panel (0 lit pixels) while the page's own instance draws (803), same image and a different firmware file, so 'the PC never entered the overlay' was measured on a radio that never reached its main loop and is inconclusive rather than a finding.

Driving the page's own /api/key and reading /api/panel reproduces what the user sees: F, 7, DOWN x3 and MENU all answer ok and the screen does not change by a pixel (ink 232 throughout). The page's firmware.bin is 109.3 KiB and the Labs build that does open the app menu is 111.9 KiB -- a different build. The radio still answers 0x0730 with four committed apps, so it supports the app region but has no F+7 entry. Testing an overlay app requires running the Labs build, and measuring through the page rather than a hand-started QEMU.
2026-10-02 09:51:58 +08:00
mckero fa0f4e7f23 Record that no overlay app runs: the copy lands, the jump never happens
Measured by polling the PC over QMP every 30 ms, finer than the model's own 100 ms probe. After launching Tetris the overlay at 0x20000280 holds Tetris's code byte for byte, so the loader's copy is correct and aligned; the earlier claim in this session that it was shifted by one byte was my own parsing dropping the first value. But 172 PC samples over 4.5 s and 64 more over 2 s were all inside firmware flash at 0x08013260, a wait loop, and none landed in the overlay. A 16-byte app that calls nothing behaves identically, and Breakout's header is field-for-field the shape of ours, so neither the app's code nor its header is the reason. The failure is upstream of the app, which makes the earlier 'Tetris runs' note stale: per this file's own rule it is treated as unverified until it reproduces.

Also carries the CI-fix branch merge and a regression test for _edit_flash creating its working-copy directory (its fixture image has to be a real size: slot 0 sits at 0x102000).
2026-10-02 09:45:33 +08:00
mckero c0c2c3230e Merge the CI fix for the unit job 2026-10-02 08:37:40 +08:00
mckero f938d0b3bd Build overlay apps without the Arm toolchain, and say how far verification got
pip install ziglang cross-compiles to thumb-freestanding-eabi, which is enough to build a .app with no arm-none-eabi-gcc and no Docker. Measured refusals: --defsym, -Ttext and --section-start come back as unsupported linker args, so the VMA is resolved into a copy of app.ld; -T is forwarded (a missing script errors); --image-base is accepted but page-aligns the segments into 0x200103C8 and 0x20020C94. tools/elf2bin.py extracts allocated sections rather than program headers, because lld maps the ELF header and phdr table as a 180-byte LOAD of its own -- following the headers starts the image at 0x20000000 and APP_ERR_VMA. One division pulled in __aeabi_uidiv, which a -nostdlib blob cannot have: Minesweeper now avoids division entirely. app_main carries the .text.entry attribute upstream's apps use, so the entry is first for the loader's jump to offset 0.

Minesweeper builds to 2408 bytes of code against a 4096-byte budget. Installed through the page, the firmware reads the slot header twice and then exactly code_size bytes from slot+0x1000 -- the only read of that size in the boot log -- so the blob shape, header, CRC, VMA and offset are all accepted. Whether control reaches the overlay is unproven: the 100 ms PC probe saw no overlay address, no APP ERROR screen appears, and the app does not draw. test_elf2bin pins the phantom-header-segment lesson; both AGENTS files record the rest.
2026-10-02 08:32:22 +08:00
mckero 61344e486e Minesweeper: verify its logic on the host, and document it in both languages
apps/minesweeper/host_test.c includes the app source with a fake app_api_t, so the real state machine runs on a PC: every string it draws is recorded and the lit pixels are counted. Measured -- a reveal/flag/new-game/digit/quit script returns normally, draws the title, the mine count and lights 24 pixels; a script that blindly reveals 85 cells reaches a terminal state eleven times, draws BOOM, and MENU starts a new game after it (terminal at record 356, title again at 4180). The win path is the one branch blind play does not reach.

Both READMEs now say what is verified (compiles clean under gcc -Wall -Wextra -Werror against upstream's real app_api.h, which is what caught the API's true member names and the absence of left/right keys; the logic runs on the host) and what is not (never built for ARM, never run on the radio -- no toolchain and no Docker here).
2026-10-01 22:31:22 +08:00
copilot-swe-agent[bot]andMCKero6423 6343c150ef Fix unit CI failures in app tests
Co-authored-by: MCKero6423 <222912022+MCKero6423@users.noreply.github.com>
2026-10-01 14:30:57 +00:00
mckero ff14dffb67 Fix the blank screen, correct the multiboot note, and add Minesweeper
The blank screen was self-inflicted and the earlier explanation was wrong. The FMP3 marker at 0x100000 is an ordinary state record -- generation, image_size, image_crc32, firmware_slot/slot_inv, config_bank/bank_inv and a matching state_crc32 (0x661286F1) -- not a pending flag. What actually happened: a firmware uploaded over a state that recorded a different identity made the firmware take the restore/adopt path, which draws nothing (panel 1024/1024 bytes zero) while the serial banner printed normally. Restoring the untouched dump fixed it at once (485/1024 bytes lit) and the three apps reinstalled and were confirmed by 0x0730.

Clearing the marker sectors was tried twice (a whole 8 KiB, then just the 24-byte headers) and is not a fix: it sends the firmware down MB_MARK_MISSING = fresh radio, which adopts the running firmware slowly and without drawing, and its own write-back restores the marker anyway. AGENTS.md and AGENTS.zh-CN.md now say so in place of the wrong claim.

apps/minesweeper/ adds our own 9x9 minesweeper for the 4 KiB overlay: no left/right keys exist on this radio (so the cursor walks with UP/DOWN and digits pick a row then a column), 81 cells need three 9-byte bit arrays rather than a uint16_t mask (the uint16_t version compiled fine and was wrong past cell 15), mines are placed after the first reveal so it cannot lose immediately, and the source compiles clean under gcc -Wall -Wextra -Werror against upstream's real app_api.h. It has not been built for ARM or run -- no toolchain here -- and the README says so.
2026-10-01 22:29:36 +08:00
copilot-swe-agent[bot] 5a8e987451 Initial plan 2026-10-01 14:27:49 +00:00
mckero 2b3222155f Record the restore loop that ate the installs
A real image whose 0x100000 marker held FMP3 next to a committed slot 0 made the factory bootloader reflash the internal flash from that slot on every power-on: the serial banner reappeared once a cycle (7 -> 8 in 25 s) while the screen never changed. Clearing the two marker sectors (0x100000..0x101FFF, stopping just before the app region at 0x102000) ended it, and the radio booted once and stayed.

It also explained why page-side installs vanished: the emulator writes its in-memory image back on exit, and the looping guest's copy was older than the file, so powering it off overwrote the installs. With the loop gone the same installs survive a power cycle -- installed, powered off, still listed, powered on, still listed, and 0x0730 answered with all three. Both notes are in AGENTS.md and AGENTS.zh-CN.md; work/app-template/ got a starter app, its README and the upstream api/ld it needs (git-ignored).
2026-10-01 22:13:00 +08:00
mckero dcf9dc00f5 Keep what the page installed when the server restarts
The page edits a working copy of the flash image on purpose -- the file it was first pointed at may be a real calibration dump -- but a restart pointed at a different image switched to that file's copy instead, and the games installed through the page looked like they had vanished (measured: the app table came back holding an older slot). work/run-webui.ps1 now prefers the working copy the page has been editing, ahead of the original, and the page's own hint says which image it is editing and that the loaded one is never modified. A test asserts the order in the script.
2026-10-01 22:03:37 +08:00
mckero 673c181d59 Fold the slot and app tables away instead of one long page
Sixteen app rows (most of them empty or holding something that is not an app) plus five slot rows made the column long enough to be annoying, so each table is now a details pane: the firmware slots start open, the app list starts folded, and the state that used to sit in its own row moved into the summary line. Same ids, so loadSlots and loadApps are untouched. A front-end test asserts both panes exist and that the app list starts folded.
2026-10-01 22:00:54 +08:00