diff --git a/AGENTS.md b/AGENTS.md index b6d1bbb..c85ba4e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -828,6 +828,27 @@ entered. The pattern is the same family as the four flash bugs already in this f transfer is over while the guest is still feeding it, so it jumps early. The suspicion is the SPI/DMA busy and complete flags, and the check is to watch them during a large read rather than to reason about them. +**The size ladder was a red herring: a 3 KiB app that calls nothing -- and then print_tiny, get_key and +display_clear in turn -- runs fine. Minesweeper is dying on its own code.** + +The ladder lumped two variables together: every small test app happened to spin first, and every big one +called the firmware immediately. Separating them, all with the 2 ms PC probe on the page's own emulator: + +* a **3028-byte** app whose body is a volatile pad array and a spin loops in the overlay indefinitely (1704 + samples, all inside its first 512 bytes) -- so blob size is not a problem; +* adding a single **print_tiny** to it changes nothing (2039 samples, and 1024 font-region reads appear in the + flash log, the first at 0x1e0000) -- so the font path, and the sector-cache worry with it, is fine; +* adding **get_key** as well: still running (2029 samples); +* adding **display_clear** too: still running (1978 of 13195 samples, 15%, inside the app). + +So the three calls no surviving app had ever made are all harmless, and the surviving apps have covered the +API surface Minesweeper uses. Its failure is in its own code. The next step is therefore an ordinary bisect: +strip pieces out of its draw() until it survives, install each variant through the page and read the same +trace. That is a much better place to be than the emulator mystery this started as. + +One measurement note: the flash probe's per-read lines were also the way the font read was spotted (1024 of +them, none of which landed on the app), which is a useful control to keep running. + Workaround, for now and marked as one: a small app is a working app. The opening spin added to Minesweeper is kept out of the repository until the real cause is fixed, because it does not work anyway. diff --git a/AGENTS.zh-CN.md b/AGENTS.zh-CN.md index 312b576..25661d6 100644 --- a/AGENTS.zh-CN.md +++ b/AGENTS.zh-CN.md @@ -675,6 +675,26 @@ flash 探针是按顺序记录每一笔事务的,而那一次启动的**最后 bug 是同一族:**传输还在进行时就告诉加载器「已完成」**,于是它提前跳转。嫌疑点是 SPI/DMA 的忙/完成标志; 验证方式是**在一次大读取期间去观察这些标志**,而不是坐着推理。 +**「体积阶梯」是假的线索:一个 3 KiB、什么都不调用的应用能一直跑;再依次加上 print_tiny、get_key、** +**display_clear,它照样跑得很好。扫雷是死在它自己的代码上。** + +那个阶梯把两个变量混在了一起:每个小测试应用碰巧都先自旋,而每个大应用都一上来就调用固件。把两者分开来, +全部在页面自己的模拟器上用 2 ms 的 PC 探针测: + +* 一个 **3028 字节**的应用(内容只是一个 volatile 数组加一个自旋)在叠加区里无限循环(1704 个样本,全部落在 + 它的前 512 字节内)—— 所以**块的大小根本不是问题**; +* 给它加一次 **print_tiny** 毫无影响(2039 个样本,而且 flash 日志里出现 1024 次字体区读取,第一条在 0x1e0000) + —— 所以**字体路径没问题,扇区缓存那个担心也一并作废**; +* 再加上 **get_key**:仍在运行(2029 个样本); +* 再加上 **display_clear**:仍在运行(13195 个样本里有 1978 个、15% 在应用内)。 + +至此,**「从未有应用成功调用过」的那三个函数全部无害**,而活下来的应用已经覆盖了扫雷用到的整个 API 面。它的失败 +**在自己的代码里**。所以下一步就是一次普通的二分:把 draw() 里的东西一块块删掉,直到它活下来,每个变体都从页面装 +进去、读同一份 trace。比起一开始那个模拟器悬案,这是好得多的处境。 + +一条测量备注:flash 探针的逐条读取记录也正是发现字体读取的方式(1024 条,且没有一条落在应用身上),这个对照值得 +一直开着。 + 临时办法(且标明只是临时办法):**小应用就是能用的应用**。给扫雷加的那个开场自旋不会进仓库 —— 它本来也没起作用。 没有按键活动的紧循环。