用户要求内置完整版模型, 不要量化版。 assets/deepcw/model.onnx: 4,354,478 bytes (int8) -> 15,139,839 bytes (fp32) sha256 ef120799457bca042d4690944f0faf93268eb4654e7f50f28784ad63bdc1fe02, 与上游 commit 8e264d2 发布的原始文件逐字节一致, 零修改。 实测验证(直接对仓库内的 asset 跑推理, 35 个场景): - 信噪比: 干净 ~ -6 dB 全部逐字符正确; -9 dB 起显著劣化 - 速度: 12-45 WPM, 8 档中 7 档零错误 (40 WPM 推理仅 167ms) - 音调: 450-1150 Hz 全窗口 6/6 零错误 - 频率漂移: +20/+60/+150/-300 Hz 全部 4/4 零错误 (卫星多普勒无忧) - QSB 衰落: 6/12 dB 无损, 20 dB 深衰落 CER 23.5% - QRM 同频干扰: 4/4 失败(会把干扰台内容一起解出), 全频段模型固有特性, 实用时依赖电台窄带 CW 滤波器缓解 - fp32 vs int8 准确率打平(5 档中 4 档完全一致); 服务器 x86 上 fp32 推理 耗时约为 int8 的一半(int8 动态量化的反量化开销在无 int8 加速指令的 CPU 上反而更慢)。手机 ARM 侧表现待装机确认。 NOTICE.md / DEEPCW.md / README.md 同步更新: 移除 int8 量化派生的记录与复现 步骤, 改为声明未修改照搬上游。 代价: APK 体积约 59MB -> 70MB, 运行内存峰值上升。此前真机闪退的根因是 R8 缺 -keep ai.onnxruntime.** 规则(已修), 与模型大小无关。
96 lines
4.8 KiB
Markdown
96 lines
4.8 KiB
Markdown
# DeepCW Decoder — module documentation
|
||
|
||
The `feature:cw` module decodes Morse code (CW) with the DeepCW neural model
|
||
instead of a classical DSP detector. This document records the architecture,
|
||
the model provenance, the pitfalls we hit on real devices, and the verification
|
||
results. It exists so the next person does not have to re-derive any of it.
|
||
|
||
## Architecture
|
||
|
||
```
|
||
microphone (44.1 kHz PCM float)
|
||
│ AudioCapture.audioFlow() → ~100 ms chunks
|
||
▼
|
||
CwDecodeScreen (feature:cw, pure Compose)
|
||
│ decoder.processBuffer(chunk) waterfall.pushSamples(chunk)
|
||
▼ ▼
|
||
CwDeepDecoder (core:data) CwWaterfallState (feature:cw)
|
||
resample → 3200 Hz resample → 3200 Hz
|
||
CwDeepBuffer (20 s rolling) CwDeepSpectrogram.compute
|
||
CwDeepSpectrogram.compute → rolling spectrogram history
|
||
ONNX Runtime session.run → Canvas waterfall
|
||
CwCtcDecoder.greedy (400–1200 Hz band, same magnitudes
|
||
→ decodedText StateFlow the model sees)
|
||
```
|
||
|
||
**Why the split.** `core:domain` must stay pure Kotlin/JVM (KMP migration
|
||
headroom, per `AGENTS.md`), so the ONNX Runtime dependency — which ships native
|
||
libraries — lives in `core:data` behind the `ICwDecoder` interface. The
|
||
spectrogram front-end (`CwDeepSpectrogram`) and CTC collapse (`CwCtcDecoder`)
|
||
are pure Kotlin and live in `core:domain`, where they are unit-tested against
|
||
the Python reference with golden vectors.
|
||
|
||
**Streaming model, not incremental.** DeepCW is a whole-utterance CTC model. It
|
||
rewrites earlier characters as more context arrives, so incremental appending is
|
||
wrong. The decoder keeps a rolling buffer capped at 20 s and re-runs the whole
|
||
window every 1.5 s, replacing the displayed text outright. 20 s is the smallest
|
||
window that hits 0 % CER on the reference clip *and* the largest that keeps
|
||
inference comfortably faster than real time (measured ~14× headroom on a server
|
||
CPU; timed per-window on device and reported via `lastInferenceMs`).
|
||
|
||
## Model
|
||
|
||
| | Value |
|
||
|---|---|
|
||
| File | `assets/deepcw/model.onnx` (fp32, as published upstream) |
|
||
| Size | 15,139,839 bytes |
|
||
| Modifications | none — vendored byte-for-byte |
|
||
| Input | `spectrogram` [1,1,T,65] float32 |
|
||
| Output | `log_probs` [1,T,42] float32 |
|
||
|
||
The full fp32 model ships inside the APK for maximum decode fidelity. An int8
|
||
`quantize_dynamic` build was trialled earlier (~4× smaller, measurably identical
|
||
at SNR ≥ −4 dB) but the shipped artifact is now the unmodified fp32 model. See
|
||
[`licenses/NOTICE.md`](licenses/NOTICE.md) for provenance, commit SHA and hash.
|
||
|
||
Audio must be packaged **uncompressed** (`noCompress += "onnx"` in the app
|
||
module): ONNX Runtime mmap's assets and refuses compressed ones.
|
||
|
||
## Pitfalls fixed on real devices (release-only)
|
||
|
||
1. **R8 minification breaks the JNI binding.** `isMinifyEnabled = true` in the
|
||
convention plugin, but `onnxruntime-android` ships **no** consumer ProGuard
|
||
rules, so `ai.onnxruntime.*` got renamed and native code crashed with no Java
|
||
stack trace. Fixed with
|
||
`-keep class ai.onnxruntime.** { *; }` in `app/proguard-rules.pro`.
|
||
Symptom was the exact "flash to home screen, empty crash log" we chased.
|
||
2. **Model load in `init{}`** took down the whole composable on failure; moved
|
||
to lazy `ensureLoaded()` with an `errorMessage` path.
|
||
3. **Default thread count** saturated all cores and starved audio capture/UI;
|
||
capped `setIntraOpNumThreads` to `cores−1` (max 4).
|
||
4. **Compose snapshot state written from the audio thread** crashes at runtime
|
||
with no log from business-class instrumentation; the waterfall now uses a
|
||
`MutableStateFlow` revision counter, and both the row deque and the pending
|
||
buffer are lock-guarded.
|
||
5. **`LaunchedEffect { launch { collect() } }`** leaked a second `AudioRecord`
|
||
when toggling; collection now runs directly in the effect body keyed on
|
||
`(isListening, permissionGranted)`.
|
||
|
||
Diagnostics (no-adb): `MainApplication` already writes uncaught exceptions to
|
||
`files/crash_log.txt`. The decoder additionally persists load/inference failures
|
||
to `files/deepcw_load_error.txt` / `files/deepcw_infer_error.txt` and step
|
||
markers to `files/probe_cw.txt` (`load_begin → load_session_ok → infer_begin →
|
||
infer_done`).
|
||
|
||
## Verification
|
||
|
||
- **Golden vectors**: `CwDeepSpectrogramTest` + `CwDeepGoldenVectorTest` pin the
|
||
Kotlin spectrogram to the Python reference across 5005 values (tolerance
|
||
1e-4). 79 domain tests pass.
|
||
- **End-to-end on server**: real 20 s noisy CW audio → resample → spectrogram →
|
||
ONNX → greedy CTC reproduces `CQ CQ DE BG7NTA BG7NTA K 5NN TU 73`
|
||
character-for-character (~1 s inference, fp32).
|
||
- **On device (release APK)**: decodes 40 WPM CW at high accuracy. Some UI
|
||
latency is expected on the first decode cycle (model load + first inference);
|
||
sustained decoding stays real-time.
|