用户要求内置完整版模型, 不要量化版。 assets/deepcw/model.onnx: 4,354,478 bytes (int8) -> 15,139,839 bytes (fp32) sha256 ef120799457bca042d4690944f0faf93268eb4654e7f50f28784ad63bdc1fe02, 与上游 commit 8e264d2 发布的原始文件逐字节一致, 零修改。 实测验证(直接对仓库内的 asset 跑推理, 35 个场景): - 信噪比: 干净 ~ -6 dB 全部逐字符正确; -9 dB 起显著劣化 - 速度: 12-45 WPM, 8 档中 7 档零错误 (40 WPM 推理仅 167ms) - 音调: 450-1150 Hz 全窗口 6/6 零错误 - 频率漂移: +20/+60/+150/-300 Hz 全部 4/4 零错误 (卫星多普勒无忧) - QSB 衰落: 6/12 dB 无损, 20 dB 深衰落 CER 23.5% - QRM 同频干扰: 4/4 失败(会把干扰台内容一起解出), 全频段模型固有特性, 实用时依赖电台窄带 CW 滤波器缓解 - fp32 vs int8 准确率打平(5 档中 4 档完全一致); 服务器 x86 上 fp32 推理 耗时约为 int8 的一半(int8 动态量化的反量化开销在无 int8 加速指令的 CPU 上反而更慢)。手机 ARM 侧表现待装机确认。 NOTICE.md / DEEPCW.md / README.md 同步更新: 移除 int8 量化派生的记录与复现 步骤, 改为声明未修改照搬上游。 代价: APK 体积约 59MB -> 70MB, 运行内存峰值上升。此前真机闪退的根因是 R8 缺 -keep ai.onnxruntime.** 规则(已修), 与模型大小无关。
4.8 KiB
DeepCW Decoder — module documentation
The feature:cw module decodes Morse code (CW) with the DeepCW neural model
instead of a classical DSP detector. This document records the architecture,
the model provenance, the pitfalls we hit on real devices, and the verification
results. It exists so the next person does not have to re-derive any of it.
Architecture
microphone (44.1 kHz PCM float)
│ AudioCapture.audioFlow() → ~100 ms chunks
▼
CwDecodeScreen (feature:cw, pure Compose)
│ decoder.processBuffer(chunk) waterfall.pushSamples(chunk)
▼ ▼
CwDeepDecoder (core:data) CwWaterfallState (feature:cw)
resample → 3200 Hz resample → 3200 Hz
CwDeepBuffer (20 s rolling) CwDeepSpectrogram.compute
CwDeepSpectrogram.compute → rolling spectrogram history
ONNX Runtime session.run → Canvas waterfall
CwCtcDecoder.greedy (400–1200 Hz band, same magnitudes
→ decodedText StateFlow the model sees)
Why the split. core:domain must stay pure Kotlin/JVM (KMP migration
headroom, per AGENTS.md), so the ONNX Runtime dependency — which ships native
libraries — lives in core:data behind the ICwDecoder interface. The
spectrogram front-end (CwDeepSpectrogram) and CTC collapse (CwCtcDecoder)
are pure Kotlin and live in core:domain, where they are unit-tested against
the Python reference with golden vectors.
Streaming model, not incremental. DeepCW is a whole-utterance CTC model. It
rewrites earlier characters as more context arrives, so incremental appending is
wrong. The decoder keeps a rolling buffer capped at 20 s and re-runs the whole
window every 1.5 s, replacing the displayed text outright. 20 s is the smallest
window that hits 0 % CER on the reference clip and the largest that keeps
inference comfortably faster than real time (measured ~14× headroom on a server
CPU; timed per-window on device and reported via lastInferenceMs).
Model
| Value | |
|---|---|
| File | assets/deepcw/model.onnx (fp32, as published upstream) |
| Size | 15,139,839 bytes |
| Modifications | none — vendored byte-for-byte |
| Input | spectrogram [1,1,T,65] float32 |
| Output | log_probs [1,T,42] float32 |
The full fp32 model ships inside the APK for maximum decode fidelity. An int8
quantize_dynamic build was trialled earlier (~4× smaller, measurably identical
at SNR ≥ −4 dB) but the shipped artifact is now the unmodified fp32 model. See
licenses/NOTICE.md for provenance, commit SHA and hash.
Audio must be packaged uncompressed (noCompress += "onnx" in the app
module): ONNX Runtime mmap's assets and refuses compressed ones.
Pitfalls fixed on real devices (release-only)
- R8 minification breaks the JNI binding.
isMinifyEnabled = truein the convention plugin, butonnxruntime-androidships no consumer ProGuard rules, soai.onnxruntime.*got renamed and native code crashed with no Java stack trace. Fixed with-keep class ai.onnxruntime.** { *; }inapp/proguard-rules.pro. Symptom was the exact "flash to home screen, empty crash log" we chased. - Model load in
init{}took down the whole composable on failure; moved to lazyensureLoaded()with anerrorMessagepath. - Default thread count saturated all cores and starved audio capture/UI;
capped
setIntraOpNumThreadstocores−1(max 4). - Compose snapshot state written from the audio thread crashes at runtime
with no log from business-class instrumentation; the waterfall now uses a
MutableStateFlowrevision counter, and both the row deque and the pending buffer are lock-guarded. LaunchedEffect { launch { collect() } }leaked a secondAudioRecordwhen toggling; collection now runs directly in the effect body keyed on(isListening, permissionGranted).
Diagnostics (no-adb): MainApplication already writes uncaught exceptions to
files/crash_log.txt. The decoder additionally persists load/inference failures
to files/deepcw_load_error.txt / files/deepcw_infer_error.txt and step
markers to files/probe_cw.txt (load_begin → load_session_ok → infer_begin → infer_done).
Verification
- Golden vectors:
CwDeepSpectrogramTest+CwDeepGoldenVectorTestpin the Kotlin spectrogram to the Python reference across 5005 values (tolerance 1e-4). 79 domain tests pass. - End-to-end on server: real 20 s noisy CW audio → resample → spectrogram →
ONNX → greedy CTC reproduces
CQ CQ DE BG7NTA BG7NTA K 5NN TU 73character-for-character (~1 s inference, fp32). - On device (release APK): decodes 40 WPM CW at high accuracy. Some UI latency is expected on the first decode cycle (model load + first inference); sustained decoding stays real-time.