feat(cw): optionally shift out-of-window CW tones into the model's range

DeepCW only analyses 400-1200 Hz - its input tensor is 65 bins wide, fixed at
training time - so a CW note outside that range is invisible to the decoder.
This adds an opt-in preprocessing step that moves such a tone to 800 Hz, the
window centre, extending the usable pitch range without touching the model.

Single-sideband mixing via a 63-tap Hilbert transformer. Plain real mixing was
measured and rejected: shifting 1500 Hz to 800 Hz left a fold-back image at
1000 Hz at 0.999 of the wanted amplitude, inside the window. Zero-stuff
upsampling plus lowpass handled downward shifts but left a 0.996 image when
shifting 300 Hz upward. The Hilbert approach measures clean on nine tones from
150 to 1550 Hz: one peak at the target, nothing above 0.3 relative amplitude.
In-window energy for a 1500 Hz input goes from 6.8% to 94.6%.

Only out-of-range audio is processed. A tone already inside 400-1200 Hz is
returned untouched (same array instance, no copy), and with the setting off the
audio path is exactly what it was before.

CwToneShifter.Streaming carries the Hilbert filter history and mixer phase
across capture chunks. Shifting each chunk in isolation left 62 of every 320
samples convolving against zeros, inflating envelope ripple to 8.7x the
whole-buffer baseline. A residual difference in the last ~3 samples of each
chunk is causal and documented: those output samples would need input that has
not been captured yet.

Detection pools chunks rather than gating on one. A capture chunk is 4410
samples at 44.1 kHz but only 320 after resampling to 3200 Hz, so requiring
1280 samples in a single chunk would have made the feature dead code - the two
independent audits both found this before it shipped. Detection now runs on a
pooled 0.4 s window, at most every 2 s.

Toggling the setting or a change in the detected shift drops the buffered
audio: the 20 s window would otherwise keep decoding samples moved by the old
amount, and the pitch readout could only be correct for one of them. The
readout itself subtracts the active shift so it shows the pitch on the radio,
not the shifted one.

Settings: OtherSettings.cwToneShiftEnabled, off by default, persisted and read
back in SettingsRepo, toggled from the Other card in Settings with a help line
explaining the 400-1200 Hz limit. Strings added to all nine locales. The
decoder reads the flag per chunk, so the toggle applies without restarting
capture.

Debug: the enabled-state transition, each detection verdict (no tone / inside
window / shifting by N Hz), and every shift change are logged, with the noisy
paths throttled to the 2 s detection interval. CwProbe records shift changes
only, keeping well inside its 1 MiB cap.

Tests: 8 shifter tests (detection sweep, noise rejection, pass-through
identity, image-free shifting across 8 tones, end-to-end spectrogram energy),
8 streaming tests (chunk continuity, history retention, reset semantics, chunk
sizes above and below the history window), and 6 gate tests including a
regression guard that a 320-sample chunk must be able to reach the detection
threshold. All 53 CW tests pass, golden vectors included.
This commit is contained in:
mckero committed 2026-08-22 01:07:38 +00:00
1 parent e3d7238721
commit 9798107d37
21 files changed
+1040 -9

No files matched your search

@@ -25,6 +25,7 @@ import android.util.Log
import com.rtbishop.look4sat.core.domain.cw.CwCtcDecoder
import com.rtbishop.look4sat.core.domain.cw.CwDeepBuffer
import com.rtbishop.look4sat.core.domain.cw.CwDeepSpectrogram
import com.rtbishop.look4sat.core.domain.cw.CwToneShifter
import com.rtbishop.look4sat.core.domain.cw.ICwDecoder
import kotlinx.coroutines.CancellationException
import kotlinx.coroutines.Dispatchers
@@ -48,9 +49,18 @@ import java.nio.FloatBuffer
* window is re-decoded every 1.5 seconds, replacing [decodedText] outright.
*
* The model's fixed 400-1200 Hz analysis window means pitch detection is built
* in; no spectral peak tracking or squelch gating is needed.
* in; no spectral peak tracking or squelch gating is needed. A tone outside that
* window is invisible to the model, so [CwToneShifter] can optionally move it in —
* see [isToneShiftEnabled].
*
* @param isToneShiftEnabled read on every chunk so toggling the setting takes effect
* without rebuilding the decoder. Defaults to disabled: with it off the audio path
* is byte-for-byte what it was before the feature existed.
*/
class CwDeepDecoder(context: Context) : ICwDecoder {
class CwDeepDecoder(
context: Context,
private val isToneShiftEnabled: () -> Boolean = { false }
) : ICwDecoder {
private companion object {
const val TAG = "CwDeepDecoder"
@@ -60,6 +70,20 @@ class CwDeepDecoder(context: Context) : ICwDecoder {
/** Evicted audio is decoded into permanent history once this much accumulates. */
const val ARCHIVE_SECONDS = 15.0
val ARCHIVE_THRESHOLD: Int = (CwDeepSpectrogram.SAMPLE_RATE * ARCHIVE_SECONDS).toInt()
/**
* Samples the detector needs for a usable estimate: 0.4 s at 3200 Hz, giving
* ~12.5 Hz resolution.
*
* A capture chunk is ~100 ms, which is 4410 samples at the 44.1 kHz capture
* rate but only 320 after resampling to 3200 Hz. Gating on a single chunk
* reaching this size would therefore never fire, so chunks are accumulated in
* [detectBuffer] until enough audio is available.
*/
const val DETECT_MIN_SAMPLES = 1280
/** Detection cadence; re-running it on every 100 ms chunk would be wasteful. */
const val DETECT_INTERVAL_MS = 2000
}
private val _decodedText = MutableStateFlow("")
@@ -95,6 +119,29 @@ class CwDeepDecoder(context: Context) : ICwDecoder {
/** Held while inference runs so slow devices skip work instead of queuing it. */
private val inferenceLock = Mutex()
/** Shift currently applied to incoming audio; 0 when the tone needs no move. */
private var activeShiftHz = 0f
/** Last detected tone, for logging and for the pitch readout while shifting. */
private var detectedToneHz: Float? = null
/** Wall clock of the last detection scan, throttling it to [DETECT_INTERVAL_MS]. */
private var lastDetectAtMs = 0L
/**
* Accumulates resampled chunks until [DETECT_MIN_SAMPLES] is reached. A single
* capture chunk is only 320 samples once resampled, so detection has to pool
* several of them.
*/
private val detectBuffer = FloatArray(DETECT_MIN_SAMPLES)
private var detectFill = 0
/** Carries Hilbert filter history and mixer phase across capture chunks. */
private val streamingShifter = CwToneShifter.Streaming()
/** Previous value of the setting, so a toggle can invalidate buffered audio. */
private var toneShiftWasEnabled = false
private var environment: OrtEnvironment? = null
private var session: OrtSession? = null
private var chars: List<String> = emptyList()
@@ -174,7 +221,8 @@ class CwDeepDecoder(context: Context) : ICwDecoder {
val resampled = CwDeepSpectrogram.resampleLinear(
samples, sampleRate, CwDeepSpectrogram.SAMPLE_RATE
)
val shouldRedecode = buffer.append(resampled)
val prepared = applyToneShift(resampled)
val shouldRedecode = buffer.append(prepared)
// Archive audio that scrolled out of the live window. It is decoded once
// when a full archive chunk has accumulated, so old text does not vanish.
@@ -237,6 +285,103 @@ class CwDeepDecoder(context: Context) : ICwDecoder {
}
}
/**
* Move an out-of-window tone into the model's analysis window when the user has
* enabled it.
*
* The detection scan is a bin-by-bin DFT, so it runs at most every
* [DETECT_INTERVAL_MS] rather than on every ~100 ms capture chunk; the decision it
* produces is cached in [activeShiftHz] and applied to the chunks in between. A
* tone already inside the window yields a zero shift, and then this returns the
* caller's array untouched.
*
* @return the audio to buffer: [resampled] itself whenever no shift applies.
*/
private fun applyToneShift(resampled: FloatArray): FloatArray {
if (!isToneShiftEnabled()) {
// Clear stale state so re-enabling starts from a fresh detection.
if (activeShiftHz != 0f || detectedToneHz != null || detectFill > 0) {
Log.i(TAG, "toneShift: disabled, clearing shift=${activeShiftHz}Hz")
activeShiftHz = 0f
detectedToneHz = null
lastDetectAtMs = 0L
detectFill = 0
streamingShifter.reset()
}
return resampled
}
accumulateForDetection(resampled)
val now = System.currentTimeMillis()
val elapsed = now - lastDetectAtMs
if (detectFill >= DETECT_MIN_SAMPLES && elapsed >= DETECT_INTERVAL_MS) {
lastDetectAtMs = now
val sample = detectBuffer.copyOf(detectFill)
detectFill = 0
runDetection(sample)
}
// Streaming keeps the Hilbert filter history and mixer phase across chunks;
// shifting each chunk in isolation distorted the 62 samples at its edges.
return streamingShifter.process(resampled, activeShiftHz, CwDeepSpectrogram.SAMPLE_RATE)
}
/** Collect resampled chunks until [DETECT_MIN_SAMPLES] is available. */
private fun accumulateForDetection(chunk: FloatArray) {
if (chunk.isEmpty()) return
// A chunk larger than the detection buffer only needs to contribute its tail.
val start = maxOf(0, chunk.size - detectBuffer.size)
for (i in start until chunk.size) {
if (detectFill == detectBuffer.size) {
// Slide the window so detection always sees the most recent audio.
detectBuffer.copyInto(detectBuffer, 0, 1, detectFill)
detectFill--
}
detectBuffer[detectFill++] = chunk[i]
}
}
/** Update [activeShiftHz] from a detection pass and log what was decided. */
private fun runDetection(sample: FloatArray) {
val analysis = CwToneShifter.analyse(sample, CwDeepSpectrogram.SAMPLE_RATE)
val previousShift = activeShiftHz
detectedToneHz = analysis.toneHz
activeShiftHz = analysis.shiftHz
when {
analysis.toneHz == null ->
Log.d(TAG, "toneShift: no tone in ${sample.size} samples, shift stays 0")
!analysis.needsShift -> Log.d(
TAG,
"toneShift: tone=${analysis.toneHz}Hz inside " +
"${CwDeepSpectrogram.MIN_FREQ_HZ}-${CwDeepSpectrogram.MAX_FREQ_HZ}Hz, no shift"
)
else -> Log.i(
TAG,
"toneShift: tone=${analysis.toneHz}Hz outside window, " +
"shifting ${analysis.shiftHz}Hz to ${CwToneShifter.TARGET_HZ}Hz"
)
}
if (previousShift != activeShiftHz) {
// The window still holds audio moved by the old amount. Mixing two shifts in
// one spectrogram smears the tone, and the pitch readout could only be right
// for one of them, so rebuild the window from the new shift.
Log.i(
TAG,
"toneShift: shift changed ${previousShift}Hz -> ${activeShiftHz}Hz, " +
"dropping buffered audio"
)
buffer.reset()
archiveSize = 0
streamingShifter.reset()
CwProbe.step("tone_shift tone=${analysis.toneHz} shift=$activeShiftHz")
}
}
private suspend fun decodeWindow(window: FloatArray) = withContext(Dispatchers.Default) {
val activeSession = session ?: return@withContext
val activeEnvironment = environment ?: return@withContext
@@ -322,7 +467,9 @@ class CwDeepDecoder(context: Context) : ICwDecoder {
val binHz = CwDeepSpectrogram.SAMPLE_RATE.toDouble() / CwDeepSpectrogram.FFT_LENGTH
// Relative bin 0 is 400 Hz; absolute bin index is 32 + bestBin.
val absoluteBin = 32 + bestBin
_estimatedPitch.value = (absoluteBin * binHz).toFloat()
// Undo the shift before reporting: the spectrogram sees the moved tone, but
// the readout must show the pitch the operator actually hears on the radio.
_estimatedPitch.value = (absoluteBin * binHz - activeShiftHz).toFloat()
val mean = total / count
_signalStrength.value = ((bestValue - mean) / bestValue).coerceIn(0f, 1f)
@@ -336,6 +483,12 @@ class CwDeepDecoder(context: Context) : ICwDecoder {
_estimatedPitch.value = null
_signalStrength.value = 0f
_lastInferenceMs.value = 0
// Re-detect from scratch: the operator may have retuned before resetting.
activeShiftHz = 0f
detectedToneHz = null
lastDetectAtMs = 0L
detectFill = 0
streamingShifter.reset()
}
override fun close() {
@@ -111,7 +111,10 @@ class MainContainer(private val context: Context) : IMainContainer {
// 每次调用返回新实例: 调用方负责 close() 释放 OrtSession, 且 Radar 内嵌
// 面板与独立 CW 页各自持有自己的解码器
override fun provideCwDecoder(): com.rtbishop.look4sat.core.domain.cw.ICwDecoder =
com.rtbishop.look4sat.core.data.cw.CwDeepDecoder(context)
com.rtbishop.look4sat.core.data.cw.CwDeepDecoder(context) {
// Read per chunk so toggling the setting applies without restarting capture.
settingsRepo.otherSettings.value.cwToneShiftEnabled
}
override fun provideSaveImage(): ISaveImage = SaveImage(context)
@@ -105,6 +105,7 @@ class SettingsRepo(
private val keyWavelogAutoUpload = "wavelogAutoUpload"
private val keyRadarCompassOffset = "radarCompassOffset"
private val keyRadarCompassOffsetElev = "radarCompassOffsetElev"
private val keyCwToneShiftEnabled = "cwToneShiftEnabled"
private val separatorComma = ","
@@ -405,6 +406,7 @@ class SettingsRepo(
putBoolean(keyWavelogAutoUpload, new.wavelogAutoUpload)
putFloat(keyRadarCompassOffset, new.radarCompassOffset)
putFloat(keyRadarCompassOffsetElev, new.radarCompassOffsetElev)
putBoolean(keyCwToneShiftEnabled, new.cwToneShiftEnabled)
}
new
@@ -431,7 +433,8 @@ class SettingsRepo(
wavelogStationId = preferences.getString(keyWavelogStationId, null) ?: "",
wavelogAutoUpload = preferences.getBoolean(keyWavelogAutoUpload, false),
radarCompassOffset = preferences.getFloat(keyRadarCompassOffset, 0f),
radarCompassOffsetElev = preferences.getFloat(keyRadarCompassOffsetElev, 0f)
radarCompassOffsetElev = preferences.getFloat(keyRadarCompassOffsetElev, 0f),
cwToneShiftEnabled = preferences.getBoolean(keyCwToneShiftEnabled, false)
)
//endregion
@@ -0,0 +1,128 @@
package com.rtbishop.look4sat.core.data.cw
import com.rtbishop.look4sat.core.domain.cw.CwDeepSpectrogram
import com.rtbishop.look4sat.core.domain.cw.CwToneShifter
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertSame
import org.junit.Assert.assertTrue
import org.junit.Test
import kotlin.math.PI
import kotlin.math.sin
/**
* The gating contract the decoder relies on: shift only when the user opted in AND the
* tone is outside the model window.
*
* [CwDeepDecoder] needs a Context and a loaded ONNX model, so it cannot be constructed
* here. What these tests do exercise is the real decision function the decoder calls -
* [CwToneShifter.analyse] - rather than a copy of it, so a wrong verdict fails here.
* The decoder's own sample accumulation and throttling are covered by the streaming
* tests in core:domain.
*/
class CwToneShiftGateTest {
private val sampleRate = CwDeepSpectrogram.SAMPLE_RATE
private fun tone(hz: Double, samples: Int = 1600): FloatArray = FloatArray(samples) { i ->
sin(2.0 * PI * hz * i / sampleRate).toFloat()
}
/**
* The enabled/disabled gate as [CwDeepDecoder.applyToneShift] applies it: when off
* the audio is returned as-is, when on the verdict comes from the real analyser.
*/
private fun gate(audio: FloatArray, enabled: Boolean): FloatArray {
if (!enabled) return audio
val analysis = CwToneShifter.analyse(audio, sampleRate)
if (!analysis.needsShift) return audio
return CwToneShifter.shift(audio, analysis.shiftHz, sampleRate)
}
@Test
fun `disabled leaves every tone untouched`() {
for (hz in listOf(150.0, 300.0, 800.0, 1200.0, 1500.0)) {
val audio = tone(hz)
assertSame(
"$hz Hz must pass through unchanged while the setting is off",
audio, gate(audio, enabled = false)
)
}
}
@Test
fun `enabled still leaves in-window tones untouched`() {
for (hz in listOf(400.0, 600.0, 800.0, 1000.0, 1200.0)) {
val audio = tone(hz)
assertSame(
"$hz Hz is inside the window; enabling the setting must not alter it",
audio, gate(audio, enabled = true)
)
}
}
@Test
fun `enabled shifts only out-of-window tones`() {
for (hz in listOf(200.0, 300.0, 1300.0, 1500.0)) {
val audio = tone(hz)
val result = gate(audio, enabled = true)
assertFalse("$hz Hz should have been shifted", result === audio)
assertEquals("shift must preserve length", audio.size, result.size)
}
}
@Test
fun `window edges count as inside`() {
val analysisLow = CwToneShifter.analyse(tone(CwDeepSpectrogram.MIN_FREQ_HZ), sampleRate)
val analysisHigh = CwToneShifter.analyse(tone(CwDeepSpectrogram.MAX_FREQ_HZ), sampleRate)
assertFalse("400 Hz is the lower edge, inside", analysisLow.needsShift)
assertFalse("1200 Hz is the upper edge, inside", analysisHigh.needsShift)
}
@Test
fun `shift target is inside the window`() {
assertTrue(
"the target must be a pitch the model can see",
CwToneShifter.isInsideWindow(CwToneShifter.TARGET_HZ.toFloat())
)
}
/**
* Regression guard for the defect that made the whole feature dead on arrival:
* the decoder gated detection on a single chunk reaching DETECT_MIN_SAMPLES, but
* AudioCapture delivers 4410 samples at 44.1 kHz, which is only 320 after
* resampling to 3200 Hz. Detection could never run.
*
* The decoder now pools chunks, so what matters is that the pooled size is
* reachable: a handful of real-sized chunks must add up to enough audio.
*/
@Test
fun `pooled capture chunks reach the detection threshold`() {
val captureRate = 44100
val captureChunk = captureRate / 10 // AudioCapture's ~100 ms read
val resampledChunk = captureChunk * CwDeepSpectrogram.SAMPLE_RATE / captureRate
assertEquals(
"a capture chunk resamples to 320 samples; if this changes revisit pooling",
320, resampledChunk
)
val threshold = 1280 // CwDeepDecoder.DETECT_MIN_SAMPLES
val chunksNeeded = (threshold + resampledChunk - 1) / resampledChunk
assertTrue(
"a single chunk ($resampledChunk) must not be expected to reach $threshold",
resampledChunk < threshold
)
assertTrue(
"pooling must reach the threshold within a second of audio, needs $chunksNeeded chunks",
chunksNeeded in 2..10
)
// And that much audio must actually be enough for the detector to work.
val pooled = tone(1500.0, samples = threshold)
val detected = CwToneShifter.detectToneHz(pooled, sampleRate)
assertEquals(
"the pooled window must be long enough to detect a tone",
1500.0, detected!!.toDouble(), 25.0
)
}
}