AprsReporter derived a passcode from the callsign whenever the operator's entry was
unusable:
passcode = cfg.passcode.toIntOrNull()?.takeIf { it >= 0 } ?: AprsPacket.passcode(cfg.callsign)
Measured against the shipped algorithm for BG7NTA, whose passcode is 21162, four
inputs produced a transmit passcode the operator never obtained: blank, whitespace,
non-numeric, and an explicit -1. That last one is the documented receive-only value,
so `takeIf { it >= 0 }` also made receive-only unreachable - and a receive-only login
is the one way to confirm a setup works without putting anything on the network,
which is exactly how this feature was supposed to be validated before release.
This is a policy question more than a bug. APRS-IS states that supplying the correct
passcode to a user is the software author's responsibility, and the passcode functions
as a licence check for transmitting. APRSdroid carries the same algorithm in the same
source file and deliberately does not use it to fill a blank, validating the operator's
entry instead. AprsPasscode follows that: it classifies an entry as Transmit,
ReceiveOnly, Mismatch or NotANumber, and anything not usable logs in as -1. The
connection still works and the operator is told separately that reports are not being
forwarded, but no packet goes out under a code the app made up.
AprsPacket.passcode stays, because validating an entry means recomputing the expected
value. Nothing substitutes it for a missing one.
Separately, and worse than the line above: the report notices never appeared at all.
onReport is invoked from AprsReporter's Dispatchers.IO scope, where constructing a
Toast throws because the thread has no Looper - and the surrounding runCatching
swallowed it. So the whole reporting path, including the unverified-login warning
added in the previous commit, was writing messages nobody could see. They now post to
the main looper. The one in startReporting is left alone: onStartCommand already runs
on the main thread.
Still outstanding for APRS, and not addressed here: the 0N 0E position fallback, the
missing TCPIP* path, the unbounded status field, the coroutine delay that does not fire
in Doze, and the dataSync foreground service type. Also unaddressed is the "Compute
passcode" button in AprsCard, which offers the operator the derived value directly and
so contradicts the policy this commit establishes - it was added by request, so it needs
a decision rather than a quiet removal.
Two defects made every APRS failure invisible. sendPacket ended with
`.getOrElse { Pair(true, "OK") }`, so a read that threw - including on a dead
socket - was reported as a successful send. And the login check threw inside a
runCatching whose result was discarded, so a server that refused to verify the
passcode could not propagate: aprsc keeps such a client connected and its writes
succeed while silently discarding every packet, which the app reported as success.
An audit put it plainly - twelve commits are all fix(aprs), none added a
socket-level test, so "it worked" was never evidence a packet had landed.
sendPacket now separates the cases. A read timeout stays a success, because
APRS-IS does not acknowledge position reports and silence is the normal outcome.
A closed stream or an IOException is a failure. The catch order matters and is
load-bearing: SocketTimeoutException extends IOException, so reversing them would
mark every normal report as failed.
The login handshake follows the spec: read the server's identification line first,
then log in, then read until a verdict arrives. AprsLogin holds that as pure logic
in core:domain with the parsing that decides it, including one trap worth naming -
"unverified" contains "verified", so the negative has to be tested first or every
refusal reads as acceptance. An explicit refusal now fails the report and shows
the operator its own message pointing at the callsign and passcode, in five
locales. A response we could not parse does not, since the packets may well be
landing and blaming the passcode would send them to fix something that works.
Three defects came out of review after that. The verdict is now bounded by a
deadline rather than a five-line budget, because a server that sent six keepalives
before its answer turned an accepted login into Unknown - telling the operator
their passcode was wrong when it had just been accepted. The greeting gets a short
two-second probe instead of the full login window, which cost eight seconds on
every connect to a server that sends none. And a refusal detected in the greeting
now aborts the connection instead of being overwritten by the next read, which had
made that branch and its comment a lie.
The worst of the three was mine: rewriting the write as print + flush dropped the
line terminator entirely. APRS-IS is a line protocol, so the server's reader never
saw a packet, while the send reported success and the read timed out into the
"silence is normal" branch. It broke healthy connections rather than dead ones and
was designed to have no symptom. Both the packet and the login line now end in an
explicit CRLF as the spec requires, rather than println's platform separator.
That defect is why this adds AprsIsClientSocketTest, which runs the client against
a stand-in server and reads the bytes back: it asserts two packets arrive as two
lines, that a login line arrives complete, that keepalive chatter does not bury the
verdict, that a greeting-less server connects promptly, and that a send to a closed
peer reports failure. Nothing in the pure-logic tests could have caught a missing
newline. Note for anyone extending it: closing the ServerSocket leaves an
established connection alive, so the dead-peer test has to close the accepted
socket - assuming otherwise made a correct implementation look broken.
AprsPacket.formatLogin is deleted, its work moved into AprsLogin.line, which also
replaces spaces in the version string because the server splits that field on
whitespace and the shipped value contained one.
The grid lookup returned String? and swallowed everything with catch { null }, so a
timeout, an expired cookie, a QRZ layout change and a station that simply has not
published a locator were one indistinguishable blank. The operator saw an empty grid
with no way to know that re-pasting their cookie would fix it. There was also no
retry at all, on a phone, mid-pass, on mobile data.
QrzGrid names the four outcomes and QrzGridParser holds the parsing, which is pure
string work and now testable without a network. The fetch moves to core:data as
QrzGridSource, using the project's own OkHttp client with three attempts and 700ms
then 2000ms of backoff. Only transport failures and 5xx are retried; a 4xx would
repeat identically. This also gets java.net.URL I/O out of core:domain, which that
module is meant to stay clear of for the KMP move.
Classifying signed-out took two goes. Keying on the detail table being absent held
for an expired cookie - QRZ genuinely serves no detail rows to an anonymous visitor,
verified against a live response - but an audit found that a callsign QRZ has never
heard of returns HTTP 200 with no detail rows either, because QRZ serves its search
form instead. That would have reported a mistyped callsign as an expired cookie and
sent the operator into settings mid-pass to re-paste one that was never broken. It
now keys on QRZ's own "Login is required for additional detail" notice, so an absent
locator degrades to the harmless outcome and only QRZ actually asking for a login
triggers the cookie prompt. All three cases are measured against live responses.
Not yet wired in: LogTab and SettingsScreen still call the old QrzGridClient, so
nothing changes for the operator yet. Cutting over needs an interface in core:domain
and a MainContainer provider, because feature modules cannot reach core:data
directly - and the cookie itself belongs in SettingsRepo rather than the separate
prefs file a composable currently reads through LocalContext.
The history box appeared to delete text. Audio leaving the 20 s live window is
decoded into the archive only once a full 15 s batch has accumulated, so until
then its characters were in neither place: not in the live decode, which had
scrolled past them, and not in the history, which had not seen them yet.
Measured on a 20 WPM timeline, the concatenated transcript held at 40 characters
from t=24 s to t=34.5 s - eleven seconds of no growth - then jumped to 70 when the
batch flushed. Up to 30 characters sat in that gap. Reading it as deletion is
reasonable; the text really was missing from the box.
The pending batch is now decoded too, on the same 1.5 s cycle as the live window,
and shown as a provisional tail after the committed text. The final archive decode
replaces it, having the whole batch for context. The transcript is monotonic
afterwards: +3 characters every cycle with no stalls.
Decoding each 100 ms capture chunk instead would have removed the gap entirely but
measured 14x the inference load - over 250% of one core across ten minutes - and a
chunk that short carries under two dot-lengths of context, so the decode would be
poor as well as expensive. One extra inference per redecode cycle costs 24.7%
against 18.3%.
Discarding buffered audio drops the provisional text with it, since that text
describes audio that no longer exists. Committed text stays: it was correct for
audio that really was archived.
The waterfall showed only the model's 400-1200 Hz window, so a tone outside it
was absent from the picture entirely. Measured on keyed audio, the brightest
column in that narrow view swings 1.01x between key-down and key-up against
13.76x for a tone in range - it carries no keying at all, so the operator could
not tell a signal was present, let alone where it was. Markers alone could not
fix that: they pointed at a frequency with nothing drawn there.
compute() now takes an optional bin range, defaulting to the model's own, so the
decoder path is byte-identical and the golden-vector test still holds. The
display asks for DC to Nyquist, 129 bins against 65. The FFT already computed
every bin - this only changes which are kept - so the cost is a wider copy.
The decoder window is framed and faintly lifted, since half the picture is now
outside what the model reads and nothing said which half.
Marker fixes found while reviewing the render: the tone marker was orange, which
is a colour the inferno ramp itself passes through, so a marker sitting on the
trace it pointed at was indistinguishable from the keying gaps in that trace -
invisible in exactly the case it existed for. It is cyan now, and both markers
are pips in a gutter above the spectrum rather than lines across it.
Also from the release audit:
- compute()'s bin-count guard was written as a three-term disjunction, which any
custom range satisfies regardless of bin count, leaving the model invariant
unenforced for the caller most able to break it. Rewritten as an implication,
with a Nyquist bound so no range can index past the FFT output.
- signalStrength was gated on a confirmed out-of-window tone, which is false when
detection fails - and it fails for a slow fist, measured at prominence 2.5
against a 4.5 threshold for 15% duty. So the meter still read half scale beside
an empty transcript. It now requires a tone confirmed decodable: 11 flow
combinations, 3 wrong before, 0 wrong after.
- detectedToneHz never expired, so after retuning into the band the hint kept
naming the frequency the operator had left, indefinitely. It now clears after
10 s without a tone, which is clear of any real gap - the longest being 1.7 s
between words at 5 WPM.
- The waterfall label read estimatedPitch while the hint read detectedToneHz, two
numbers up to 800 Hz apart both claiming to be the tone. Both read the latter.
- Removed a redundant toFloat() that the compiler warned about.
Accessibility, untouched until now: the waterfall was a bare Canvas and the AMSAT
day cells bare Boxes, so both announced nothing at all - on the status page that
is the entire content of the screen. Both now carry a contentDescription naming
the tone or the day's worst status and report count. The AMSAT tap target goes
from 28 dp to 48 dp with the coloured tile still 28 dp, so the grid keeps its
density. Strings in all nine locales for both modules.
Opinion split on the stripes, so Settings > Other now has a switch. On by
default, since the flat tile it replaced hid intra-day outages, which is the
problem the stripes were introduced to solve.
Flat mode is deliberately not the old behaviour. The old cell took its colour
from the first slot with a report and its count from that same slot, so a day
that worked in the morning and failed all afternoon read as "worked" - measured
across eight representative day shapes, two of them had their failure hidden
outright, and the count reported 1 where the day held 24 reports. Flat mode now
takes the day's worst status and the day's total count, so the summary can
understate detail but not hide bad news. The help text says so, in case someone
turns the switch off expecting the tile they remember.
The count is drawn in black or white by relative luminance rather than always
white: on the telemetry amber, white measured 1.83:1 against WCAG's 3:1 for
large text, and that cell does carry a count whenever a day held nothing but
telemetry reports. All six status colours now clear 3:1, the worst being 3.03.
SatStatusViewModel collects the setting rather than reading it once - the switch
is on another screen, so the operator is always elsewhere when they change it
and would otherwise return to the old style.
Strings in all nine locales.
With tone shift off and the operator tuned outside 400-1200 Hz, the page did not
go quiet - it went confidently wrong. Three measurements, all reproduced against
the real spectrogram path:
estimatedPitch is (32 + loudestBin) * 12.5 - shiftHz with the bin confined to
0..64, so with no shift applied it can only ever report 400-1200 Hz. It cannot
express 1500 Hz, and it does not try: it publishes whichever window edge the
leakage piles against. For a 1500 Hz tone that is 1200 Hz.
That leakage is not faint. The waterfall normalises to the loudest value on
screen, so 50 of 65 bins clear the 0.06 draw threshold and the picture shows a
keyed-looking column pinned to the right edge - the 1200 Hz column runs 25 times
the 400 Hz one.
signalStrength is prominence over the window mean, so the same leakage scores
0.78 and paints the meter to 78% of full width.
So the operator got a strong-signal bar, a plausible 1200 Hz readout, a picture
that looked like a signal, and an empty transcript, with nothing saying why.
The scan that can see past the window now runs whether or not shifting is
enabled - it is the only measurement that can - and publishes through a new
detectedToneHz flow kept separate from estimatedPitch. Overloading the latter is
what let the 1200 Hz claim out in the first place, so the two meanings stay in
two flows. The shift decision still only happens when the setting is on. Cost is
one 121-bin scan every 2 s.
The meter now reads zero when a tone is out of range and not being shifted in: it
is a claim that something decodable is present, and in that state nothing is.
A line under the waterfall says which case the operator is in - the tone was
moved in, or it is out of range and tone shift is off, naming the frequency and
the remedy. Strings in all nine locales; feature:cw only had five, so values-es,
values-ru, values-si and values-uk are new, with the Turkish apostrophe escaped.
CwToneShifterTest pins the premise the hint rests on: that the scan reports tones
the model window excludes, at 120, 250, 1400 and 1500 Hz.
When the tone-shift feature moves a tone into the model's 400-1200 Hz window,
the waterfall now shows two visual markers so the operator can see what is
happening: a green dashed line at the target (800 Hz) and an orange frequency
label at the top-left showing the original pitch.
The waterfall draws the RAW audio, not the shifted audio, so a 1500 Hz tone was
always invisible regardless of the shift setting. The markers close the gap:
the operator can now see that a tone was detected and where it was moved, even
when the original pitch is outside the visible band.
activeShiftHz is now a StateFlow exposed through ICwDecoder so the UI can
observe it without polling.
All three AMSAT endpoint calls still declared Look4Sat/4.5.5 while the project
has been at 4.5.7 for several releases. The API does not appear to validate the
header, but it misrepresents the client version in server logs.
The API caps at 500 records regardless of the hours requested. With 88 catalog
satellites, eight of them more active than 50 reports per 72 hours, quieter
satellites get crowded out. Measured live: the global pull returned 500 reports
covering 36 satellites, while the summary endpoint reported 743 reports across 38
satellites. 26 of 38 satellites had incomplete data, and two (PO-101_[FM] and
TEVEL2-6_[FM]) had zero reports in the global pull despite having reports in the
summary.
The summary endpoint (api/v1/summary.php) returns per-satellite report counts in
one request, so the fix adds one extra call rather than the 88-request
alternative of per-satellite pulls. A satellite whose global pull is incomplete
gets a subdued "68 / 116" marker next to its name, telling the operator the page
knows there is more data it could not fetch. The marker is silent when the
summary is unavailable or the counts match, so the feature degrades gracefully.
The earlier no-data grey (0xFFE8E8E8) already prevented the worst case: slots
crowded out of the global pull were marked as "we never looked" rather than
claiming "nobody reported". The marker now closes the remaining gap: the page
can honestly say "we know there are 116 reports for this satellite but we could
only show you 68 of them".
Also fixed a subagent mutation-testing residue: the coverage floor had been
moved from global (reports.minOfOrNull) to per-satellite (satReports.minOfOrNull)
and left in the tree. One test caught it (coverage is judged from all reports,
not one satellite's), proving the test has teeth.
Adds getAmSatSummary to IRemoteSource and RemoteSource, parseSummary to
AmSatRepository, and summaryCount to SatStatus. All eight test-file
implementations of IRemoteSource were updated for the new method.
Grey meant two different things. The API caps at 500 records however many hours
are requested: measured against the live endpoint, a 72-hour request returned 500
reports spanning only 49 hours, so the oldest 9.5 hours of the third day had no
data at all. Those cells were painted the same grey as "nobody reported", which
claimed knowledge we did not have - 352 of 3168 cells on a real page, a third of
the third day's column.
Slots entirely older than the earliest report in the response now use a lighter
grey. Coverage is judged from all reports rather than per satellite: a quiet
satellite has no reports of its own, but the slots it shares with the rest of the
response were still covered, so it must read as "not heard" rather than "unknown".
The two greys are now in the legend, which previously listed only the four active
states. That matters more than it sounds: on the live page 81% of cells are
"nobody reported" and 11% are outside our data, so a user looking at a mostly-grey
row had no way to tell a dead satellite from a gap in what we fetched. The legend
chips use a solid dot, so the two greys stay distinguishable despite the 25%
alpha background. Strings added to all nine locales.
Three tests cover it: a day entirely before the data starts, a day straddling the
boundary, and an empty response marking nothing as covered.
The UTC alignment landed with tests covering the normal cases; these cover the
ones that would have made it wrong quietly.
Midnight arithmetic is exercised at exactly midnight, a second either side,
every leap-day combination around 2028-02-29, both year boundaries, and the
first of all twelve months in a leap and a non-leap year. Since the code steps
back a day by subtracting 86400 rather than using Calendar arithmetic, those
dates are where a naive step would drift.
Every slot edge across all three days is probed at the boundary and one second
either side, asserting each instant occupies exactly one cell and that the cell's
day matches the report's UTC date - `until` versus `..` on the slot range is a
one-character mistake that would double-count edge reports.
Also pinned: the shared Calendar is not re-read after the labels loop (it points
at the oldest day by then), repeated calls are idempotent, duplicate catalogue
names produce duplicate rows carrying the same report, reports for names absent
from the catalogue are dropped, and the build stays linear in reports rather than
quadratic.
Adds a comment recording why reusing that Calendar is safe: each pass assigns
timeInMillis outright instead of adjusting fields.
235 tests pass.
Two defects in our own AMSAT page, both found by auditing the change that exposed
them.
The day columns claimed to be dates but were a rolling window anchored on the
fetch time. Fetching at 06:07 UTC put 17.9 hours of yesterday into the cell
labelled today; measured against a live amsat.org page of 1021 reports, 73% of
them landed in the wrong day column and none matched the official cell. Days are
now UTC calendar days and slots are fixed UTC bands - slot 0 is 22:00-24:00, slot
11 is 00:00-02:00 - so a cell's contents match its label whenever it is fetched.
The day cell painted one colour for the whole day, taken from the first slot that
had a report, so a satellite that worked all morning and failed all afternoon
looked identical to one that worked once - the reported symptom. It now draws one
stripe per two-hour slot in the same 64x28 dp footprint. Twelve stripes are about
5 dp each, roughly 15 px at 440 dpi, and runs of the same status merge visually,
so a day reads as a few blocks rather than twelve lines. Every density from ldpi
up allocates all twelve without dropping one, and the 4 dp corner radius leaves
95% of the end stripes visible. The report count text is gone; tapping a day
still lists every report from it, which was already the richer view.
buildStatuses and ApiReport are internal rather than private so the grid contract
can be tested. AmSatSlotBuildTest drives it directly: fetchStatus cannot be
tested here because the parsing around it uses Android's JSONObject, a JVM stub
that makes every call return null - eight of nine tests written against it failed
for that reason before being rewritten.
Also corrects three KDoc comments claiming 5 days when the code builds 3, and
records in AGENTS.md that the status colours are ARGB literals in core:data,
duplicated in MainTheme, which anything needing themeable or colour-blind-safe
colours has to fix first.
Mutation testing found the decision rule was effectively untested. Four defects
injected into it - removing the silence guard, comparing shifts instead of
tones, never setting the hysteresis anchor, and inverting the comparison - all
left the entire suite green. The rule lived inside CwDeepDecoder, which needs an
Android Context and a loaded ONNX session, so tests could only restate it, and a
restated rule cannot fail when the real one is wrong.
CwShiftDecider now holds the rule as a pure class that both the decoder and the
tests drive. Its outcome is reported as an enum so the decoder's logging is a
presentation concern rather than a second copy of the logic. CwShiftDeciderTest
targets each of the four surviving mutants directly.
MIN_PROMINENCE lowered from 8.0 to 4.5. Raising it to 8.0 last round overshot:
measured on 400 ms windows of keyed CW in noise, a comfortably copyable signal
reaches only 7.6-9.0 at 0 dB SNR and 5.2-6.7 at -3 dB, so 8.0 silently refused
to shift weak out-of-window signals - the exact failure the feature exists to
prevent. Pure noise peaks at 2.2-3.4, so 4.5 keeps zero false positives across
40 noise windows while retaining the weak end. A false tone is worse than a
missed one: it moves a good signal out of range, whereas a miss leaves the audio
alone until a stronger window arrives. Windows dominated by keying gaps measure
2.4 and are indistinguishable from noise at any threshold; those are skipped.
Test files reorganised to match: the decision rule is covered by
CwShiftDeciderTest against real code, signal-level properties by
CwToneShiftSignalTest, and the restated-logic file it replaces is gone.
80 CW tests pass, golden vectors included.
Two audit findings, both measured, both able to silently disable the feature.
A detection window landing in a keying gap used to collapse an established
shift to zero. CW is keyed, so gaps are normal: over 180 s of keyed audio at
1400 Hz, 11 of 90 detections saw no tone, and each one wiped the decode window
and left the next ~2 s buffered unshifted - outside the model's range and
therefore invisible to it. Absence of a tone is now absence of evidence and the
active shift is retained.
Hysteresis moved from shift space to tone space, anchored on the pitch that
produced the active shift. The old rule required a non-zero previous shift and
a needed shift, so it lapsed exactly where the jump is largest: at the 1200 Hz
edge one 12.5 Hz estimate hop flips between "inside" (shift 0) and "outside"
(a large shift). Measured 35 window drops in 60 detections for a 1205 Hz tone,
and 10 in 10 for a bare one-bin hop. A shift of zero is a real state, not the
absence of one. Slow drift still catches up, since the anchor bounds staleness
at the margin rather than letting it accumulate.
Detection prominence raised from 3.0 to 8.0. Pure noise peaks at 2.0-3.3 times
its own spectral mean, so 3.0 admitted roughly one noise window in five as a
"tone" - and a false tone is worse than none, since it moves a good signal out
of range. Keyed CW measures 47-51, so the gap is wide.
Shifted output is clamped to the +/-1.0 range the spectrogram assumes. The
Hilbert kernel's L1 gain is 2.51, so mixing overshoots: a full-scale square
wave measured 2.35 and even a plain sine 1.05.
The detection pool moved to core:domain as CwDetectionPool so its ring
behaviour can be tested directly - mutation testing showed the previous private
implementation was unreachable from any test. Its chronological-order contract
now has 11 tests driving the real class.
Removed the write-only detectedToneHz field.
74 CW tests pass, golden vectors included.
The detection pool shifted its whole array down one slot per incoming sample
once full. Detection is throttled to 2 s but the pool fills in 400 ms, so for
the remaining 1.6 s of every cycle each chunk arrived at a full buffer: 320
copies of 1280 floats per chunk, measured at 24320 whole-array moves per 10 s
of audio, all on the capture thread.
Writing to a ring index is O(1) per sample. Draining walks the ring from the
oldest slot so the analyser still receives the most recent audio in
chronological order - a test feeds a ramp past capacity and asserts the exact
contents, since getting the wrap wrong would splice the waveform and corrupt
every estimate silently.
Follow-up to the tone-shift feature, closing gaps the audits surfaced.
Toggling the setting, or the detector settling on a materially different shift,
now discards the buffered audio. Without it the 20 s decode window kept feeding
the model samples moved by the old amount for up to 20 s after the user acted,
and updateSignalMetrics corrected the pitch readout by an offset that no longer
matched the window. Text already committed to the history is kept: it was
correct when it was decoded.
The previous-state flag is nullable and seeded from the current setting on the
first chunk, so a decoder created while the setting is already on does not
report a spurious change and wipe an empty buffer. reset() clears it back to
null for the same reason. Two decoders can be live at once (the CW screen and
the Radar panel) and each tracks its own state.
Re-shifting is now gated by a 40 Hz hysteresis. Detection resolution is 12.5 Hz
and a real tone wanders, so without it an estimate hopping between adjacent
scan bins would drop the window every 2 s - costing far more decoding context
than re-centring gains. 40 Hz absorbs two bins of jitter while still following
a genuine retune; a test pins both halves of that trade-off.
DeepCW only analyses 400-1200 Hz - its input tensor is 65 bins wide, fixed at
training time - so a CW note outside that range is invisible to the decoder.
This adds an opt-in preprocessing step that moves such a tone to 800 Hz, the
window centre, extending the usable pitch range without touching the model.
Single-sideband mixing via a 63-tap Hilbert transformer. Plain real mixing was
measured and rejected: shifting 1500 Hz to 800 Hz left a fold-back image at
1000 Hz at 0.999 of the wanted amplitude, inside the window. Zero-stuff
upsampling plus lowpass handled downward shifts but left a 0.996 image when
shifting 300 Hz upward. The Hilbert approach measures clean on nine tones from
150 to 1550 Hz: one peak at the target, nothing above 0.3 relative amplitude.
In-window energy for a 1500 Hz input goes from 6.8% to 94.6%.
Only out-of-range audio is processed. A tone already inside 400-1200 Hz is
returned untouched (same array instance, no copy), and with the setting off the
audio path is exactly what it was before.
CwToneShifter.Streaming carries the Hilbert filter history and mixer phase
across capture chunks. Shifting each chunk in isolation left 62 of every 320
samples convolving against zeros, inflating envelope ripple to 8.7x the
whole-buffer baseline. A residual difference in the last ~3 samples of each
chunk is causal and documented: those output samples would need input that has
not been captured yet.
Detection pools chunks rather than gating on one. A capture chunk is 4410
samples at 44.1 kHz but only 320 after resampling to 3200 Hz, so requiring
1280 samples in a single chunk would have made the feature dead code - the two
independent audits both found this before it shipped. Detection now runs on a
pooled 0.4 s window, at most every 2 s.
Toggling the setting or a change in the detected shift drops the buffered
audio: the 20 s window would otherwise keep decoding samples moved by the old
amount, and the pitch readout could only be correct for one of them. The
readout itself subtracts the active shift so it shows the pitch on the radio,
not the shifted one.
Settings: OtherSettings.cwToneShiftEnabled, off by default, persisted and read
back in SettingsRepo, toggled from the Other card in Settings with a help line
explaining the 400-1200 Hz limit. Strings added to all nine locales. The
decoder reads the flag per chunk, so the toggle applies without restarting
capture.
Debug: the enabled-state transition, each detection verdict (no tone / inside
window / shifting by N Hz), and every shift change are logged, with the noisy
paths throttled to the 2 s detection interval. CwProbe records shift changes
only, keeping well inside its 1 MiB cap.
Tests: 8 shifter tests (detection sweep, noise rejection, pass-through
identity, image-free shifting across 8 tones, end-to-end spectrogram energy),
8 streaming tests (chunk continuity, history retention, reset semantics, chunk
sizes above and below the history window), and 6 gate tests including a
regression guard that a 320-sample chunk must be able to reach the detection
threshold. All 53 CW tests pass, golden vectors included.
The merged upstream AMSAT implementation fetches status data from AMSAT's JSON
endpoints (getAmSatCatalog / getAmSatReports), so the fork's HTML scraping path
no longer has a caller:
- core/data/.../source/AmSatParser.kt (136 lines): parsed the amsat.org status
table, deriving state from the page's inline colour codes.
- IRemoteSource.getStatusHtml() plus its RemoteSource implementation and the
DatabaseRepoTest fake override.
Verified zero references repo-wide before removing, and again afterwards.
Request / CancellationException imports in RemoteSource remain in use by the
other fetchers. compileReleaseKotlin plus core:domain / core:data /
feature:map / feature:roaming unit tests stay green.
Post-merge audit found code the merge left unreferenced:
- SettingsRepo: keySatelliteUrls / keyTransceiversUrls / separatorUrl were
upstream's list-shaped data-source keys. The merge kept the fork's map-shaped
DataSourcesSettings, so these three had a definition and zero uses.
- feature/status/res/drawable/ic_refresh.xml: SatStatusScreen imports
core.presentation.R only, so its R.drawable.ic_refresh resolves to the
core copy; the feature-local copy was never addressable. It was the only
file under feature/status/src/main/res, so the directory goes with it.
Verified zero references with a repo-wide grep before removing each symbol.
compileReleaseKotlin plus core:domain / core:data / feature:map /
feature:roaming unit tests stay green.
AudioCapture.audioFlow's finally ran recorder.stop() then release() naked. If
startRecording() threw - permission revoked mid-request, audio device error -
the finally's stop() threw IllegalStateException (stop on an uninitialized
recorder), which replaced the original error AND skipped release(), leaking
the AudioRecord. The flow's caller saw "recorder failure" instead of "no
permission" and the native recorder was never freed.
Wrapping each cleanup step in runCatching preserves the original exception
while guaranteeing release() runs. Probe: a start failure previously surfaced
as RuntimeError with released=false; it now surfaces as the original
PermissionError with released=true.
RadioTrackingService wrote lastSetTxFreq/lastSetRxFreq unconditionally after
calling setFrequency, ignoring its Boolean result. When the radio rejected the
frequency - the FT-817 CAT limit added in the previous commit, a dropped
Bluetooth link, or a failed ack - the remembered value no longer matched what
the radio actually holds. The manual-tuning detector then saw a phantom dial
change on the next read-back (read is the real frequency, lastSet is the one
that never landed) and entered tuning mode: it locked onto the wrong base and
kept rewriting the radio.
Probe of the state machine: before the fix, a rejected 1.26 GHz write against
a radio sitting on 145.5 MHz left lastSet at 1.26 GHz, so every subsequent
cycle read a 1.1 GHz gap and flagged manual tuning forever. After the fix the
lastSet is only updated on success, so the detector sees no change and the
loop keeps applying the next valid frequency. Same fix applied to the split
IC-705 path (setWorkingFrequency/setTxVfoFrequency).
RemoteSource's five suspend functions and SatStatusViewModel's two fetch paths
caught bare Exception, which also swallows CancellationException. When the
owning scope is cancelled (screen leaves, app closes) a cancelled network call
was reported as a null/error result instead of stopping: the caller kept
running until the next suspension point, and SatStatusViewModel wrote state
updates into an already-cancelled scope. Correct coroutine hygiene is to let
cancellation propagate - rethrow CancellationException before the generic
catch. Verified semantically with an asyncio probe: a swallowed cancel returns
a normal-looking null and the caller continues; a propagated cancel stops the
coroutine immediately.
No behaviour change for real errors; :core:data and :feature:status compile.
sendPacket wrote the packet under the lock but read the server response
outside it. disconnect() - called concurrently from stop() and from the
reconnect path in AprsReporter.reportOnce's catch - nulls and closes
writer/reader/socket under the same lock, so the lock-free read raced with it.
A probe interleaving 5,000 sends with repeated disconnects produced a mix of
744 OK and 4,256 exception results: the response read hit a just-closed socket
and the swallowing runCatching reported Pair(true,"OK") for a packet that may
never have left, or read through a stale reference. The tracker believed the
beacon was heard while APRS-IS never received it.
Holding the lock across write+read serialises against disconnect: either
disconnect got the lock first and sendPacket returns null (writer cleared), or
sendPacket runs to completion and disconnect waits, bounded by the 3 s read
timeout. Re-ran the interleaving probe: 3,000 sends, zero inconsistent
results. Compiles and :core:data tests stay green.
The FT-817 CAT frequency field is 4 BCD bytes at 10 Hz resolution, so the
largest representable value is 999,999,990 Hz. encodeFrequencyBcd is exact
below that, but for anything above it the %08d formatting silently drops the
leading digit: 1,267.6 MHz encodes as 126.76 MHz. Verified against the release
bytecode - 1,000,000,000 Hz -> [10 00 00 00] -> 100,000,000 Hz, ten times
lower - and the SatNOGS catalogue has 17 transmitters with uplinks over 1 GHz
(QO-100 at 2400.05 MHz, several 23 cm links), so the wrong value is reachable
via RadioTrackingService when an FT-817 is mis-configured as the TX radio.
The tracking loop's read-back then locks onto the wrong band with no warning.
Reject out-of-range frequencies at setFrequency with a log and return false
instead of sending a corrupted command. In-range values are unaffected
(probe: 7.074/145.5/435.1 MHz and both 999,999,98x/99x MHz round-trip exactly;
every value above the limit is refused before the encoder runs).
Two leaks in NetworkReporter, the same family as the Bluetooth/APRS/radio
socket leaks fixed earlier:
1. ensureRotatorConnected/ensureFrequencyConnected assigned the field directly,
so a channel that opened but threw during the rest of setup was never
closed and remained referenced. Use a local `opened` and close it in the
catch, matching the pattern used in AprsIsClient/Ic705Controller/
Ft817Controller/BluetoothReporter.
2. write() only flipped connected=false on failure. The broken channel stayed
in the field, the next ensure* reconnected and overwrote it, and the old
channel was never closed. Now a failed write closes the channel and nulls
the field. The null check is identity-based (socket === field) so a stale
reference from a concurrent report can never close a newer channel.
State-machine probe: normal write keeps the socket, failed write closes and
nulls, next report reconnects fresh, and passing a stale reference does not
close the newer socket.
CwProbe.step() appended one line per call with no size limit, rotation or
cleanup, and it runs on every build: CwDeepDecoder is the only ICwDecoder
implementation and writes infer_begin + infer_done every 1.5 s inference tick.
Measured against the actual line format that is ~170 KB/hour, ~4 MB/day of
unbounded growth in files/probe_cw.txt while CW audio is monitored, plus
synchronous disk I/O on every inference.
Truncate when the file exceeds 1 MiB instead of deleting, so the probe keeps
the most recent diagnostics (the reason it exists: the last lines show where a
flash-crash died). Simulated 10 h of continuous use: 2.7 MB written in total,
file stays bounded around ~640 KB; previously it would have kept all 2.7 MB
and grown without limit.
The status grid always drew six day columns, but the AMSAT reports endpoint
cannot supply six days for the full catalogue. Measured against the live API:
limit=500 -> meta.count=500, covers 4 days (Aug 11..Aug 14)
limit=1000 -> meta.count=500, same 4 days (server clamps the limit)
hours=336 -> meta.count=500, same 4 days (window size does not help)
before/offset/page -> ignored, same 500 newest rows
With ~90 catalogued satellites the 500 newest rows only reach about four days
back, so the two oldest columns were guaranteed to be uniformly gray. Gray means
"no report" in this UI, so the screen asserted nobody reported those days when
the truth was that the data was never fetched.
Derive the column count from the oldest report actually received, capped at six.
On live data that yields four columns labelled Aug 14..Aug 11 instead of six with
Aug 10 and Aug 9 blank. The UI already renders whatever days it is given, so no
UI change is needed.
Also name the request constants and record what was measured about the endpoint,
so the 500 is not mistaken for an arbitrary choice that can simply be raised.
Note for a future change: the per-satellite form of the endpoint
(reports.php?name=...) is not affected by the cap - sampling eight satellites
returned 926 rows spanning eight days, i.e. full six-day coverage - but it needs
one request per satellite (~0.8 s each, ~68 s for the whole catalogue), so
switching to it is a deliberate trade-off rather than a bug fix.
:core:data:compileReleaseKotlin, :core:data:testDebugUnitTest and
:core:domain:test all BUILD SUCCESSFUL.
AmSatRepository labelled columns by calendar date but filled them by slicing a
rolling 72-slot window ending at fetch time. The two timelines coincide only
near 23:59 UTC. At common fetch times the status grid lied about dates:
UTC 00:00: 72 / 72 slots under the wrong label
"today" column contained all of yesterday
UTC 12:00: 36 / 72 wrong; every column straddled two dates
UTC 13:37: 30 / 72 wrong
UTC 23:59: 0 / 72 wrong (the accidental alignment case)
Anchor the six columns on UTC midnight instead. Every SatDay now covers exactly
[day 00:00, next day 00:00), split into twelve 2-hour slots newest-first so the
UI's existing first-non-gray lookup still chooses the latest daily report.
A standalone Java probe porting the old arithmetic reproduced the 72/72,
36/72 and 30/72 mismatches. Porting the new formula gives 0/72 mismatches at
00:00, 12:00, 13:37 and 23:59 UTC.
Also restore core:data's unit-test compilation. DatabaseRepoTest's fakes were
stale after IRemoteSource gained AMSAT methods and ISettingsRepo's zero-arg GPS
setter became suspend; the whole data test suite previously could not compile,
so data-layer regressions were untestable. Updated the fake members and verified
:core:data:testDebugUnitTest plus :core:data:compileReleaseKotlin BUILD
SUCCESSFUL. The product code does not use org.json in JVM tests because Android
org.json stubs throw there, so the date math remains verified by the standalone
same-JVM probe rather than a misleading mocked parser test.
A full-domain round-trip probe found three related boundary bugs.
1. Exact positive limits wrapped the square/subsquare terms to zero
positionToQth clamped only the A-R field index. At +90 latitude / +180
longitude the field saturated at R, but all later terms used modulo and wrapped
to square 0 / subsquare a:
(90, 180) -> RR00aa00 -> (80.002083, 160.004167)
error: -9.998 deg latitude, -19.996 deg longitude
The existing test incorrectly asserted RR00aa00 and had therefore fossilised
the defect. Clamp shifted coordinates just inside the half-open upper bound so
the limits land in the final cell RR99xx99.
2. isValidPosition allowed longitude through +360
Maidenhead covers -180..180, but 181..360 was accepted and produced plausible
locators that decoded 20-200 degrees away:
lon 181 -> decoded 161.004167 (error -19.996)
lon 270 -> decoded 170.004167 (error -99.996)
lon 360 -> decoded 160.004167 (error -199.996)
Restrict the converter contract to -180..180.
3. Locator validation allowed S-X as field letters
The first pair has 18 fields A-R, while only the later subsquare pairs use
A-X. The shared [A-X]{2} regex accepted SS00aa / XX99xx and decoded them past
the poles (up to lat 149.98, lon 299.96). Use A-R for the field pair.
The SettingsRepo caller had a separate wrapping bug that masked part of this:
it mapped longitude>180 by subtracting 180 (270 -> +90, wrong hemisphere)
instead of modulo 360 (270 -> -90). Fix that at the writer too.
Verification:
- standalone JVM sweep: 519,841 points, old code had 1,441 large-error points
with max drift 9.997917 deg lat / 19.995833 deg lon
- new Kotlin regression sweep requires every 8-char round trip <=0.01 deg
- QthConverterTest BUILD SUCCESSFUL
- full :core:domain:test + :core:data:compileReleaseKotlin BUILD SUCCESSFUL
All five connect paths opened a socket, completed the TCP/RFCOMM handshake,
and only afterwards stored it in a field. Any exception in between leaked the
socket: the catch block just flipped a boolean, and disconnect() can only close
what already reached the fields.
Leak windows (statements that can throw after the handshake succeeded):
AprsIsClient.connect soTimeout / tcpNoDelay / getOutputStream / getInputStream
Ic705Controller.connect outputStream / inputStream / sendAndWaitAck
Ft817Controller.connect outputStream / inputStream
BluetoothReporter x2 outputStream
AprsIsClient is the worst case because AprsReporter retries on a timer
(intervalMin, minimum 1 minute) and nulls out the client after each failure,
so every failed attempt permanently loses one fd:
failure rate leaked fds/hour time to exhaust 1024 fds
5% 3.0 ~14.2 days
20% 12.0 ~3.6 days
50% 30.0 ~1.4 days
100% 60.0 ~17 hours
Typical trigger is a weak link where the TCP handshake succeeds but the peer
immediately RSTs (overloaded or rate-limiting APRS-IS server). Once fds run out
nothing in the process can open a socket or file any more: TLE updates, AMSAT
status and WaveLog uploads all start failing with no obvious cause.
Each path now keeps a local reference to the socket it opened and closes it in
the catch block, also clearing the stream/socket fields so a half-initialised
connection is not mistaken for a live one.
Verified: :core:data:compileReleaseKotlin BUILD SUCCESSFUL; grep confirms all
five close calls are present.
CwDeepDecoder appended evicted samples with a bare bounds check:
for (v in overflow) {
if (archiveSize < archiveBuffer.size) archiveBuffer[archiveSize++] = v
}
if (archiveSize >= ARCHIVE_THRESHOLD) { flush() }
Once archiveBuffer (64000 samples / 20 s) filled up mid-batch the remaining
samples were silently discarded, because the flush only ran after the loop.
Worst measured case: 47999 samples already accumulated (just under the 48000
flush threshold, so no flush) plus a 64000-sample overflow batch means 111999
samples pushed into a 64000 buffer -> 47999 dropped, i.e. 15 s of audio missing
from the permanently archived CW history.
Now the buffer is flushed as soon as it is full and before appending, so every
sample reaches archiveDecode. Simulation over five batch patterns: dropped
count goes 47999 -> 0 for the worst case and all 111999 samples are archived.
Bounds: single append() can evict at most capacity samples, so drainOverflow()
returns at most 64000 - archiveBuffer never needs to grow.
The previous guard (if (_isCalculating.value) return) silently dropped
concurrent calls. Every call carries filter settings the user just applied,
so a dropped one left the list showing results for the previous filter:
User clicks 'Apply' with elevation>=5
-> UI updates to show elevation>=5
-> calculatePasses(elevation>=5) called
-> but if _isCalculating=true, return immediately
-> list still shows elevation>=30 results
The guard window is wide: delay(1000) + real calculation time (hundreds
of ms to seconds), exactly when the progress indicator spins and users
naturally interact again.
Mutex serializes calls instead: the second one queues and eventually runs
with its own parameters. This also fixes the original concurrency issue
(duplicate parallel calculations) and adds finally {} so a thrown exception
cannot leave isCalculating stuck at true (frozen progress indicator).
Reverts the regression introduced in the previous attempt to add concurrency
protection.
Addresses @AlanCui4080's review feedback about the manual date calculation.
Uses SimpleDateFormat to eliminate the hand-written calendar math (regex +
days-from-epoch calculation). Avoids java.time since it requires desugaring
on minSdk 24, keeping dependencies minimal.
Before: 23 lines of manual day-from-epoch calculation
After: 7 lines using SimpleDateFormat
Ref: https://github.com/rt-bishop/Look4Sat/pull/233#discussion_r1868599947