radiocli Home Blog GitHub

All posts

Ask Again

In “Ask the Screen” I wrote that the screen, unlike the API, has to be right. That line has been the working theory of this whole project: when the SDS150’s serial protocol won’t answer, drive the menus, read the glass, trust what’s drawn. This week the theory picked up an asterisk. Nothing about the screen stopped being true. What I forgot is that a screen is only right once it has finished drawing, and there is no signal, none, that says when that is.

“Which layout is on screen” sounds cosmetic, and it isn’t. The scanner draws with seven layouts, and the soft keys along the bottom row are the only screen areas not in the built-in map, because their widths follow whatever labels the current mode is showing. So the tool reads them off the live screen, and only when the layout it’s reporting on is the one actually being drawn. Call the layout wrong and five areas come back unplaced. That’s a wrong answer dressed as a missing one, and downstream nothing distinguishes them. The full working notes are in layout-detection.md.

Three questions, cheapest first

Working out the layout takes up to three questions, asked in order of cost. First, ask what the scanner is doing, which is a plain status read: conventional scanning and trunk scanning each name a pair of layouts, one simple and one detail, while every other mode names exactly one. Second, if it named a pair, count the rows on the screen: simple draws 14 lines and detail draws 17, and counting is another plain read that opens nothing. Third, only if the count is neither, open the display menu and read the highlighted entry, which is the actual authority and the only step that costs anything. Circumstantial evidence first, authority last.

There’s one deliberate wait in that path. The scanner reports its screen as plain text for a moment after being moved, one of the quirks cataloged in the oddities file, so the candidate check gives the screen up to three seconds to settle. And there’s one deliberate refusal to wait: if the screen says it’s a menu, the check returns immediately, on the reasoning that a menu will stay a menu and waiting would only delay a good error message. Keep that reasoning in mind. It’s sound where I wrote it, and it’s also the prime suspect in the bug I haven’t fixed yet.

The bug I fixed

A full read of a layout walks the scanner’s menus for about thirty seconds, and the code used to read the live screen for the soft keys the instant it climbed back out. At that instant the scanner is still drawing the menu it just left. The bottom row isn’t three runs of reverse video yet, the soft key reader correctly declines to guess at a row that doesn’t look like soft keys, and all three came back with no position.

The tell was two commands disagreeing about reality seconds apart: a full colors read said the soft keys had no positions, and a cached read of the same layout on the same scanner placed them fine. The bug hid for a long time behind exactly that asymmetry. Any path that never enters a menu reads a settled screen and gets the right answer, so the fast paths all worked and the slow path quietly didn’t.

The fix asks until the row looks like soft keys, on the same three-second budget the settle logic already uses. A layout that genuinely draws no soft keys spends the whole budget and still reports none, which is the right answer arrived at slowly, and slowly is fine. The lesson generalizes past this radio: anything reading a live display right after you’ve moved the device is reading a screen in transition, and with no “redraw finished” signal your only defense is to ask until the answer looks like the thing you asked about. Then ask again.

The bug I haven’t

The test suite drives the commands back to back with no pause, and in one run a plain colors placed the soft keys at line 16 with lengths 9, 1, 10, 1, 9, then two named reads seconds later reported dashes for all five. Same layout, same scanner, same minute. That has happened once in five runs of the suite, and nine attempts to reproduce it in isolation, five pairing a scan with a cached named read and four pairing a weather read with the same, all passed. Whatever the window is, it’s narrow, and it closes when I aim instruments at it.

This is the shape Jim Gray was writing about at Tandem back in 1985, the transient timing-dependent fault the industry came to call a Heisenbug: it disappears the moment you arrange to watch it. Gray’s observation was that most production faults look like this, and his prescription was the unglamorous one, retry. Forty years on, my settle loops are the same medicine in a smaller bottle.

So which of my suspects is it? There are two, both consistent with everything seen, and they need different fixes. The first: the screen was still reporting a menu when the check ran, because the previous command had just climbed out of one. The candidate check returns immediately with no candidates on a menu screen, and the code reads “no candidates” as “the named layout is not among them.” That’s a false pretending to be an unknown, and the timing fits exactly. The second: the row count was read mid-redraw. The count comes from a single display read matched against 14 or 17, and a partially drawn screen that happens to show exactly 14 rows would confidently answer with the wrong layout. It declines safely on any other count, so only the value 14 is dangerous, and only on a scanning screen. I don’t know which it is yet, and I’m not going to guess with code. One debug line carrying the screen name and the row count, and a loop of the test until it fails, will separate them. Observation first, fix second.

False is not unknown

The reason this bug is silent is three lines I wrote on purpose:

want, _ := lookup(wanted)
current, _ := isCurrent(ctx, client, want)
return want, current, nil

That second underscore throws away the error from the currency check, and the reasoning was defensible: whether the named layout happens to be on screen is worth mentioning, not worth failing a whole read over. Go made me sign for the decision with a blank identifier, so I can’t even claim I did it by accident. The cost only became visible later: the one signal that would explain a wrong answer is discarded, and a scanner that could not be asked becomes indistinguishable from a scanner that answered no.

Bottom line: false and unknown are different values, and any interface that collapses them will eventually lie to you. The report should say “not the current layout” when the scanner answered and “could not tell” when it didn’t, because the second case is the one where the soft keys are missing rather than absent. That fix is coming regardless of which suspect the debug line convicts. So is a measurement I’ve been putting off: the three-second settle budget was picked as “enough” and never measured against an actual menu exit, and a measured number would let every settle loop in the command share one honest constant.

The screen still has to be right. It just doesn’t have to be right yet.