Refusal

Why we say indeterminate

Two things stop an epoch before it is given a state, and the number behind the second was not chosen by anyone. It falls out of the five weights the scorer already had, and it commits Delve in advance to refusing at most one genuine reading in a thousand.

The standard

Impossible, not merely unusual

Two things stop an epoch before it is given a state. The first is a single rule about the spectrum: if the 0.5–4 Hz band holds less than one percent of the power, there is nothing to score. That is not a quiet brain. Scalp EEG is never empty down there — the minimum across 140 epochs of real human recording is 0.108 — while 568 of the 615 epochs Delve had stored from its own broken decoder reported delta as exactly 0.000000. It is a failed integration, and it is refused.

One rule, where there used to be four. The bar a rule has to clear to sit here is that it describes something impossible, not something unusual, and three rules that had been inherited from an earlier calibration failed that bar the moment there was real signal to test them against:

  • delta > 0.60
  • gamma > 0.25
  • 1/f exponent < 0.5

All three came from the broken sessions, because that was the only data anyone had. Applied to 140 epochs of ordinary awake resting EEG they refuse 72, 5 and 19 epochs respectively — 91 of 140, sixty-five percent, of perfectly healthy signal. A gate with a sixty-five percent false-refusal rate is not a gate. They were removed, and the reasoning written down so they cannot come back.

That leaves the second stop, which is the interesting one.

The number nobody chose

Each state is a Gaussian in five weighted coordinates, so an epoch has a distance from the nearest one, measured in units of that state’s own spread. Past some distance, the honest answer is that this is not that state. The question is which distance, and the answer is not a matter of taste.

When an epoch really does come from a state, its score Q is a weighted sum of five independent χ²(1) terms. That is not itself a χ², but the Welch–Satterthwaite approximation matches its first two moments with a scaled one, and the whole derivation is three lines of the weights the scorer already had:

h = (Σw)²/Σw² = 25/6.5 = 3.8462
g = Σw²/Σw = 6.5/5 = 1.3
Q_REJECT = 1.3 × χ²₀.₉₉₉(3.8462) = 23.5817

Which is a distance of 2.1717 per coordinate. An epoch whose best-fitting state fits worse than the 99.9th percentile of what that state itself produces is not that state, and is refused.

Deriving it rather than picking it fixes the false-refusal rate in advance: at most one genuine epoch in a thousand is turned away on distance, committed to before the threshold was ever run on anything. Change one of the five weights and the number moves with it. It is a consequence of the model, not a dial on it.

what a state itself producesQ_REJECT · 0.999g·χ²(h)

Q, the weighted squared distance · shaded: the 99.9% that is read

The distribution the threshold is a quantile of, computed from Σw and Σw² rather than typed in. Almost the whole of what a state produces lies below ten; the refusal sits at 23.58.

The ceiling

0.84, borrowed from two humans disagreeing

The epochs that are read carry a cap on what may be claimed for them, and it is the same derivation problem one step along. This judgement has no ground truth, so its accuracy cannot be measured. What can be quoted is how well trained experts agree with each other on the nearest comparable task — scoring graded drowsiness from EEG.

In that literature, agreement on the frankest, most clear-cut cases reaches κ = 0.80, and it collapses to chance in the grey zone. Take the best of it, because no epoch can be easier to read than the easiest case on record. Cohen’s κ relates to raw agreement by pₒ = κ + (1−κ)·pₑ, and for five states at roughly uniform marginals that is 0.80 + 0.20 × 0.20 = 0.84. Four seconds of four-channel dry-electrode EEG is not entitled to claim more than the best two humans manage on an easier question.

The cap falls from there to 0.50 at the refusal boundary, and that end is not chosen either — it is forced. The reject class is scored at exactly the point where belief stops, so at the boundary it and the winning state have identical scores. The epoch is, by construction, a coin flip between “this state” and “nothing I recognise”.

Niruddha is capped at 0.50 everywhere, including at its own profile centre — by design rather than by measurement. It is the least evidenced of the five states and should require agreement across consecutive epochs, and this classifier is per-epoch and stateless. Until a layer exists that can watch a whole sitting, half is the honest maximum.

And two things that no longer refuse

A broken electrode and a face full of jaw tension both used to withhold the reading. Neither does now; both discount the confidence instead, on a floor of 0.50/0.84 = 0.5952 — the ratio of the same two anchors. A caveat is actionable and a blank panel, arriving every four seconds, is indistinguishable from a page that has not loaded. That change is set out in full in the methodology post, and the gamma half has a post of its own.

One consequence is visible on screen: because the ceiling, the discounts and the reject class all withhold probability rather than redistributing it, the five states shown do not add up to one. The gap is the part of the answer nobody is claiming.

In a sitting

What this means for your practice

“Indeterminate” is not an error and it is not a comment on you. It is the system declining to name a state it cannot defend — roughly once in a thousand epochs from the distance rule alone, and more often when the band is genuinely not reading. A sitting with a scattering of them is normal.

If they arrive in a long unbroken run, that is worth acting on, and the thing to act on is the hardware rather than the attention: reseat the band, wet the contacts behind the ears, push the hair clear. And if the five bars in front of you do not sum to a hundred percent, nothing is broken. That gap is the honest part.

The limits of this

One epoch in a thousand is the false-refusal rate under the classifier’s own model — that is, if the five profiles are right and an epoch really is drawn from one of them. It is a property of the arithmetic, not a measured field rate, and it says nothing about how often the model is wrong in ways it cannot see.

The 0.84 ceiling is borrowed from a different judgement, made by different people, on different equipment. It is quoted as a bound on what may be claimed, not as evidence that Delve agrees with anybody at that rate. Delve is a research and education tool. It is not a medical device and makes no health claims.

Sources: docs/classifier-spec.md §3, §3.1, §6, §8 · docs/real-eeg-validation-fixture.md §6

© 2026 Delve

Delve is a research and education tool. It is not a medical device and is not intended to diagnose, treat, or monitor any condition. Mental-state labels are estimates from EEG signal analysis, not clinical assessments, and their accuracy varies with the individual, the session and electrode contact. If you have concerns about your mental health, speak to a qualified healthcare professional.