The symptom
Six hundred epochs, almost all of them certain
The first version of the Chitta Bhūmi classifier scored an epoch by closed-form Gaussian likelihood against five published profiles and softmaxed the result. Replayed over the 615 epochs then stored in the database, 599 of them came back above 0.95 confidence, with a median of exactly 1.000. In one 65-epoch export, nineteen epochs reported 100.0%.
Certainty on that scale is a claim, and it was not being earned. Over the same 615 epochs the classifier returned Kṣipta 608 times and Ekāgra seven; Mūḍha, Vikṣipta and Niruddha were assigned to no epoch at all. Three of the five states it advertised were unreachable.
Worse than either: much of that signal was not EEG. A fault in the headband’s Bluetooth decoder had been inflating gamma about sixty-fold and suppressing alpha, and delta read exactly 0.000000 in every epoch of every real-hardware session. So the system was at its most confident about recordings that were, in the arithmetic sense, nothing at all.
That combination is the diagnosis. Measured directly, the Spearman correlation between an epoch’s distance from the nearest state profile and the confidence reported for it was +0.34. The further an epoch sat from anything the model recognised, the surer the model became.
It was the arithmetic, not the tuning
For a Gaussian classifier, the log-odds between any two states is
(μ₁ − μ₂)(x − (μ₁ + μ₂)/2) / σ²
which is linear in the epoch and has no upper bound. Push x far enough away from both profiles and one of them is always slightly less astronomically unlikely than the other, and that ratio is all a softmax has to work with. On the shipped delta profiles a relative delta of 0.60 already gives 0.908; 0.75 gives 0.997; 0.90 gives 0.99998 — from one band, before the other five multiply in.
There is a sharper way to say it. In the old posterior, a term common to every state cancels: multiply every state’s distance by ten and the reported probabilities do not move at all. The model could tell you which profile was least bad. It had no way to tell you whether any of them was any good.
Lowering the softmax temperature or widening the standard deviations moves where the saturation happens. Neither removes it.
distance from the nearest profile, in RMS z per coordinate
Both curves are computed from the two scorers’ own arithmetic along a line leading away from every profile, not drawn to be persuasive. They are the two-state form; a real epoch is scored against five states and a reject class, so the heights are illustrative and the directions are not.
The rebuild
A class that does not move
The fix is older than the problem. Chow’s rule: give the classifier a sixth option whose score does not depend on how far away the epoch is, and the cancellation breaks. Delve’s reject class is scored as exactly as likely as the winning state would be at the precise distance where the model stops believing it — so as an epoch drifts away from everything, the reject class comes to dominate the denominator instead of scaling with it.
That is one half. The other is a ceiling on what may be claimed for the epochs that are read, falling from 0.84 at a perfect fit to 0.50 at the refusal boundary. Both anchors are derived rather than chosen, and the next post is about where they come from.
Measured the same way as before, on the same 615 epochs, the correlation between distance and confidence is now −0.761. On the 140 epochs of real, independently recorded human EEG it is −0.193, from +0.075 under a flat ceiling — still the wrong sign, on real signal, until the ceiling was allowed to fall. And no accepted epoch can report above 0.84 at all any more, because 0.84 is the ceiling: the old headline of 599 epochs above 0.95 is now arithmetically unreachable rather than merely rarer.
- ρ +0.34 → −0.761
- 615 epochs
- ceiling 0.84
- reject at d = 2.17
In a sitting
What this means for your practice
The number beside your state is now a smaller number, and that is the improvement. Eighty-four percent is the top of the scale, not a disappointing score; a reading in the forties or fifties is the system telling you those four seconds did not clearly resemble any one state, which is a fact about the epoch and not a verdict on your sitting.
Do not read a falling confidence as a falling quality of attention. Confidence measures how well the signal matched a profile. It is perfectly ordinary for a good sitting to spend long stretches between two states, and Delve now says so rather than picking one and insisting.
The limits of this
The 615 epochs this post quotes are not good data, and that is the reason they can be quoted here. They demonstrate a defect in a scorer — a scorer being confidently wrong about signal that was not EEG — and nothing whatever about anybody’s meditation. No practitioner recording appears in this post, and none is used to support any claim in it.
The direction of confidence is now correct and measured. That is not the same as the confidence being calibrated: this judgement has no ground truth to be calibrated against, so 0.84 is a bound on what may honestly be claimed rather than an accuracy figure. Delve is not a medical or diagnostic instrument and makes no health claims.
