Evidence

How a sitting is read, and what the reading is worth

Everything technical about Delve, in the order it happens: four electrodes to six bands to one of five states, the two rules that make the classifier decline to answer, the ceiling on what it is allowed to claim when it does, and the checks that would go red if any of it drifted.

Numbers here are quoted from the specification the three implementations answer to, not from memory. Where the site previously published a figure this page contradicts, the contradiction is stated rather than tidied away.

Working prototype · Aug 2026

The premise

Attention leaves a trace

Not a thought, and not a mood. What a headband can see is the rhythm of the cortex underneath it, and the rhythm changes with what attention is doing. That is the whole of the bet this project makes.

What is actually measured

Four electrodes on a consumer headband — two at the temples, two on the forehead — sampling a few hundred times a second. Every four seconds that stream is cut into an epoch, decomposed into six frequency bands, and scored against five profiles. The reading that comes back names a state and says how far from it the epoch sat.

The strongest evidence that this is measuring anything at all is the plainest one in the field: close your eyes and the 8–13 Hz alpha rhythm rises sharply. This pipeline recovers that on real human recordings it did not produce and cannot tune — three times over, on five subjects out of five.

What is not measured

Nothing here reads content. There is no thought, no image, no memory and no intent anywhere in the signal chain. Four channels give a whole-head average, not a map of the brain, and the software says so in its own interface rather than only here.

The pipeline

How a sitting is read

Five steps from a headband to a sitting you can read back. The first four are arithmetic: every number they produce can be exported and checked by hand. The fifth only says those numbers back to you.

01Wear
ARTIFACT-SCREENED · MEDIAN ± MADCH 1CH 2CH 3CH 41024 samples · ~4 sflagged above 2% bad samples4TH-ORDER BUTTERWORTH · ZERO-PHASEδ0.5–4θ4–8α8–13β₁13–18β₂18–30γ30–50WELCH PSD · HzGAUSSIAN LIKELIHOOD · PUBLISHED PROFILESMūḍhaoff the ladderKṣiptaSurfaceVikṣiptaEmergingEkāgraDeepNiruddhaProfoundΔ 0.34refuses past 2.17 σ from every stateSTORED WITH THE SITTINGstatedepthswaraα / θREAD BACK AS ONE PARAGRAPH
4 channelsWeb Bluetoothno install
01A headband, and a browser tab

Wear

A consumer EEG headband across the forehead: nothing implanted, nothing that needs a lab. It pairs to the browser over Web Bluetooth.

021024 samples, roughly four seconds

Sense

Every epoch is screened before it is trusted: a per-channel median and MAD, a 6.0 robust-sigma threshold, a channel flagged past 2% bad samples, and its own signal-quality figure.

03Six bands, filtered without smearing time

Decompose

A 4th-order Butterworth band-pass runs zero-phase, forwards then backwards, so nothing shifts in time. Welch's method splits the spectrum into six bands. Low and high beta stay apart: averaged, settled alertness and agitation collapse into one number.

04A state, or an honest shrug

Interpret

Band powers are scored by closed-form Gaussian likelihood, giving one of five Chitta Bhūmi: four rungs of a depth ladder, and Mūḍha, dullness, which sits off it. Neuromarkers from the same epoch read underneath as signed evidence; when they dissent, the reading is downgraded rather than defended.

05The record, said back to you

Describe

Step 04's state is the first figure kept with the sitting. Ask for a summary and Delve writes the stored figures back as a paragraph, then answers questions about that sitting. A language model does the writing, handed those figures and nothing else: no new measurement, nothing diagnosed, and asked about anything but your own session it says so and stops.

Forty fields are stored per epoch, all of them available as TXT, CSV or JSON.

Refusal

It says when it doesn't know

Two things stop an epoch before it is given a state, and both are derived rather than picked. Two others used to and no longer do: they lower what may be claimed instead of withholding the reading altogether.

A spectrum with nothing in it

total power ≤ 0 → no-spectrum

An epoch the transform could not turn into band powers has nothing to score. This is arithmetic rather than judgement, and it is checked before any state is looked at.

An empty delta band

δ < 0.010 → delta-empty

0.5–4 Hz is never empty in scalp EEG. Across 140 epochs of real human recording the lowest relative delta observed is 0.108, while 568 of 615 epochs stored by an earlier, broken signal chain report exactly 0.000000. A band that is exactly zero is a broken integration, not a quiet brain.

The rule is deliberately the only one in this gate. “Delta is high” says nothing — real awake EEG at these sites has a median relative delta of 0.605 and reaches 0.92 — so the classifier does not score delta at all. It reads it for this one thing it can honestly say.

Too far from every state to be any of them

distance > 2.17 σ per coordinate → no-confident-reading

Each state is a Gaussian in five coordinates, so an epoch can be asked how far it sits from the nearest one in units of that state’s own spread. Past a point, the honest answer is that it is not that state.

The threshold is derived, not chosen. It is the 99.9th percentile of the scoring statistic under the classifier’s own model, so at most one genuine epoch in a thousand is declined on distance — a false-refusal rate stated in advance rather than discovered afterwards. Change one of the five weights and the threshold moves with it.

A broken electrode: no longer a refusal

Every channel is still screened on its own — for a flat trace, for saturation against the amplifier’s limit, for a frozen stream, for a slow drift that has swamped everything else, and for an amplitude no scalp produces. None of that detection changed. What changed is what is done with it.

The verdict, the list of failing checks and the signal-quality figure are all still computed, stored and shown. They now discount the confidence rather than replace the reading, because the screen is per-channel and the refusal was per-epoch, and because a caveat is actionable where a blank panel is not. On four channels a clean epoch is untouched, one dead electrode costs about a tenth of the claim, and a wholly unusable epoch costs four tenths.

Muscle in the gamma band: no longer a refusal

30–50 Hz power sits on top of the frontalis and temporalis muscles, whose own peak falls inside the band, so high gamma is usually a clenched jaw rather than cortex. That physiology is real. The threshold that acted on it was not: it was the 99th percentile of a laboratory recording made with wet electrodes and a clinical amplifier, and this product runs on a dry band worn on a real head, where the band is never clean at any level.

Published comparisons of consumer and research hardware do not even agree on the direction of the gamma bias. A number that unreliable may lower a confidence; it should not be allowed to withhold a reading, which is the strongest thing this system can say. Elevated gamma now discounts every state instead, on a ramp anchored to the profile table’s own measured means.

The flag this replaced, and what it was really measuring.

What may be claimed, at best

0.84 at a perfect fit → 0.50 at the refusal distance

Accepted epochs carry a ceiling. At a perfect fit it is 0.84, and it falls to 0.50 at the distance where the reading would be refused outright — the point where, by construction, “this state” and “nothing I recognise” are a coin flip.

The 0.84 is borrowed rather than measured, and that is stated because it matters. This judgement has no ground truth, so its accuracy cannot be tested; what can be quoted is how well trained experts agree with each other on the nearest comparable task, scoring graded drowsiness from EEG. Their agreement on the easiest cases in that literature works out at 0.84 across five categories, and it collapses to chance in the grey zone. Four seconds of four-channel dry-electrode EEG is not entitled to claim more than the best two humans manage on an easier question.

Niruddha is capped at 0.50 whatever the fit, by design rather than by measurement. It is the deepest state in the vocabulary and this classifier is stateless — it cannot see whether the neighbouring epochs agree — so a single four-second window is not allowed to claim it confidently. Earning the rest of that claim needs a session-level layer, and that layer does not exist yet.

Neither discount ever reaches zero. Both bottom out at the ratio of the two anchors above, 0.50 / 0.84, because an epoch read off nothing usable and an epoch sitting exactly on the refusal boundary are the same statement — the evidence has run out — and it would be arbitrary for them to disagree.

The withheld probability stays withheld

What the ceiling and the discounts take is left visibly unclaimed rather than redistributed, so the five states on screen do not sum to one. A sitting has long stretches where the signal genuinely does not separate two states, and naming one anyway would be the wrong kind of confident. A system that always has an answer is not more capable, only less informative about which answers to believe.

An earlier version of this argument, published here and elsewhere on the site, said the reading came back indeterminate when the margin between the top two states fell below 0.10. That rule is no longer the refusal path. It survives as a display flag and nothing more, and it is recorded as a correction rather than quietly dropped.

Where the distance threshold comes from, in full — and what the classifier this replaced got wrong.

Validation

Checked against things that could disagree

Two checks, neither of them an accuracy figure. One asks whether three separate implementations compute the same thing. The other asks whether the pipeline finds the plainest effect in the field on real human recordings nobody here made.

  • 3Independent implementationsA Python reference, the C# backend and the JavaScript that runs in the browser. The specification is the authority; when they disagree, all three are wrong.
  • 39/39Golden cases in agreementGenerated from the Python reference, asserted against C# to twelve decimal places and against the browser by a separate harness.
  • 0.000e+00Max |Δ confidence|Bit-identical rather than merely within tolerance. The harness tolerance is 1e-12 and nothing needs it.

Three implementations, one answer

The same classifier is written three times over, in three languages, for three different runtimes. A generated fixture holds all of them to the reference implementation’s output case by case, and they currently agree on every case with a maximum confidence difference of 0.000e+00 — not close, identical.

That number is worth what it is because of what it replaced. Before the harness existed the C# and JavaScript classifiers agreed on 52.3% of the stored epochs, with mean confidences of 0.895 and 0.453, and nothing in the build noticed for months.

And the arithmetic underneath it

Below the classifier, the signal processing answers to a NumPy / SciPy reference oracle. The zero-phase Butterworth band-pass, the Welch PSD, the six band powers, frontal alpha asymmetry, the Hilbert phase-locking value and the complexity stack each have a counterpart written against the scientific-Python implementations, and production code has to reproduce the oracle rather than something plausible. If a change moves a value, the tests go red.

That is a narrower claim than an accuracy figure and much harder to fake: the arithmetic between the electrode and the band power is not something anyone has to take on trust.

Both checks in full, and the signal-processing argument with the filter response and the golden-test plot.

Real human EEG

Close your eyes and alpha rises. The pipeline finds it.

  • 3.16×Eyes-closed alpha over eyes-openRelative alpha rises from 0.0496 to 0.1569 on eye closure, a difference of +0.1072.
  • 1.70Cohen's dA large effect, on 140 epochs — 70 per condition — from a public dataset this project did not produce.
  • 5 / 5Subjects it holds onPer-subject ratios run from 2.06× to 4.83×, so it is not one person's anatomy.

The fixture

PhysioNet eegmmidb v1.0.0

One minute of eyes-open and one minute of eyes-closed resting EEG from five subjects in a public motor-imagery dataset recorded in 2009. Nobody on this project produced that data and nobody on this project can tune it. It is cut into 140 four-second epochs and run through the shipped feature extractor unmodified.

The dataset has no temporal electrodes at the headband’s own positions, so the two nearest available sites were substituted by measured scalp distance. That choice was made in the conservative direction on purpose: the sites actually used give 3.16×, a pair 9 mm further away would have given 4.41×, and going occipital — where the alpha rhythm is largest — would have given 7.42×. The fixture takes the smallest of the three numbers available to it.

What this proves, and what it does not

It proves that given real human EEG at headband-equivalent positions, this pipeline recovers the largest and most reliable effect in the field, in the right direction, at better than three to one, on every subject — and that it does not report that alpha as delta or call an alpha-rich epoch dullness. That is a floor, not a ceiling.

It is not meditation data. Those five people were not meditators and were not practising anything, so the fixture validates none of the five states individually or as a set, none of the profile means the classifier scores against, and nothing about what a practised meditator’s spectrum looks like. There is no ground-truth labelled data for that task anywhere in this project, which is why no accuracy figure appears on this page.

The taxonomy

A state, not a score

Patañjali’s Yoga Sūtras name five grounds of the mind, the chitta bhūmi. Four of them form a ladder from restless to restrained. The fifth, Mūḍha, is dullness, and it sits off that ladder rather than at the foot of it.

01

Kṣipta

thrown, scattered

Surface

Restless and flung outward. Attention lands, slips, and goes.

02

Vikṣipta

distracted, but partly gathered

Emerging

Attention returns more often than it leaves. Most of a practice life is spent here.

03

Ekāgra

one-pointed

Deep

A single object holds, and returning to it stops being work.

04

Niruddha

restrained, stilled

Profound

The movements of the mind have quieted.

off

Mūḍha

dull, inert

Deep Inertia

From outside it can pass for depth. Alpha collapsed rather than merely quiet, and complexity low enough to resemble drowsiness rather than absorption.

Kṣipta shallow, Niruddha deep

The bhūmi is not the only reading. Delve also reports the guṇa triad: sattva (clarity), rajas (activity) and tamas (inertia), including triguṇa-sāmya, where none of the three dominates.

Unlike the bhūmi, the triad rests on no published EEG profile. A systematic search of the meditation-neuroimaging literature turns up no peer-reviewed primary research quantifying sattva, rajas or tamas against band power, and no validated instrument for them that has been correlated with EEG at all. It is reported as the tradition’s reading of a sitting. It is not a measurement of one.

Try it

How are you, right now?

Not in general. Today, in the last hour. Choose the one that is nearest.

Waiting on you

Nothing chosen yet

Pick one above and Delve will name it: the ground it would call that, the balance of the three guṇas that goes with it, and which way the tradition would point from there.

What it will not do is tell you what to practise. That is not modesty, it is the shape of the software: the prescribing screen has not been built, and you have not been measured.

Delve would call this

Kṣipta

thrown, scattered

SurfaceRajasic

rung 1 of 4

Restless and flung outward. Attention lands, slips, and lands somewhere else, and the sitting is mostly noticing it has left. In the guṇa triad this is where rajas predominates: activity rather than clarity or inertia.

Direction of travel

the framework leans toward settling

The classical answer to an over-roused mind is slow, even breathing. The flag Delve raises for high-beta agitation names Nāḍī Śodhana, which this site lists as a settling practice. That is the framework talking, and a teacher deciding.

Delve would call this

Mūḍha

dull, inert

Deep InertiaTamasic

off the ladder

Heavy rather than agitated: clouded instead of scattered, and the difficulty is beginning anything at all. Of the five grounds this is the one that carries the most tamas, which is inertia rather than activity. Mūḍha is the fifth ground, and it sits off the ladder the other four form.

Direction of travel

the framework leans toward waking

The classical answer to a dull, heavy ground is stimulating breath work rather than stillness, and Kapālabhāti and Bhastrikā are listed here as rousing rather than calming. That is the framework talking, and a teacher deciding. Delve itself raises no note here — the flag that used to recommend one fired on most of a healthy adult’s ordinary resting EEG, so it was removed rather than re-tuned.

Delve would call this

Vikṣipta

distracted, but partly gathered

EmergingBalanced

rung 2 of 4

Attention returns more often than it leaves, though it still leaves. Fourteen places at once is the ordinary condition of a working mind, and most of a practice life is spent here. The three guṇas sit near equilibrium, with none of them clearly on top.

Direction of travel

the framework does not lean

Neither, and that is the honest answer rather than a missing one. Nothing here says settle and nothing says wake; on a balanced reading none of the flags that carry a suggestion fire at all. This is the ground where the choice belongs to a teacher rather than to a rule.

Delve would call this

Ekāgra

one-pointed

DeepSattvic

rung 3 of 4

A single object holds without being held, and returning to it stops being work. Sattva predominates: clarity rather than activity or inertia. The corroboration Delve runs on this ground looks for retained complexity, so that genuine stillness is told apart from drowsiness.

Direction of travel

nothing to correct

There is no direction of travel to name. The framework’s own reading of this ground is that it is the good one to be practising from, so the useful response is to carry on sitting.

You tapped a word. That is a self-description, and it is not a reading. Delve reaches these names from EEG, about one reading every four seconds, and you have not worn the band: nothing above has been measured.

Delve does not hand out practice plans. It names the ground and reports the guṇa balance, and a small set of tattva flags can surface a pranayama suggestion, but only ever from measured signal. The prescribing screen in the app is an empty state that says as much. The direction of travel above belongs to the framework and to whoever teaches you, never to a measurement of you.

Svara · नाडी

Where svara is actually read from

Not from your breath, and not from your nostrils. Svara and nāḍī come out of the EEG, from the difference between the two hemispheres’ alpha power and from nothing else.

Frontal alpha asymmetry

faa = ln(alpha right) − ln(alpha left)

Alpha power is averaged across the two right-hand channels and the two left-hand ones, and the reading is the log of their ratio. Below −0.15 Delve reports Iḍā, above +0.15 Piṅgalā, and the band between them Suṣumnā. The axis is a binary; the reading is not.

It is not evidence for the state

The bhūmi classifier does not consume this figure at all. Between- subject variance in frontal alpha asymmetry runs three to ten times the between-condition effect, so on a single session it is noise, and a number that noisy is not allowed to move which state is named. It is reported as svara, and it informs the guṇa balance, and that is its whole remit.

And the correspondence is the tradition’s, not a finding

Everything above is about where the figure comes from. Whether a left-right alpha difference is the thing the texts call Iḍā and Piṅgalā is a separate question, and the honest answer is that the chain it rests on has not replicated. Frohlich et al. 2025 (PLoS ONE, high-density EEG, N = 19) found no significant lateralisation between left and right unilateral nostril breathing: t = 1.00, p = 0.33. Earlier work reporting the coupling has failed replication elsewhere, ipsilaterally or not at all.

Delve keeps reporting svara because it is the vocabulary the practice is taught in, and because the figure underneath it is a real measurement of something. It is not offered as evidence that the tradition’s account of that something is correct.

Zero is not balance

The failure mode worth stating: when both hemispheres fall under the extractor’s floor, the log ratio comes out at exactly zero — which looks like perfect equilibrium and is in fact an absent measurement. It travels with a validity flag for that reason, and the readings that depend on it are gated on the flag rather than on the value.

The record

Thirty-seven fields, every four seconds

Band power is not all four seconds of signal will tell you. Everything the pipeline computes for an epoch is kept with the sitting and can be taken out of the software whole.

What one epoch carries

Exactly thirty-seven fields per epoch in the export, from the state and its confidence down to the raw band shares and the sensors underneath them. Nothing here is a summary of something else: these are the numbers the reading was computed from.

epochNumelapsedSecondsrecordedAtdataQualitychittaBhumichittaConfidencecontemplativeDepthswaraswaraConfidenceswaraNotesattvarajastamasgunaLabeldeltathetaalphalowBetahighBetabetagammafaafaaValidplvplvValidplvInformativevrittiIndexnirodhaStatelzivhiguchiFdsampleEntropypermEntropyaperiodicExponentaperiodicOffsettattvaFlagsheartRatestillnessScore

Structure and connectedness

Several of those fields are measures of how organised the four seconds were rather than how much power sat in a band. They are what the corroboration layer reads when it agrees or dissents with the state above it.

Lempel-Ziv complexityHiguchi fractal dimensionsample entropypermutation entropyaperiodic 1/f fit · 2–40 HzwSMI connectivityindividual alpha frequency

Out of the software, in three formats

TXT · CSV · JSON

TXT is a narrative report laid out for hand-annotation, so a teacher can mark up a sitting the way they always have. CSV and JSON carry the same source data for anyone who would rather plot it themselves.

A field that was not measured is written as an empty cell rather than a zero, which is the difference between “no reading” and “a reading of nothing” — and the difference between a plot with a gap in it and a plot with a lie in it.

The product

A teacher’s instrument, not a wellness dashboard

Built for someone running a cohort: watching a sitting unfold in real time, then reading it back without flattering it.

delve · live monitor
Delve's live monitor, running in the app's own Demo mode: Chitta Bhūmi reading Vikshipta at 34% confidence, the full state-probability spread, Svara Nāḍī balance, a 'what the signals say' corroboration panel, spectral band powers, the triguṇa triad, inner texture, heart rate and physical stillness. The waveform is chipped 'simulated — not a real signal'.

01 / Live reading

Every four seconds a fresh epoch lands: the contemplative state, how confident the classifier is, and underneath that, whether the Western neuromarkers agree. In demo mode the waveform says, on its face, that it is a demonstration and that nothing is being recorded.

What is not claimed

No accuracy figure appears anywhere on this page, and that is not modesty. There is no ground-truth labelled data for this task anywhere in this project, so every number above is a property of the design or a measurement of behaviour, and none of them is a claim about being right.

Four channels give a whole-head average, not a per-electrode map. The deepest three states are close to unresolvable from one another on this hardware, and the classifier reports that as low confidence rather than as a confident guess. The profile table the states are scored against comes from published work on 103 people; the dullness row in it is extrapolated rather than measured, and it is marked as such in the specification.

Delve is a research and education tool. It is not a medical device and is not intended to diagnose, treat, or monitor any condition. Mental-state labels are estimates from EEG signal analysis, not clinical assessments, and their accuracy varies with the individual, the session and electrode contact. If you have concerns about your mental health, speak to a qualified healthcare professional.