# EDF+ annotations and gaps

> How edf2csv reads the EDF+ annotations channel, exports events, and preserves the real timing of discontinuous recordings

edf2csv 0.9.51. Canonical HTML version: https://edf2csv.vercel.app/docs/edf-plus-annotations

Plain EDF stores nothing but numbers. EDF+ adds a way to store text alongside the
signals: sleep stages, stimulus markers, technician notes, seizure onsets, and the
timestamps that tell you where in the recording each data record actually sits.
edf2csv reads that channel in full and writes it to `annotations.csv`, and it uses
the same information to make sure the `time_s` column in `signals.csv` reflects
when the data was really measured.

## The annotations channel

An EDF+ file declares one or more channels with the reserved label
`EDF Annotations`. BDF+ files, produced by BioSemi hardware, spell the same thing
`BDF Annotations`. edf2csv recognises both.

These channels occupy space in every data record exactly like a signal does, but
the bytes are UTF-8 text rather than samples — where the writer managed UTF-8. Where
it did not, they are read as latin1 instead, one character per byte, which is how
this parser reads every other text field in the file. Decoding them strictly as
UTF-8 and taking what comes back means a description written `café` by an older
recorder arrives as `caf�`: a character the file does not contain, in a text
column. Bytes that decode as UTF-8 are UTF-8, including a replacement character the
file really holds; bytes that do not decode are not UTF-8, and nothing is invented
either way. That's why the annotations channel
never appears in the channel table, never gets a column in `signals.csv`, and is
never listed in `channels.csv`. `--info` reports it separately:

```bash
edf2csv sleep-study.edf --info
```

```
Channels   5 signals + 1 annotation channel
```

The text is stored as a run of Time-stamped Annotation Lists (TALs). Each TAL
begins with a signed onset, optionally followed by a duration, then zero or more
text strings, and is terminated by a NUL byte. The separators are control
characters: `0x15` between onset and duration, `0x14` before and after each text
string. Written out with the control bytes made visible, one TAL looks like this:

```
+1.25<0x15>0.5<0x14>Seizure onset<0x14><0x00>
```

That says: at 1.25 seconds from the start of the recording, an event lasting 0.5
seconds, described as `Seizure onset`.

## What an annotation carries

Three things, and only three things:

| Field | Meaning |
| --- | --- |
| onset | Seconds from the start of the recording. Always present. |
| duration | Seconds. Optional. A TAL may omit it entirely. |
| text | The description. A single TAL may carry several. |

`annotations.csv` has one row per annotation, with a fourth column recording
which data record the annotation was physically stored in:

```
onset_s,duration_s,description,record_index
0.5,1,Sleep stage W,0
1.25,,Lights off,1
2,0.5,Seizure onset,2
```

An omitted duration is written as an empty cell, not as `0`. The two mean different
things: a duration of zero states that the event was instantaneous, while an empty
cell records that no duration could be read. In pandas the empty cell reads back as
`NaN`, which is the right value for something that was never recorded.

The empty cell covers two cases the file distinguishes and the column cannot: a TAL
that stated no duration, and one that stated something which is not a number. The
event is kept either way, and the second raises an `ANNOTATION_DECODE_FAILED`
warning saying how many rows it happened to — the cell itself cannot tell them
apart, so the warning is the only place the difference survives.

Descriptions are escaped by normal CSV rules, so an annotation containing a comma,
a quote, or a newline is quoted rather than allowed to break the column count.

A single TAL can carry several text strings sharing one onset and duration. Each
becomes its own row, all with the same `onset_s` and `duration_s`.

## EDF+C and EDF+D

Bytes 192 to 235 of the header hold a reserved field. EDF+ files write one of two
markers there:

- **EDF+C**, continuous. Every data record follows the one before it with no gap.
  Record `n` starts at `origin + n * record_duration` seconds, where the origin is
  what the first record's timekeeping TAL says — usually `+0`, and not always. A
  continuous recording whose first TAL reads `+0.5` writes its first row as `0.500`
  and its records still sit end to end. What continuity fixes is the spacing, not
  the starting point.
- **EDF+D**, discontinuous. The records are still in file order, but they aren't
  adjacent in time. There are holes between them.

BDF+ writes `BDF+C` and `BDF+D` for the same two states. edf2csv normalises both
spellings to one internal marker, so a BioSemi discontinuous file behaves exactly
like an EDF+D one.

Discontinuity isn't an exotic case. It's what you get when an ambulatory
recorder is paused and resumed, when a system drops out and reconnects, when a
long recording is segmented and the segments are concatenated, or when a vendor
tool exports only the annotated epochs of a much longer study.

The important consequence: in an EDF+D file, the arithmetic `record_index *
record_duration` is no longer the record's position in time. It's only a count of
how much data came before. The real position is stored in the first TAL of every
data record, the mandatory timekeeping annotation. That TAL is the only place a
discontinuous file records where its data sits.

It usually carries an onset and nothing else, and it is allowed to carry event text
after the onset as well — writers do. Those events are exported like any other, which
is why an unreadable first TAL costs both the record's position and whatever events
came with it, and why the warning for one counts them separately from the warning for
the other.

`--info` names the format directly:

```
Format     EDF+ (discontinuous)
```

and conversion warns before it starts:

```
warning: This recording is marked discontinuous (EDF+D): its data records need not be contiguous in time.
         Each row carries its true recording time, so gaps stay visible instead
         of being closed.
```

## Gaps stay visible

edf2csv reads the timekeeping TAL of every record and uses it as the base time for
every sample in that record. Sample `i` of a record that declares itself at
`t` seconds is written at `t + i / sampling_rate`. Nothing is inserted to bridge a
gap, and nothing is shifted to close one. A gap therefore appears in the output as
a jump in `time_s` between two consecutive rows.

The test fixture `discontinuous.edf` demonstrates this in miniature. It holds
three one-second records of a single 10 Hz channel, and its timekeeping TALs place
those records at 0 s, 1 s and 10 s. The first two are adjacent; the third sits
after a nine-second hole.

```bash
edf2csv discontinuous.edf --out ./converted
head -1 ./converted/signals.csv
sed -n '20,23p' ./converted/signals.csv
```

If the gap were ignored and the records were laid end to end, the output would run
0.0 s to 2.9 s with nothing to indicate that anything was missing. What edf2csv
actually writes is:

```
time_s,EEG Fpz-Cz
...
1.800,2.259
1.900,2.381
10.000,2.503
10.100,2.625
...
```

That's thirty rows, one per recorded sample, and the jump from `1.900` to `10.000`
is the gap. Anything that computes a sampling interval by differencing `time_s`
sees the discontinuity rather than a smooth ramp that misplaces every sample after
the hole by nine seconds.

This has three practical consequences.

**The recording's span isn't the amount of data it holds.** This file contains
three seconds of samples but covers eleven seconds of wall time. Time windows are
resolved against the real span, so a window can legitimately reach past where the
data would have ended if it were contiguous:

```bash
edf2csv discontinuous.edf --start 5 --out ./converted
```

That isn't an error, and the record at 10 s isn't clipped away. It writes the ten
samples from `10.000` to `10.900`. The same window against a naively flattened
version of this file would have been rejected as starting past the end.

**`--info` reports both the amount of data and the span.** For this file it prints
`Duration 3s (3 records of 1s)`, because that's how many seconds of signal exist, and
`Time span 11s (includes discontinuities)` on the line under it, because that's how
long the recording covers. The gap is the difference between them, and is reported
again by the EDF+D warning. A span *shorter* than the duration is the other way the two
can disagree — records that overlap, which is what a device does when it re-sends a
buffer — and the line says `(records overlap in time)` instead, since a recording
covering less time than its own records account for has no gaps in it at all.

**Rows are written in file order.** If a file's timekeeping TALs are themselves out
of order, so that a record claims to start before the one preceding it, edf2csv
writes the rows anyway and warns that `time_s` won't increase monotonically. It
doesn't silently reorder your data.

## The whole annotation channel is always read

When you request a time window, edf2csv still scans the annotation channel of the
entire file, from the first record to the last.

This is deliberate. Nothing in the EDF+ specification obliges a writer to store an
annotation in the data record whose time span contains its onset. Some tools write
every annotation in the file into the first record. Others batch them. Reading only
the records that fall inside the requested window would silently drop events that
belong in that window but happen to be stored elsewhere, and the resulting
`annotations.csv` would look complete.

The fixture `annotations-front-loaded.edf` is exactly this shape: ten one-second
records, with all three annotations stored in record 0, at onsets 0.5 s, 5.5 s and
8.5 s. Asking for the window from 5 s to 7 s produces the event that belongs there:

```bash
edf2csv annotations-front-loaded.edf --start 5 --duration 2 --out ./converted
cat ./converted/annotations.csv
```

```
onset_s,duration_s,description,record_index
5.5,,middle,0
```

Note `record_index` is `0` while `onset_s` is `5.5`. That column exists precisely
so you can see when a writer has done this.

The scan is cheap. edf2csv seeks straight to the annotation channel's byte range
inside each record rather than reading whole records through memory, so on a
multi-gigabyte recording the cost is a few kilobytes of I/O rather than all of it.

Annotations are filtered to the requested window by their onset, using the
half-open interval `[start, end)`. An annotation whose onset falls inside a gap in
a discontinuous file is still exported, because the window covers the recording's
real span rather than only the parts that contain samples.

## Exporting only the events

Some analyses need the event list and nothing else: building a hypnogram, counting
stimulus markers, checking that a scoring pass covers the whole night. Converting
an eight-hour 256 Hz study to get a few hundred rows of text is wasteful.

```bash
edf2csv sleep-study.edf --annotations-only --out ./events
```

This skips signal conversion entirely. No `signals.csv` is written and no samples
are read. The output directory contains `annotations.csv`, `channels.csv`
describing the channels that weren't converted, and `metadata.json`.

`--start`, `--duration` and `--end` still apply, so you can pull the events from a
single hour.

Asking for annotations from a file that has no annotation channel isn't an error,
but it doesn't pass silently either:

```
warning: --annotations-only was requested but this recording has no annotation channel, so there are no events to export.
         Plain EDF files carry no annotations. Convert without
         --annotations-only to get the signals.
```

A file whose only content is annotations, with no signal channels at all, converts
fine. It produces `annotations.csv` and an empty `channels.csv`, plus a warning
saying the file carries no signal channels.

## When timekeeping is missing or unreadable

The timekeeping TAL is the only record of where a discontinuous file's data sits.
When it's absent or can't be parsed, that position is simply not knowable from
the file.

edf2csv doesn't stop, and it doesn't guess quietly. The affected record is timed
as if it were contiguous with the records around it, at `origin + index *
record_duration` — where `origin` comes from the first record that does state a
time, since a recording need not begin at zero. Writing it as `index *
record_duration` would put the record at the wrong instant on every file whose
first record says anything but `+0`. The substitution is reported by name:

```
warning: 1 of 3 data records carries no readable timekeeping annotation (record 1), so its true position in time is unknown.
         That record is timed as if it were contiguous; treat its timestamp as
         unreliable.
```

Up to eight record indices are listed, and the rest are counted rather than dropped:
`... 8, 9 and 2 more`. The same warning appears in `metadata.json` under `notes`,
and in the `warnings` array of `--json` output, so a scripted pipeline can detect
it without parsing stderr.

Two related cases get their own warnings.

A file marked EDF+D that has no annotation channel at all is self-contradictory: it
claims gaps and provides nowhere to record them. Times are written as if the
records were contiguous, and edf2csv says so plainly, noting that any gaps are
lost.

Individual TALs that can't be decoded are skipped rather than aborting the
conversion, and the count is reported:

```
warning: 2 annotation entries were unreadable and could not be exported.
         Every entry that could be read is exported. The file may have been
         written by a non-conforming tool.
```

One malformed annotation shouldn't cost you an entire conversion, but dropping it
silently would leave you with an event list you had no reason to question.

## How other tools handle this

Discontinuity is where EDF readers differ most, so it's worth knowing what your
existing tooling does.

**pyEDFlib** refuses EDF+D files outright, raising rather than returning data. That
guarantees it never gives you wrong timestamps, but it also means discontinuous
recordings give you nothing at all. Converting with edf2csv first is one way to get
the data into a form you can work with.

**mne.io.read_raw_edf** reads EDF+D files and presents them as continuous. The
gaps are closed. Samples from either side of a hole end up adjacent in the array,
and the returned time vector counts uniformly from zero as though no interruption
occurred. Downstream, every sample after the first gap carries a timestamp that's
wrong by the accumulated gap length, and nothing in the object marks which samples
those are. Annotations are read from the file with their original onsets, so on a
discontinuous recording the event times and the sample times refer to different
clocks.

edf2csv takes a third position: read the file, keep the real times, and report the
structure. If a gap matters for your analysis, deciding what to do about it is
yours to make with the gap in front of you.

## Reading the output

Loading the two files together in pandas is enough to line events up against
signal:

```python
import pandas as pd

signals = pd.read_csv("converted/signals.csv")
events = pd.read_csv("converted/annotations.csv")

# Find gaps: any interval larger than the sampling period.
dt = signals["time_s"].diff()
gaps = signals.loc[dt > dt.median() * 1.5, "time_s"]

# Samples covered by one annotation.
event = events.iloc[0]
end = event.onset_s + (event.duration_s if pd.notna(event.duration_s) else 0)
window = signals[(signals.time_s >= event.onset_s) & (signals.time_s < end)]
```

Because `time_s` carries the true recording time in both files, the comparison is
valid across a gap. `duration_s` is `NaN` wherever the file gave no duration — or gave
one that is not a number, which the run warns about — which is why the check above is
explicit rather than assuming zero. A duration below zero is written out as the file gave
it and warned about too: added to `onset_s` it ends the window before the event starts, so
`end` above would come out less than the onset and select nothing.

## Programmatic access

The annotation decoder is exported, so you can read events without producing CSV
at all:

```javascript
import { EdfFile } from 'edf2csv';

const file = await EdfFile.open('sleep-study.edf');
const { annotations, recordStarts, malformed } = await file.readAnnotations();

for (const a of annotations) {
  console.log(a.onset, a.duration, a.text, a.recordIndex);
}

// recordStarts[i] is the declared start of record i, or null when the
// timekeeping annotation was missing or unreadable.
console.log(recordStarts[0], recordStarts.at(-1), malformed);

await file.close();
```

`readAnnotations` returns annotations sorted by onset, then by record index for
ties. `recordStarts` has one entry per data record actually present in the file.
`decodeRecordAnnotations` is also exported for decoding a single record's
annotation bytes directly.
