Pipeline details¶
This section walks through each stage of the ffrprep pipeline
— BIDS validation, preprocessing, and analysis — covering the
nodes that make up each Nipype workflow and the file layout each
stage writes to disk.
Pipeline Overview¶
The ffrprep pipeline consists of three main stages:
BIDS Validation - Ensures dataset compliance with BIDS standards
Preprocessing - Filters, re-references, and epochs the EEG data
Analysis - Computes evoked responses, time-frequency representations, and FFR metrics
Each stage is implemented as a modular Nipype workflow. The CLI parallelizes
across (task, run) iterations within a subject via a ProcessPoolExecutor
sized by --n_procs; for cross-subject scaling on clusters, run one CLI
invocation per subject (e.g. via slurm job arrays). See the
Parallelization section under Usage for details.
Stage 1: BIDS Validation¶
The first stage validates that your input dataset follows the Brain Imaging Data Structure (BIDS) specification. This ensures reproducibility and compatibility with other neuroimaging tools.
Purpose: Verify dataset structure, file naming conventions, and required metadata files before processing begins.
Implementation:
The validation uses the bids-validator tool with a custom configuration that ignores warnings not relevant to EEG/FFR data.
Key Functions: - validate_input_dir() - Main validation function - Custom validator configuration for EEG-specific requirements
Validation Steps:
Directory Structure Check
Verifies presence of required BIDS directories (
sub-*/,derivatives/)Checks for
dataset_description.jsonand other required metadata filesValidates subject/session/task naming conventions
EEG-Specific Validation
Confirms presence of EEG data files (
.edf,.bdf,.vhdr,.fif,.set)Validates channel description files (
*_channels.tsv)Checks event files (
*_events.tsv) for proper formattingVerifies EEG-specific metadata in JSON sidecars
Participant Selection
Validates requested participant labels exist in dataset
Checks for required EEG data for specified participants
Reports any missing or incomplete data
Error Handling:
If validation fails, ffrprep provides detailed error messages indicating specific BIDS compliance issues and suggestions for resolution.
Skip Option:
Validation can be bypassed using --skip_bids_validation (not recommended for production analyses).
Stage 2: Preprocessing¶
The preprocessing stage converts raw EEG data into clean, epoched data suitable for FFR analysis. This stage implements standard electrophysiological preprocessing steps optimized for frequency-following responses.
Purpose: Transform raw continuous EEG into clean, filtered, and epoched data while preserving FFR-relevant neural signals.
Implementation: Implemented as a Nipype workflow (create_preprocessing_workflow()) with the following nodes:
Preprocessing Workflow Nodes¶
1. Data Loading Node
Function: load_data()
Purpose: Robustly load EEG data from BIDS datasets using pybids and MNE-BIDS.
- Sub-steps:
Create
BIDSLayoutobject for dataset queryingQuery for EEG files matching participant/session/task/run criteria
Try multiple file extensions (
.edf,.bdf,.vhdr,.fif,.set)Load data using
mne_bids.read_raw_bids()Extract and validate channel information
Load associated event data and metadata
Outputs: Raw EEG data object, BIDS path information, original filename
2. Re-referencing Node
Function: reference_data()
Purpose: Apply appropriate reference scheme to reduce common-mode noise and artifacts.
- Sub-steps:
Parse reference channel specification (average, single channel, or channel list)
Validate reference channels exist in data
Apply re-referencing using MNE’s
set_eeg_reference()Update channel information and provenance
- Reference Options:
Average reference (
--ref_channels average): Uses all EEG channelsSingle channel (
--ref_channels Cz): References to one electrodeMultiple channels (
--ref_channels "Cz,Fz"): Average of specified channels
Outputs: Re-referenced EEG data
3. Filtering Node
Function: filter_data()
Purpose: Apply temporal filtering to remove noise while preserving FFR signals.
- Sub-steps:
Apply high-pass filter to remove slow drifts and DC offsets
Apply low-pass filter to remove high-frequency noise
Use zero-phase FIR filters to avoid temporal distortions
Log filter parameters and transition bands
- Default Parameters:
High-pass: 70.0 Hz (removes slow drifts and low-frequency noise, preserves FFR frequencies)
Low-pass: 1000.0 Hz (preserves FFR harmonics while removing high-frequency noise)
Filter design: Zero-phase FIR with automatic transition bandwidth.
--filter-method iirswitches to zero-phase first-order Butterworth high-pass + low-pass passes (12 dB/octave overall).
Outputs: Filtered EEG data
4. Epoching Node
Function: epoch_data()
Purpose: Segment the continuously-filtered EEG into time-locked epochs around stimulus events, applying baseline correction and amplitude-based rejection.
- Sub-steps:
Load events from the BIDS
*_events.tsv(or use annotations embedded in the raw recording when no sidecar is present)Build
mne.Epochswith the requestedtmin/tmax, baseline window, picks, andrejectthresholdsApply
epochs.drop_bad()to materialize amplitude-based rejection; keep the resulting Epochs object as the workflow output
- Default Parameters (FFR-typical, all overridable from the CLI):
Epoch window:
--tmin -0.2to--tmax 0.6seconds around stimulus onsetBaseline:
--baseline -0.2 0seconds (pre-stimulus)Rejection:
--reject-eeg 75e-6(75 µV peak-to-peak); pass--no-auto-rejectto disable, or--reject-mode absto drop epochs whose absolute amplitude reaches the threshold at any sample
Outputs: Epoched EEG data and the post-rejection drop log.
5. Save Preprocessing Outputs Node
Function: save_preprocessing_outputs()
(invoked through the internal save_preprocessing_node wrapper,
which fans the Epochs out by trial type before calling it).
Purpose: Persist the epoched data plus a self-describing BIDS
sidecar — one file per trial type by default, or a single
combined file under --no-split-by-trial-type.
- Sub-steps:
When
--split-by-trial-typeis on (the default), partition the input Epochs byevent_idand write one_desc-preproc{Cond}_epo.fifper trial type with a matching sidecar carrying aConditionfield. Under--no-split-by-trial-type, write a single bare_desc-preproc_epo.fifinstead.Reconstruct a fresh
EpochsArrayfrom the data + events +event_idso trial-type metadata survives the save / load round-trip (and the upstream Epochs object’s internal state doesn’t leak through.save()).Write each sibling
.jsonsidecar withEpochCount/EpochCountTotal/EpochCountRejected/RejectionMode/RejectionThresholds/Filtering/SamplingFrequency/EpochTmin/EpochTmax/Channelsplus run / session /ConcatenatedRuns/Conditionprovenance.SourcesandRawSourceslist the raw EEG file(s) the epochs came from, when they can be found.Initialize the per-derivatives
dataset_description.jsonif missing.
Reporting runs after the workflow drains, in the CLI rather than as a workflow node — see reporting below.
Output Structure (default split-by-trial-type):
derivatives/ffrprep-preprocessing/
├── sub-XX/
│ └── eeg/
│ ├── sub-XX_task-YY_run-ZZ_desc-preprocPositive_epo.fif
│ ├── sub-XX_task-YY_run-ZZ_desc-preprocPositive_epo.json
│ ├── sub-XX_task-YY_run-ZZ_desc-preprocNegative_epo.fif
│ ├── sub-XX_task-YY_run-ZZ_desc-preprocNegative_epo.json
│ ├── sub-XX_preprocessing_report.html
│ └── sub-XX_preprocessing.log
Under --no-split-by-trial-type the per-trial-type files are
replaced by a single _desc-preproc_epo.fif + its sidecar.
The sidecar JSON carries provenance, EpochCount /
EpochCountTotal / EpochCountRejected, RejectionMode
(peak-to-peak or absolute-amplitude) with
RejectionThresholds, Filtering (high-pass and low-pass
cut-offs), sampling frequency, run / session identifiers, the raw
file(s) in Sources / RawSources, and Condition (for
per-trial-type files only). The epoch counts cover only the trial
types being written: in a per-trial-type file they are that trial
type’s counts, and events outside event_id (for example markers)
are not counted. Epochs dropped by --reject-mode abs count as
rejected.
Outputs: File paths, processing metadata.
Stage 3: Analysis¶
The analysis stage averages each preprocessed epochs group into a structured set of evoked responses (per-trial-type + combined + optional difference) and saves them to BIDS-derivatives.
Purpose: Produce per-(task, run) evoked responses suitable for downstream statistical analysis or visualization, persisted in MNE-readable format with self-describing BIDS sidecars.
Implementation:
Implemented as a Nipype workflow (create_analysis_workflow())
with two nodes. The CLI worker collects all per-trial-type
_desc-preproc{Cond}_epo.fif files for one (task, run) group
via _collect_analysis_groups, stitches them back together with
mne.concatenate_epochs (which preserves event_id), and
passes the resulting Epochs directly into the workflow’s
inputnode. The per-(task, run) granularity means each group
runs the workflow once regardless of how many trial types it
holds.
Analysis Workflow Nodes¶
1. Build Analysis Payload Node
Function: build_analysis_payload()
Purpose: From a single Epochs object, produce a structured
{by_type, combined, diff} payload covering every evoked the
analysis stage emits.
- Sub-steps:
Average all events into a single combined Evoked via
make_combined_evoked(always emitted).When
split_by_trial_typeis True (the default), partition byevent_idviamake_evoked(by_event_type=True)into a per-trial-type dict.When at least two trial types are present and either (a) there are exactly two types (auto-paired) or (b) the user passed
--difference-pairs A:B [C:D …], compute one difference Evoked per pair viamake_difference_evokeds(which wrapsmne.combine_evoked([A, B], weights=[1, -1])).
Outputs: dict with keys "by_type" (dict trial_type → Evoked),
"combined" (Evoked), and optionally "diff" (dict (A, B) →
Evoked).
2. Save Analysis Node
Function: save_analysis_outputs()
Purpose: Persist every Evoked in the payload plus a self-describing sidecar per file.
- Sub-steps:
For each entry in
by_type: write_desc-evoked{Cond}.fif+ matching sidecar withCondition: <cond>.For
combined: write the bare_desc-evoked.fif+ sidecar (noConditionfield).For each entry in
diff: write_desc-evokedDiff{A}Vs{B}.fif+ sidecar carryingDifferenceOf: [A, B].Every sidecar also carries
AverageCount/Baseline/SamplingFrequency/Tmin/Tmax/Channels/TaskName/AnalysisTypeplus run / session /ConcatenatedRunsprovenance.Initialize the per-derivatives
dataset_description.jsonif missing.
Output Structure (default split-by-trial-type, 2-trial-type dataset):
derivatives/ffrprep-analysis/
├── sub-XX/
│ ├── sub-XX_task-YY_run-ZZ_desc-evokedPositive.fif
│ ├── sub-XX_task-YY_run-ZZ_desc-evokedPositive.json
│ ├── sub-XX_task-YY_run-ZZ_desc-evokedNegative.fif
│ ├── sub-XX_task-YY_run-ZZ_desc-evokedNegative.json
│ ├── sub-XX_task-YY_run-ZZ_desc-evoked.fif
│ ├── sub-XX_task-YY_run-ZZ_desc-evoked.json
│ ├── sub-XX_task-YY_run-ZZ_desc-evokedDiffPositiveVsNegative.fif
│ ├── sub-XX_task-YY_run-ZZ_desc-evokedDiffPositiveVsNegative.json
│ ├── sub-XX_analysis_report.html
│ └── sub-XX_analysis.log
Under --no-split-by-trial-type, only the combined
_desc-evoked.fif + sidecar are emitted per (task, run).
Outputs: Saved file paths.
Reporting (post-workflow)¶
The single-file HTML reports are built in the CLI after the
workflow drains, not as a workflow node. The CLI’s
_build_preproc_report and _build_analysis_report glob the
saved _desc-preproc*_epo.fif and _desc-evoked*.fif files
via _collect_analysis_groups / _collect_evoked_groups,
group them by (task, run), and build sections via the
ffrprep.reports builders (build_raw_section /
build_epoch_section / build_evoked_section /
build_phase_consistency_section / make_group); the
subject-level HTML is rendered via build_subject_report /
build_analysis_report.
Per (task, run) group layout in the analysis report:
one Evoked section per per-trial-type file (e.g. Positive, Negative). Each carries:
waveform / PSD / TFR / autocorrelation / pitch-track figures (via
ffrprep.reports.evoked_qa());scalar metrics:
RMS SNR (<window>),Mean power 90-110 Hz, <window>, where<window>defaults to 100-200 ms and is configurable via--response-window START END(seconds);when the BIDS
stim_filecolumn is populated inevents.tsv, aStim correlation (peak r)+Stim correlation (lag, ms)row plus a stim ↔ response cross-correlation lag plot (computed by_stim_correlation_datainffrprep_cli.py).
one combined Evoked section (the across-events average) with the same plots + scalars, plus a second pair of rows / plot for the envelope correlation (combined ≈ ENV proxy in FFR, so
|hilbert(stim)|is the natural reference).one difference Evoked section per
(A, B)pair (auto for the 2-trial-type case; opt-in via--difference-pairs) with the standard plots + a single raw-waveform stim correlation row + plot (diff ≈ TFS proxy).when exactly two per-condition preproc files exist for the group, one Phase Consistency section combining
ffrprep.analysis.compute_phase_consistency()withffrprep.analysis.plot_phase_consistency()(or, when the caller passesmask=True, the masked variant). Single-subject reports default to unmasked; group-level builders passmask=True. Uses seaborn’sflare_rcolormap; subplot titles surface the trial-type names plus auto-derivedA + B/A − Bfor the sum and difference panels.
Per (task, run) group layout in the preprocessing report:
one Raw section loaded from the original BIDS recording (with run concatenation when the preproc output sourced multiple runs).
one Epoched section per per-trial-type file — each carries the standard epoch metadata plus a
Mean trial-to-trial rrow (viaffrprep.analysis.response_consistency()) when there are at least 10 trials.
Figures and scalar metrics are computed at report time from
the loaded Epochs / Evoked objects — they are not
separately persisted to disk. To recompute them yourself, see the
Working with Outputs in Python section of the
Tutorial walkthrough.
The single-file *_report.html embeds all figures as inline
base64 PNGs — there is no sibling figures/ directory.
- File Formats:
MNE format (.fif): Epoched / Evoked data, loadable in MNE-Python.
WAV (under
<dataset>/stimuli/): stimulus audio referenced byevents.tsv’sstim_filecolumn. Fetched viaffrprep-download example --with-stimuli.JSON: BIDS sidecar (human-readable, machine-parseable).
HTML: Self-contained single-file per-subject report.
Pipeline Integration and Quality Control¶
Workflow Management:
Each stage implemented as a Nipype workflow for per-iteration dependency tracking
Outer-loop parallelism: the CLI dispatches per-(task, run) iterations to a
ProcessPoolExecutorsized by--n_procsFail-fast on any iteration error; the exception propagates to the CLI entry point
nipype caches per-iteration intermediates under
work/so re-runs that already have a saved_desc-preproc_epo.fifskip the workflow re-execution--skip-existingskips a (task, run) iteration whose sidecar outputs already exist, so an interrupted run can be repeated with the same command. The check uses the.jsonsidecars rather than the epochs files, so it also works after--no-keep-epochs. An analysis iteration counts as done when its_desc-evoked.jsonexists, and no report is rebuilt for iterations that are skipped.--clean-work-dirdeletes an iteration’s Nipype working files once it has succeeded. Log files are kept, and the files of a failed iteration are left in place.--no-keep-epochsdeletes the_epo.fiffiles after the analysis stage and its report have finished. The JSON sidecars stay. With--stage preprocessingthe flag is ignored and a warning is printed, because the analysis stage still needs the epochs.
Quality Control Checkpoints:
BIDS validation before processing
Data quality assessment after loading
Preprocessing quality metrics and reports
Analysis validation and statistical checks
Output Organization:
BIDS-compatible directory structure
Per-file BIDS sidecars with run / session / condition provenance
Standardized file formats for interoperability
Version-controlled processing parameters
Customization:
All preprocessing and analysis parameters are exposed as CLI flags (see Usage).
The Nipype workflows can be imported and reused programmatically from
ffrprep.preproc(create_preprocessing_workflow/create_analysis_workflow).Section builders in
ffrprep.reportsacceptextra_summaryandextra_figureskwargs so downstream code can fold caller-computed scalars or figures into a section’s table or figure gallery.
Group-Level Aggregation¶
The group analysis level (ffrprep.group.run_group_level)
aggregates participant-level derivatives already written under
output_dir — it does not read raw BIDS data. For each (task,
session, run) with derivatives from at least two subjects, it:
discovers each subject’s combined (
_desc-evoked.fif) and, when present, difference (_desc-evokedDiff{A}Vs{B}.fif) Evoked files viaffrprep.group.discover_group_inputs(), and computes a grand average withmne.grand_average(ffrprep.group.compute_grand_average()). If the subjects’ single channel has different names (for exampleA32andCzat different sites), it is renamed to a common name first. Multi-channel sets that differ raise an error;recomputes the same scalar FFR metrics the participant-level report shows (RMS SNR, band power) directly from each subject’s saved combined Evoked, plus trial-to-trial response consistency from that subject’s saved preprocessing Epochs (
ffrprep.group.compute_subject_metrics()) — nothing is recomputed from raw data;adds further columns when asked.
--f0addsrms_snr_polarity_sum,f0_uvandupper_harmonics_uv;--stimulusaddsstim2resp_r,stim2resp_zandstim2resp_lag_ms(plusstim2resp_lim_*with--xcorr-lag-range);--n-trials-presentedaddsusable_pct. These are computed from the sum of the two per-trial-type evoked files.f0_uvandupper_harmonics_uvare FFT amplitudes averaged in a band around each harmonic of F0 (ffrprep.analysis.harmonic_amplitudes()), andstim2resp_ris the maximum normalized cross-correlation between the stimulus and the response (ffrprep.analysis.stim_to_resp_xcorr());joins subject-level covariates onto the table by
participant_id, from<bids_dir>/participants.tsvand any--covariatesfiles (ffrprep.group.merge_covariates()). A column already in the table is not overwritten;with
--min-usable-pctand/or--min-snr, addsqc_*columns that mark recordings below a threshold (ffrprep.group.add_qc_flags()). Rows are flagged, not removed, and the grand averages still include every subject;writes the grand-average Evoked(s), a per-subject metrics TSV with a
_metrics.jsondata dictionary next to it (ffrprep.group.build_metrics_dictionary()), and a singlegroup_report.html(reusing the same section builders and template as the participant-level reports) underoutput_dir/ffrprep-group/.
Output Structure:
derivatives/ffrprep-group/
├── dataset_description.json
├── task-YY_run-ZZ_desc-grandAverage_ave.fif
├── task-YY_run-ZZ_desc-grandAverage_ave.json
├── task-YY_run-ZZ_desc-grandAverageDiffPositiveVsNegative_ave.fif
├── task-YY_run-ZZ_desc-grandAverageDiffPositiveVsNegative_ave.json
├── task-YY_run-ZZ_metrics.tsv
├── task-YY_run-ZZ_metrics.json
└── group_report.html
This step is deliberately scoped to aggregation, not inference: it summarizes what participant-level ffrprep already computed and performs no group-level statistics (no hypothesis tests, no GLM). Statistical analysis is left to the user, e.g. directly in MNE-Python against the saved grand-average / per-subject derivatives.