Pipeline details

This section walks through each stage of the ffrprep pipeline — BIDS validation, preprocessing, and analysis — covering the nodes that make up each Nipype workflow and the file layout each stage writes to disk.

Pipeline Overview

The ffrprep pipeline consists of three main stages:

  1. BIDS Validation - Ensures dataset compliance with BIDS standards

  2. Preprocessing - Filters, re-references, and epochs the EEG data

  3. Analysis - Computes evoked responses, time-frequency representations, and FFR metrics

Each stage is implemented as a modular Nipype workflow. The CLI parallelizes across (task, run) iterations within a subject via a ProcessPoolExecutor sized by --n_procs; for cross-subject scaling on clusters, run one CLI invocation per subject (e.g. via slurm job arrays). See the Parallelization section under Usage for details.

Stage 1: BIDS Validation

The first stage validates that your input dataset follows the Brain Imaging Data Structure (BIDS) specification. This ensures reproducibility and compatibility with other neuroimaging tools.

Purpose: Verify dataset structure, file naming conventions, and required metadata files before processing begins.

Implementation: The validation uses the bids-validator tool with a custom configuration that ignores warnings not relevant to EEG/FFR data.

Key Functions: - validate_input_dir() - Main validation function - Custom validator configuration for EEG-specific requirements

Validation Steps:

  1. Directory Structure Check

    • Verifies presence of required BIDS directories (sub-*/, derivatives/)

    • Checks for dataset_description.json and other required metadata files

    • Validates subject/session/task naming conventions

  2. EEG-Specific Validation

    • Confirms presence of EEG data files (.edf, .bdf, .vhdr, .fif, .set)

    • Validates channel description files (*_channels.tsv)

    • Checks event files (*_events.tsv) for proper formatting

    • Verifies EEG-specific metadata in JSON sidecars

  3. Participant Selection

    • Validates requested participant labels exist in dataset

    • Checks for required EEG data for specified participants

    • Reports any missing or incomplete data

Error Handling: If validation fails, ffrprep provides detailed error messages indicating specific BIDS compliance issues and suggestions for resolution.

Skip Option: Validation can be bypassed using --skip_bids_validation (not recommended for production analyses).

Stage 2: Preprocessing

The preprocessing stage converts raw EEG data into clean, epoched data suitable for FFR analysis. This stage implements standard electrophysiological preprocessing steps optimized for frequency-following responses.

Purpose: Transform raw continuous EEG into clean, filtered, and epoched data while preserving FFR-relevant neural signals.

Implementation: Implemented as a Nipype workflow (create_preprocessing_workflow()) with the following nodes:

Preprocessing Workflow Nodes

1. Data Loading Node

Function: load_data()

Purpose: Robustly load EEG data from BIDS datasets using pybids and MNE-BIDS.

Sub-steps:
  • Create BIDSLayout object for dataset querying

  • Query for EEG files matching participant/session/task/run criteria

  • Try multiple file extensions (.edf, .bdf, .vhdr, .fif, .set)

  • Load data using mne_bids.read_raw_bids()

  • Extract and validate channel information

  • Load associated event data and metadata

Outputs: Raw EEG data object, BIDS path information, original filename

2. Re-referencing Node

Function: reference_data()

Purpose: Apply appropriate reference scheme to reduce common-mode noise and artifacts.

Sub-steps:
  • Parse reference channel specification (average, single channel, or channel list)

  • Validate reference channels exist in data

  • Apply re-referencing using MNE’s set_eeg_reference()

  • Update channel information and provenance

Reference Options:
  • Average reference (--ref_channels average): Uses all EEG channels

  • Single channel (--ref_channels Cz): References to one electrode

  • Multiple channels (--ref_channels "Cz,Fz"): Average of specified channels

Outputs: Re-referenced EEG data

3. Filtering Node

Function: filter_data()

Purpose: Apply temporal filtering to remove noise while preserving FFR signals.

Sub-steps:
  • Apply high-pass filter to remove slow drifts and DC offsets

  • Apply low-pass filter to remove high-frequency noise

  • Use zero-phase FIR filters to avoid temporal distortions

  • Log filter parameters and transition bands

Default Parameters:
  • High-pass: 70.0 Hz (removes slow drifts and low-frequency noise, preserves FFR frequencies)

  • Low-pass: 1000.0 Hz (preserves FFR harmonics while removing high-frequency noise)

  • Filter design: Zero-phase FIR with automatic transition bandwidth. --filter-method iir switches to zero-phase first-order Butterworth high-pass + low-pass passes (12 dB/octave overall).

Outputs: Filtered EEG data

4. Epoching Node

Function: epoch_data()

Purpose: Segment the continuously-filtered EEG into time-locked epochs around stimulus events, applying baseline correction and amplitude-based rejection.

Sub-steps:
  • Load events from the BIDS *_events.tsv (or use annotations embedded in the raw recording when no sidecar is present)

  • Build mne.Epochs with the requested tmin / tmax, baseline window, picks, and reject thresholds

  • Apply epochs.drop_bad() to materialize amplitude-based rejection; keep the resulting Epochs object as the workflow output

Default Parameters (FFR-typical, all overridable from the CLI):
  • Epoch window: --tmin -0.2 to --tmax 0.6 seconds around stimulus onset

  • Baseline: --baseline -0.2 0 seconds (pre-stimulus)

  • Rejection: --reject-eeg 75e-6 (75 µV peak-to-peak); pass --no-auto-reject to disable, or --reject-mode abs to drop epochs whose absolute amplitude reaches the threshold at any sample

Outputs: Epoched EEG data and the post-rejection drop log.

5. Save Preprocessing Outputs Node

Function: save_preprocessing_outputs() (invoked through the internal save_preprocessing_node wrapper, which fans the Epochs out by trial type before calling it).

Purpose: Persist the epoched data plus a self-describing BIDS sidecar — one file per trial type by default, or a single combined file under --no-split-by-trial-type.

Sub-steps:
  • When --split-by-trial-type is on (the default), partition the input Epochs by event_id and write one _desc-preproc{Cond}_epo.fif per trial type with a matching sidecar carrying a Condition field. Under --no-split-by-trial-type, write a single bare _desc-preproc_epo.fif instead.

  • Reconstruct a fresh EpochsArray from the data + events + event_id so trial-type metadata survives the save / load round-trip (and the upstream Epochs object’s internal state doesn’t leak through .save()).

  • Write each sibling .json sidecar with EpochCount / EpochCountTotal / EpochCountRejected / RejectionMode / RejectionThresholds / Filtering / SamplingFrequency / EpochTmin / EpochTmax / Channels plus run / session / ConcatenatedRuns / Condition provenance. Sources and RawSources list the raw EEG file(s) the epochs came from, when they can be found.

  • Initialize the per-derivatives dataset_description.json if missing.

Reporting runs after the workflow drains, in the CLI rather than as a workflow node — see reporting below.

Output Structure (default split-by-trial-type):

derivatives/ffrprep-preprocessing/
├── sub-XX/
│   └── eeg/
│       ├── sub-XX_task-YY_run-ZZ_desc-preprocPositive_epo.fif
│       ├── sub-XX_task-YY_run-ZZ_desc-preprocPositive_epo.json
│       ├── sub-XX_task-YY_run-ZZ_desc-preprocNegative_epo.fif
│       ├── sub-XX_task-YY_run-ZZ_desc-preprocNegative_epo.json
│       ├── sub-XX_preprocessing_report.html
│       └── sub-XX_preprocessing.log

Under --no-split-by-trial-type the per-trial-type files are replaced by a single _desc-preproc_epo.fif + its sidecar.

The sidecar JSON carries provenance, EpochCount / EpochCountTotal / EpochCountRejected, RejectionMode (peak-to-peak or absolute-amplitude) with RejectionThresholds, Filtering (high-pass and low-pass cut-offs), sampling frequency, run / session identifiers, the raw file(s) in Sources / RawSources, and Condition (for per-trial-type files only). The epoch counts cover only the trial types being written: in a per-trial-type file they are that trial type’s counts, and events outside event_id (for example markers) are not counted. Epochs dropped by --reject-mode abs count as rejected.

Outputs: File paths, processing metadata.

Stage 3: Analysis

The analysis stage averages each preprocessed epochs group into a structured set of evoked responses (per-trial-type + combined + optional difference) and saves them to BIDS-derivatives.

Purpose: Produce per-(task, run) evoked responses suitable for downstream statistical analysis or visualization, persisted in MNE-readable format with self-describing BIDS sidecars.

Implementation: Implemented as a Nipype workflow (create_analysis_workflow()) with two nodes. The CLI worker collects all per-trial-type _desc-preproc{Cond}_epo.fif files for one (task, run) group via _collect_analysis_groups, stitches them back together with mne.concatenate_epochs (which preserves event_id), and passes the resulting Epochs directly into the workflow’s inputnode. The per-(task, run) granularity means each group runs the workflow once regardless of how many trial types it holds.

Analysis Workflow Nodes

1. Build Analysis Payload Node

Function: build_analysis_payload()

Purpose: From a single Epochs object, produce a structured {by_type, combined, diff} payload covering every evoked the analysis stage emits.

Sub-steps:
  • Average all events into a single combined Evoked via make_combined_evoked (always emitted).

  • When split_by_trial_type is True (the default), partition by event_id via make_evoked(by_event_type=True) into a per-trial-type dict.

  • When at least two trial types are present and either (a) there are exactly two types (auto-paired) or (b) the user passed --difference-pairs A:B [C:D …], compute one difference Evoked per pair via make_difference_evokeds (which wraps mne.combine_evoked([A, B], weights=[1, -1])).

Outputs: dict with keys "by_type" (dict trial_type → Evoked), "combined" (Evoked), and optionally "diff" (dict (A, B) → Evoked).

2. Save Analysis Node

Function: save_analysis_outputs()

Purpose: Persist every Evoked in the payload plus a self-describing sidecar per file.

Sub-steps:
  • For each entry in by_type: write _desc-evoked{Cond}.fif + matching sidecar with Condition: <cond>.

  • For combined: write the bare _desc-evoked.fif + sidecar (no Condition field).

  • For each entry in diff: write _desc-evokedDiff{A}Vs{B}.fif + sidecar carrying DifferenceOf: [A, B].

  • Every sidecar also carries AverageCount / Baseline / SamplingFrequency / Tmin / Tmax / Channels / TaskName / AnalysisType plus run / session / ConcatenatedRuns provenance.

  • Initialize the per-derivatives dataset_description.json if missing.

Output Structure (default split-by-trial-type, 2-trial-type dataset):

derivatives/ffrprep-analysis/
├── sub-XX/
│   ├── sub-XX_task-YY_run-ZZ_desc-evokedPositive.fif
│   ├── sub-XX_task-YY_run-ZZ_desc-evokedPositive.json
│   ├── sub-XX_task-YY_run-ZZ_desc-evokedNegative.fif
│   ├── sub-XX_task-YY_run-ZZ_desc-evokedNegative.json
│   ├── sub-XX_task-YY_run-ZZ_desc-evoked.fif
│   ├── sub-XX_task-YY_run-ZZ_desc-evoked.json
│   ├── sub-XX_task-YY_run-ZZ_desc-evokedDiffPositiveVsNegative.fif
│   ├── sub-XX_task-YY_run-ZZ_desc-evokedDiffPositiveVsNegative.json
│   ├── sub-XX_analysis_report.html
│   └── sub-XX_analysis.log

Under --no-split-by-trial-type, only the combined _desc-evoked.fif + sidecar are emitted per (task, run).

Outputs: Saved file paths.

Reporting (post-workflow)

The single-file HTML reports are built in the CLI after the workflow drains, not as a workflow node. The CLI’s _build_preproc_report and _build_analysis_report glob the saved _desc-preproc*_epo.fif and _desc-evoked*.fif files via _collect_analysis_groups / _collect_evoked_groups, group them by (task, run), and build sections via the ffrprep.reports builders (build_raw_section / build_epoch_section / build_evoked_section / build_phase_consistency_section / make_group); the subject-level HTML is rendered via build_subject_report / build_analysis_report.

Per (task, run) group layout in the analysis report:

  • one Evoked section per per-trial-type file (e.g. Positive, Negative). Each carries:

    • waveform / PSD / TFR / autocorrelation / pitch-track figures (via ffrprep.reports.evoked_qa());

    • scalar metrics: RMS SNR (<window>), Mean power 90-110 Hz, <window>, where <window> defaults to 100-200 ms and is configurable via --response-window START END (seconds);

    • when the BIDS stim_file column is populated in events.tsv, a Stim correlation (peak r) + Stim correlation (lag, ms) row plus a stim ↔ response cross-correlation lag plot (computed by _stim_correlation_data in ffrprep_cli.py).

  • one combined Evoked section (the across-events average) with the same plots + scalars, plus a second pair of rows / plot for the envelope correlation (combined ≈ ENV proxy in FFR, so |hilbert(stim)| is the natural reference).

  • one difference Evoked section per (A, B) pair (auto for the 2-trial-type case; opt-in via --difference-pairs) with the standard plots + a single raw-waveform stim correlation row + plot (diff ≈ TFS proxy).

  • when exactly two per-condition preproc files exist for the group, one Phase Consistency section combining ffrprep.analysis.compute_phase_consistency() with ffrprep.analysis.plot_phase_consistency() (or, when the caller passes mask=True, the masked variant). Single-subject reports default to unmasked; group-level builders pass mask=True. Uses seaborn’s flare_r colormap; subplot titles surface the trial-type names plus auto-derived A + B / A − B for the sum and difference panels.

Per (task, run) group layout in the preprocessing report:

  • one Raw section loaded from the original BIDS recording (with run concatenation when the preproc output sourced multiple runs).

  • one Epoched section per per-trial-type file — each carries the standard epoch metadata plus a Mean trial-to-trial r row (via ffrprep.analysis.response_consistency()) when there are at least 10 trials.

Figures and scalar metrics are computed at report time from the loaded Epochs / Evoked objects — they are not separately persisted to disk. To recompute them yourself, see the Working with Outputs in Python section of the Tutorial walkthrough.

The single-file *_report.html embeds all figures as inline base64 PNGs — there is no sibling figures/ directory.

File Formats:
  • MNE format (.fif): Epoched / Evoked data, loadable in MNE-Python.

  • WAV (under <dataset>/stimuli/ ): stimulus audio referenced by events.tsv’s stim_file column. Fetched via ffrprep-download example --with-stimuli.

  • JSON: BIDS sidecar (human-readable, machine-parseable).

  • HTML: Self-contained single-file per-subject report.

Pipeline Integration and Quality Control

Workflow Management:

  • Each stage implemented as a Nipype workflow for per-iteration dependency tracking

  • Outer-loop parallelism: the CLI dispatches per-(task, run) iterations to a ProcessPoolExecutor sized by --n_procs

  • Fail-fast on any iteration error; the exception propagates to the CLI entry point

  • nipype caches per-iteration intermediates under work/ so re-runs that already have a saved _desc-preproc_epo.fif skip the workflow re-execution

  • --skip-existing skips a (task, run) iteration whose sidecar outputs already exist, so an interrupted run can be repeated with the same command. The check uses the .json sidecars rather than the epochs files, so it also works after --no-keep-epochs. An analysis iteration counts as done when its _desc-evoked.json exists, and no report is rebuilt for iterations that are skipped.

  • --clean-work-dir deletes an iteration’s Nipype working files once it has succeeded. Log files are kept, and the files of a failed iteration are left in place.

  • --no-keep-epochs deletes the _epo.fif files after the analysis stage and its report have finished. The JSON sidecars stay. With --stage preprocessing the flag is ignored and a warning is printed, because the analysis stage still needs the epochs.

Quality Control Checkpoints:

  • BIDS validation before processing

  • Data quality assessment after loading

  • Preprocessing quality metrics and reports

  • Analysis validation and statistical checks

Output Organization:

  • BIDS-compatible directory structure

  • Per-file BIDS sidecars with run / session / condition provenance

  • Standardized file formats for interoperability

  • Version-controlled processing parameters

Customization:

  • All preprocessing and analysis parameters are exposed as CLI flags (see Usage).

  • The Nipype workflows can be imported and reused programmatically from ffrprep.preproc (create_preprocessing_workflow / create_analysis_workflow).

  • Section builders in ffrprep.reports accept extra_summary and extra_figures kwargs so downstream code can fold caller-computed scalars or figures into a section’s table or figure gallery.

Group-Level Aggregation

The group analysis level (ffrprep.group.run_group_level) aggregates participant-level derivatives already written under output_dir — it does not read raw BIDS data. For each (task, session, run) with derivatives from at least two subjects, it:

  • discovers each subject’s combined (_desc-evoked.fif) and, when present, difference (_desc-evokedDiff{A}Vs{B}.fif) Evoked files via ffrprep.group.discover_group_inputs(), and computes a grand average with mne.grand_average (ffrprep.group.compute_grand_average()). If the subjects’ single channel has different names (for example A32 and Cz at different sites), it is renamed to a common name first. Multi-channel sets that differ raise an error;

  • recomputes the same scalar FFR metrics the participant-level report shows (RMS SNR, band power) directly from each subject’s saved combined Evoked, plus trial-to-trial response consistency from that subject’s saved preprocessing Epochs (ffrprep.group.compute_subject_metrics()) — nothing is recomputed from raw data;

  • adds further columns when asked. --f0 adds rms_snr_polarity_sum, f0_uv and upper_harmonics_uv; --stimulus adds stim2resp_r, stim2resp_z and stim2resp_lag_ms (plus stim2resp_lim_* with --xcorr-lag-range); --n-trials-presented adds usable_pct. These are computed from the sum of the two per-trial-type evoked files. f0_uv and upper_harmonics_uv are FFT amplitudes averaged in a band around each harmonic of F0 (ffrprep.analysis.harmonic_amplitudes()), and stim2resp_r is the maximum normalized cross-correlation between the stimulus and the response (ffrprep.analysis.stim_to_resp_xcorr());

  • joins subject-level covariates onto the table by participant_id, from <bids_dir>/participants.tsv and any --covariates files (ffrprep.group.merge_covariates()). A column already in the table is not overwritten;

  • with --min-usable-pct and/or --min-snr, adds qc_* columns that mark recordings below a threshold (ffrprep.group.add_qc_flags()). Rows are flagged, not removed, and the grand averages still include every subject;

  • writes the grand-average Evoked(s), a per-subject metrics TSV with a _metrics.json data dictionary next to it (ffrprep.group.build_metrics_dictionary()), and a single group_report.html (reusing the same section builders and template as the participant-level reports) under output_dir/ffrprep-group/.

Output Structure:

derivatives/ffrprep-group/
├── dataset_description.json
├── task-YY_run-ZZ_desc-grandAverage_ave.fif
├── task-YY_run-ZZ_desc-grandAverage_ave.json
├── task-YY_run-ZZ_desc-grandAverageDiffPositiveVsNegative_ave.fif
├── task-YY_run-ZZ_desc-grandAverageDiffPositiveVsNegative_ave.json
├── task-YY_run-ZZ_metrics.tsv
├── task-YY_run-ZZ_metrics.json
└── group_report.html

This step is deliberately scoped to aggregation, not inference: it summarizes what participant-level ffrprep already computed and performs no group-level statistics (no hypothesis tests, no GLM). Statistical analysis is left to the user, e.g. directly in MNE-Python against the saved grand-average / per-subject derivatives.