Pipeline Details#
MEGFlow is a Nextflow workflow that combines structural MRI processing, continuous MEG preprocessing, artifact detection, ICA cleaning, optional epoching, covariance estimation, MEG-MRI coregistration, forward modeling, source reconstruction, and static quality-control reporting.
The main workflow is implemented in nextflow/megflow.nf.
Configuration is supplied through nextflow.config and can be overridden by
selected command-line options. See Configuration Reference for the
complete configuration reference.
Execution Modes#
The workflow is controlled by params.megflow.defaults.steps and optional
dataset- or recording-level steps overrides:
Mode |
Processing scope |
|---|---|
|
MRI import, FreeSurfer or DeepPrep reconstruction, head surface, and BEM. |
|
MEG import, continuous preprocessing, bad-channel and bad-segment detection, then static report. |
|
|
|
|
|
Full MEG workflow using an existing |
|
Structural MRI workflow plus full MEG workflow. |
|
Static HTML report only, using existing outputs. |
Aliases meg, artifacts, ica, and epochs map to meg_all,
meg_artifacts, meg_ica, and meg_epochs. The optional
with_anatomy modifier can accompany meg_artifacts, meg_ica, or
meg_epochs; skip_ica is valid with meg_epochs. A recording-level
stage may reduce an enabled dataset MEG path but cannot add anatomy or exceed
the dataset stage.
High-Level Flow#
The complete meg_all or all dependency graph is:
MEG import
-> NormMEG-QC scoring (when enabled)
-> NMDQ min_score gate
-> continuous preprocessing
-> artifact detection
-> ICA fit
-> ICA labeling
-> ICA application
-> epoching
|-> noise covariance -------------------------|
|-> LCMV data covariance (only for LCMV) -----|-> source reconstruction
|-> forward solution <- coregistration -------|
-> static HTML report
When megqc.enabled is false, imported recordings bypass both the scoring
process and its gate. When enabled, an unscored recording or one below
megqc.min_score does not enter continuous preprocessing.
Covariance and coregistration may execute concurrently after their own inputs are ready. Forward modeling waits for both the recording’s epoch file and its coregistration transform. Source reconstruction is a strict keyed join of that recording’s exact forward file, noise covariance, and conditional LCMV data covariance. The covariance branch also carries the hash of the exact source Raw/Epochs used to compute it; an unmatched, duplicate, or inconsistent key is an error rather than a silently skipped source result.
With covariance.type = "raw", the recording selected by
raw_covariance_task_id follows the same continuous preprocessing and ICA
path, then feeds the experimental recording’s noise-covariance branch. It does
not create its own epochs, forward model, or source estimate.
When anatomy is enabled, its branch can run concurrently with MEG preprocessing. The branches join only when coregistration, forward modeling, or source reconstruction requires anatomy:
MRI import
-> FreeSurfer or DeepPrep reconstruction
-> head surface
-> BEM model
-> MEG-MRI coregistration dependencies
When anatomy.method = "pseudomri", MEGFlow first imports MEG
files to access digitization/headshape points, generates a pseudo T1 image, and
then reuses the normal FreeSurfer and BEM stages:
MEG import
-> pseudo-MRI generation
-> FreeSurfer reconstruction
-> BEM model
-> MEG-MRI coregistration dependencies
Continuous Core Preprocessing#
The continuous MEG core is task independent and applies to both resting-state and task-based recordings. NormMEG-QC, when enabled, is a preflight gate before this core rather than one of its signal-processing operations.
import_meg_datasetdiscovers input recordings. BIDS input is filtered bymeg_importentities. Raw input is selected byfile_suffixand optionalraw_include_keywords/raw_exclude_keywords. Raw discovery matches both files and directories, so CTF.dsfolders are supported.score_meg_qualityruns NormMEG-QC when enabled, writes the NMDQ score sidecars, and appliesmegqc.min_scorebefore downstream processing.meg_basic_preproccallsmeg_preproc_osl.py, which passes the effectivepreprocblock to OSL-Ephysrun_proc_batch. The listed preprocessing steps are executed in order. Common steps include Maxwell/tSSS for Elekta/MEGIN data, band-pass filtering, notch filtering, and resampling. Resampling is the current configurable downsampling mechanism.detect_artifactscallsmeg_detect_artifacts.py. It detects bad channels and bad time spans using the configured PyPREP, PSD, OSL, MNE, and optional DeepReject methods. Within DeepReject, BadChnNet runs first; its bad channels are masked before BadSegNet predicts bad time windows. Results from all enabled detectors are merged into*_bad_channels.txtand*_bad_segments.txt. The process also writes detector provenance and a recording-wide artifact-mask heatmap. Detailed waveform images are optional.run_icaloads the preprocessed raw file plus the artifact sidecars. Bad channels are excluded from picks, and bad annotations are ignored during ICA fitting throughreject_by_annotation=True. Before signal samples are loaded, a fast preflight reads only the FIF header and sidecars. It checks that at least one requested MEG channel and some unannotated samples remain, and that an integer component request fits the available input. Overlapping or touchingBAD...annotations are counted once using the same sample boundaries as MNE ICA. The result is saved asica_input_validation.jsonwith sample, second, channel, component, and sidecar details.fit_required,ica_cache_exists, andica_cache_pathshow whether this run will fit a new ICA or reuse an existing one. A cached ICA still requires usable channels and samples, but a new component request is checked only when a new fit is required. This file is also written when validation fails whenever the output directory is available.run_ic_labellabels artifact-related ICA components using configured MNE ECG/EOG detection, the original MNE-ICALabel MEGNet model, the independently optional retrained MEGNet model, and rule-based settings. Method switches determine which detectors may run;ic_ecg,ic_eog, andic_outlierare category master switches applied to every method. The automatic union therefore contains only detections whose method and category are both enabled. A failure isolated to the optional retrained model is recorded without blocking the other methods. Category results and automatic/manual component provenance are stored inecg_eog_scores.json.apply_icaloads the marked components, applies the ICA solution, and saves*_clean_raw.fif. The cleaned continuous file keeps the bad-channel and bad-segment metadata.
Interactive Edits and Resume#
Some sidecar files can be edited after a Nextflow run through the interactive
reports. MEGFlow includes content hashes of those files in the relevant
Nextflow task inputs so -resume can invalidate only the affected downstream
tasks:
Adding or removing a raw input invalidates dataset import. Existing unchanged recordings remain cacheable, while newly discovered recordings create new task branches.
Changing one raw recording invalidates QC and all processing for that recording without invalidating other recordings.
Changing a BIDS
events.tsvsidecar invalidates epoching, epoch-based covariance, forward modeling, and source reconstruction for that recording; continuous preprocessing and ICA remain cacheable.Editing
artifact_report/*/*_bad_channels.txtorartifact_report/*/*_bad_segments.txtinvalidates ICA fitting and later steps for that recording.Editing
ica_report/*/marked_components.txtinvalidates ICA application and later steps for that recording.Editing
trans/*/coreg-trans.fifinvalidates forward modelling and source reconstruction for that recording.Changing a T1 input or reconstructed anatomy invalidates the relevant structural/BEM lineage and downstream coregistration or forward tasks.
Changing the MEGFlow processing implementation invalidates cached tasks through an implementation fingerprint; Python cache files are deliberately excluded.
This downstream hash mechanism is separate from published-output deletion.
Each required external result has a task-local symbolic-link cache guard that
Nextflow records as a path output. This includes both
marked_components.txt and ecg_eog_scores.json for ICA labeling. An
unchanged target keeps the owner task cacheable. If the published target is
deleted, the guard becomes dangling and -resume reruns the owner task; any
consumer whose effective input fingerprint changes then reruns as well. Other
recordings remain cacheable.
Editable bad-channel, bad-segment, and marked-component sidecars use both
mechanisms deliberately. Deleting one reruns its owner so the required file is
restored. Editing its content leaves the owner cached, preserves the manual
edit, and invalidates only the downstream consumers listed above. Static report
processes use cache false and therefore regenerate on every completed or
lenient run.
If configuration, labeling code, or the retrained model changes, Nextflow
safely refreshes the labeling task. A marked_components.txt file that still
matches the previous automatic result is replaced by the new automatic union.
If its content differs, MEGFlow treats it as a manual edit and leaves it
unchanged while refreshing detector details in ecg_eog_scores.json. The
JSON marked_components.mode field records auto or
preserved_manual. Its auto_indices field records the newly detected
category union, while written_indices always matches the exact contents of
marked_components.txt consumed by apply_ica. Direct script users can
request an unconditional reset with run_ica_label.py --overwrite-existing.
Bad Segments: Marking vs Exclusion#
Artifact detection marks bad segments as MNE annotations. This does not cut samples out of the continuous raw file. Downstream steps decide whether the annotations should exclude data:
ICA fitting ignores annotated spans when estimating ICA components.
ICA application writes a cleaned raw file with annotations attached.
With
skip_ica, epoching loads the detected bad-channel and bad-segment sidecars directly into the preprocessed raw before constructing epochs.Epoching drops epochs overlapping annotations only when
epochs.epochs.reject_by_annotationis true.Additional epoch rejection can come from
epochs.epochs.rejector optionalautoreject.
By default, artifacts.find_bad_segments.keep_existing_annotations is
false. Artifact detection therefore starts from an empty annotation set and
writes only annotations produced during the current run. Set it to true
when existing input annotations are trusted and should be retained alongside
new detector annotations. This setting controls annotation merging, not sample
deletion.
DeepReject intervals use the annotation description BAD_deepreject. See
DeepReject Artifact Detection for the BadChnNet and BadSegNet decision rules,
mode thresholds, and provenance output.
If ICA stops with ICA_INPUT_ALL_BAD, open
ica_input_validation.json and check bad_coverage_fraction and the
recorded bad-segment sidecar path. A value of 1.0 means the annotations
cover every sample, so ICA has no data from which to learn components. Inspect
the bad-segment report, correct or regenerate an over-broad sidecar, and resume
the workflow. Do not work around this error by merely increasing the component
count. no_eligible_meg_channels similarly means the bad-channel sidecar
excluded every MEG channel. invalid_bad_segment_sidecar means the sidecar
could not be aligned with this recording and should be regenerated from it.
invalid_bad_channel_sidecar identifies an unreadable bad-channel file, and
invalid_ica_modality identifies a modality other than meg, eeg, or
meeg. When fitting a new ICA, invalid_component_request means the
request must be an integer of at least 2 or a finite fractional variance target
strictly between 0 and 1.
Resting-State and Task-Based Epochs#
Epoching is optional and happens after the continuous core. The effective
epochs block
selects how epochs are built:
task_type: restingcreates fixed-length events withresting.fixed_length_duration.task_type: taskwithevent_source: find_eventsuses MNEfind_eventsand thefind_eventsconfig block.task_type: taskwithevent_source: event_filereads BIDS*_events.tsvfiles and applies theevent_filefilters or label-to-id mappings. If the inferred event path is not a tabular.tsvevent file (for example a non-BIDS raw.fifinput), MEGFlow falls back tomne.find_eventsso raw datasets are not parsed as text.exclude_event_idcan be set to one id or a list of ids to remove those events before epoching. Withepochs.event_id: null, MEGFlow keeps all remaining event ids.event_time_shift_seccan be set in theepochsblock to shift all task events before epoching. Positive values move event samples later, for example to compensate for a stable auditory or visual stimulus delivery delay after the hardware trigger.epochs.preproccan optionally filter, notch-filter, or resample the cleaned continuous Raw recording before events are converted into epochs. The default is empty and preserves the existing continuous data unchanged.
The resulting epoch FIF file and rejection log are written under
preprocessed/epochs/<recording>/. When epochs.preproc is configured,
the same directory also contains *_analysis-raw.fif; epoch-based covariance
uses this exact analysis-ready continuous recording.
Covariance and Empty-Room Style Records#
Covariance is computed only in the full MEG stage. Two modes are available:
covariance.type = "epochs"estimates noise covariance from baseline epochs created from each cleaned experimental recording.covariance.type = "raw"estimates noise covariance from a continuous raw recording selected bycovariance.raw_covariance_task_id.
Both modes write bl-cov.fif. The same compute_covariance task writes
lcmv-data-cov.fif only when the effective source.source_methods contains
LCMV. That matrix is computed from the exact *-epo.fif when
source.type = "epochs" or the exact analysis-ready Raw when
source.type = "raw". dSPM and other minimum-norm-only runs do not perform
this extra calculation.
When analysis preprocessing is configured for epochs, raw covariance applies
the same operations in memory to the paired baseline recording. Empty
epochs.preproc configurations retain the original covariance behavior.
When covariance is estimated from epochs, covariance.event_time_shift_sec
uses the same sign convention as the epoching stage so baseline epochs remain
aligned with the corrected event timing.
For raw covariance, MEGFlow pairs experimental recordings with a noise or
baseline recording by replacing the BIDS task-... part of the filename with
task-${covariance.raw_covariance_task_id}. This is the current mechanism for
empty-room or empty-room-like recordings. For example, if
covariance.raw_covariance_task_id = "emptyroom", an experimental file with
task-aef is paired with a file whose matching name contains
task-emptyroom while every other filename entity remains unchanged.
The pairing is performed between current-run clean-recording channels. It waits for the noise branch instead of checking whether a predicted output path happens to exist, and its clean-file fingerprint participates in the covariance cache lineage. A single noise recording may feed multiple experimental recordings. Reference recordings complete continuous preprocessing and ICA, then skip their own epoch/source branches. If a requested pair is absent, the strict downstream join fails the run instead of producing an incomplete source set.
Before covariance estimation, target and noise inputs are restricted to common
good channels in target-channel order. rank_policy is resolved from the
final experimental target and provides the default rank for noise covariance,
LCMV data covariance, and source reconstruction. For raw noise, MEGFlow also
requires its empirical rank to be at least the target rank. This detects an
insufficient empty-room input, but equal ranks do not prove that independently
fitted ICA operators describe the same linear subspace. See
Rank, Covariance, and Source Imaging for configuration precedence and this
compatibility boundary.
The covariance task writes the resolved dictionary and its ordered target
channels to resolved-rank.json. Source reconstruction consumes that file,
verifies the channel order after alignment with covariance and forward inputs,
and passes the stored dictionary to the configured MNE functions. It does not
derive a second default rank.
Coregistration, Forward Model, and Source Reconstruction#
coregistration or coregistrations aligns MEG sensor space to the
subject anatomy. The process uses fiducial fitting, ICP, and a fine-tuned ICP
stage controlled by the effective coreg block. It writes coreg-trans.fif,
coregistration figures, and distance summaries.
forward_solution builds the forward model using the keyed epoch file,
transform, anatomy fingerprint, FreeSurfer subject directory, and the effective
forward block. The emitted tuple carries the exact generated forward FIF
path rather than reconstructing that path later from directory names.
source_imaging consumes either epochs or raw data according to
source.type.
It receives and loads the exact forward model, noise covariance, and optional
LCMV data covariance selected by the workflow. It verifies that covariance
channel names and order match the source data and forward model, consumes the
routed default-rank artifact, and then applies the configured source methods.
LCMV never recomputes a covariance inside source_localization.py. Deterministic
rank, routing, channel-contract, or missing-output errors terminate rather than
being retried and ignored as a successful partial run.
Static Processing and QC Report#
At the end of each selected MEG milestone, generate_static_html_report scans
the existing outputs and writes a portable report under
<dataset_output_dir>/static_html_report. The report includes a workflow
manifest, a config snapshot when available, subject pages, dataset summaries,
and evidence files. See Reports and
Quality Control Metrics for details.
Multi-Dataset Profiles#
params.megflow.corpus_root can point to a directory containing multiple
datasets. dataset_include and dataset_exclude select the children to
run. Each dataset can override any default module block under
params.megflow.datasets. A dataset can also define recordings entries
with BIDS-entity match rules for task- or run-specific settings. This allows
one workflow run to give WAND, SMN4Lang, and MEG-MASC different event
definitions, preprocessing, artifact settings, and source labels. The runnable
example is documented in Runnable Three-Dataset Source Demo.
Primary Outputs by Step#
Step |
Output location |
Main outputs |
|---|---|---|
NormMEG-QC |
|
|
Continuous preprocessing |
|
|
Artifact detection |
|
Bad-channel text and provenance JSON, bad-segment annotations,
|
ICA |
|
ICA FIF, source FIF, marked components, ECG/EOG scores, plots. |
ICA-clean raw |
|
|
Epochs |
|
|
Covariance |
|
|
Coregistration |
|
|
Forward model |
|
Forward solution FIF and head-model figures. |
Source reconstruction |
|
Source estimate files and visualization figures. |
Static report |
|
Dataset dashboard, subject pages, JSON/CSV summaries. |