Skip to content

Core Data Classes

Design Patterns covered the contract shapes (ParticipantConfig, RawBioData/RawBehaviourData, PipelineOutputData) in passing. This page is a closer look at three classes you'll touch directly while writing or debugging a pipeline: TrialIntervals, PipelineStatus, and PipelineOutputData.

Declaring a class (and a child class) — the basics

If classes are still new: a class is a blueprint for objects that bundle data and behaviour together. A child class (or "subclass") reuses everything a parent class already has, and only needs to state what's different — new fields, or a method it wants to replace:

class Dish:
    def __init__(self, name):
        self.name = name

    def describe(self):
        return f"{self.name}"


class Salad(Dish):                    # Salad IS-A Dish — inherits __init__ and describe()
    def __init__(self, name, dressing):
        super().__init__(name)        # reuse the parent's setup
        self.dressing = dressing

    def describe(self):               # override: same method name, different behaviour
        return f"{self.name} with {self.dressing}"

Salad("Greek Salad", "olive oil") gets self.name for free from Dish; it only had to add self.dressing and change what describe() returns. Every class below (and every Protocol in Design Patterns) is a variation on this same idea, written with @dataclass instead of a hand-written __init__ — see below for what that changes.

TrialIntervals

A named set of (start, end) time pairs for one participant — e.g. "Baseline": (0.0, 5.0). Produced by a GetTrialIntervalsStartegy, consumed by every processing step that needs to slice data per trial. See Interval QC Plot for the conceptual picture (what a trial interval means, and the QC plot that checks one was built correctly) — this section is the code-level detail behind it.

@dataclass
class TrialIntervals:
    """
    Trial intervals are named time periods in the experiment at the subject level, encoded
    as (start, end) time pairs keyed by a str name.

    For example:
    "Baseline": (0, 5) — the "Baseline" trial spans from 0 to 5 seconds.
    """

    intervals: dict[str, tuple[float, float]] = field(default_factory=dict)

    def __len__(self):
        return len(self.intervals)

Strip away the class, and it's really just a dict: {"Baseline": (0, 5), "Stress": (5, 65), ...}. The class wraps that dict and adds behaviour on top — keeping it sorted by start time, filling gaps, relabelling, comparing two sets of intervals — which is why the pipeline passes TrialIntervals objects around instead of bare dictionaries.

Toy example: creating one yourself

You can build one directly, with made-up values, to get a feel for the shape — no recording data needed:

from vrlab_toolbox.processing.trial_intervals import TrialIntervals

my_intervals = TrialIntervals(
    intervals={
        "Baseline": (0.0, 60.0),
        "Stress": (60.0, 120.0),
        "Recovery": (120.0, 180.0),
    }
)

len(my_intervals)       # 3 — one per named interval, via __len__
my_intervals.intervals  # the underlying dict, always kept sorted by start time

Real example 1: unlabelled, straight from trigger pulses

TrialIntervals.from_raw_interval_pairs is a shortcut constructor for when all you have is a plain list of (start, end) pairs, with no meaningful names yet — exactly the situation right after detecting raw trigger pulses, before anything's matched to behaviour:

from vrlab_toolbox.processing.trial_intervals import TrialIntervals

raw_pairs = [(0.0, 60.0), (60.0, 120.0), (120.0, 180.0)]
unlabelled = TrialIntervals.from_raw_interval_pairs(raw_pairs)
# {"TP0": (0.0, 60.0), "TP1": (60.0, 120.0), "TP2": (120.0, 180.0)}

That TP0, TP1, … naming is exactly what labels the blue "Raw biopac" row in Interval QC Plot's example plot — this constructor is what produces it.

Real example 2: labelled, straight from behaviour data

Compare that to get_crane_trigger_behav_intervals (crane_trial_intervals.py), which builds a TrialIntervals from a participant's behaviour dataframe instead, using real trial names:

def get_crane_trigger_behav_intervals(validated_behav_df: pd.DataFrame) -> TrialIntervals:
    intervals_out = {}
    for _, row in validated_behav_df.iterrows():
        intervals_out.update(
            {
                f"{row['BlockType']}_{row['TrialType']}_{row['TrialNr']}": tuple(
                    [row["TrialStartTime"], row["TrialEndTime"]]
                )
            }
        )
    return TrialIntervals(intervals=intervals_out)

This is what produces names like NonStressBlock_NonSlipTrial_1 — the orange "Raw behav" row on the same plot. Matching this object against the unlabelled one above (correcting for clock drift along the way) is exactly what CraneGetTrialIntervalStrategyStep (below) does.

Real example 3: what the matched result actually contains

Running the full matching step (CraneGetTrialIntervalStrategyStep) on the synthetic DUMMY000 participant's raw, pre-BIDS output in examples/ (see Testing) produces a TrialIntervals with 23 entries — this is .intervals, first five shown:

{
    'NonStressBlock_NonSlipTrial_1_Training': (4.9845, 65.3135),
    'ITI_0': (65.3135, 94.522),
    'NonStressBlock_SlipTrial_2_Training': (94.522, 155.1415),
    'ITI_1': (155.1415, 205.774),
    'NonStressBlock_NonSlipTrial_3_Training': (205.774, 266.104),
    ...
}

This is exactly the object drawn as the red "Matched with Behav" row in Interval QC Plot's example plot — real trial names, timestamps corrected onto the trigger channel's clock.

The two strategies behind the QC plot

crane_trial_intervals.py defines two strategy steps (see Design Patterns for what "strategy step" means here):

  • CraneGetTrialIntervalStrategyStep — the main path. Needs both physiology and behaviour data, and produces all four rows of the QC plot, including a filled-in "Matched with Behav" row.
  • CraneGetTrialIntervalStrategyFallbackStep — set as the main step's fallback_strategy, and used when matching against behaviour data fails (see Pipeline Rules). It only has trigger data to work with, so its version of the plot only fills rows 1–3 — the "Matched with Behav" row stays empty, and physiology is processed against unlabelled trigger intervals instead. See Interval QC Plot: when behaviour data is missing or doesn't match for what that looks like on the plot itself.

PipelineStatus

A type → ProcessingStatus mapping (OK, PARTIAL, CORRECTED, ERROR, NOT_RUN) — one entry per data type the pipeline touched for a participant. Every Sequential*Steps.run() returns one alongside its data, and PipelineTemplate.run() merges them all into a single status for the participant.

@dataclass
class PipelineStatus:
    status: dict[type, ProcessingStatus] = field(default_factory=dict)

    def set(self, data_type, status: ProcessingStatus):
        self.status[data_type] = status

    def merge(self, other: "PipelineStatus") -> "PipelineStatus":
        ...  # worst status per type wins, across both
from vrlab_toolbox.processing.processing_status import PipelineStatus, ProcessingStatus
from vrlab_toolbox.processing.crane_behaviour import RawCraneBehaviourData

status = PipelineStatus()
status.set(RawCraneBehaviourData, ProcessingStatus.OK)
status.get_as_text()   # "RawCraneBehaviourData=ok" — this is what ends up in the output row

PipelineOutputData — and why it has child classes

One participant's output: a one-row subject_df_out DataFrame, a status (PipelineStatus), and any QC figure_data_out. Every processing step returns one; .merge() combines two into one, always staying one row — see One participant at a time for why that invariant matters.

@dataclass
class PipelineOutputData:
    subject_id: str
    subject_df_out: pd.DataFrame = field(init=False)
    figure_data_out: dict[str, Figure] = field(default_factory=dict, init=False)
    validation_schema: pa.DataFrameSchema = field(default_factory=build_base_pipeline_output_schema)
    status: PipelineStatus = field(default_factory=PipelineStatus, init=False)

    def validate_participant_output(self) -> pd.DataFrame:
        return self.validation_schema.validate(self.subject_df_out)

Every experiment ends up with its own output columns — Crane's row looks nothing like the Long Walk pipeline's row. Rather than rewrite .merge(), .append_dataframe(), and the row-shape rules per experiment, each experiment defines a child class that changes only the one thing that actually differs — which Pandera schema validates the row:

@dataclass
class CranePipelineOutputData(PipelineOutputData):
    def validate_participant_output(self) -> pd.DataFrame:
        return build_crane_participant_output_schema().validate(self.subject_df_out)

That's the whole class. merge(), append_dataframe(), figure_data_out, the one-row rule — all inherited unchanged from PipelineOutputData, same as Salad inheriting Dish.__init__ above. You'll see the same pattern for CraneBehaviourOutputData, CraneDebriefPipelineOutput, EdaPhysiologyOutputData, LongWalkPipelineOutputData, and FohPipelineOutputData (foh_pipeline.py) — one override each, nothing more, all still exactly one row per subject.

Going further

FohPipelineOutputData.validate_participant_output() currently validates against a bare pa.DataFrameSchema() — the override exists, but build_foh_participant_output_schema() hasn't been filled in with real columns yet, so validation is a no-op today. Once it is, it'll follow the same "build from constants" shape build_crane_participant_output_schema() already uses. See Lab Streaming for the rest of FOH's anticipated pipeline shape.

Declaring fields: = default vs. field(default_factory=...)

All three classes above use @dataclass, and all three mix both styles of default. The difference matters, and Python enforces it rather than just recommending it:

@dataclass
class Broken:
    intervals: dict[str, tuple[float, float]] = {}     # in-place literal default
ValueError: mutable default <class 'dict'> for field intervals is not
allowed: use default_factory

A plain = {} (or = [], = SomeDataclass()) would mean every instance starts out sharing the exact same dict object — mutate one participant's intervals, and every other TrialIntervals ever created without an explicit value would see the change too. @dataclass refuses to let you write that by accident:

@dataclass
class TrialIntervals:
    intervals: dict[str, tuple[float, float]] = field(default_factory=dict)

field(default_factory=dict) says "call dict() fresh, once per instance" instead — the fix, not just a longer way to write the same thing. The same reasoning is why PipelineOutputData.status is field(default_factory=PipelineStatus, init=False) rather than = PipelineStatus(): a fresh PipelineStatus() per participant, not one PipelineStatus silently shared across every PipelineOutputData ever built.

Plain, immutable defaults (str, int, float, bool, None) don't have this problem — they can't be mutated in place, so there's nothing to accidentally share. That's why subject_id: str above never needs field(...) at all; only the mutable ones do.


See also: Design Patterns for the Protocol contracts these classes flow through, and Data Stores for where RawBioData/RawBehaviourData instances (a related but separate pair of classes) get held between import and processing.