Core Data Classes
Design Patterns covered the contract shapes
(ParticipantConfig, RawBioData/RawBehaviourData, PipelineOutputData)
in passing. This page is a closer look at three classes you'll touch
directly while writing or debugging a pipeline: TrialIntervals,
PipelineStatus, and PipelineOutputData.
Declaring a class (and a child class) — the basics
If classes are still new: a class is a blueprint for objects that bundle data and behaviour together. A child class (or "subclass") reuses everything a parent class already has, and only needs to state what's different — new fields, or a method it wants to replace:
class Dish:
def __init__(self, name):
self.name = name
def describe(self):
return f"{self.name}"
class Salad(Dish): # Salad IS-A Dish — inherits __init__ and describe()
def __init__(self, name, dressing):
super().__init__(name) # reuse the parent's setup
self.dressing = dressing
def describe(self): # override: same method name, different behaviour
return f"{self.name} with {self.dressing}"
Salad("Greek Salad", "olive oil") gets self.name for free from Dish;
it only had to add self.dressing and change what describe() returns.
Every class below (and every Protocol in Design Patterns)
is a variation on this same idea, written with @dataclass instead of a
hand-written __init__ — see below
for what that changes.
TrialIntervals
A named set of (start, end) time pairs for one participant — e.g.
"Baseline": (0.0, 5.0). Produced by a GetTrialIntervalsStartegy,
consumed by every processing step that needs to slice data per trial. See
Interval QC Plot for the
conceptual picture (what a trial interval means, and the QC plot that
checks one was built correctly) — this section is the code-level detail
behind it.
@dataclass
class TrialIntervals:
"""
Trial intervals are named time periods in the experiment at the subject level, encoded
as (start, end) time pairs keyed by a str name.
For example:
"Baseline": (0, 5) — the "Baseline" trial spans from 0 to 5 seconds.
"""
intervals: dict[str, tuple[float, float]] = field(default_factory=dict)
def __len__(self):
return len(self.intervals)
Strip away the class, and it's really just a dict:
{"Baseline": (0, 5), "Stress": (5, 65), ...}. The class wraps that dict
and adds behaviour on top — keeping it sorted by start time, filling gaps,
relabelling, comparing two sets of intervals — which is why the pipeline
passes TrialIntervals objects around instead of bare dictionaries.
Toy example: creating one yourself
You can build one directly, with made-up values, to get a feel for the shape — no recording data needed:
from vrlab_toolbox.processing.trial_intervals import TrialIntervals
my_intervals = TrialIntervals(
intervals={
"Baseline": (0.0, 60.0),
"Stress": (60.0, 120.0),
"Recovery": (120.0, 180.0),
}
)
len(my_intervals) # 3 — one per named interval, via __len__
my_intervals.intervals # the underlying dict, always kept sorted by start time
Real example 1: unlabelled, straight from trigger pulses
TrialIntervals.from_raw_interval_pairs is a shortcut constructor for
when all you have is a plain list of (start, end) pairs, with no
meaningful names yet — exactly the situation right after detecting raw
trigger pulses, before anything's matched to behaviour:
from vrlab_toolbox.processing.trial_intervals import TrialIntervals
raw_pairs = [(0.0, 60.0), (60.0, 120.0), (120.0, 180.0)]
unlabelled = TrialIntervals.from_raw_interval_pairs(raw_pairs)
# {"TP0": (0.0, 60.0), "TP1": (60.0, 120.0), "TP2": (120.0, 180.0)}
That TP0, TP1, … naming is exactly what labels the blue "Raw
biopac" row in Interval QC Plot's example
plot — this constructor is what produces it.
Real example 2: labelled, straight from behaviour data
Compare that to get_crane_trigger_behav_intervals
(crane_trial_intervals.py), which builds a TrialIntervals from a
participant's behaviour dataframe instead, using real trial names:
def get_crane_trigger_behav_intervals(validated_behav_df: pd.DataFrame) -> TrialIntervals:
intervals_out = {}
for _, row in validated_behav_df.iterrows():
intervals_out.update(
{
f"{row['BlockType']}_{row['TrialType']}_{row['TrialNr']}": tuple(
[row["TrialStartTime"], row["TrialEndTime"]]
)
}
)
return TrialIntervals(intervals=intervals_out)
This is what produces names like NonStressBlock_NonSlipTrial_1 — the
orange "Raw behav" row on the same plot. Matching this object
against the unlabelled one above (correcting for clock drift along the
way) is exactly what CraneGetTrialIntervalStrategyStep (below) does.
Real example 3: what the matched result actually contains
Running the full matching step (CraneGetTrialIntervalStrategyStep) on
the synthetic DUMMY000 participant's raw, pre-BIDS output in examples/
(see Testing) produces a TrialIntervals
with 23 entries — this is .intervals, first five shown:
{
'NonStressBlock_NonSlipTrial_1_Training': (4.9845, 65.3135),
'ITI_0': (65.3135, 94.522),
'NonStressBlock_SlipTrial_2_Training': (94.522, 155.1415),
'ITI_1': (155.1415, 205.774),
'NonStressBlock_NonSlipTrial_3_Training': (205.774, 266.104),
...
}
This is exactly the object drawn as the red "Matched with Behav" row in Interval QC Plot's example plot — real trial names, timestamps corrected onto the trigger channel's clock.
The two strategies behind the QC plot
crane_trial_intervals.py defines two strategy steps (see Design
Patterns for what "strategy step" means
here):
CraneGetTrialIntervalStrategyStep— the main path. Needs both physiology and behaviour data, and produces all four rows of the QC plot, including a filled-in "Matched with Behav" row.CraneGetTrialIntervalStrategyFallbackStep— set as the main step'sfallback_strategy, and used when matching against behaviour data fails (see Pipeline Rules). It only has trigger data to work with, so its version of the plot only fills rows 1–3 — the "Matched with Behav" row stays empty, and physiology is processed against unlabelled trigger intervals instead. See Interval QC Plot: when behaviour data is missing or doesn't match for what that looks like on the plot itself.
PipelineStatus
A type → ProcessingStatus mapping (OK, PARTIAL, CORRECTED,
ERROR, NOT_RUN) — one entry per data type the pipeline touched for a
participant. Every Sequential*Steps.run() returns one alongside its
data, and PipelineTemplate.run() merges them all into a single status
for the participant.
@dataclass
class PipelineStatus:
status: dict[type, ProcessingStatus] = field(default_factory=dict)
def set(self, data_type, status: ProcessingStatus):
self.status[data_type] = status
def merge(self, other: "PipelineStatus") -> "PipelineStatus":
... # worst status per type wins, across both
from vrlab_toolbox.processing.processing_status import PipelineStatus, ProcessingStatus
from vrlab_toolbox.processing.crane_behaviour import RawCraneBehaviourData
status = PipelineStatus()
status.set(RawCraneBehaviourData, ProcessingStatus.OK)
status.get_as_text() # "RawCraneBehaviourData=ok" — this is what ends up in the output row
PipelineOutputData — and why it has child classes
One participant's output: a one-row subject_df_out DataFrame, a
status (PipelineStatus), and any QC figure_data_out. Every
processing step returns one; .merge() combines two into one, always
staying one row — see One participant at a time
for why that invariant matters.
@dataclass
class PipelineOutputData:
subject_id: str
subject_df_out: pd.DataFrame = field(init=False)
figure_data_out: dict[str, Figure] = field(default_factory=dict, init=False)
validation_schema: pa.DataFrameSchema = field(default_factory=build_base_pipeline_output_schema)
status: PipelineStatus = field(default_factory=PipelineStatus, init=False)
def validate_participant_output(self) -> pd.DataFrame:
return self.validation_schema.validate(self.subject_df_out)
Every experiment ends up with its own output columns — Crane's row looks
nothing like the Long Walk pipeline's row. Rather than rewrite .merge(),
.append_dataframe(), and the row-shape rules per experiment, each
experiment defines a child class that changes only the one thing that
actually differs — which Pandera schema validates the row:
@dataclass
class CranePipelineOutputData(PipelineOutputData):
def validate_participant_output(self) -> pd.DataFrame:
return build_crane_participant_output_schema().validate(self.subject_df_out)
That's the whole class. merge(), append_dataframe(), figure_data_out,
the one-row rule — all inherited unchanged from PipelineOutputData, same
as Salad inheriting Dish.__init__ above. You'll see the same pattern
for CraneBehaviourOutputData, CraneDebriefPipelineOutput,
EdaPhysiologyOutputData, LongWalkPipelineOutputData, and
FohPipelineOutputData (foh_pipeline.py) — one override each, nothing
more, all still exactly one row per subject.
Going further
FohPipelineOutputData.validate_participant_output() currently
validates against a bare pa.DataFrameSchema() — the override exists,
but build_foh_participant_output_schema() hasn't been filled in with
real columns yet, so validation is a no-op today. Once it is, it'll
follow the same "build from constants" shape
build_crane_participant_output_schema() already uses. See
Lab Streaming for the rest of FOH's anticipated
pipeline shape.
Declaring fields: = default vs. field(default_factory=...)
All three classes above use @dataclass, and all three mix both styles of
default. The difference matters, and Python enforces it rather than just
recommending it:
@dataclass
class Broken:
intervals: dict[str, tuple[float, float]] = {} # in-place literal default
ValueError: mutable default <class 'dict'> for field intervals is not
allowed: use default_factory
A plain = {} (or = [], = SomeDataclass()) would mean every instance
starts out sharing the exact same dict object — mutate one participant's
intervals, and every other TrialIntervals ever created without an
explicit value would see the change too. @dataclass refuses to let you
write that by accident:
@dataclass
class TrialIntervals:
intervals: dict[str, tuple[float, float]] = field(default_factory=dict)
field(default_factory=dict) says "call dict() fresh, once per
instance" instead — the fix, not just a longer way to write the same
thing. The same reasoning is why PipelineOutputData.status is
field(default_factory=PipelineStatus, init=False) rather than
= PipelineStatus(): a fresh PipelineStatus() per participant, not one
PipelineStatus silently shared across every PipelineOutputData ever
built.
Plain, immutable defaults (str, int, float, bool, None) don't
have this problem — they can't be mutated in place, so there's nothing to
accidentally share. That's why subject_id: str above never needs
field(...) at all; only the mutable ones do.
See also: Design Patterns for the Protocol
contracts these classes flow through, and Data Stores
for where RawBioData/RawBehaviourData instances (a related but
separate pair of classes) get held between import and processing.