BIDS Crosscheck: Architecture
BIDS Crosscheck Plan covers the design decisions behind this tool — why it exists, what's deliberately out of scope, the mockup it was built from. This page is the code map: what's actually built, where it lives, and what to touch to change or extend it. If you're looking to use the tool instead, see FOH Crosscheck or Crane Crosscheck.
Philosophy
- Human-in-the-loop, decisions recorded not applied. The tool never guesses which file is right. Every choice a person makes is written to a JSON file before anything on disk changes, so the record survives even if the app crashes mid-operation.
- One tool per dataset, not one app with a dataset switcher — matches
this repo's existing pattern of one CLI/entry point per pipeline
(
vrlab_crane_process,vrlab_foh_assess_data, …). - No automatic collision resolution. If a rename ever leaves two files looking like duplicates again, that's treated as an ordinary duplicate, routed through the same review — deliberately one mechanism, not two.
- Additive-only. Nothing here changes
ParticipantConfigor any existing pipeline code — see the plan's own guardrail section for the full reasoning.
Code layout
processing/bids_crosscheck.py Qt-free engine: scanning, JSON I/O, corrections
gui/bids_crosscheck_common.py shared PySide6 window (BidsCrosscheckWindow)
gui/crane_bids_crosscheck_gui.py thin entry point: crane's DatasetConfig
gui/foh_bids_crosscheck_gui.py thin entry point: FOH's DatasetConfig + CandidateExtras
tests/test_bids_crosscheck.py tests for the engine (no Qt dependency)
Everything Qt-specific stays out of processing/bids_crosscheck.py on
purpose — it's plain Python + pathlib, testable without a display, and in
principle reusable from a future non-GUI entry point.
Key types
| Type | Lives in | What it is |
|---|---|---|
DatasetConfig |
processing/bids_crosscheck.py |
One dataset's scan types, glob patterns, and whether/how task-tagging (FOH's "Tag with foh BIDS tags") applies -- task_tag_task/task_tag_acq/task_tag_suffix (the BIDS entities/suffix a tag writes) plus an optional task_tag_folder_name, which also renames a tagged file's parent folder (e.g. FOH's eeg/ -> beh/), carrying along anything else still in it. |
ScanTypeConfig |
same | One scan type's name + the glob patterns that find its candidate files. |
SubjectScan |
same | One subject/scan-type's candidate files, with a status property ("ok" / "missing" / "duplicate", derived purely from file count). |
BidsFolderScan |
same | The result of scanning a whole folder — every subject × scan-type, plus has_issues(). |
CandidateExtras |
gui/bids_crosscheck_common.py |
The hook dataset-specific GUIs override to add per-candidate info (see FOH below). Crane uses the no-op default. |
BidsCrosscheckWindow |
same | The actual master-detail window: subject table (Subject / Tag / Datatype / Info columns) on the left, per-scan-type detail panel on the right. |
BidsCrosscheckWindow never imports anything crane- or FOH-specific — it
only knows DatasetConfig and the CandidateExtras protocol. That's what
lets crane_bids_crosscheck_gui.py and foh_bids_crosscheck_gui.py stay a
few dozen lines each.
The subject table's Tag and Datatype columns
_build_subject_tag_widget shows, per tag-eligible scan type (i.e.
extras.task_tag_available(scan_type) is true), whether the
currently-effective candidate carries the task-<task> marker yet -- the
task value itself if tagged, "not tagged" in PLEASE_SELECT_COLOR if not.
Blank for a scan type that doesn't support tagging at all (e.g. every one of
Crane's), or with nothing effective yet. Same underlying check as
_needs_task_tag, just surfaced as its own column instead of only
the 🏷 icon.
_build_subject_datatype_widget shows, per scan type, the BIDS datatype
folder (file.parent.name) the currently-effective candidate lives in right
now -- blank if there's nothing effective yet (missing, or an unresolved
duplicate). _datatype_tooltip explains the value on hover: what the current
folder is, and -- if task_tag_folder_name is configured and differs
from the current folder -- what it becomes once tagged. BIDS_DATATYPE_NAMES
(a small, dataset-agnostic dict of BIDS's own datatype abbreviations, e.g.
"beh": "behavioural") glosses both in plain English wherever it recognizes
the folder name; an unrecognized one is shown without a gloss rather than
guessed. Deliberately doesn't assert anything about the current folder's
accuracy (only the target one, which the dataset config author chose on
purpose) -- FOH's eeg/ is a real example of a current folder whose name is
simply wrong for what it contains.
Column order is Subject / Tag / Datatype / Info, left to right.
The artifacts a BIDS folder ends up with
The tool never touches anything outside the BIDS folder it's pointed at (see the plan's "BIDS folder only" decision). Inside it, these appear as you use the tool:
crosscheck.json— finalized decisions:selected_run,date_correction,id_correction,task_tag/task_tag_removed,crosschecked. Keyed by{subject_id}_{scan_type}(_decision_key), exceptcrosscheckeduses a different key (_crosschecked_key, same base plus_crosschecked) so marking something crosschecked can never overwrite an existingselected_runentry for that same subject/scan-type. Update (2026-09-01): each key's value is a list of decisions, oldest first (_append_decision/_latest_decision), not one dict — a later correction on the same key no longer erases an earlier one. An old single-entrycrosscheck.jsonis transparently upgraded to[entry]on load, so nothing already on disk needs migrating by hand. Most callers (crosschecked_scan_types,existing_subject_ids) only need_latest_decision, i.e. current state;revert_all_decisionsandrebuild_from_raw(below) are the two that walk the full history.crosscheck_pending.json— radio picks made but not yet committed (load_pending_selections/save_pending_selections). Kept in a separate file fromcrosscheck.jsonon purpose: these aren't decisions yet, just in-progress GUI state, autosaved so closing the app before clicking "commit" doesn't lose the picks. Filenames only, not full paths — re-resolved against a fresh scan's actual candidates on load, so a stale entry (file renamed/deleted outside the tool) is silently dropped rather than crashing.excluded_subjects.json—{subject_id: reason}for every subjectrecord_subject_excludedhas removed from BIDS. Unlike the earliercrosscheck_junk//crosscheck_review/folders (removed 2026-08-27), a removed subject'ssub-XXX/folder is deleted outright, not moved anywhere -- safe only because neither importer/converter ever touches the raw folder (see "BIDS folder only" in the plan), so it's always the real recoverable copy.existing_subject_idsunions this file's keys with the folders actually on disk, so an excluded subject doesn't look "new" again to the next raw-to-BIDS refresh.record_selected_run(a duplicate's non-selected candidate) deletes the file directly instead and does not write here -- it's still recorded incrosscheck.json'sselected_runentry, which is enough for restore (below) to find it.restore_all_from_bidsreverses both: it clears this file entirely, and for everyselected_rundecision it finds, deletes that subject's wholesub-XXX/folder and the decision itself, so the next raw-to-BIDS refresh re-derives them fresh. There's no more permanent-delete operation in this module -- deleting was already the only operation, so a separate "confirm forever" action on top of it would be redundant.
Both JSON writes go through _write_json_atomic — write to a .tmp file,
then os.replace() — so a crash mid-write can't corrupt either file.
Backup and disaster recovery
Everything in a BIDS folder is either raw data (already safe — the raw folder
is never touched, see "BIDS folder only" in the plan) or one of the JSON
files above; only crosscheck.json/excluded_subjects.json/
crosscheck_pending.json actually need backing up (crosscheck_info_cache.json
is a re-derivable performance cache, not a decision record). Two functions
handle this:
backup_decisions(bids_folder, backup_folder)— copies those three files tobackup_folder. Wired to the GUI's "Backup crosscheck data..." button.rebuild_from_raw(bids_folder, decisions, excluded)— the recovery side. Assumesbids_folderhas just been freshly re-imported from raw (the GUI runs the existingraw_converter/"Refresh BIDS" hook first, unchanged), then replays every decision's filesystem effect against it — the same rename/delete eachrecord_*function performs, but computed directly from what the entry already recorded (original_filename→corrected_filename, etc.) rather than recomputed from scratch, so there's one definition of what each decision type means, not two. The one exception isid_correction, which re-applies the same original/corrected id token-replacement rulerecord_id_correctionuses, since its entry stores the original relative paths rather than each file's new name.
Every entry (across every key) is retried pass after pass until a pass
makes no further progress, rather than walked once in a fixed order — two
things can make an entry temporarily unresolvable: it isn't the oldest
step in its own key's chain yet (its original_filename doesn't exist
until an earlier entry in the same list produces it), or it depends on an
id_correction recorded under a different key finishing first. Both
resolve themselves once whatever they were waiting on succeeds on an
earlier pass — no explicit ordering/timestamp needed, and notably not
derivable from crosscheck.json's on-disk key order anyway, since
_write_json_atomic writes with sort_keys=True. Whatever still can't
resolve once no pass makes progress is reported unresolved rather than
guessed at (see "no automatic collision resolution" in the plan for why).
Wired to the GUI's "Rebuild from backup..." button, which loads
decisions/excluded from a chosen backup folder (via the same
load_decisions/load_excluded_subjects already used for the live BIDS
folder — both take a bare folder path) before calling this.
Not every dataset-specific correction fits crosscheck.json's "decision
about an already-converted file" shape, though — crane's
debrief_id_corrections.json/raw_filename_id_corrections.json
(cli/crane_convert_to_bids.py, surfaced via its own extra_raw_actions
dialogs) are inputs the converter itself reads on the next "Refresh BIDS"
run, not something to replay afterward. BidsCrosscheckWindow's
extra_backup_filenames constructor parameter (set by
crane_bids_crosscheck_gui.py, empty for FOH) names these so
backup_decisions includes them too, and a new restore_backup_files
function puts them back into bids_folder before raw_converter runs
during a rebuild — the window itself doesn't know what these files mean,
only that they need to travel with a backup the same way crosscheck.json
does (the same "this window doesn't know what the callback does" contract
extra_raw_actions already uses).
Extending: adding a new dataset
A new dataset needs exactly two things: a DatasetConfig, and a main()
that calls run_bids_crosscheck_app. crane_bids_crosscheck_gui.py is the
minimal template (no CandidateExtras override):
from vrlab_toolbox.gui.bids_crosscheck_common import CandidateExtras, run_bids_crosscheck_app
from vrlab_toolbox.processing.bids_crosscheck import DatasetConfig, ScanTypeConfig
CRANE_DATASET_CONFIG = DatasetConfig(
dataset_name="crane",
scan_types=(
ScanTypeConfig(name="physiology", glob_patterns=("*physiology*",)),
ScanTypeConfig(name="behaviour", glob_patterns=("*behaviour*",)),
ScanTypeConfig(name="debrief", glob_patterns=("*redcap*",)),
),
)
def main() -> None:
run_bids_crosscheck_app(
CRANE_DATASET_CONFIG, "Crane BIDS Crosscheck", CandidateExtras(),
settings_app_name="CraneBidsCrosscheck",
)
foh_bids_crosscheck_gui.py shows the richer path: a CandidateExtras
subclass (FohCandidateExtras) that overrides describe() to show a
recording's date/duration/stream-presence (colored HTML, parsed via
pyxdf + processing/lsl.gather_xdf_data_streams/get_start_time), and
task_tag_available() to enable the "Tag with foh BIDS tags" button.
Every parse result is cached per-file in self._info_cache for the life
of the window — see the caching TODO below for the next step (persisting
that across restarts too).
Add settings_app_name (used as the QSettings application name — see
below) unique per dataset, and register a console-script entry in
pyproject.toml's [project.scripts], matching the existing
vrlab_crane_bids_crosscheck / vrlab_foh_bids_crosscheck pattern.
Libraries used
- PySide6 — the GUI framework itself (official Python bindings for Qt).
QSettings(part ofPySide6.QtCore) — remembers the last-opened BIDS folder across restarts, one registry/plist entry persettings_app_name, so crane and FOH don't share state.rich.progress.Progress— the same progress-bar library the CLIs already use (vrlab_crane_process.py,mobi_FOH_assess_data.py). Used for exactly one case: the very first auto-restore on startup happens beforewindow.show(), so the Qt progress bar would be invisible —BidsCrosscheckWindowfalls back to a terminal bar viaself.isVisible(), then switches to the normal Qt one for every subsequent rescan.pyxdf— FOH-only, for reading stream/timestamp info out of.xdffiles (already a toolbox dependency viaprocessing/lsl.py).
Known gaps / TODOs
- FOH info caching isn't persistent yet — logged in BIDS Crosscheck Plan.
- Bulk FOH-rename silently skips unpicked subjects — no warning is shown when a subject is skipped because no recording had been picked yet. See BIDS Crosscheck Plan.
- Crane parity with FOH's
CandidateExtras— FOH is ahead (richdescribe()info,task_tag); crane still uses the no-op base. Deliberately one dataset at a time — bring crane's GUI up to match FOH's once a good crane-specific info source is identified, not scoped yet. - Crane's glob patterns are still placeholders — see the note at the
top of
crane_bids_crosscheck_gui.py; there's no real crane BIDS output to check them against yet (see BIDS Converter Plan).
Testing
tests/test_bids_crosscheck.py covers processing/bids_crosscheck.py
directly — no Qt, no display needed, plain unittest.TestCase with
tempfile.mkdtemp() BIDS-folder fixtures (see TestScanBidsFolder for the
pattern). GUI-level behaviour has been verified with ad hoc headless
scripts (QT_QPA_PLATFORM=offscreen) during development rather than a
committed GUI test suite — worth formalizing if this tool keeps growing.
Also see: FOH Crosscheck and Crane Crosscheck for the user-facing walkthroughs, and BIDS Crosscheck Plan for the original design rationale.