Skip to content

Crane Pipeline Milestones

Running log of architectural/skill milestones in this repo, for REVIEW mode (see AGENTS.md) to compare current work against. Not a full changelog — only entries that mark a real shift in approach (new pattern, new tooling, new discipline), so progress over months is visible at a glance.

Date Commit Milestone Why it matters
2026-03-11 70e2d40 First EDA/XDF processing script (check_plux_data.py) Flat exploratory script: hardcoded paths, print() diagnostics, no tests, no modules. The baseline to measure everything else against.
2026-04-14 a6a8d52 Refactor check_plux_data into focused processing modules First split from one script into separate processing modules.
2026-04-16 fa18788 / c55501b Added basic logging, then improved it print() → logging module. Diagnostics become structured and filterable.
2026-06-10 fa5e52d Added basic unit test to crane pipeline First test in the crane path — testing becomes part of the workflow, not an afterthought.
2026-06-19 d444a5c Implemented status class Processing outcomes become a typed Status/PipelineStatus object instead of ad-hoc bools/strings.
2026-07-01 189389a Added template and strategy base classes. Refactored key data types Strategy/Template pattern introduced — the architectural turning point from "pipeline as a sequence of function calls" to "pipeline as composable, typed steps."
2026-07-06 1b700df / 11177da Added Ruff settings, initial ruff reformat Static analysis and consistent formatting adopted project-wide.
2026-07-17 032d3dd First implementation of pipeline.py for crane. Tests passed First green run of the new pipeline architecture end-to-end.
2026-07-22 2d91140 Tidied up a bit. Added deprecation warnings @deprecated used to retire flawed logic explicitly instead of silently deleting or leaving it live — keeps the migration honest.
2026-07-24 4ca8143 (WIP tip at last REVIEW) Logging added; pipeline runs through alignment; edge-case tests in place but 3 still failing Most recent REVIEW checkpoint. Trigger/behaviour alignment refactored to return real PipelineStatus (was previously a stub); remaining gaps are documented edge cases (PID4572, PID16230, PID9188), not silent failures.
2026-07-24 cc639d9 (WIP tip) Placeholder pass tests replaced with real assertions; align_biopac_trigger_drift_from_behav_file stopped discarding its own result The 6 edge-case tests that were silent pass stubs now assert and fail loudly where the algorithm genuinely can't yet handle the subject's data — test red is now signal, not a gap in coverage. Also added a remove_crane_delayed_start heuristic and switched fill_in_gaps to return a new sorted TrialIntervals instead of mutating in place.
2026-07-28 669645e PipelineStatus rewritten to type-keyed tracking; two UnboundLocalError classes of bug root-caused and fixed via captured-output-type-before-try; None-input guard added instead of widening an except tuple; dead code (CraneDebriefOutputData) deleted rather than left to rot The two SequentialBehaviourImportSteps/SequentialPhysiolgyImportSteps bugs and the AttributeError-swallowing fix show a specific, reusable diagnostic move — noticing that catching a broad exception type is "working by luck," not by design — and choosing the narrower, more correct fix over the shortcut, even after the shortcut already made tests pass. All three previously-@unittest.skipped tests are unskipped and green.

| 2026-07-30 | af1d483 (WIP tip) | LSL/FOH begins moving onto shared raw-data contracts | LslPhysiologyDataImportStrategy now returns RawBioData, and raw physiology labels are normalized at the contract boundary; this shows the pipeline architecture starting to become reusable toolbox infrastructure rather than Crane-only structure. | | 2026-08-07 | f720a47 (WIP tip) | Batch CLI (mobi_FOH_process_batch.py) migrated from deprecated run_lsl_pipeline to run_pipeline/PipelineTemplate; test_basic_foh_pipeline strengthened from a single truthiness check to per-datatype status + figure assertions | First real entry point proving the FOH PipelineTemplate end-to-end through an actual CLI, not just direct unit tests. mobi_FOH_process.py (the single-file CLI) was deliberately left on the deprecated path rather than migrated in the same commit — matches this project's own "one structural issue at a time" working rule, not an oversight. |

| 2026-08-14 | 192385f (WIP tip) | BIDS Crosscheck tool built end-to-end: bids_crosscheck.py module (atomic JSON writes, cross-platform case-insensitive collision handling, full audit trail of renames/junking/reverts), a shared PySide6 GUI driving both a Crane and a FOH variant, and 810 lines of passing tests | The design doc referenced by a prior REVIEW (bids_crosscheck_plan.md) is now real, working infrastructure, not just a plan. Architecturally significant beyond the tool itself: ParticipantConfig.from_lsl_data (input_data.py) now raises on zero or multiple matching physiology files instead of silently guessing "last one" — safe only because the crosscheck tool now guarantees exactly one file per subject/scan-type upstream. That's a whole-system contract change reasoned from a sibling tool's guarantees, not a local fix. | | 2026-08-24 | a5dfd37 (PR #5) | First outside contribution merged: a student (suzeschulenburg) adds the graphomotor/spiral pipeline (spiral.py, graphomotor_pipeline.py, graphomotor_qc.py, graphomotor_xdf.py, 382 lines of passing tests) and a standalone pull_redcap.py puller | First real multi-contributor workflow, not just solo commits — and the first evidence of what needs to be taught explicitly rather than absorbed from the codebase: the pipeline code picks up the project's own conventions well (dataclasses, type hints, logging, decomposed modules, self-written tests), but pull_redcap.py is a flat top-level script with print()-debugging, no main()/click wrapper, and a committed API_TOKEN = "" placeholder — a credential-shaped landmine if ever filled in and committed live. Both PRs (#5, #6) were merged by the 28866630 account, not reviewed/merged by stefan — worth tightening given the mixed quality. | | 2026-08-28 | WIP tip (uncommitted) | Packaging grows from "build one exe" into a real distribution pipeline: all 11 PyInstaller specs moved into specs/ and rewritten around REPO_ROOT = SPECPATH/.. so they're relocatable regardless of invocation directory; a new vrlab_toolbox_launcher.py PySide6 window fronts every GUI tool; toolbox_installer.iss (Inno Setup) bundles all built exes into one no-admin, per-user installer that manages its own PATH entry; build.ps1 and release.yml rewritten in lockstep to drive the same sequence locally and in CI | First time the project reasons about deployment as a system rather than "does this one script produce an exe": explicit no-admin rationale for lab machines, DiskSpanning pre-empting a size limit before it's hit, WM_SETTINGCHANGE broadcast so PATH updates apply without a logoff, and docs (packaging.md) that name the remaining gap (build_mac.sh still only builds two tools, plainly) rather than paper over it. | | 2026-08-28 | b5e71e0 (WIP tip) | REVIEW confirms test-backed triage after the packaging push: full suite now isolates one live FOH batch regression while Crane/BIDS/packaging work stays broadly coherent with the docs | Shows a more competent maintenance loop: compare docs to code, run the suite, distinguish real red tests from stale notes, and preserve a short punch list instead of rewriting architecture on impulse. | | 2026-09-01 | 3a7fe55 (WIP tip) | Root-caused the installer's silent upgrade-overwrite failure with two real bugs, not one guess: GetUninstallString's registry key was missing the braces Inno actually wraps its AppId in (always querying a key that could never exist), then RemoveQuotes mangling QuietUninstallString's trailing /SILENT into the exe path once the first bug was fixed. Confirmed via temporary MsgBox instrumentation in [Code] rather than re-guessing, then removed once both were verified fixed. Installer filenames and AppVersion now carry the git tag (MooiToolboxSetup-v0.7.2.exe), aligned between build.ps1 and release.yml | Shows debugging a black-box (Pascal Script inside a compiled installer) by adding your own instrumentation and reading the evidence back, instead of pattern-matching to the first plausible cause — the first fix (registry key) looked complete but wasn't; only the second debug run surfaced the quoting bug underneath it. | | 2026-09-01 | f596c44 (HEAD) | Crosscheck decisions storage rewritten from one-entry-per-key to full history-per-key (_append_decision/_latest_decision, transparent old-shape migration in load_decisions), unlocking per-subject "Restore from raw..." and a revert_all_decisions that now walks a whole correction chain instead of only its latest step; paired with a new disaster-recovery pair (backup_decisions/rebuild_from_raw) that replays a backed-up crosscheck.json onto a fresh raw re-import pass-by-pass until nothing more resolves. Crane's raw-filename and debrief-record_id ambiguity (a "(N)" duplicate-copy marker, two subjects sharing one literal record_id) is now surfaced to a human instead of guessed either way. REVIEW's own full-suite run (8 failed / 133 passed / 3 skipped) caught a real regression from this window: crane_debrief_behaviour.py dropped the High_at_start_end column from both schema and dummy-data generator, but the checked-in examples/crane_bids_dummy/ fixture was never regenerated to match, so pandera's now-actually-enforced validation (get_group_debrief_data previously computed and discarded its own .validate() result -- a real bug e8e012b fixed) rejects the stale extra column across 5 dummy-data tests; a 6th failure against real committed data (sub-00011) is that same enforcement correctly catching a genuinely incomplete debrief export, not a code bug; a 7th is a one-line test-authoring gap (missing sub-002/ folder in a new backup test); the 8th (test_crane_batch_processing, FOH) is the same pre-existing failure b5e71e0's REVIEW already isolated, not new. | Shows the payoff of making pandera validation actually run (previously silent no-op): it immediately started catching a real stale-fixture regression and a real incomplete-data subject in the same run it was fixed in -- the fixture-regen step just hadn't caught up yet, which is a process gap (verify after a schema trim), not a design flaw. |

How this file grows

After each REVIEW, append one new row: date, commit hash (or "WIP tip" if uncommitted), a short milestone label, and a one-line note on what it demonstrates skill-wise. Keep entries to real shifts in approach, not every commit — this file is meant to stay short enough to read in one pass years from now.