A high-rigor cocktail-party paradigm probing the neural dynamics of endogenous attention switching between a Familiar (English) and an Unfamiliar (Spanish) talker, read out offline with CCA / envelope tracking.
Two concurrent talkers (one male, one female; spatially separated) play for ~60 s. The subject attends one, and at a jittered mid-trial cue re-orients to the other. A subtle acoustic-anomaly counting task keeps attention engaged without a semantic confound between languages.
stream A ─┐ equiluminant cue (25–35 s, jittered)
attend A ────────▶ ├──▶ 60 s ──┼──▶ attend B ──────────────▶
stream B ─┘ │
markers: TRIAL_START ANOMALY… CUE / SWITCH ANOMALY… TRIAL_END RESPONSE
LSL env: ├───────────── envL, envR, attended_side @128 Hz ──────────────────┤
diode: ▉ onset edge ▉ cue edge (occluded corner)
| Factor | Levels | Balancing |
|---|---|---|
| Attention transition | FF, FU, UF, UU | 10 each (exact) |
| Attended-first side | L / R | 5 / 5 within every condition |
| Attended-first sex | M / F | 5 / 5 within every condition |
| Target answer (# anomalies) | 0,1,2,3 | 10 each, decorrelated from condition |
| Switch time | U(25, 35) s | per-trial jitter (anti-anticipation) |
| ITI | U(3, 5) s | per-trial jitter (baseline recovery) |
| Breaks | after trials 13 & 26 | fixed |
F = English (familiar), U = Spanish (unfamiliar). Everything is deterministic
from (config.seed, subject_id).
run.py drives the whole pipeline; config.yaml controls everything.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python run.py S01 --build-speech en,es # build real speech, design, stimuli, then run
python run.py S01 --synthetic # placeholder speech (timing dry-run, no downloads)
python run.py S01 --prepare-only # stop after rendering stimuli
python run.py S01 --run-only # just run a prepared subject
python run.py S01 --config my_config.yaml # use an alternative config# 1. Real speech bank -> data/raw/{en,es}/{M,F}/*.wav (streams HuggingFace, once)
python -m src.build_speech_bank --all # en: LibriSpeech, es: VoxPopuli
python -m src.build_speech_bank --lang en --split dev_clean --max-speakers 2 # quick test
# (or bring your own single-talker files into data/raw/<lang>/<sex>/)
# 2. Balanced design for a subject
python -m src.build_design S01
# 3. Render 40 trial stimuli (LUFS, anomalies, spatialise, envelopes)
python -m src.prepare_stimuli S01 # add --synthetic to dry-run without audio
# 4. Calibrate audio latency/jitter (physical loopback cable OUT->IN)
python -m src.measure_latency --list
python -m src.measure_latency --device <idx> --reps 100
# 5. Start LabRecorder, record EEG + SpeechEnvelope + AttnSwitchMarkers, then:
python run.py S01 --run-only # or: python -m src.run_experiment S01
# 6. Offline
python -m src.analysis.cca_align recording.xdf --plotSpeech source: English (Familiar) streams gilkeyio/librispeech-alignments
(real LibriSpeech read speech, with sex); Spanish (Unfamiliar) streams
facebook/multilingual_librispeech config spanish (Spanish audiobook speech,
register-matched to the English). MLS has no gender label, so speaker sex is
assigned from median F0 (validated 100% on the English speakers, with a guard
band that skips ambiguous pitches). Both are config-driven (speech_sources: in
config.yaml) and swappable. Note: facebook/voxpopuli is script-based and
stalls on datasets>=5 — don't use it; Common Voice es works but is gated.
Per-trial reports: after each trial the subject answers the anomaly count
(0–3) and rates switch difficulty 1–5 (rating.switch_difficulty in the
config; logged + DIFFICULTY LSL marker).
- Backend: PsychPortAudio, aggressive latency (
latency_class=3). On Windows select an ASIO device; on macOS/Linux CoreAudio/ALSA under PortAudio. - True onset: we schedule each buffer to a future vsync
(
getFutureFlipTime('ptb')) and read the card's realStartTime/ElapsedOutSamples. Markers/envelope are anchored to that DAC onset, never to "now". Seesrc/audio_engine.py,src/run_experiment.py. - One clock:
ClockMapmeasures theGetSecs → local_clock()offset (median of interleaved reads) so all hardware event times map exactly to LSL. - Photodiode: pulses on the audio-onset frame and the cue-onset frame. Mount the sensor over the corner patch so the subject never sees the luminance edge. Combined with the measured audio latency it yields audio onset in EEG time independent of software timestamps.
- Equiluminant cue: the grey fixation cross performs a smooth 0→45→0° rotational wobble — no net luminance/area change → minimal VEP.
- Envelope for CCA: the continuous 3-ch LSL stream (
envL, envR, attended_side) is redundancy; the ground truth is the saved per-stream envelope arrays + the onset marker.analysis/cca_align.pyinterpolates EEG onto the envelope's timestamps (shared LSL clock) for sample-accurate alignment.
- Loopback latency (
measure_latency.py) covers DAC+ADC+cable. For the true acoustic onset add the headphone/air constant via a StimTracker or a microphone tap feeding a TTL / spare EEG channel. - Use hard-ish L/R panning (default) for clean per-stream envelopes, or set
audio.spatialisation: hrtfwith a SOFA HRIR for binaural spatial release.
run.py top-level launcher (build speech -> design -> prepare -> run)
config.yaml all timing/signal/balancing parameters (single source of truth)
src/config.py loader (honours LINGU_CONFIG for an alternative file)
src/build_speech_bank.py stream real speech from HuggingFace -> data/raw/<lang>/<sex>/
src/download_stimuli.py archive.org / LibriVox fetch (alternative source)
src/build_design.py balanced, jittered, validated 40-trial design
src/preprocessing.py LUFS/RMS, pan/HRTF, anomaly gating, envelope extraction
src/prepare_stimuli.py renders per-trial stereo WAV + envelopes + metadata
src/lsl_utils.py clock map, marker outlet, continuous envelope outlet
src/audio_engine.py PsychPortAudio wrapper (onset time, playout position)
src/run_experiment.py main flip-locked presentation loop
src/measure_latency.py loopback latency/jitter calibration
src/analysis/cca_align.py XDF load, EEG↔envelope alignment, CCA + AAD decoding
Recorded sessions live under data/sessions/EEG/**/*_eeg.xdf (LabRecorder: EEG +
SpeechEnvelope + AttnSwitchMarkers). Run the whole pipeline over all of them:
python -m src.analysis.run_analysis # -> data/analysis/report.html (+ results.json)Each session is loaded (pyxdf clock-sync) → preprocessed → aligned → epoched;
then a subject's sessions are merged (trials pooled, ERP epochs concatenated) and a
self-contained HTML report (data/analysis/report.html) is produced with:
-
AAD — leave-one-trial-out attended-talker decoding (overall + per transition).
-
Envelope tracking — attended vs unattended canonical correlation (paired, per trial)
- the CCA forward-pattern scalp topography.
-
Switch-locked ERP — grand-average butterfly + GFP with topomaps (endogenous re-orientation; the cue is equiluminant, so it is not a VEP).
-
Re-orientation dynamics — switch-locked pre- vs post-talker tracking (mean ± SEM).
-
Switch-time decoding — the switch time estimated from EEG alone (zero-crossing of corr(EEG,post)−corr(EEG,pre)), with the true-vs-decoded scatter and error distribution.
-
Preprocessing (
src/analysis/preprocess.py): 1 Hz high-pass → bad-channel detection (impedance stream + robust log-variance) → STAR — Sparse Time Artifact Removal (de Cheveigné 2016,meegkit), the dry-EEG glitch cleaner → interpolate bads → common-average reference → FastICA ocular removal → 32 Hz low-pass. -
Alignment (
src/analysis/decode.py): EEG interpolated onto the envelope's DAC-anchored 128 Hz LSL grid; attended/unattended from theattended_sidechannel. -
Decoding: the de Cheveigné CCA (the project's
aud_cca, vendored undersrc/analysis/cca_decoding/) — FIR-lagged EEG ~ FIR-lagged attended envelope, PCA-truncated whitening, leave-one-trial-out CV; a trial is decoded as attended = the stream with the larger leading canonical correlation. A switch-locked sliding window traces the re-orientation. All knobs live underanalysis:inconfig.yaml.