MusWM: A Capacity-Limited Architecture for Musical Working Memory
This work advances an adaptive theory that complements prevailing statistical-learning accounts: those accounts address WHAT cognitive processes predict, whereas MusWM addresses HOW. It also introduces a real-time tool that synchronizes musical, physiological, and behavioral information to validate or falsify its hypotheses.
For improved readability, view on a mobile device or use your browser’s zoom controls on desktop.
01 WHY A WORKING-MEMORY ARCHITECTURE?
Prevailing expectation models in statistical learning are powerful—but cognitive capacity, attention, and multimodal integration remain implicit. Their black-box character also limits explainability, auditability, and accountable use—core requirements of Responsible AI. While statistical learning addresses WHAT cognitive processes predict, Musical Working Memory (MusWM) specifies HOW these processes operate through biologically and ecologically valid mechanisms.
STATISTICAL LEARNING
- • Offline symbolic analysis
- • Unimodal
- • Expert annotation
- • Memory argument is analogy and left implicit
MusWM
- • Real-time symbolic analysis
- • Multimodal
- • Automatic Annotation
- • Memory argument is explicit
PITCH-CLASS TOKENIZATION
→ bass-relative [0,4,7] → I (C major)
Duplicates and octave doublings collapse to unique pitch classes.
COGNITIVELY MOTIVATED CONTROLS
- Capacity-limited pitch-class buffer Tempo-relative decay window
- Bar-synchronous inhibitory reset Tonal reorientation
- Interval-based chord classification Functional label co-indexed with key
02 BADDELEY → MUSICAL WORKING MEMORY
Widely recognized statistical models describe machine-learning processes rather than human memory mechanisms, leaving cognition analogical and memory processes implicit. MusWM instead operationalizes Baddeley’s working-memory mechanisms (Baddeley & Hitch, 1974; Baddeley, 2000) as explicit and inspectable.
MusWM outperforms black-box models in real-time Roman-numeral analysis, as demonstrated in the companion SMPC poster, “A Real-Time Multimodal Architecture for Musical Expectation.”
1 BADDELEY’S SYSTEM
2 MusWM · MUSIC APPLICATION
SLAVE SYSTEMS
In MusWM, each slave system is modality specific: auditory, visual–spatial, olfactory, gustatory, and somatosensory information is maintained in a separate time-limited buffer. Controlled decay preserves each stream independently before cross-modal integration.
Each modality-specific subsystem contributes a distinct sensory representation; episodic integration binds them into a coherent percept of the world.
EPISODIC BUFFER
The episodic buffer binds the independently maintained modality-specific streams into a unified, temporally coherent musical episode. It makes integrated states available for awareness and analysis while admitting only persistent, structurally coherent information.
03 LIVE MULTIMODAL INTEGRATION
Every musical, physiological, and behavioral event is timestamped at millisecond resolution and broadcast across active streams.
MIDI EVENTS
Ableton Live
EEG · 8 CH
250 Hz · band power
CARDIOVASCULAR
HR · HRV / RMSSD
BEHAVIORAL
face · valence · Go/NoGo
shared boundaries
REAL-TIME HARMONIC LAYER
O(1) = constant-time lookup, enabling reliable real-time use.
Currently used: 41 of 2,048 possible voice combinations for Roman-numeral generation and Western music scales.
PRELIMINARY FEASIBILITY
Continuous harmonic annotations synchronized with EEG band-power, cardiovascular indices, and behavioral responses across varied harmonic contexts.
Successful cross-domain binding in real performance conditions.WHY IT MATTERS
The world’s first research tool for investigating harmonic expectation, melodic affect, executive load, and affect–structure coupling in real time in music cognition research and experiments.