AuralPrimer turns any song you own into a playable practice surface. Drop in an audio file or a Suno stem export, and the pipeline does the rest: stem separation, beat/tempo analysis, per-instrument transcription, and a falling-note in-game view that scrolls toward a "play here" line you can adjust on the fly.
The shot above is the in-game Band Setup view of the Keys/Synth lane for a real Suno piano stem the pipeline transcribed. Notes fall toward the "PLAY HERE" line; the bright cap at the bottom of each pill is the attack, the fading tail above is the sustain. The detected key signature (
F# minor, 3 sharps) is read straight off the note distribution by the in-process Krumhansl–Schmuckler analyzer inviz-tab. Press [ / ] in-game to spread / compress the falling notes StepMania-style.
This is a two-app desktop suite plus the pipelines, plugins, and research that feed them.
- AuralPrimer — the gameplay app. Loads
.auralsongpacks, plays them, drives the visualizer plugins, and routes live MIDI for practice. - AuralStudio — the authoring app. Imports raw audio / stem folders,
runs the pipeline, and emits
.auralsongpacks the game can load. - Python sidecar pipeline (
python/ingest/) — extraction, stem-separation, and per-instrument transcription tooling shipped as PyInstaller binaries so users don't need a Python install. - Visualizer SDK + plugins (
packages/viz-sdk,visualizers/) — pluggable canvas renderers (drum highway, beats, tab, piano-roll, chord-lane, lyrics) that all consume the same canonical transport state.
git clone https://github.com/wirelessdreamer/AuralPrimer.git
cd AuralPrimer
npm ci # installs the JS workspace + downloads model packs that ship in CI
cargo test --workspace
npm test
python -m pip install -e python/ingest # only needed if you want to call the pipeline from source
# Build a portable bundle (Windows): produces AuralPrimerPortable/{AuralPrimer.exe, AuralStudio.exe, ...}
pwsh ./create_portable.ps1 -PortableRoot ./AuralPrimerPortableSee BUILDING.md for the full install / test / package
matrix and docs/local-dev-prereqs.md for
OS-level prerequisites.
- Test-driven development. Tests first, implementation second; CI stays green.
- AuralSong-first runtime. AuralPrimer consumes
.auralsongpacks as canonical content. The folder watcher pulls in new packs without a restart. - Deterministic imports. Cacheable, reproducible, versioned outputs.
- Local-first shipping. Required tooling lives in the desktop
artifact. ML model weights are NOT bundled in the installer — they
download/import into
assets/models/on first use. - Plugin-first visualization. Visualizers stay decoupled from the
runtime via the
viz-sdkplugin SDK. - Rights-neutral importer scope. PRs for source-specific proprietary game/DLC archive importers are out of scope.
AuralPrimer stands on open-source music-ML research and tooling. Because it
ships commercially, every component below is commercially licensed — and for
the neural models we verify the trained weights license separately from the
code license (an open-code / non-commercial-weights split is exactly what
disqualifies a model here; see the license gate in the research docs). Current
stack: permissive throughout (MIT / Apache-2.0 / BSD / ISC / MPL-2.0 / CC0 /
Unlicense), with LGPL present only as dynamically-linked binaries (ffmpeg,
libsndfile) and GPL only in a build-time tool (PyInstaller, used under its
bootloader exception — its output is not GPL-encumbered). No GPL runtime
dependencies; no non-commercial model weights.
One row per maker, ordered by layer — models first, then ML runtime, app shell, and native audio/MIDI:
| Layer | Organization / maintainer | What AuralPrimer uses from them | License |
|---|---|---|---|
| Model | Meta AI · FAIR — Défossez et al. | Demucs — stem separation | MIT (code + weights) |
| Model | Sony AI — Tan · Cheuk · Mitsufuji | MR-MT3 — neural drum transcription (+ mt3-infer wrapper) |
MIT (code + weights) |
| Model | Google Research · Magenta | MT3 — the architecture MR-MT3 is fine-tuned from | Apache-2.0 |
| Model | Spotify · Audio Intelligence Lab | Basic Pitch — piano / polyphonic melodic transcription | Apache-2.0 |
| Model | CPJKU · JKU Linz — Foscarin · Schlüter · Widmer | Beat This! — beat / downbeat / meter (drives the editor grid) | MIT (code + weights) |
| Model | NYU MARL · Northwestern — Kim · Salamon · Bello · Morrison | CREPE + torchcrepe — bass / guitar pitch | MIT |
| Runtime | Meta Platforms — PyTorch project | PyTorch (torch · torchaudio · torchvision) — ML runtime |
BSD |
| Runtime | Hugging Face | transformers — model architectures | Apache-2.0 |
| Runtime | Lightning AI — William Falcon | PyTorch Lightning — inference scaffolding | Apache-2.0 |
| Runtime | Phil Wang (lucidrains) | x-transformers · rotary-embedding-torch | MIT |
| Runtime | NumFOCUS community — Brian McFee et al. | NumPy · SciPy · scikit-learn · librosa (onset / CQT / tempo) | BSD · ISC |
| Runtime | Bastian Bechtold · libsndfile team | soundfile — audio I/O | BSD (LGPL backend, dynamic) |
| Runtime | The FFmpeg project | ffmpeg — decode binary | LGPL-2.1+ (dynamic) |
| App | Tauri WG · Commons Conservancy | Tauri v2 — desktop shell + dialog/shell plugins | MIT / Apache-2.0 |
| App | VoidZero — Evan You · Anthony Fu | Vite · Vitest — build & test | MIT |
| App | Microsoft | TypeScript · Playwright | Apache-2.0 |
| Native | RustAudio + independent Rust | cpal · rtrb · symphonia · midir · midly — native audio/MIDI | Apache / MIT / MPL-2.0 |
Lineage: Google/Magenta's MT3 → Sony's MR-MT3 (fine-tuned for drums). Full per-component detail — including build/test-only deps — is in the tables below.
| Component | Role in AuralPrimer | Made by | License (code / weights) |
|---|---|---|---|
Demucs (htdemucs_6s) |
Stem separation (vocals/drums/bass/guitar/keys/other) | Meta AI / FAIR — Alexandre Défossez | MIT / MIT |
MR-MT3 (mr_mt3 ckpt) + mt3-infer |
Neural drum transcription (which drum + velocity) | Sony AI — Hao Hao Tan, Kin Wai Cheuk, Yuki Mitsufuji et al.; wrapper by openmirlab | MIT / MIT |
| MT3 (lineage) | Base multi-track transcription architecture MR-MT3 derives from | Google Research · Magenta | Apache-2.0 / Apache-2.0 |
| Basic Pitch | Piano / polyphonic melodic transcription | Spotify — Audio Intelligence Lab | Apache-2.0 / Apache-2.0 |
| Beat This! | Beat / downbeat / meter tracking (drives the editor grid) | CPJKU, JKU Linz — Foscarin, Schlüter, Widmer | MIT / MIT |
| torchcrepe + CREPE | Monophonic bass / guitar pitch tracking | Max Morrison (Northwestern) · CREPE by NYU MARL (Kim, Salamon, Bello) | MIT / MIT |
| librosa | DSP: onset detection, CQT/spectrogram, tempo analysis | librosa dev team (Brian McFee, NYU) | ISC |
| Component | Made by | License |
|---|---|---|
PyTorch (torch, torchaudio, torchvision) |
Meta Platforms (PyTorch project) | BSD-2/3-Clause |
| PyTorch Lightning | Lightning AI (William Falcon) | Apache-2.0 |
| transformers | Hugging Face, Inc. | Apache-2.0 |
| x-transformers, rotary-embedding-torch | Phil Wang (lucidrains) | MIT |
| NumPy, SciPy, scikit-learn | NumFOCUS-sponsored communities | BSD-3-Clause |
| soundfile (+ libsndfile) | Bastian Bechtold · libsndfile team | BSD-3-Clause (LGPL backend, dynamic) |
| beartype | Cecil Curry | MIT |
| ffmpeg (bundled binary) | The FFmpeg project | LGPL-2.1+ (dynamic) |
| PyInstaller (build tool) | PyInstaller Development Team | GPL-2.0+ w/ bootloader exception |
| Component | Made by | License |
|---|---|---|
Tauri v2 (+ @tauri-apps/api, dialog/shell plugins) |
Tauri Working Group / Commons Conservancy | MIT / Apache-2.0 |
| Vite, Vitest | VoidZero (Evan You / Anthony Fu) | MIT |
| TypeScript, Playwright | Microsoft | Apache-2.0 |
| ajv (JSON Schema) | Evgeny Poberezkin | MIT |
| fflate (in-browser zip) | Arjun Barrett | MIT |
| jsdom (test DOM) | jsdom project (Domenic Denicola et al.) | MIT |
| Component | Made by | License |
|---|---|---|
| cpal (native audio out) | RustAudio | Apache-2.0 |
| rtrb (realtime ring buffer) | Matthias Geier | MIT / Apache-2.0 |
| symphonia (pure-Rust decode) | Philip Deljanov | MPL-2.0 |
midir / midly (MIDI I/O + .mid parse) |
Patrick Reisert · Martín Andrighetti | MIT · Unlicense |
| tokio (async runtime) | tokio-rs | MIT |
serde(+json/yaml) |
David Tolnay | MIT / Apache-2.0 |
| zip, sha2, hex, notify | zip-rs · RustCrypto · rust-hex · notify-rs | MIT / Apache-2.0 / CC0 |
License-gate note. Two widely-repeated misconceptions were checked against primary sources and cleared: (1) the Demucs
htdemucs_*weights are MIT (the CC-BY-NC claim circulating in third-party repackagings conflates them with the 2019 Conv-TasNet research models); (2)madmom's trained models are CC-BY-NC-SA and were therefore rejected for meter tracking in favor of Beat This! Pin exact model revisions (HF/checkpoint) so a future card relicense can't silently breach the gate.
We treat transcription quality as a research problem and publish the numbers, the corpora, the algorithms, and the reproducibility commands alongside the code. The work directly informs production defaults — nothing in the ingest pipeline is shipped without head-to-head benchmark evidence captured in a document below.
Ingest auto-selects a primary transcriber per instrument — the
gameplay_default profile (transcription.py) — backed by a fail-safe
fallback chain so machines without the neural checkpoints still import.
The current production picks:
| Instrument | Top-tier transcriber | Why it leads | Fallback chain |
|---|---|---|---|
| Drums | mr_mt3_drums — neural MT3 ADT (GPU/CPU auto) |
catches the dense hi-hats / ghost notes the DSP engines miss (their E-GMD recall collapses to F1 ≈ 0.13) | beat-conditioned multiband → spectral-flux multiband → adaptive-beat-grid → DSP bandpass — used when the MT3 checkpoint/runtime is absent |
| Bass | torchcrepe — neural monophonic pitch (MIT) |
octave-clean (~0.3% octave jumps vs ~45% for Basic Pitch), tighter low register, ~9× faster | YIN octave+HPS fix → adaptive → YIN-bass80 → octave-fix |
| Guitar — lead | torchcrepe — monophonic |
octave-clean, ~6–8× faster than the DSP chain | melodic-adaptive → octave-fix → combined → Basic Pitch |
| Guitar — rhythm | melodic_hpss_combined — HPSS + onset (polyphonic) |
keeps the chord voices a monophonic tracker drops | melodic-adaptive → octave-fix → combined |
| Keys / piano | piano_auto — scored gate: Basic Pitch → PTI (Edwards/Kong) cleanup |
picks the best-scoring engine per stem; the tuned piano_chord_supplement path benchmarks F1 0.928 / precision 0.981 on the synthetic piano corpus |
PTI-consensus-clean → PTI-clean → polyphonic-clean |
| Vocals | (no pitch transcription) | vocals drive lyric alignment, not a note chart | — |
Alternate profiles are selectable per import: fidelity_midi (denser
symbolic output for A/B review) and research_ab (every local candidate,
defaults unchanged). Distorted/electric guitar is a known frontier
(rendered-tone onset-F1 ceiling ~0.78–0.84) — see the amp-tone research doc.
Per-case breakdowns, JSON reports, reproducibility commands, and the "what didn't work and why" notes live in the docs below.
-
Piano transcription cleanup deep-dive — 2026-06-20 Standalone narrative of the keys F1=0.706 → 0.928 climb. Tells each of the four steps end-to-end: the failure mode found by inspecting the previous step's output, the fix's gating contract, per-case numbers, and the "what didn't work" log (naive ensembles, onset- threshold sweeps, other engines). The doc to read first if you want the story behind the headline number.
-
Ground-truth benchmarks — 2026-06-14 Full per-instrument deep dive across all four instruments. Builds the annotated-corpus benchmark harness, ships dataset adapters for E-GMD / GuitarSet / Guitar-TECHS, documents every tuned variant we tried (the wins AND the failures), and pins reproducibility commands so anyone can re-run a sweep with one
aural_ingest gt-benchmarkinvocation. Covers drums, bass, guitar, and the round summary for keys (deep-dive doc above expands the keys section). -
ADT architecture deep-dive — 2026-05-07 2024–2025 ADT / transcription literature scan that revised 10 architectural assumptions baked into the original pipeline. Each assumption is checked against published work, the resulting paths- forward list is the source of the current production-default trail.
-
Electric-guitar amp-tone transcription — 2026-06-27 Adversarially-verified research on the distorted/electric-guitar frontier: why general models collapse on real amp tone, the amp-tone augmentation lever (the dominant fix), license-clean datasets/assets, and a ranked plan. Sets the ~0.78–0.84 onset-F1 ceiling cited above.
-
Research decision gates The locked-in production defaults the rest of the codebase reads from (beat/tempo backend, stem-separator policy, benchmark thresholding stance). Updated whenever a benchmark round flips a decision.
-
docs/CLAUDE_CODE_RESUME_PLAN.mdResumable v0.2 task plan (synth corpus, ADTOF integration, Demucs gate, real Basic Pitch, multi-label CRNN, real-audio fixtures). -
DRUM_TRANSCRIPTION_ALGORITHM_NOTES.md,TRANSCRIPTION_RECOVERY_NOTES.md,TRANSCRIPTION_REGRESSION_HISTORY.mdPre-rewrite recovery context (lost-tree recovery from 2026-03-03); preserved because the algorithm choices inpython/ingest/src/aural_ingest/algorithms/trace back through this material.
- Pick an annotated corpus and pick / write a dataset adapter under
python/ingest/src/aural_ingest/dataset_adapters/. - Run
aural_ingest gt-benchmark --dataset <name> --algorithm <one> --algorithm <other> ...The runner registers the requested algorithms, scores each against the corpus's reference events with greedy onset matching, and emits a JSON report underbenchmarks/{drums,melodic}/gt_runs/. - Read the per-case breakdown. If a variant Pareto-dominates the production default (every case strictly improves or stays unchanged on F1 / precision / recall), promote it; otherwise ship it as a workspace candidate.
- Append the round's results table + reproducibility command + the
"what didn't work" log to
docs/research-ground-truth-benchmarks-<date>.mdand update the summary scoreboard above.
The harness lives at
python/ingest/src/aural_ingest/ground_truth_benchmark.py
and the CLI subcommand is documented in
docs/ingest-pipeline.md.
Authoritative requirements + planning:
spec.md— app boundaries, hard constraints, MIDI/audio ruleswip.md— living implementation tracker (milestones, in-flight tasks, decisions)docs/roadmap.md— milestones from MVP to v1docs/risk-register.md— technical risks and mitigations
Architecture and contracts:
docs/architecture.md— system overview, module boundaries, runtime flowsdocs/auralsong-spec.md—.auralsongformat, event model, versioning/migrationsdocs/auralsong-deliverable.md— deterministic pack build contractdocs/ingest-pipeline.md— Python pipeline DAG, stage plugins, caching, CLI contractdocs/visualization-plugins.md— visualization plugin API + loading modeldocs/audio-codec-policy.md— host playback uses Rust/Symphonia; FFmpeg stays in the ingest sidecardocs/midi-keyboard-testing.md— hardware MIDI input verification path
Process, tooling, packaging:
docs/local-dev-prereqs.md— OS-level prerequisitesBUILDING.md— install / test / build / portable-package instructionsdocs/testing-strategy.md— TDD layers, fixtures, golden testsdocs/packaging-ci.md— bundling sidecars / decoders / models, CI build matrixdocs/performance-baselines.md— hardware profiles backing benchmark thresholds
/AuralPrimer
/apps
/game # AuralPrimer gameplay app (Tauri)
/desktop # AuralStudio authoring app (Tauri)
/packages
/core-music # shared schema + utilities (TS + Rust)
/viz-sdk # visualization plugin SDK (TS)
/auralsong # AuralSong reader/writer/validator (TS + Rust)
/python
/ingest # Python extraction pipeline (built into sidecars)
/src/aural_ingest
/algorithms # per-instrument transcribers + the new piano cleanup family
/dataset_adapters # E-GMD / GuitarSet / Guitar-TECHS / piano_synthetic
/ground_truth_benchmark.py
/visualizers
/viz-beats # beat/section grid
/viz-drum-highway # drum lanes from host-provided MIDI events
/viz-fretboard # fretboard cursor placeholder
/viz-lyrics # data-driven karaoke lyrics
/viz-nashville # chord-lane placeholder
/viz-tab # piano-roll + tab renderer for keys/bass/guitar
/benchmarks # frontend/python/rust benches + thresholds.yml + reports
/scripts # Node + PowerShell launchers for build/bench/portable
/assets
/models # downloaded on first use; NOT bundled in installer
/test_fixtures
/docs
/assets/screenshots
At runtime, users should not need a separate Python / FFmpeg / runtime install:
- ingest tools ship as PyInstaller sidecar executables
- decoder binaries are bundled when needed
- ML model packs download / import post-install into
assets/models/(never bundled in the installer)
See docs/packaging-ci.md for the full CI build
matrix.
