Evolutionary neural architecture search for automatic chord estimation: a reproducible, vocabulary-aware study of structured harmonic sequence models
Abstract
Automatic chord estimation (ACE) takes a musical recording and produces a time-aligned sequence
of labels. These labels, such as C major, A minor 7, or G7/B, provide a compact description
of a song’s harmony and support transcription, teaching, accompaniment, harmonic search, and
large-scale musical analysis. Automating this process is difficult because the same harmony can
be voiced in different ways, chords often share pitches, chord labels are contextual to surrounding
harmonic information, detailed chord types may be rare, and human annotators may disagree.
This thesis presents a reproducible evolutionary neural architecture search (NAS) framework
for large-vocabulary ACE. Starting from a validated convolutional–recurrent baseline, the framework
searches for improved convolutional and recurrent architectures while keeping the musical
representation, chord vocabulary, training procedure, and decoder fixed. Multiple independent
searches are conducted using a validation objective that balances recognition performance with
model complexity, after which selected architectures are independently retrained and evaluated.
On the McGill Billboard benchmark test set, the proposed model achieved a Weighted Chord-
Symbol Recall (WCSR) of 63 49 0 36%, compared with 60 09 0 16% for the baseline, with a
mean paired improvement of 3 39 0 42 percentage points across three initialization seeds. After
topology freeze, the same model was retrained from scratch on a genre-stratified, song-disjoint
split of Chordonomicon, with the corresponding chord progressions rendered as synthetic audio; it
achieves 96 50 0 12% WCSR compared with 95 26 0 14% for the baseline, an improvement of
1 23 0 23 percentage points.
Description
Thesis is embargoed until September 15 2027.
Keywords
Musical analysis, Machine-learning, Automation
