Ear training from zero: what actually works, and what most people get wrong
Relative pitch is trainable at any age. Perfect pitch mostly is not, and the research openly disagrees with itself about whether adults can get it at all.
Ear training that works is mostly relative pitch, not perfect pitch. It is the skill of hearing how a note functions against a key centre — that it is the fifth, that it is the flat seventh, that it is the note leaning hard on the tonic and wanting to fall — rather than identifying its absolute frequency in isolation. Relative pitch is genuinely trainable at any age, and it is built through interval drills, movable-do solfège, and above all transcription: writing down real music by ear, repeatedly, across months and years. Absolute pitch is a separate thing. It is rare, strongly associated with very early exposure to a tone language, and not something a method or an app reliably gives an adult. The research itself is split on whether it can be given to an adult at all.
That last point is where most ear training conversations go wrong, so it is worth doing properly before we get to the useful part.
Perfect pitch: what the research shows, and where it argues with itself
The single most interesting finding in this field is a 2006 study by Diana Deutsch and colleagues, published in the Journal of the Acoustical Society of America. They tested 203 conservatory students — 88 native Mandarin speakers at Beijing's Central Conservatory and 115 non-tone-language speakers at Eastman — and found absolute pitch far more prevalent in the Chinese cohort at every age-of-training-onset band. Among students who began training at four or five, the approximate split was around 74 per cent against 14 per cent; among those who started at eight or nine, roughly 42 per cent against effectively zero. Those exact percentages come from a secondary rendering of the data rather than the abstract, so hold them loosely — but the shape is solid, and the shape is the argument. The gap persisted at every starting age, which is what rules out "they simply started younger" as the explanation and points instead towards early tone-language exposure and a speech-related critical period.
Note carefully what that does not say. It does not say that a Mandarin or Cantonese speaker will develop absolute pitch. It is a population-level prevalence difference, not an individual prediction.
You will also see the claim that absolute pitch occurs in about one person in ten thousand. That number is repeated everywhere and is flagged in the literature as poorly evidenced; a 2019 review instead put prevalence at at least four per cent among music students specifically, using a strict scoring threshold. Four per cent of a conservatory cohort and one in ten thousand of the general public are not measuring the same thing, and neither figure should be quoted with confidence.
Then there is the contradiction that nobody selling ear training courses seems keen to mention. One widely-circulated summary states that despite a century of attempts, no adult has ever been documented acquiring genuine absolute listening ability. Against that, Van Hedger, Heald and Nusbaum at the University of Chicago published a paper in PLOS ONE on 24 September 2019 titled, unambiguously, "Absolute pitch can be learned by some adults". They trained six adults aged 18 to 26 — pre-screened for strong auditory working memory — for eight weeks, roughly four hours a week, about 32 hours in total, on speed- and accuracy-focused pitch-naming drills. Two of the six reached absolute-pitch-qualifying accuracy, scoring 94.44 and 97.22 per cent on a standardised test, and were still there at a four-month follow-up. The other four improved only modestly.
Six participants, pre-selected for an unusual cognitive trait, with a one-in-three hit rate: a real peer-reviewed result, and a very thin plank to stand a worldview on. In the strangest corner of this literature, a 2013 randomised, double-blind, placebo-controlled crossover trial by Gervain and colleagues gave 24 healthy adult men, median age 23, either valproate or placebo for 15 days while training them to name six pitches. The valproate group scored significantly above chance — a mean of 5.09 out of 9, against a chance level of 3 and placebo's 3.50 — in the arm where they took the drug first. Genuine evidence that a childhood-like plastic window can be pharmacologically nudged open. Also nowhere near full absolute pitch, and it involves taking an anticonvulsant to learn six notes.
My honest read: the field has not settled this, and I am suspicious of anyone in either direction who says it has. The practical consequence is the same either way. Absolute pitch is not what makes musicians good, and chasing it is a poor use of the hour you have today.
Functional hearing and interval hearing are not rivals
There is a fashionable framing that interval training is the outdated way and functional training is the enlightened one. In a teaching room that falls apart within a term, because the two do different jobs.
Functional ear training teaches you to hear each note's role against an established tonal centre — that is the third of the key, that is the flat seventh — rather than the bare distance between two notes. Interval training teaches the distance: perfect fifth, minor sixth. The functional approach is what transfers to playing, because real music has a key and your ear is always orienting to it. But interval names are the vocabulary that lets you say what you just heard, check it, and write it down.
The moment this clicks happens in the same small back room with the upright piano at least twice a year. A guitarist, usually early twenties, usually good at intervals from an app, gets asked to sing the fourth degree over a droning C chord. They cannot do it reliably, even though they can name a perfect fourth in a blind test at ninety per cent. The drill trained a comparison between two isolated events; it did not train orientation inside a key. So we drop the app, hold a C with the sustain pedal down, and sing degrees against it: do, then sol, then fa, then back. Four weeks of that and the same player starts finding melodies on the fretboard on the first or second attempt instead of the fifth — no longer searching, but hearing "that is the fifth" and going straight there. Which is why learning by ear is not a shortcut around reading but its counterpart.
So: both, in that order. Functional first, because it is load-bearing. Intervals alongside, because they are how you name and verify what functional hearing gives you. If you are working through modes by their actual sound, functional hearing is doing the work there too.
Movable do, fixed do, and a schoolteacher from Norwich
Which solfège system you learn is mostly a question of where you live, and almost nobody explains this.
In fixed do, do always means C, whatever key you are in. That is the norm in France, Italy, Spain, Portugal, Belgium, Romania, Russia, Turkey, Ukraine, Bulgaria, Israel, Brazil, much of Latin America and French-speaking Canada. In movable do, do is the tonic of whatever key you are in — so the syllables carry function rather than pitch names. That is the norm in Australia, the UK, Ireland, the US, English-speaking Canada, China, Japan and Hong Kong.
The Anglophone version has a specific history worth knowing. Tonic sol-fa was devised by Sarah Anna Glover (1786–1867), a schoolteacher in Norwich, who developed her Norwich Sol-fa Ladder from 1812 and published it in 1845 and 1850. It was popularised by John Curwen (1816–1880), who was commissioned by a conference of Sunday school teachers in 1841 to find a workable method for teaching singing, published his Grammar of Vocal Music in 1843, founded the Tonic Sol-Fa Association in 1853 and released his Standard Course of Lessons in 1858. A century later, Zoltán Kodály built the Hungarian method that still bears his name on movable-do solfège adapted from Curwen, along with Curwen's hand signs and rhythm syllables drawn from Dalcroze.
Two hundred years, one idea: syllables that name function, not frequency. If you are Australian, use movable do — it is the local convention, and AMEB permits sol-fa as an alternative to scale-degree numbers from Grade 1. That is not a verdict on fixed do, which produces superb musicians across continental Europe. It is a verdict on swimming against your own examination system for no reason.
Transcription is the engine
Everything above is scaffolding. Transcription is the building.
By transcription I mean: take a recording, work out what is played, and write it down — melody, bass line, chord changes, a solo, a horn part. No tab, no chord sheets, no looking it up until you have committed to an answer.
I will be straight about the evidence. There is no controlled trial establishing an effect size for transcription as an ear training method. What exists is pedagogical consensus, mostly out of the jazz tradition, where transcription has been the central discipline for the better part of a century and where players who agree about nothing else agree about this. Consensus is weaker than a randomised trial. But a method that survives that long, across that many teachers, usually survives for a reason.
The reason, I think, is that transcription loads every component at once. You hear a phrase. You hold it in working memory. You sing it. You find it on your instrument. You decide what it is called. You write it. Then the recording tells you immediately whether you were right. Interval apps train one narrow link in that chain and give you a score. Transcription trains the whole chain and gives you a piece of music.
What I actually set: eight bars. Not a solo, not a whole tune — eight bars the student chose because they like them. A week to bring back a written version. The first attempt is usually rhythmically wrong and harmonically vague, which is fine, because the second one is not. What changes over a term is not the accuracy. It is that the student stops asking me what the chords are. They hear a change coming before it lands, and they make that face — the one where they have predicted something correctly and are slightly annoyed at how satisfying it is. That is functional hearing arriving, and it arrives through the writing.
Practical notes. Slow the audio down; do not shift the pitch. Sing the line before you play it — if you cannot sing it you have not heard it, you have guessed it. Write the rhythm before the pitches. Take the bass line first on anything harmonically busy, because the bass gives you the root and the root gives you the function. Keep the finished pages: they are the record of a skill that improves too slowly to notice day to day, which is the same argument for keeping a practice journal.
How long this actually takes
I will not give you an hours figure, because nobody has one that survives contact with evidence, and every "40 hours to relative pitch" claim you have seen is marketing.
What can be quoted are institutional benchmarks. Oberlin Conservatory requires Bachelor of Music and double-degree students to complete or exempt out of four sequential semesters of Aural Skills, alongside a matched four-semester theory sequence, with aural classes meeting 100 minutes a week. That is a top conservatory's view of the runway, for students who arrived already able to play.
Closer to home, the AMEB Musicianship Syllabus 2022 does not introduce a dedicated Aural Section until Grade 4, at 30 minutes, expanding to 40 minutes by Grades 5 and 6. A Grade 5 melodic dictation is more modest than "aural exam" suggests: a six-crotchet phrase in a key of up to two sharps or flats, played six times, with a one-minute gap to finish writing before a final check playing. Six notes, six hearings. By Associate and Licentiate diploma level, candidates handle two-part and three-part dictation respectively. The syllabus is deliberately silent on how many years that takes, and so am I — but the increments are small and the ladder is long. That structure differs between boards, which is part of choosing between AMEB, ABRSM and Trinity.
The best evidence on what predicts dictation ability is a 2025 study by Nichols and Springer in Frontiers in Psychology. Across 34 undergraduate music majors, three factors independently predicted melodic dictation performance: tonal working memory (β = 0.406), years of private piano lessons (β = 0.410) and semesters of aural skills coursework completed (β = 0.353). Together they explained 36.5 per cent of the variance. Attentional flexibility and vividness of auditory imagery predicted nothing.
A small study with a clear message: the skill is built cumulatively, from several directions at once — coursework, keyboard familiarity, memory capacity — not by one clever hack. Read it as encouraging, too. Two of those three are things you can go and accumulate.
Tools worth using
Free first. musictheory.net's browser-based Interval Ear Trainer is unglamorous, configurable and instant — no account, no install, good for five minutes between things. GNU Solfege is a GPL-licensed cross-platform program distributed as part of the GNU Project, drilling intervals, chords, scales and rhythm; broad coverage, and an interface that looks its age.
On the paid side, judge by approach rather than price, since pricing moves. EarMaster is the comprehensive-curriculum option, a graded course across intervals, chords, rhythm and dictation. Functional Ear Trainer takes the narrower, opinionated route: notes presented against an established key context, which is exactly the functional skill described above. Pick according to which gap you are filling.
Whatever you use, drilling only compounds if it happens most days, and daily transcription is the thing most likely to quietly fall off the list — logging it alongside scales and repertoire in a practice tracker like Crescender at least makes the gap visible when it opens.
Common questions
Can adults develop relative pitch at any age?
Yes. Unlike absolute pitch, relative pitch has no documented critical period, and the evidence points to accumulation rather than innate wiring. Nichols and Springer's 2025 study found melodic dictation performance predicted by semesters of aural skills coursework, years of piano lessons and tonal working memory — three things that are acquired or trainable, not fixed at birth. Progress is slow and cumulative, but it is available to anyone.
Is perfect pitch trainable in adulthood?
The literature genuinely contradicts itself. One widely-repeated summary holds that no adult has ever been documented acquiring absolute listening ability. A 2019 PLOS ONE study by Van Hedger, Heald and Nusbaum reports the opposite in a small pre-screened sample: two of six adults reached absolute-pitch accuracy after about 32 hours, sustained at four months. Both are real. Be suspicious of certainty in either direction.
What is the difference between functional and interval ear training?
Interval training identifies the distance between two notes in isolation — a perfect fourth, a minor seventh. Functional training identifies a note's role inside an established key — the fourth degree, the leading note. Functional hearing transfers directly to playing, improvising and predicting chord changes; interval naming gives you the vocabulary to describe and verify what you heard. Learn functional first, keep interval work running alongside.
Fixed do or movable do — which should I learn?
If you are in Australia, movable do. It is the Anglophone convention, shared with the UK, Ireland, the US, Canada, China and Japan, and AMEB permits sol-fa from Grade 1. Fixed do, where do always means C, is standard in France, Italy, Spain, Russia, Brazil and much of Latin America, and produces excellent musicians. It is a regional convention, not a correctness question.
How long does ear training actually take?
Nobody has a defensible number, and app marketing that offers one is guessing. The available anchors are institutional: Oberlin requires four semesters of Aural Skills at 100 minutes a week, and AMEB does not introduce a dedicated aural section until Grade 4, with Grade 5 dictation asking for just six crotchets across six playings. Small increments, long ladder, measured in years.
What percentage of people have absolute pitch?
Unclear, and the commonly quoted "one in ten thousand" is flagged in the literature as poorly evidenced. A 2019 review put prevalence at at least four per cent among music students under a strict scoring threshold, which is a different population from the general public. Deutsch et al.'s 2006 conservatory study found rates varying enormously by cohort and age of training onset.
If you take one instruction from all of this, take the eight bars. Pick a recording you actually love, choose eight bars of it, and write them down this week without looking anything up. It will take longer than you expect and the first attempt will be wrong in ways that are slightly embarrassing. Do it again next week with eight different bars, and again the week after. Everything else here — the solfège, the intervals, the apps, the exam syllabus — is support structure for that one habit, and in twelve months the difference between musicians who kept it up and musicians who did not will not be subtle.
Put the idea into practice
Crescender helps musicians, teachers, and families organise the work around music without scattering it across disconnected tools.
Start now