A note’s pitch is how often its shape repeats. Watch two detectors below on a note built to fool them: the first stops at half the period, an octave too high, while the second, in the bottom panel, finds 220 Hz.
A strong second harmonic fakes an octave
A synthetic 220 Hz note, harmonics 0.3, 1, 0.15, 0.6, … (even ones strong), 64 ms at 16 kHz, noise seed 291 (RMS 0.05, measured 0.0495).
One 64 ms frame of a 220 Hz note whose even harmonics are strong.
Describe this picture
A synthetic 220 Hz note with harmonics 0.3, 1, 0.15, 0.6, …, the even ones strong, 64 ms at 16 kHz, with noise from seed 291 (RMS 0.05, measured 0.0495). Three stacked panels. The first shows the frame as a thin solid line, from −3 to 3 against time from 0 to 64 ms. The second, the normalised autocorrelation, draws as a solid line from −0.5 to 1 against the lag from 0 to 400 samples, with a dotted level at 0.5 for “first peak above 0.5” and filled circles at the refined peaks. The third, YIN, draws as a solid line from 0 to 2 on the same lag axis, with values above 2 drawn at 2, a dotted threshold at 0.1, and a filled diamond at the chosen lag. The readouts are the rule, the lag in samples with two decimals and the pitch in Hz with one decimal. There is no control. The 13 s clip opens on the frame alone. From 3 s the autocorrelation draws, and the rule marks its first peak above 0.5: lag 36.35, half the period, 440.1 Hz, an octave too high; the true period’s peak, 0.925, comes later. From 8 s YIN’s curve draws: it stays at 0.156 at the half period, above 0.1, and first falls below it at lag 72; refined, the dip sits at lag 72.72, 220.0 Hz.
A strong second harmonic fakes an octave
Pluck a guitar string and you hear one note, even though the string shakes in a complicated way. The shape it makes repeats, over and over. How often it repeats is what your ear calls the note’s pitch.
A tuner has to measure that repeat rate from a few hundredths of a second of sound. This page builds three ways to do it from earlier lessons, shows the classic way they go wrong, and ends with a tuner you can turn.
Every sound on this page is synthetic. Each one is a sum of harmonics from a stated formula, plus noise from the site’s seeded generator, so the numbers are the same on every visit.
Pitch is a period
In “Repeats per second” of Sinusoids (3.2), a wave that repeats every seconds has frequency . In “A second arrow, three times as fast” of Signals as sums of sinusoids (7.1), a repeating sound is a sum of harmonics. The th harmonic turns times as fast as the slowest one, the fundamental .
The whole sum repeats at the fundamental’s rate. So the pitch is , and finding it means finding the period.
Sampled at samples per second, the period is samples. I use Hz on this page. A note at 220 Hz then repeats every samples, which is not a whole number.
Two notes whose frequencies differ by a factor of 2 are an octave apart. They sound so alike that a choir singing in octaves sounds like one tune. The note an octave above 220 Hz is 440 Hz, with half the period.
The autocorrelation finds the period
You met the tool in “A period hidden in noise” of Correlation (24.3). The autocorrelation slides the signal along itself, as in “Slide, multiply, add: no flip” there. A periodic signal lands on itself at a lag of one period, so the autocorrelation peaks there again.
As in 24.3, I divide by the value at lag 0, the frame’s energy. That gives the normalised autocorrelation
which is 1 at lag 0 and at most 1 everywhere else. It is the of 24.3 with one signal in both places. For a frame of samples, the sum at lag has only the products where the frame overlaps its shifted copy.
The obvious detector is a rule: take the first peak of above 0.5, call its lag the period, and divide it into . A peak can fall between two lags. “A parabola through three bins” of Zero-padding and resolution (15.3) fixes that, and I come back to it below.
A note built to fool it
Here is the test note. It is one frame of 1024 samples at 16 kHz, 64 ms of sound, with eight harmonics of 220 Hz:
for from 0 to 1023. The amplitudes to are 0.3, 1, 0.15, 0.6, 0.1, 0.35, 0.05 and 0.2, and the phases are radians. The noise is white noise from the site’s seeded generator, seed 291, scaled to RMS 0.05. This draw measures RMS 0.0495.
Look at the pattern: the odd harmonics are weak and the even ones strong. The even harmonics, 2, 4, 6 and 8 times 220 Hz, are all harmonics of 440 Hz. So most of this note’s power sits in a 440 Hz note, and the weak odd harmonics are all that makes it repeat at 220 Hz.
The picture at the top of the page runs the rule on this frame. Its third panel shows a second detector, YIN, which I build properly in the next section. For now, read its curve, called D′, as “how different the frame is from its shifted copy”. Near 0 means a near-perfect match, and YIN takes the first lag where D′ falls below 0.1.
Watch the rule stop at lag 36.35 and report 440.1 Hz, an octave too high, while YIN goes on to lag 72.72 and 220.0 Hz.
Notice the two heights in the middle panel. The peak at half the period reaches 0.818, and the one at the true period 0.925. The rule never gets that far: it stops at the first peak above 0.5.
Why half a period looks like a whole one
Let’s predict the height at half a period. In 24.3, one sine of amplitude gave an autocorrelation per sample. Each harmonic here adds its own cosine, and the products of two different harmonics average out, as the term did there.
With noise of power and the overlap counted, that gives
At half a period, , the cosine of the th harmonic is . That is for even and for odd . So the even harmonics push the peak up and the odd ones pull it down.
The even harmonics hold 0.761 of the power and the odd ones 0.0625, out of 0.824. With the noise and the overlap of 988 products out of 1024, the formula gives 0.816 at half a period. The instrument measures 0.818.
At the true period every cosine is . The formula gives 0.926 there, and the instrument measures 0.925. The note does repeat better at its true period, but only by a little.
Reporting a multiple of the pitch, usually twice it, is an octave error. It is the classic failure of pitch detectors. Any rule that takes the first good peak fails on a note whose even harmonics are strong enough.
Between two lags
The period, 72.73 samples, falls between lags 72 and 73. 15.3’s parabola finds the top between them. Take the three values around the highest sample, , and at lags , and . The top sits at
In 15.3 the three values were bins of a spectrum; here they are lags. The pitch is then . Around the true period the parabola puts the peak at lag 72.72, which gives 220.0 Hz, as the worked example below shows step by step.
Tune the string to the needle
The autocorrelation asks where the frame is most alike with itself. YIN, from A. de Cheveigné and H. Kawahara (2002), asks where it is least different. The two questions sound the same. The difference is in how YIN decides what counts as “small enough”.
YIN’s difference
Take the first 512 samples of the frame, 32 ms, and subtract the same stretch moved samples later. Add up the squared differences:
At a lag of one period the two stretches match, and is close to 0. I compute it for lags 0 to 400, so the lowest pitch it can find is Hz.
There is a catch. At lag 0 the difference is exactly 0, and at small lags it is small too, because neighbouring samples are alike. So “the smallest ” would pick a lag near 0.
YIN fixes this by comparing each value with the average of all the values up to it:
with . This is the cumulative-mean-normalised difference. At lag 1 it is 1, because the average of one value is that value. It falls well below 1 only where the frame matches itself far better than it did at the smaller lags.
YIN’s rule has three steps:
- Find the first lag where drops below a threshold, here 0.1.
- Walk down from there to the bottom of that dip.
- Refine the bottom with the parabola, and divide the lag into .
On the octave trap, is 0.156 at lag 36, the half period, so step 1 skips it. It first drops below 0.1 at lag 72, and the bottom of that dip is at lag 73. The parabola refines it to lag 72.72: 220.0 Hz.
Why “the first dip”, and not the deepest? On this frame the deepest sampled value is at lag 291, four periods: 0.0034, against 0.0062 at lag 73. Taking the deepest would report 55.0 Hz, two octaves too low.
YIN is not immune to the trap. Halve the odd harmonics of this note, and at the half period falls to 0.047, below the threshold. Then YIN stops at lag 36 and reports an octave too high as well.
Notes and cents
A tuner names the note and says how far off it is. Western music splits each octave into 12 equal steps, the semitones. Each step multiplies the frequency by , so 12 of them double it.
The notes are pinned to A4 = 440 Hz, so they are Hz for whole numbers . Two octaves below A4 is A2, at Hz, the open fifth string of a guitar.
For finer steps, each semitone is split into 100 cents, 1200 per octave. The offset of a frequency from the nearest note is
It lies between −50 and +50, because a note further away than half a semitone has a nearer neighbour. A negative offset means the note is flat, below the target, and a positive one sharp.
One cent is a ratio of 1.00058, far finer than the octave trap’s mistake. YIN’s 220.03 Hz on the trap was 0.2 cents above 220 Hz, while the rule’s octave error was 1200 cents off.
The plucked string
The tuner’s string is a synthetic plucked A2. It has eight harmonics, the th with amplitude , and each dies away after the pluck, the higher ones faster:
The first harmonic falls to , 37 % of its start, in 0.8 s, and the eighth in 0.26 s. The noise is scaled to RMS 0.01 and comes from the seeded generator, seed 2910, from the pluck on.
The peg sets the string’s frequency in cents from A2: , with from −45 to 45. YIN runs on one 64 ms frame, = 2000 to 3023, which starts 125 ms after the pluck. At 110 Hz the period is 145.45 samples, well inside the 400 lags.
Tune the string to the needle
A synthetic plucked A2 (110 Hz nominal): harmonics 1/k, each decaying, plus noise (seed 2910); YIN on a 64 ms frame from 125 ms after the pluck.
A2, −30.0 cents: the string is flat; the needle sits left.
Describe this picture
A synthetic plucked A2, 110 Hz nominal: harmonics , each decaying, plus noise (seed 2910), with YIN on a 64 ms frame from 125 ms after the pluck. Two panels. The first is a half-dial from −50 to +50 cents, with ticks every 10 cents and a needle pointing at the reading, the note name above it. The second shows the frame as a thin solid line, from −2 to 2 against time from 0 to 64 ms. The readouts are the note, the pitch in Hz with two decimals and the cents with one decimal. The 12 s clip opens with the peg at −30 cents: A2, 108.11 Hz, −30.0 cents, the string flat and the needle left. From 3 s the peg turns up to −12 cents, 109.24 Hz. From 7.5 s it turns to 0: in tune, 110.00 Hz, 0.1 cents. When the clip ends, a full-width slider, “Peg”, turns the peg from −45 to 45 cents, starting at 0 (arrow keys 1 cent, Page Up and Page Down 10), and the caption gives the cents and the pitch. A “Hear it” button plays the current string for 1 s. The setting is kept in the link, as tuner.c.
Watch the needle swing to the centre as the peg turns up from −30 cents to 0. When the clip ends, turn the peg yourself.
Notice how closely the needle follows the peg. For every whole-cent setting from −45 to 45, the reading is within 0.18 cents of the peg. The largest miss is 0.172 cents, at 38.
Even in tune, the needle reads 0.1, not 0.0. A clean string that does not decay would read 0.02 cents high, from the parabola alone. The noise and the decay bring it to 0.06.
The harmonic product spectrum
The third detector looks at the spectrum instead. In 7.1, a note’s spectrum has lines at , , and so on. Squeeze the spectrum by 2, and the line at lands on . Squeeze it by 3, and the line at lands there too.
So take , the size of the DFT of the frame, and multiply five copies, squeezed by 1 to 5:
Each copy puts one of the note’s harmonics on the true pitch, so the product piles up there. This is the harmonic product spectrum.
For the octave trap, I taper the frame with a Hann window, as in “Taper the ends, and the leakage falls” of Windowing & spectral leakage (15.1). Then I pad it to 16 384 points, as in “Zeros add points, not detail” of 15.3, so the bins are 0.9766 Hz apart. The product peaks at 219.7 Hz: right as well.
Without the fifth copy it would be wrong. The spectrum on its own peaks at 440.4 Hz, the strong second harmonic, and with four copies the product still peaks there. At 440 Hz the first four copies land on harmonics 2, 4, 6 and 8, whose amplitudes multiply to 0.042. At 220 Hz they land on harmonics 1 to 4, which give only 0.027.
The fifth copy decides. At 440 Hz it lands on 2200 Hz, the tenth harmonic, which this note does not have. At 220 Hz it lands on the fifth harmonic. So the harmonic product spectrum struggles when harmonics are missing: set the fifth harmonic to 0, and it reports 439.5 Hz.
The maths behind it · inner products
is the squared length of the difference between the frame and its shifted copy :
The last term is an autocorrelation. So YIN and the autocorrelation are one inner product, seen two ways.
The maths behind it · priors
An octave error is a choice between two explanations. A pitch and a pitch both explain this frame almost equally well. YIN’s threshold works like a prior that favours the lower pitch unless the data clearly rule it out.
Worked example
1. Refining a lag. Around the true period, is 0.9068, 0.9227 and 0.8672 at lags 72, 73 and 74. Then
and the peak is at lag .
2. Lag to pitch. Hz. The rule’s first peak, at lag 36.35, gives Hz from the rounded lag; the instrument’s unrounded lag gives 440.1 Hz.
3. Note name. For 108.11 Hz, semitones. The nearest whole number is −24, two octaves below A4: A2 at 110 Hz.
4. Cents. cents: the string is flat by 30 cents, as the tuner’s first frame shows.
5. A bin to hertz. The harmonic product spectrum peaks at bin 225 of 16 384, so Hz.
Where you’ll meet this
A clip-on guitar tuner or a tuning app does what the second instrument does. It measures the pitch of a frame, names the nearest note and shows the offset in cents. Real tuners track the pitch over many frames in a row, frame by frame as in Spectrograms & the STFT (15.5), and smooth the reading.
Pitch correction, of the kind made famous by Auto-Tune, tracks a singer’s pitch the same way and shifts each note towards the nearest semitone. Query by humming finds a song from a tune you hum, by comparing the hummed pitch track with stored melodies.
In speech, the pitch of the voice rises and falls as we talk: that is intonation. It can turn a statement into a question in many languages. The source–filter model of speech (30.1) treats the voice’s pitch as the source of its sound.
I left out three harder problems: several notes at once (polyphonic pitch), deciding whether a frame has a pitch at all (voicing), and tracking the pitch smoothly over time. For more, see A. de Cheveigné and H. Kawahara, “YIN, a fundamental frequency estimator for speech and music” (J. Acoust. Soc. Am., 2002); M. Müller, Fundamentals of Music Processing (2nd ed., 2021), chapter 8; and L. Rabiner and R. Schafer, Theory and Applications of Digital Speech Processing (2011), chapter 10.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Normalised autocorrelation | peaks at the period | |
| Half-period height | even harmonics add, odd subtract | strong even ones: octave error |
| YIN difference | 0 at a perfect match | |
| YIN, normalised | first dip below 0.1 | |
| Refine | , | parabola, as in 15.3 |
| Pitch | lags to 400: down to 40 Hz at 16 kHz | |
| Harmonic product spectrum | fails when a harmonic is missing | |
| Note | Hz, whole | A2 = 110 Hz |
| Cents | 100 per semitone, −50 to 50 |