Skip to content

Pitch detection

Find a note's period with autocorrelation and YIN, see a strong second harmonic fake an octave, and read the pitch as a note and cents.

Before thisZero-padding and resolution (15.3), Correlation (24.3)

2 more before it

Sinusoids (3.2), Signals as sums of sinusoids (7.1)

Before this15.3 · 24.3 · 2 more
Chapter 29 · Lesson 1 of 3

First, the picture

A note’s pitch is how often its shape repeats. Watch two detectors below on a note built to fool them: the first stops at half the period, an octave too high, while the second, in the bottom panel, finds 220 Hz.

A strong second harmonic fakes an octave

A synthetic 220 Hz note, harmonics 0.3, 1, 0.15, 0.6, … (even ones strong), 64 ms at 16 kHz, noise seed 291 (RMS 0.05, measured 0.0495).

One 64 ms frame of a 220 Hz note whose even harmonics are strong.

rule
—
lag
—
pitch
—
0.00 / 13.00 s
Describe this picture

A synthetic 220 Hz note with harmonics 0.3, 1, 0.15, 0.6, …, the even ones strong, 64 ms at 16 kHz, with noise from seed 291 (RMS 0.05, measured 0.0495). Three stacked panels. The first shows the frame as a thin solid line, from −3 to 3 against time from 0 to 64 ms. The second, the normalised autocorrelation, draws ρx[ℓ]\rho_{x}[\ell] as a solid line from −0.5 to 1 against the lag from 0 to 400 samples, with a dotted level at 0.5 for “first peak above 0.5” and filled circles at the refined peaks. The third, YIN, draws Dx′[ℓ]D_{x}'[\ell] as a solid line from 0 to 2 on the same lag axis, with values above 2 drawn at 2, a dotted threshold at 0.1, and a filled diamond at the chosen lag. The readouts are the rule, the lag in samples with two decimals and the pitch in Hz with one decimal. There is no control. The 13 s clip opens on the frame alone. From 3 s the autocorrelation draws, and the rule marks its first peak above 0.5: lag 36.35, half the period, 440.1 Hz, an octave too high; the true period’s peak, 0.925, comes later. From 8 s YIN’s curve draws: it stays at 0.156 at the half period, above 0.1, and first falls below it at lag 72; refined, the dip sits at lag 72.72, 220.0 Hz.

A strong second harmonic fakes an octave

Pluck a guitar string and you hear one note, even though the string shakes in a complicated way. The shape it makes repeats, over and over. How often it repeats is what your ear calls the note’s pitch.

A tuner has to measure that repeat rate from a few hundredths of a second of sound. This page builds three ways to do it from earlier lessons, shows the classic way they go wrong, and ends with a tuner you can turn.

Every sound on this page is synthetic. Each one is a sum of harmonics from a stated formula, plus noise from the site’s seeded generator, so the numbers are the same on every visit.

Pitch is a period

In “Repeats per second” of Sinusoids (3.2), a wave that repeats every TT seconds has frequency f=1/Tf=1/T. In “A second arrow, three times as fast” of Signals as sums of sinusoids (7.1), a repeating sound is a sum of harmonics. The kkth harmonic turns kk times as fast as the slowest one, the fundamental f0f_{0}.

The whole sum repeats at the fundamental’s rate. So the pitch is f0f_{0}, and finding it means finding the period.

Sampled at fsf_s samples per second, the period is P=fs/f0P=f_s/f_{0} samples. I use fs=16 000f_s=16\,000 Hz on this page. A note at 220 Hz then repeats every 16 000/220=72.7316\,000/220=72.73 samples, which is not a whole number.

Two notes whose frequencies differ by a factor of 2 are an octave apart. They sound so alike that a choir singing in octaves sounds like one tune. The note an octave above 220 Hz is 440 Hz, with half the period.

The autocorrelation finds the period

You met the tool in “A period hidden in noise” of Correlation (24.3). The autocorrelation Rx[ℓ]=∑nx[n] x[n−ℓ]R_{x}[\ell]=\sum_n x[n]\,x[n-\ell] slides the signal along itself, as in “Slide, multiply, add: no flip” there. A periodic signal lands on itself at a lag of one period, so the autocorrelation peaks there again.

As in 24.3, I divide by the value at lag 0, the frame’s energy. That gives the normalised autocorrelation

ρx[ℓ]=Rx[ℓ]Rx[0],\rho_x[\ell]=\frac{R_x[\ell]}{R_x[0]},

which is 1 at lag 0 and at most 1 everywhere else. It is the ρxy[ℓ]\rho_{xy}[\ell] of 24.3 with one signal in both places. For a frame of NN samples, the sum at lag ℓ\ell has only the N−ℓN-\ell products where the frame overlaps its shifted copy.

The obvious detector is a rule: take the first peak of ρx[ℓ]\rho_{x}[\ell] above 0.5, call its lag the period, and divide it into fsf_s. A peak can fall between two lags. “A parabola through three bins” of Zero-padding and resolution (15.3) fixes that, and I come back to it below.

A note built to fool it

Here is the test note. It is one frame of 1024 samples at 16 kHz, 64 ms of sound, with eight harmonics of 220 Hz:

x[n]=∑k=18akcos⁡(2πk 22016 000 n+0.7k)+v[n],\begin{aligned} x[n]&=\sum_{k=1}^{8}a_k\cos\Big(2\pi k\,\frac{220}{16\,000}\,n\\ &\qquad+0.7k\Big)+v[n], \end{aligned}

for nn from 0 to 1023. The amplitudes a1a_{1} to a8a_{8} are 0.3, 1, 0.15, 0.6, 0.1, 0.35, 0.05 and 0.2, and the phases are 0.7k0.7k radians. The noise v[n]v[n] is white noise from the site’s seeded generator, seed 291, scaled to RMS 0.05. This draw measures RMS 0.0495.

Look at the pattern: the odd harmonics are weak and the even ones strong. The even harmonics, 2, 4, 6 and 8 times 220 Hz, are all harmonics of 440 Hz. So most of this note’s power sits in a 440 Hz note, and the weak odd harmonics are all that makes it repeat at 220 Hz.

The picture at the top of the page runs the rule on this frame. Its third panel shows a second detector, YIN, which I build properly in the next section. For now, read its curve, called D′, as “how different the frame is from its shifted copy”. Near 0 means a near-perfect match, and YIN takes the first lag where D′ falls below 0.1.

Watch the rule stop at lag 36.35 and report 440.1 Hz, an octave too high, while YIN goes on to lag 72.72 and 220.0 Hz.

Notice the two heights in the middle panel. The peak at half the period reaches 0.818, and the one at the true period 0.925. The rule never gets that far: it stops at the first peak above 0.5.

Why half a period looks like a whole one

Let’s predict the height at half a period. In 24.3, one sine of amplitude AA gave an autocorrelation (A2/2)cos⁡(2πℓ/P)(A^2/2)\cos(2\pi\ell/P) per sample. Each harmonic here adds its own cosine, and the products of two different harmonics average out, as the cos⁡(a+b)\cos(a+b) term did there.

With noise of power σv2\sigma_v^2 and the overlap counted, that gives

ρx[ℓ]≈N−ℓN×∑k12ak2cos⁡(2πkℓ/P)∑k12ak2+σv2.\begin{aligned} \rho_x[\ell]&\approx\frac{N-\ell}{N}\\ &\quad\times\frac{\sum_k\tfrac12a_k^2\cos(2\pi k\ell/P)}{\sum_k\tfrac12a_k^2+\sigma_v^2}. \end{aligned}

At half a period, ℓ=P/2\ell=P/2, the cosine of the kkth harmonic is cos⁡(πk)\cos(\pi k). That is +1+1 for even kk and −1-1 for odd kk. So the even harmonics push the peak up and the odd ones pull it down.

The even harmonics hold 0.761 of the power and the odd ones 0.0625, out of 0.824. With the noise and the overlap of 988 products out of 1024, the formula gives 0.816 at half a period. The instrument measures 0.818.

At the true period every cosine is +1+1. The formula gives 0.926 there, and the instrument measures 0.925. The note does repeat better at its true period, but only by a little.

Reporting a multiple of the pitch, usually twice it, is an octave error. It is the classic failure of pitch detectors. Any rule that takes the first good peak fails on a note whose even harmonics are strong enough.

Between two lags

The period, 72.73 samples, falls between lags 72 and 73. 15.3’s parabola finds the top between them. Take the three values around the highest sample, aa, bb and cc at lags ℓ−1\ell-1, ℓ\ell and ℓ+1\ell+1. The top sits at

ℓ+d,d=(a−c)/2a−2b+c.\ell+d,\qquad d=\frac{(a-c)/2}{a-2b+c}.

In 15.3 the three values were bins of a spectrum; here they are lags. The pitch is then f0=fs/(ℓ+d)f_{0}=f_s/(\ell+d). Around the true period the parabola puts the peak at lag 72.72, which gives 220.0 Hz, as the worked example below shows step by step.

Tune the string to the needle

The autocorrelation asks where the frame is most alike with itself. YIN, from A. de Cheveigné and H. Kawahara (2002), asks where it is least different. The two questions sound the same. The difference is in how YIN decides what counts as “small enough”.

YIN’s difference

Take the first 512 samples of the frame, 32 ms, and subtract the same stretch moved ℓ\ell samples later. Add up the squared differences:

Dx[ℓ]=∑n=0511(x[n]−x[n+ℓ])2.D_x[\ell]=\sum_{n=0}^{511}\big(x[n]-x[n+\ell]\big)^2.

At a lag of one period the two stretches match, and Dx[ℓ]D_{x}[\ell] is close to 0. I compute it for lags 0 to 400, so the lowest pitch it can find is 16 000/400=4016\,000/400=40 Hz.

There is a catch. At lag 0 the difference is exactly 0, and at small lags it is small too, because neighbouring samples are alike. So “the smallest Dx[ℓ]D_{x}[\ell]” would pick a lag near 0.

YIN fixes this by comparing each value with the average of all the values up to it:

Dx′[ℓ]=Dx[ℓ]1ℓ∑j=1ℓDx[j],D_x'[\ell]=\frac{D_x[\ell]}{\frac1\ell\sum_{j=1}^{\ell}D_x[j]},

with Dx′[0]=1D_{x}'[0]=1. This is the cumulative-mean-normalised difference. At lag 1 it is 1, because the average of one value is that value. It falls well below 1 only where the frame matches itself far better than it did at the smaller lags.

YIN’s rule has three steps:

  1. Find the first lag where Dx′[ℓ]D_{x}'[\ell] drops below a threshold, here 0.1.
  2. Walk down from there to the bottom of that dip.
  3. Refine the bottom with the parabola, and divide the lag into fsf_s.

On the octave trap, Dx′D_{x}' is 0.156 at lag 36, the half period, so step 1 skips it. It first drops below 0.1 at lag 72, and the bottom of that dip is at lag 73. The parabola refines it to lag 72.72: 220.0 Hz.

Why “the first dip”, and not the deepest? On this frame the deepest sampled value is at lag 291, four periods: 0.0034, against 0.0062 at lag 73. Taking the deepest would report 55.0 Hz, two octaves too low.

YIN is not immune to the trap. Halve the odd harmonics of this note, and Dx′D_{x}' at the half period falls to 0.047, below the threshold. Then YIN stops at lag 36 and reports an octave too high as well.

Notes and cents

A tuner names the note and says how far off it is. Western music splits each octave into 12 equal steps, the semitones. Each step multiplies the frequency by 21/12=1.05952^{1/12}=1.0595, so 12 of them double it.

The notes are pinned to A4 = 440 Hz, so they are 440⋅2m/12440\cdot2^{m/12} Hz for whole numbers mm. Two octaves below A4 is A2, at 440/4=110440/4=110 Hz, the open fifth string of a guitar.

For finer steps, each semitone is split into 100 cents, 1200 per octave. The offset of a frequency ff from the nearest note fnotef_\text{note} is

cents=1200log⁡2ffnote.\text{cents}=1200\log_2\frac{f}{f_\text{note}}.

It lies between −50 and +50, because a note further away than half a semitone has a nearer neighbour. A negative offset means the note is flat, below the target, and a positive one sharp.

One cent is a ratio of 1.00058, far finer than the octave trap’s mistake. YIN’s 220.03 Hz on the trap was 0.2 cents above 220 Hz, while the rule’s octave error was 1200 cents off.

The plucked string

The tuner’s string is a synthetic plucked A2. It has eight harmonics, the kkth with amplitude 1/k1/k, and each dies away after the pluck, the higher ones faster:

y[n]=∑k=181k e−(1+0.3(k−1)) n/(0.8fs)×sin⁡(2πk ffs n)+v[n].\begin{aligned} y[n]&=\sum_{k=1}^{8}\tfrac1k\,e^{-(1+0.3(k-1))\,n/(0.8f_s)}\\ &\quad\times\sin\Big(2\pi k\,\frac{f}{f_s}\,n\Big)+v[n]. \end{aligned}

The first harmonic falls to e−1e^{-1}, 37 % of its start, in 0.8 s, and the eighth in 0.26 s. The noise v[n]v[n] is scaled to RMS 0.01 and comes from the seeded generator, seed 2910, from the pluck on.

The peg sets the string’s frequency in cents from A2: f=110⋅2c/1200f=110\cdot2^{c/1200}, with cc from −45 to 45. YIN runs on one 64 ms frame, nn = 2000 to 3023, which starts 125 ms after the pluck. At 110 Hz the period is 145.45 samples, well inside the 400 lags.

Tune the string to the needle

A synthetic plucked A2 (110 Hz nominal): harmonics 1/k, each decaying, plus noise (seed 2910); YIN on a 64 ms frame from 125 ms after the pluck.

A2, −30.0 cents: the string is flat; the needle sits left.

note
A2
pitch
108.11 Hz
cents
−30.0
0.00 / 12.00 s
Describe this picture

A synthetic plucked A2, 110 Hz nominal: harmonics 1/k1/k, each decaying, plus noise (seed 2910), with YIN on a 64 ms frame from 125 ms after the pluck. Two panels. The first is a half-dial from −50 to +50 cents, with ticks every 10 cents and a needle pointing at the reading, the note name above it. The second shows the frame as a thin solid line, from −2 to 2 against time from 0 to 64 ms. The readouts are the note, the pitch in Hz with two decimals and the cents with one decimal. The 12 s clip opens with the peg at −30 cents: A2, 108.11 Hz, −30.0 cents, the string flat and the needle left. From 3 s the peg turns up to −12 cents, 109.24 Hz. From 7.5 s it turns to 0: in tune, 110.00 Hz, 0.1 cents. When the clip ends, a full-width slider, “Peg”, turns the peg from −45 to 45 cents, starting at 0 (arrow keys 1 cent, Page Up and Page Down 10), and the caption gives the cents and the pitch. A “Hear it” button plays the current string for 1 s. The setting is kept in the link, as tuner.c.

Watch the needle swing to the centre as the peg turns up from −30 cents to 0. When the clip ends, turn the peg yourself.

Notice how closely the needle follows the peg. For every whole-cent setting from −45 to 45, the reading is within 0.18 cents of the peg. The largest miss is 0.172 cents, at 38.

Even in tune, the needle reads 0.1, not 0.0. A clean string that does not decay would read 0.02 cents high, from the parabola alone. The noise and the decay bring it to 0.06.

The harmonic product spectrum

The third detector looks at the spectrum instead. In 7.1, a note’s spectrum has lines at f0f_{0}, 2f02f_{0}, 3f03f_{0} and so on. Squeeze the spectrum by 2, and the line at 2f02f_{0} lands on f0f_{0}. Squeeze it by 3, and the line at 3f03f_{0} lands there too.

So take ∣X[k]∣\lvert X[k]\rvert, the size of the DFT of the frame, and multiply five copies, squeezed by 1 to 5:

∣X[k]∣⋅∣X[2k]∣⋅∣X[3k]∣⋅∣X[4k]∣⋅∣X[5k]∣.\begin{aligned} &\lvert X[k]\rvert\cdot\lvert X[2k]\rvert\cdot\lvert X[3k]\rvert\\ &\quad\cdot\lvert X[4k]\rvert\cdot\lvert X[5k]\rvert. \end{aligned}

Each copy puts one of the note’s harmonics on the true pitch, so the product piles up there. This is the harmonic product spectrum.

For the octave trap, I taper the frame with a Hann window, as in “Taper the ends, and the leakage falls” of Windowing & spectral leakage (15.1). Then I pad it to 16 384 points, as in “Zeros add points, not detail” of 15.3, so the bins are 0.9766 Hz apart. The product peaks at 219.7 Hz: right as well.

Without the fifth copy it would be wrong. The spectrum on its own peaks at 440.4 Hz, the strong second harmonic, and with four copies the product still peaks there. At 440 Hz the first four copies land on harmonics 2, 4, 6 and 8, whose amplitudes multiply to 0.042. At 220 Hz they land on harmonics 1 to 4, which give only 0.027.

The fifth copy decides. At 440 Hz it lands on 2200 Hz, the tenth harmonic, which this note does not have. At 220 Hz it lands on the fifth harmonic. So the harmonic product spectrum struggles when harmonics are missing: set the fifth harmonic to 0, and it reports 439.5 Hz.

The maths behind it · inner products

Dx[ℓ]D_{x}[\ell] is the squared length of the difference between the frame x\mathbf{x} and its shifted copy xℓ\mathbf{x}_{\ell}:

∥x−xℓ∥2=∥x∥2+∥xℓ∥2−2⟨x,xℓ⟩.\begin{aligned} \lVert\mathbf{x}-\mathbf{x}_\ell\rVert^2&=\lVert\mathbf{x}\rVert^2+\lVert\mathbf{x}_\ell\rVert^2\\ &\quad-2\langle\mathbf{x},\mathbf{x}_\ell\rangle. \end{aligned}

The last term is an autocorrelation. So YIN and the autocorrelation are one inner product, seen two ways.

The maths behind it · priors

An octave error is a choice between two explanations. A pitch ff and a pitch 2f2f both explain this frame almost equally well. YIN’s threshold works like a prior that favours the lower pitch unless the data clearly rule it out.

Worked example

1. Refining a lag. Around the true period, ρx[ℓ]\rho_{x}[\ell] is 0.9068, 0.9227 and 0.8672 at lags 72, 73 and 74. Then

d=(0.9068−0.8672)/20.9068−2⋅0.9227+0.8672=0.0198−0.0714=−0.277,\begin{aligned} d&=\frac{(0.9068-0.8672)/2}{0.9068-2\cdot0.9227+0.8672}\\ &=\frac{0.0198}{-0.0714}=-0.277, \end{aligned}

and the peak is at lag 73−0.277=72.7273-0.277=72.72.

2. Lag to pitch. 16 000/72.72=220.016\,000/72.72=220.0 Hz. The rule’s first peak, at lag 36.35, gives 16 000/36.35=440.216\,000/36.35=440.2 Hz from the rounded lag; the instrument’s unrounded lag gives 440.1 Hz.

3. Note name. For 108.11 Hz, 12log⁡2(108.11/440)=−24.3012\log_2(108.11/440)=-24.30 semitones. The nearest whole number is −24, two octaves below A4: A2 at 110 Hz.

4. Cents. 1200log⁡2(108.11/110)=−30.01200\log_2(108.11/110)=-30.0 cents: the string is flat by 30 cents, as the tuner’s first frame shows.

5. A bin to hertz. The harmonic product spectrum peaks at bin 225 of 16 384, so 225×16 000/16 384=219.7225\times16\,000/16\,384=219.7 Hz.

Where you’ll meet this

A clip-on guitar tuner or a tuning app does what the second instrument does. It measures the pitch of a frame, names the nearest note and shows the offset in cents. Real tuners track the pitch over many frames in a row, frame by frame as in Spectrograms & the STFT (15.5), and smooth the reading.

Pitch correction, of the kind made famous by Auto-Tune, tracks a singer’s pitch the same way and shifts each note towards the nearest semitone. Query by humming finds a song from a tune you hum, by comparing the hummed pitch track with stored melodies.

In speech, the pitch of the voice rises and falls as we talk: that is intonation. It can turn a statement into a question in many languages. The source–filter model of speech (30.1) treats the voice’s pitch as the source of its sound.

I left out three harder problems: several notes at once (polyphonic pitch), deciding whether a frame has a pitch at all (voicing), and tracking the pitch smoothly over time. For more, see A. de Cheveigné and H. Kawahara, “YIN, a fundamental frequency estimator for speech and music” (J. Acoust. Soc. Am., 2002); M. Müller, Fundamentals of Music Processing (2nd ed., 2021), chapter 8; and L. Rabiner and R. Schafer, Theory and Applications of Digital Speech Processing (2011), chapter 10.

Reference card

QuantityFormulaNotes
Normalised autocorrelationρx[ℓ]=Rx[ℓ]/Rx[0]\rho_x[\ell]=R_x[\ell]/R_x[0]peaks at the period
Half-period heighteven harmonics add, odd subtractstrong even ones: octave error
YIN differenceDx[ℓ]=∑n=0511(x[n]−x[n+ℓ])2D_x[\ell]=\sum_{n=0}^{511}(x[n]-x[n+\ell])^20 at a perfect match
YIN, normalisedDx′[ℓ]=Dx[ℓ]/(1ℓ∑j=1ℓDx[j])D_x'[\ell]=D_x[\ell]\big/\big(\frac1\ell\sum_{j=1}^{\ell}D_x[j]\big)first dip below 0.1
Refineℓ+d\ell+d, d=(a−c)/2a−2b+cd=\dfrac{(a-c)/2}{a-2b+c}parabola, as in 15.3
Pitchf0=fs/(ℓ+d)f_0=f_s/(\ell+d)lags to 400: down to 40 Hz at 16 kHz
Harmonic product spectrum∣X[k]∣⋯∣X[5k]∣\lvert X[k]\rvert\cdots\lvert X[5k]\rvertfails when a harmonic is missing
Note440⋅2m/12440\cdot2^{m/12} Hz, mm wholeA2 = 110 Hz
Cents1200log⁡2(f/fnote)1200\log_2(f/f_\text{note})100 per semitone, −50 to 50

End of lesson 29.1

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look