Skip to content

Correlation

Slide one signal past another and add the products: the peak gives the delay, and a signal's correlation with itself reveals a hidden period.

Before this24.2 · 4 more
Chapter 24 · Lesson 3 of 4

First, the picture

Slide one signal past another, multiply the samples that meet, and add. Here xx = 1, 2, 1, −1 slides past yy, the same four values 2 samples later. Watch where the tallest sum lands: at lag 2, the delay.

Slide, multiply, add: no flip

x = 1, 2, 1, −1 and y, the same values 2 samples later.

Lag −3: x shifted left by 3; no sample meets a sample of y, so the sum is 0.

lag ℓ
−3
R_yx[ℓ]
0
0.00 / 14.00 s
Describe this picture

Three stacked panels. The first holds yy still, as stems with dot heads, for samples −4 to 9 and heights −1.5 to 2.5. The second, x[n−ℓ]x[n-\ell], has the same axes: it shows xx at its shifted place as stems with square heads, and writes the product of each pair of samples that meet in a row of products under the panel. The third, Ryx[ℓ]R_{yx}[\ell], runs from −2 to 8 against lag ℓ from −3 to 5, and each lag adds one stem with a diamond head. The readouts are the lag and Ryx[ℓ]R_{yx}[\ell], both whole numbers. There is no control. At lag −3 no sample meets a sample of yy, so the sum is 0. The lag then steps from −2 to 5: at each lag, shift xx, multiply where samples meet, add. At the end all lags read 0, 0, −1, −1, 3, 7, 3, −1, −1, and the largest, 7, at ℓ = 2, is circled and labelled “peak”. It equals xx‘s energy, 1 + 4 + 1 + 1.

Slide, multiply, add: no flip

Lay a transparency of one fingerprint over another, and slide it around. At one position the ridges match, and you know the two prints line up. Correlation does the same with two signals. You slide one along the other and look for the place where they agree.

Here is the recipe. Shift xx later by ℓ\ell samples (earlier when ℓ\ell is negative), multiply it sample by sample with yy, and add the products. The shift ℓ\ell is the lag, and the sum is the cross-correlation of yy with xx:

Ryx[ℓ]=∑ny[n] x[n−ℓ].R_{yx}[\ell]=\sum_n y[n]\,x[n-\ell].

Read it the way you read the convolution sum in “Flip and slide: a mechanical way to compute it” of Discrete convolution (5.2). Here x[n−ℓ]x[n-\ell] is xx moved ℓ\ell samples later, the shift of “Sliding in time: shift” in Shifting, reversing and scaling time (2.1). But nothing is flipped: xx keeps its order and only slides.

Let’s try it on four numbers. Take xx = 1, 2, 1, −1 at nn = 0 to 3. Let yy be the same four values two samples later, at nn = 2 to 5. So y[n]=x[n−2]y[n]=x[n-2], and yy is xx delayed by 2. The picture at the top of the page slides xx past yy one lag at a time.

Notice where the tallest stem lands: at ℓ=2\ell=2, the delay between the two signals.

Why is the peak 7? At ℓ=2\ell=2 the shifted xx sits exactly on yy, so every product is a sample times itself: 1+4+1+1=71+4+1+1=7. That is the energy of xx, E=∑nx[n]2E=\sum_n x[n]^2, from “Adding it all up” in How big is a signal (1.3). At every other lag the samples meet out of step, and the products partly cancel.

Which signal slides

In general I write Rxy[ℓ]=∑nx[n] y[n−ℓ]R_{xy}[\ell]=\sum_n x[n]\,y[n-\ell]: the first signal stays still and the second slides. With this order, a peak at ℓ=n0\ell=n_0 means the first signal lags the second by n0n_0 samples. Here yy lags xx by 2, so I put yy first, and the peak lands at +2+2.

Swap the two and the peak moves to −2-2, because Rxy[ℓ]=Ryx[−ℓ]R_{xy}[\ell]=R_{yx}[-\ell]. Some books slide the other signal, so their peaks have the opposite sign. SciPy’s correlate(y, x) uses this page’s order, and correlation_lags gives the lag of each output.

A convolution with one signal reversed

Convolution flips one signal and slides it; correlation only slides. So reverse xx first, and the flip inside the convolution undoes the reversal. Write x[−n]x[-n] for xx played backward, as in “Playing it backward: reversal” of 2.1. Convolving yy with it gives back the correlation:

y[ℓ]∗x[−ℓ]=∑ny[n] x[−(ℓ−n)]=∑ny[n] x[n−ℓ]=Ryx[ℓ].\begin{aligned} y[\ell]*x[-\ell]&=\sum_n y[n]\,x[-(\ell-n)]\\ &=\sum_n y[n]\,x[n-\ell]\\ &=R_{yx}[\ell]. \end{aligned}

In the example, convolve yy with xx reversed, −1, 1, 2, 1, and the same nine numbers come out.

This means every fast way to convolve also correlates. Fast convolution (14.3) multiplies two FFTs; reverse one signal first, and a long correlation costs no more than a long convolution.

A signal against itself

The autocorrelation is a signal’s correlation with itself:

Rx[ℓ]=∑nx[n] x[n−ℓ].R_x[\ell]=\sum_n x[n]\,x[n-\ell].

At lag 0 it is the energy. It is symmetric, Rx[−ℓ]=Rx[ℓ]R_x[-\ell]=R_x[\ell], because sliding a copy forward or back by the same amount matches the same pairs. For our xx it is −1, −1, 3, 7, 3, −1, −1 for ℓ\ell = −3 to 3.

These are the clip’s numbers again, two lags earlier. Since y[n]=x[n−2]y[n]=x[n-2], the cross-correlation is the autocorrelation moved: Ryx[ℓ]=Rx[ℓ−2]R_{yx}[\ell]=R_x[\ell-2].

You met the autocorrelation of a random process in “White noise forgets, coloured noise remembers” of Random processes (24.2). There it was an average, E{x[n] x[n−ℓ]}\mathbb{E}\{x[n]\,x[n-\ell]\}. For one recording I use the plain sum. Divide it by the number of products, and you have 24.2’s estimate.

Two microphones, one delay

Clap once in a room with two microphones. The sound reaches the nearer one first, and the other a moment later. If I can measure that moment, I know how much further the sound travelled to the second microphone. With the microphones’ spacing, that tells me the direction the sound came from.

Real recordings are not four clean numbers. Each microphone adds its own hiss, and the sound itself may be noise, such as a rush of air. The correlation still finds the delay, until the noise is too strong. To see when that happens, I first need a scale for the correlation.

Between −1 and 1

The size of RxyR_{xy} depends on how loud the signals are: double both, and it grows four times. So divide by the square root of the two energies:

ρxy[ℓ]=Rxy[ℓ]Rx[0] Ry[0].\rho_{xy}[\ell]=\frac{R_{xy}[\ell]}{\sqrt{R_x[0]\,R_y[0]}}.

This is the normalised correlation. It lies between −1 and 1. It reaches 1 only when the sliding signal, moved by ℓ\ell, is a positive multiple of the other, and −1 for a negative multiple. In the example, ρyx[2]=7/7⋅7=1\rho_{yx}[2]=7/\sqrt{7\cdot7}=1, because yy is xx moved by 2.

It is the correlation coefficient ρxy\rho_{xy} of “A cloud that leans: correlation” in Random variables for signals (24.1), with the samples as the draws. Treat each nn as one draw of the pair x[n]x[n] and y[n−ℓ]y[n-\ell]. Then ρxy[ℓ]\rho_{xy}[\ell] says how close those pairs come to a straight line. 24.1 subtracted the means first; for signals of mean 0, like the noise in the next instrument, that step changes almost nothing.

This section assumes one thing from statistics: noises from different sources are independent, so their products average towards 0.

The next instrument records a burst of noise sampled at 8 kHz on two microphones. Microphone 2 hears it 23 samples after microphone 1, and each adds its own noise. The SNR of “Signal and noise, in decibels” in 1.3 is the same at both.

I correlate the two recordings with microphone 2 still and microphone 1 sliding, and call the result ρ21[ℓ]\rho_{21}[\ell]. Microphone 2 lags by 23 samples, so the peak should land at ℓ=23\ell=23. Each sum runs over the samples where both recordings exist, and the denominator uses all 2000 samples of each.

Two microphones, one delay

A noise burst at 8 kHz reaches microphone 2 23 samples after microphone 1; each microphone adds its own noise; 2000 samples (seed 243).

SNR +10 dB at each microphone: a clear peak at lag 23, ρ = 0.903. 23 samples at 8 kHz is 2.875 ms, 0.986 m of extra path for sound.

SNR
+10.0 dB
peak at
lag 23
ρ at the peak
0.903
0.00 / 14.00 s
Describe this picture

Three stacked panels for a noise burst at 8 kHz that reaches microphone 2 23 samples after microphone 1; each microphone adds its own noise, over 2000 samples (seed 243). The first two show the first 200 samples of each recording as a thin line, each divided by its own RMS so that both stay inside −5 to 5 at every SNR. The third shows ρ21[ℓ]\rho_{21}[\ell] as stems, from −0.2 to 1 against lag ℓ from −40 to 40 samples. A dotted vertical line labelled “true delay” marks lag 23, and a filled diamond labelled “peak” marks the highest stem. The readouts are the SNR in dB, the lag of the peak, and ρ at the peak, to 3 decimals. At +10 dB there is a clear peak at lag 23, ρ = 0.903: 23 samples at 8 kHz is 2.875 ms, 0.986 m of extra path for sound. At 0 dB, noise as strong as the sound, the peak falls to 0.501 but stays at lag 23, and the rest stays below 0.089. At −10 dB the peak, 0.086, still sits at lag 23, only just above the noise floor’s highest value, 0.081. After the clip a slider named “SNR at each microphone” runs from +10 to −20 dB in steps of 0.5 dB. At other settings the caption reads like “−5.0 dB: peak at lag 23, ρ = 0.238.” When the highest stem is not at lag 23, it reads like “−15.0 dB: the highest value, 0.073, is at lag −30: the delay is lost in the noise.”

Watch the peak at lag 23 shrink as the noise grows, while the stems elsewhere stay low.

Notice the end frame: at −10 dB the peak, 0.086, beats the highest other stem, 0.081, by very little. Move the slider one step, to −10.5 dB, and the stem at lag −30 wins, 0.080 against 0.077 at lag 23. From there down to −20 dB the delay stays lost.

When does a peak stand out?

Why does the peak shrink as it does? Give the sound power 1, and each microphone’s noise power σ2\sigma^2. At the true lag the products of the two sound parts add up. Every product that involves a noise averages towards 0, because the noises are independent of the sound and of each other.

Each recording then has power 1+σ21+\sigma^2, so the normalised peak is about

ρ21[23]≈11+σ2.\rho_{21}[23]\approx\frac{1}{1+\sigma^2}.

From 1.3’s SNR, with the sound’s RMS 1, σ=10−SNRdB/20\sigma=10^{-\mathrm{SNR}_\text{dB}/20}. At +10 dB, σ2=0.1\sigma^2=0.1, and the formula gives 0.909 against 0.903 measured. At 0 dB it gives 0.5 (0.501 measured), and at −10 dB, 0.091 (0.086).

Away from lag 23 the true correlation is 0, but 2000 noisy products do not add to exactly 0. The variances of independent draws add, as 24.1 showed, so the sum of NN products has an RMS N\sqrt N times one product’s. After dividing, the stems spread by about 1/N=1/2000=0.0221/\sqrt N=1/\sqrt{2000}=0.022. Here they measure 0.021 to 0.025.

With 80 such stems, the highest one lands 2.7 to 3.8 spreads up. So my rule of thumb is to trust a peak only when it stands at least 4 spreads tall, about 4/N=0.0894/\sqrt N=0.089 here. At −10 dB the formula’s 0.091 barely clears that, and half a decibel later the delay is lost.

Two things lift a weak peak. A longer recording lowers the floor: four times the samples halves 1/N1/\sqrt N. A louder sound, or quieter microphones, raises the peak itself.

The maths behind it · angles between vectors

Treat each recording as a vector of 2000 numbers. The normalised correlation at one lag is the cosine of the angle between the two vectors, with one of them shifted. The peak lag is the shift that makes them most parallel.

From lag to distance

Once the peak’s lag n0n_0 is known, the delay in seconds is n0/fsn_0/f_s. Here that is 23/8000=2.87523/8000=2.875 ms. Sound travels about 343 m/s in air at 20 °C, so the extra path is 343×0.002875=0.986343\times0.002875=0.986 m.

At 8 kHz one sample is 4.3 cm of path. So the lag’s whole-number steps limit how finely the direction can be found. Faster sampling helps, and so does fitting a smooth curve through the peak and its two neighbours.

Rooms add echoes, and each echo makes a smaller peak of its own. Practical systems often weight the spectrum before correlating, to sharpen the main peak. That is the generalised cross-correlation of C. H. Knapp and G. C. Carter (1976), and this page leaves it there.

A period hidden in noise

The autocorrelation has a second use. A periodic signal lands on itself when you move it by one period. So its autocorrelation peaks again at every multiple of the period. White noise does not: as 24.2 showed, it is alike with itself only at lag 0.

Take a sine of period PP samples, Asin⁡(2πn/P)A\sin(2\pi n/P), in white noise of power σ2\sigma^2. Multiply the sine by itself moved ℓ\ell samples, and use sin⁡a sin⁡b=12[cos⁡(a−b)−cos⁡(a+b)]\sin a\,\sin b=\tfrac12[\cos(a-b)-\cos(a+b)]. The term cos⁡(a+b)\cos(a+b) swings through whole periods and averages to 0. The term cos⁡(a−b)=cos⁡(2πℓ/P)\cos(a-b)=\cos(2\pi\ell/P) does not depend on nn at all.

So, per sample, the sine contributes (A2/2)cos⁡(2πℓ/P)(A^2/2)\cos(2\pi\ell/P), and the noise adds σ2\sigma^2 at lag 0 only. Divided by the value at lag 0:

Rx[ℓ]Rx[0]≈A2/2A2/2+σ2cos⁡2πℓP(ℓ≠0).\begin{aligned} \frac{R_x[\ell]}{R_x[0]}&\approx\frac{A^2/2}{A^2/2+\sigma^2}\cos\frac{2\pi\ell}{P}\\ &\quad(\ell\ne0). \end{aligned}

With A=0.7A=0.7 and σ2=1\sigma^2=1, the factor is 0.245/1.245=0.1970.245/1.245=0.197. The SNR is 10log⁡100.245=−6.110\log_{10}0.245=-6.1 dB. This draw’s noise measures RMS 1.010, which makes it −6.2 dB.

x[n]20−20sample n199Rx[ℓ] / Rx[0]100204060lag ℓ
Fig. A sine of period 20 at −6.2 dB SNR: invisible in the trace, but its autocorrelation rises to 0.202, 0.201 and 0.217 at lags 20, 40 and 60 and dips to −0.244, −0.205, −0.188 in between, a cosine of period 20; the formula gives 0.197 for the peaks.

The six values in the caption sit within 0.05 of 0.197 or −0.197, about two of the floor’s spreads of 1/2000=0.0221/\sqrt{2000}=0.022. Nothing in the trace shows the sine, but the ripple’s period, 20 lags, is the sine’s period.

A pitch detector works this way. The first tall peak of a voice’s autocorrelation after lag 0 gives the period of the voice, even when the waveform looks ragged. Pitch detection (29.1) builds one.

The maths behind it · the cross-correlation function

In statistics, the cross-correlation function (CCF) between two time series is this ρxy[ℓ]\rho_{xy}[\ell]. A lagged relation between two series, such as rainfall and a river’s level days later, shows as a CCF peak at the lag.

Worked example

1. By hand. Take xx = 1, 2, 1, −1 at nn = 0 to 3, and y[n]=x[n−2]y[n]=x[n-2]. At each lag, multiply the samples that meet and add:

lagproducts where samples meetsum
−3, −2none0
−11·(−1)−1
01·1 + 2·(−1)−1
11·2 + 2·1 + 1·(−1)3
21·1 + 2·2 + 1·1 + (−1)·(−1)7
32·1 + 1·2 + (−1)·13
41·1 + (−1)·2−1
5(−1)·1−1

The peak, 7, is at ℓ=2\ell=2, the delay. Normalised, it is 7/7⋅7=17/\sqrt{7\cdot7}=1. Convolving yy with −1, 1, 2, 1, the reversed xx, gives the same nine numbers.

2. Delay to distance. The peak at lag 23, at fs=8f_s=8 kHz, is 23/8000=2.87523/8000=2.875 ms. At 343 m/s that is 0.9860.986 m of extra path.

3. Will the peak stand out? At −5 dB, σ2=100.5=3.162\sigma^2=10^{0.5}=3.162, so the peak should be about 1/4.162=0.2401/4.162=0.240. The instrument measures 0.238, well above 4/2000=0.0894/\sqrt{2000}=0.089, and the peak sits at lag 23.

Where you’ll meet this

A GPS receiver correlates what it hears with each satellite’s known code. The peak lag is the signal’s travel time, and travel times from several satellites give the position. Radar and sonar do the same with an echo of their own pulse: the lag is the round trip, so the distance is half the path.

Microphone arrays in conference phones and smart speakers estimate the delay between microphones to find who is talking. Modems find the start of each packet by correlating with a known pattern at its start, the preamble. In images, sliding a small template over a picture and correlating finds where the template appears.

To find a known pulse in white noise, correlating with the pulse gives the highest SNR any linear filter can reach; Matched filters and detection (26.1) shows why. And the DTFT of the autocorrelation is the spectrum of a random signal, the subject of Power spectral density (24.4).

Reference card

QuantityFormulaNotes
Cross-correlationRxy[ℓ]=∑nx[n] y[n−ℓ]R_{xy}[\ell]=\sum_nx[n]\,y[n-\ell]slide, multiply, add
As convolutionRxy[ℓ]=x[ℓ]∗y[−ℓ]R_{xy}[\ell]=x[\ell]*y[-\ell]FFT route (14.3)
AutocorrelationRx[ℓ]=∑nx[n] x[n−ℓ]R_x[\ell]=\sum_nx[n]\,x[n-\ell]symmetric; Rx[0]R_x[0] is the energy
Normalisedρxy[ℓ]=Rxy[ℓ]/Rx[0]Ry[0]\rho_{xy}[\ell]=R_{xy}[\ell]/\sqrt{R_x[0]R_y[0]}between −1 and 1
Delaypeak at ℓ=n0\ell=n_0time n0/fsn_0/f_s
Noise floorspread about 1/N1/\sqrt Ntrust a peak above about 4/N4/\sqrt N
Hidden periodRx[ℓ]R_x[\ell] peaks at multiples of the periodnoise averages out

End of lesson 24.3

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look