Cut one long record of noise into shorter segments, and average their periodograms. Here 4096 samples are cut into 1, 4, 16 and then 64 segments. Watch the estimate smooth out, while its peak sinks below the dashed truth.
Average the pieces: less scatter, less detail
4096 samples of AR(2) noise (seed 252; driving noise of variance 1, measured 0.994), cut into K segments of length L; rectangular window, no overlap.
One segment, the whole record: the raw periodogram, spread 5.51 dB.
Describe this picture
One panel: 4096 samples of AR(2) noise (seed 252; driving noise of variance 1, measured 0.994), cut into segments of length , with a rectangular window and no overlap. PSD in dB from −45 to 35 against Ω from 0 to π rad/sample. The estimate is a solid line through its bins, and the true PSD a dashed line; a key names them “estimate” and “true PSD”. The readouts are , , the spread and the average peak, both in dB. The clip starts with one segment, the whole record: the raw periodogram, spread 5.51 dB, average peak 22.0 dB. Then segments of 1024: spread 2.48 dB, peak 22.0 dB. Then of 256: spread 1.15 dB, and the peak starts to blur, 21.7 dB on average. It ends at of 64: a smooth estimate, spread 0.50 dB, but the peak averages only 20.6 dB against the true 22.1 dB. Each 4× in roughly halves the spread and costs resolution. When the clip has finished, a slider named “Segment length L” sets to a power of two from 64 to 4096.
Average the pieces: less scatter, less detail
The periodogram (25.1) left us with two faults. In “More data, same scatter” its values scattered about the true PSD by as much as the PSD itself, however long the record was. In “Short records blur the peak” a short record smeared a sharp peak and lifted the valleys.
This page cures the first fault with an average, and pays for it with the second. Here is the recipe.
Take a record of samples and cut it into pieces, called segments, each samples long, so . Earlier chapters used for a harmonic and for a gain, and for an interpolation factor; on this page counts segments and is a length.
Take each segment’s periodogram, add the of them, and divide by . That average is the estimate ; the hat means “estimated”, as in 25.1. Segment is the stretch to , counting from , and its periodogram is 25.1’s with samples:
This average is Bartlett’s method. Why should it help? One segment’s periodogram scatters about its own average by about that average (25.1).
Suppose the segments are nearly independent of each other. Then “A longer average: less noise, more delay” in Simple smoothing filters (18.2) gives the answer: averaging unrelated values divides their spread by . Four segments halve the scatter, and sixteen leave a quarter.
I assume this on the whole page. It is fair here: for this page’s noise, samples more than 57 apart have a correlation coefficient under 0.05, and every segment is at least 64 long.
The price is detail. Each periodogram sees only samples, so it blurs like 25.1’s short record. In “Five windows, one trade” of Window functions compared (15.2), the rectangular window’s main lobe is 2 bins wide, and one bin is . So the estimate cannot show details much narrower than about .
It is like several quick glances at a scene against one long look: steadier, but less detailed. With fixed, more segments means shorter ones, and you cannot have both.
To measure the scatter alone, I compare the estimate with its own average, bin by bin. The spread is the standard deviation, over the bins, of , in dB. For each the average is 25.1’s average periodogram with in place of , so it already contains the blur, and the spread shows only the scatter.
For one periodogram, , the exponential law of 25.1 gives a spread of 5.57 dB. Statistics derives that number; I only use it.
The record is 4096 samples of 25.1’s AR(2) noise,
with white Gaussian noise of variance 1 (seed 252; measured 0.994). I run the recursion for 500 samples first and drop them, so the record starts settled. Its true PSD has a sharp peak of 22.1 dB at and falls to dB at .
The picture at the top of the page cuts this record into segments of length , with a rectangular window and no overlap. Its spread reads 5.51, 2.48, 1.15 and 0.50 dB for = 1, 4, 16 and 64. Its average peak is the average periodogram at the true peak, so it shows the blur without the scatter.
After the clip, drag the slider to 128 and watch the two readouts that matter. becomes 32, the spread 0.86 dB, and the average peak 21.3 dB: steadier than at 256, and blurrier.
Look at the end frame. The estimate is smooth, but its peak is only 20.6 dB on average, 1.5 dB under the truth. That is the blur. One bin at is rad, about the width of the true peak itself, 0.103 rad at half its height.
At a bin is rad, a quarter of the peak’s width. There the average peak, 21.7 dB, is within 0.4 dB of the truth. The valley fills too: at the rectangular window’s leakage lifts the average at from dB to dB.
Now the scatter. In 25.1’s terms, spread ÷ mean of is 0.967, 0.515, 0.273 and 0.112 for = 1, 4, 16 and 64. That is close to 18.2’s 1, 1/2, 1/4 and 1/8.
In dB the rule bends a little. A small change by a fraction moves a level by about dB; 10 % is 0.41 dB. So once is large the spread is about dB: 1.09 dB at and 0.54 dB at , against the measured 1.15 and 0.50.
So choose from the detail you need. Make one bin, , a few times narrower than the narrowest peak you want to see; then is what the record gives you. In hertz, one bin is .
The maths behind it · the variance of a mean
Averaging independent exponential variables gives a gamma law, a scaled chi-square with degrees of freedom. Its spread relative to its mean is : the variance of a mean, from statistics.
Overlap buys segments
Welch’s method makes two changes to Bartlett’s recipe. First, it multiplies each segment by a window before the DFT. A Hann window has side lobes from −31.5 dB instead of the rectangular window’s −13.3 dB, so it leaks far less. Its main lobe is 4 bins instead of 2, so it blurs a little more (15.2).
That window costs data, though. Hann falls to 0 at both ends, so the samples near each join between segments barely count. “Hop and overlap” in Spectrograms & the STFT (15.5) found the same for frames.
So the second change is to overlap the segments. Each one starts a hop after the last one, with . The last segment must fit inside the record, so the number of segments is
where rounds down. With no overlap the hop is , and as before.
Welch’s method averages windowed periodograms, as Bartlett’s averages plain ones. Segment now starts at sample , and its windowed periodogram is
The window takes away some of the noise’s power, and dividing by instead of puts it back: for white noise the average comes out right. With the sum is , and Bartlett’s recipe returns.
SciPy computes exactly this: scipy.signal.welch(x, window='hann', nperseg=256, noverlap=128, detrend=False, return_onesided=False) gives the instrument’s estimate at 50 % overlap.
The instrument keeps the same 4096 samples and segments of 256, and changes only the overlap. Picture a relay, where each runner starts before the last one finishes.
Overlap buys segments
The same 4096 samples, Hann-windowed segments of 256.
No overlap: 16 Hann segments, spread 1.08 dB.
Describe this picture
Two stacked panels for the same 4096 samples, cut into Hann-windowed segments of 256. The first shows every Hann window the record holds against sample from 0 to 4095, each a hop after the last: faint outlines across the whole record, with the first six drawn bold so that you can see how they overlap. The second has the axes and key of the first instrument: the estimate is a solid line, and the true PSD a dashed one. The readouts are the overlap, the number of segments and the spread. With no overlap there are 16 Hann segments and the spread is 1.08 dB. At 50 % overlap there are 31 segments from the same data, spread 0.76 dB. At 75 % there are 61 segments, spread 0.75 dB: the segments now repeat each other, so the gain stops. When the clip has finished, three buttons, “0 %”, “50 %” and “75 %”, in a group named “Overlap”, switch between the three.
Watch the spread as the windows move closer: 1.08 dB, then 0.76 dB, then 0.75 dB.
Notice the last step. Going from 50 % to 75 % doubles the segments, but the spread only falls from 0.76 dB to 0.75 dB.
The reason is that overlapping segments share samples, so their periodograms are no longer independent. For white noise, two Hann segments that overlap by half give periodograms with a correlation coefficient of only 0.03: nearly independent. At 75 % overlap, the periodograms of next-door segments correlate by 0.43.
Welch worked out what that does to the scatter. At 50 % overlap the 31 segments scatter like 29.4 independent ones; at 75 % the 61 scatter like only 32.0. So with Hann, 50 % overlap already gets most of what overlap can give, and more overlap adds little.
The same trade from the lag side
There is a second road to the same place, through the autocorrelation. Random processes (24.2) estimated from one recording, and Power spectral density (24.4) showed that the PSD is the DTFT of . So take the DTFT of the estimate.
Over all lags, that DTFT is called the correlogram. With the division by that 24.2 mentions in place of , it equals the periodogram exactly, and it scatters just as much.
The trouble is in the far lags. 24.2’s estimate at lag averages only products, so the far lags are noisy. The Blackman–Tukey method multiplies the estimate by a lag window , which is 1 at , falls to 0 by a chosen largest lag and stays 0 beyond it. Then it takes the DTFT:
A short lag window gives a smooth spectrum and a long one a sharp, noisy spectrum: the same trade. Bartlett’s method sits on the same road. By 25.1, its average is the DTFT of tapered by the triangle , a lag window that reaches lag .
Multitaper estimates are a third relative: they average the periodograms of the whole record seen through several different, carefully chosen windows.
The maths behind it · correlation matrices
Write segment as a vector . Its periodogram is a squared coordinate of along a DFT row of The DFT as a matrix (13.5). Averaging over segments averages the rank-one matrices into an estimate of the correlation matrix, before that projection.
Worked example
1. Bartlett. Cutting 4096 samples into segments of 4096, 1024, 256 and 64 gives = 1, 4, 16 and 64. The spread is 5.51 dB, 2.48 dB, 1.15 dB and 0.50 dB: each 4× in multiplies it by 0.45, 0.46 and 0.44 in turn.
2. Welch. Segments of 256 with hops of 256, 128 and 64 give
The spread is 1.08 dB, 0.76 dB and 0.75 dB.
3. Choosing for this record. The true peak is 0.103 rad wide at half its height. At one bin, 0.098 rad, is about as wide as the peak, and the peak loses 1.5 dB.
At a bin is about a quarter of the peak’s width, the average peak is 21.7 dB, and Hann with 50 % overlap gives a spread of 0.76 dB. At kHz, resolves about Hz.
Where you’ll meet this
Welch’s method is the usual PSD estimate in measurement software, vibration analysis and EEG, and Hann with 50 % overlap is the usual setting. SciPy’s welch uses it by default: a Hann window, segments of 256 and an overlap of half a segment.
Two of its defaults differ from this page. It removes each segment’s mean first. It also returns the positive frequencies only, doubled except at 0 and : the one-sided density of Reading a spectrum: scaling and units (15.4). Its fs argument turns the result into units per hertz.
No average beats the limit of about . When that limit is too coarse for the record you have, other estimates help. Parametric models and linear prediction (25.3) fits a model to the data, and High-resolution frequency estimation (25.4) separates tones closer than one bin.
For more, see P. D. Welch, “The use of fast Fourier transform for the estimation of power spectra”, IEEE Trans. Audio and Electroacoustics (1967); J. G. Proakis and D. G. Manolakis, Digital Signal Processing, ch. 14; and P. Stoica and R. Moses, Spectral Analysis of Signals (2005), ch. 2.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Bartlett | average of periodograms of length | spread roughly |
| Resolution | about | shorter segments blur; Hann’s main lobe is twice as wide |
| Welch | windowed, overlapping segments | Hann, 50 %: the usual choice |
| Segments | ||
| Blackman–Tukey | DTFT of | lag window |