Skip to content

Averaged periodograms: Bartlett and Welch

Cut one record into segments and average their periodograms: the scatter shrinks, the peaks blur, and overlapping Hann-windowed segments buy extra averages.

Before this15.5 · 25.1 · 3 more
Chapter 25 · Lesson 2 of 4

First, the picture

Cut one long record of noise into shorter segments, and average their periodograms. Here 4096 samples are cut into 1, 4, 16 and then 64 segments. Watch the estimate smooth out, while its peak sinks below the dashed truth.

Average the pieces: less scatter, less detail

4096 samples of AR(2) noise (seed 252; driving noise of variance 1, measured 0.994), cut into K segments of length L; rectangular window, no overlap.

One segment, the whole record: the raw periodogram, spread 5.51 dB.

L
4096
K
1
spread
5.51 dB
average peak
22.0 dB
0.00 / 14.00 s
Describe this picture

One panel: 4096 samples of AR(2) noise (seed 252; driving noise of variance 1, measured 0.994), cut into KK segments of length LL, with a rectangular window and no overlap. PSD in dB from −45 to 35 against Ω from 0 to π rad/sample. The estimate is a solid line through its L/2+1L/2+1 bins, and the true PSD a dashed line; a key names them “estimate” and “true PSD”. The readouts are LL, KK, the spread and the average peak, both in dB. The clip starts with one segment, the whole record: the raw periodogram, spread 5.51 dB, average peak 22.0 dB. Then K=4K=4 segments of 1024: spread 2.48 dB, peak 22.0 dB. Then K=16K=16 of 256: spread 1.15 dB, and the peak starts to blur, 21.7 dB on average. It ends at K=64K=64 of 64: a smooth estimate, spread 0.50 dB, but the peak averages only 20.6 dB against the true 22.1 dB. Each 4× in KK roughly halves the spread and costs resolution. When the clip has finished, a slider named “Segment length L” sets LL to a power of two from 64 to 4096.

Average the pieces: less scatter, less detail

The periodogram (25.1) left us with two faults. In “More data, same scatter” its values scattered about the true PSD by as much as the PSD itself, however long the record was. In “Short records blur the peak” a short record smeared a sharp peak and lifted the valleys.

This page cures the first fault with an average, and pays for it with the second. Here is the recipe.

Take a record of NN samples and cut it into KK pieces, called segments, each LL samples long, so K=N/LK=N/L. Earlier chapters used KK for a harmonic and for a gain, and LL for an interpolation factor; on this page KK counts segments and LL is a length.

Take each segment’s periodogram, add the KK of them, and divide by KK. That average is the estimate S^x(ejΩ)\hat S_x(e^{j\Omega}); the hat means “estimated”, as in 25.1. Segment ii is the stretch x[iL]x[iL] to x[iL+L−1]x[iL+L-1], counting from i=0i=0, and its periodogram is 25.1’s with LL samples:

1L∣∑n=0L−1x[iL+n] e−jΩn∣2.\frac1L\Bigl\lvert\sum_{n=0}^{L-1}x[iL+n]\,e^{-j\Omega n}\Bigr\rvert^2.

This average is Bartlett’s method. Why should it help? One segment’s periodogram scatters about its own average by about that average (25.1).

Suppose the segments are nearly independent of each other. Then “A longer average: less noise, more delay” in Simple smoothing filters (18.2) gives the answer: averaging KK unrelated values divides their spread by K\sqrt K. Four segments halve the scatter, and sixteen leave a quarter.

I assume this on the whole page. It is fair here: for this page’s noise, samples more than 57 apart have a correlation coefficient under 0.05, and every segment is at least 64 long.

The price is detail. Each periodogram sees only LL samples, so it blurs like 25.1’s short record. In “Five windows, one trade” of Window functions compared (15.2), the rectangular window’s main lobe is 2 bins wide, and one bin is 2π/L2\pi/L. So the estimate cannot show details much narrower than about 2π/L2\pi/L.

It is like several quick glances at a scene against one long look: steadier, but less detailed. With NN fixed, more segments means shorter ones, and you cannot have both.

To measure the scatter alone, I compare the estimate with its own average, bin by bin. The spread is the standard deviation, over the bins, of 10log⁡10(S^x/E{S^x})10\log_{10}\bigl(\hat S_x/\mathbb{E}\{\hat S_x\}\bigr), in dB. For each LL the average E{S^x}\mathbb{E}\{\hat S_x\} is 25.1’s average periodogram with LL in place of NN, so it already contains the blur, and the spread shows only the scatter.

For one periodogram, K=1K=1, the exponential law of 25.1 gives a spread of 5.57 dB. Statistics derives that number; I only use it.

The record is 4096 samples of 25.1’s AR(2) noise,

x[n]=1.1168 x[n−1]−0.9025 x[n−2]+v[n],\begin{aligned} x[n]={}&1.1168\,x[n-1]\\ &-0.9025\,x[n-2]+v[n], \end{aligned}

with v[n]v[n] white Gaussian noise of variance 1 (seed 252; measured 0.994). I run the recursion for 500 samples first and drop them, so the record starts settled. Its true PSD has a sharp peak of 22.1 dB at Ω=0.2998π\Omega=0.2998\pi and falls to −9.6-9.6 dB at π\pi.

The picture at the top of the page cuts this record into KK segments of length LL, with a rectangular window and no overlap. Its spread reads 5.51, 2.48, 1.15 and 0.50 dB for KK = 1, 4, 16 and 64. Its average peak is the average periodogram at the true peak, so it shows the blur without the scatter.

After the clip, drag the slider to 128 and watch the two readouts that matter. KK becomes 32, the spread 0.86 dB, and the average peak 21.3 dB: steadier than at 256, and blurrier.

Look at the end frame. The estimate is smooth, but its peak is only 20.6 dB on average, 1.5 dB under the truth. That is the blur. One bin at L=64L=64 is 2π/64=0.0982\pi/64=0.098 rad, about the width of the true peak itself, 0.103 rad at half its height.

At L=256L=256 a bin is 2π/256=0.0252\pi/256=0.025 rad, a quarter of the peak’s width. There the average peak, 21.7 dB, is within 0.4 dB of the truth. The valley fills too: at L=64L=64 the rectangular window’s leakage lifts the average at π\pi from −9.6-9.6 dB to −7.1-7.1 dB.

Now the scatter. In 25.1’s terms, spread ÷ mean of S^x/E{S^x}\hat S_x/\mathbb{E}\{\hat S_x\} is 0.967, 0.515, 0.273 and 0.112 for KK = 1, 4, 16 and 64. That is close to 18.2’s 1, 1/2, 1/4 and 1/8.

In dB the rule bends a little. A small change by a fraction ee moves a level by about 4.34e4.34e dB; 10 % is 0.41 dB. So once KK is large the spread is about 4.34/K4.34/\sqrt K dB: 1.09 dB at K=16K=16 and 0.54 dB at K=64K=64, against the measured 1.15 and 0.50.

So choose LL from the detail you need. Make one bin, 2π/L2\pi/L, a few times narrower than the narrowest peak you want to see; then K=N/LK=N/L is what the record gives you. In hertz, one bin is fs/Lf_s/L.

The maths behind it · the variance of a mean

Averaging KK independent exponential variables gives a gamma law, a scaled chi-square with 2K2K degrees of freedom. Its spread relative to its mean is 1/K1/\sqrt K: the variance of a mean, from statistics.

Overlap buys segments

Welch’s method makes two changes to Bartlett’s recipe. First, it multiplies each segment by a window w[n]w[n] before the DFT. A Hann window has side lobes from −31.5 dB instead of the rectangular window’s −13.3 dB, so it leaks far less. Its main lobe is 4 bins instead of 2, so it blurs a little more (15.2).

That window costs data, though. Hann falls to 0 at both ends, so the samples near each join between segments barely count. “Hop and overlap” in Spectrograms & the STFT (15.5) found the same for frames.

So the second change is to overlap the segments. Each one starts a hop after the last one, with hop=L(1−overlap)\text{hop}=L(1-\text{overlap}). The last segment must fit inside the record, so the number of segments is

K=⌊N−Lhop⌋+1,K=\left\lfloor\frac{N-L}{\text{hop}}\right\rfloor+1,

where ⌊⋅⌋\lfloor\cdot\rfloor rounds down. With no overlap the hop is LL, and K=N/LK=N/L as before.

Welch’s method averages KK windowed periodograms, as Bartlett’s averages plain ones. Segment ii now starts at sample i hopi\,\text{hop}, and its windowed periodogram is

∣∑n=0L−1w[n] x[i hop+n] e−jΩn∣2∑n=0L−1w[n]2.\frac{\Bigl\lvert\sum_{n=0}^{L-1}w[n]\,x[i\,\text{hop}+n]\,e^{-j\Omega n}\Bigr\rvert^2}{\sum_{n=0}^{L-1}w[n]^2}.

The window takes away some of the noise’s power, and dividing by ∑w[n]2\sum w[n]^2 instead of LL puts it back: for white noise the average comes out right. With w[n]=1w[n]=1 the sum is LL, and Bartlett’s recipe returns.

SciPy computes exactly this: scipy.signal.welch(x, window='hann', nperseg=256, noverlap=128, detrend=False, return_onesided=False) gives the instrument’s estimate at 50 % overlap.

The instrument keeps the same 4096 samples and segments of 256, and changes only the overlap. Picture a relay, where each runner starts before the last one finishes.

Overlap buys segments

The same 4096 samples, Hann-windowed segments of 256.

No overlap: 16 Hann segments, spread 1.08 dB.

overlap
0 %
segments K
16
spread
1.08 dB
Overlap
0.00 / 12.00 s
Describe this picture

Two stacked panels for the same 4096 samples, cut into Hann-windowed segments of 256. The first shows every Hann window the record holds against sample nn from 0 to 4095, each a hop after the last: faint outlines across the whole record, with the first six drawn bold so that you can see how they overlap. The second has the axes and key of the first instrument: the estimate is a solid line, and the true PSD a dashed one. The readouts are the overlap, the number of segments KK and the spread. With no overlap there are 16 Hann segments and the spread is 1.08 dB. At 50 % overlap there are 31 segments from the same data, spread 0.76 dB. At 75 % there are 61 segments, spread 0.75 dB: the segments now repeat each other, so the gain stops. When the clip has finished, three buttons, “0 %”, “50 %” and “75 %”, in a group named “Overlap”, switch between the three.

Watch the spread as the windows move closer: 1.08 dB, then 0.76 dB, then 0.75 dB.

Notice the last step. Going from 50 % to 75 % doubles the segments, but the spread only falls from 0.76 dB to 0.75 dB.

The reason is that overlapping segments share samples, so their periodograms are no longer independent. For white noise, two Hann segments that overlap by half give periodograms with a correlation coefficient of only 0.03: nearly independent. At 75 % overlap, the periodograms of next-door segments correlate by 0.43.

Welch worked out what that does to the scatter. At 50 % overlap the 31 segments scatter like 29.4 independent ones; at 75 % the 61 scatter like only 32.0. So with Hann, 50 % overlap already gets most of what overlap can give, and more overlap adds little.

The same trade from the lag side

There is a second road to the same place, through the autocorrelation. Random processes (24.2) estimated R^x[ℓ]\hat R_x[\ell] from one recording, and Power spectral density (24.4) showed that the PSD is the DTFT of Rx[ℓ]R_x[\ell]. So take the DTFT of the estimate.

Over all lags, that DTFT is called the correlogram. With the division by NN that 24.2 mentions in place of N−ℓN-\ell, it equals the periodogram exactly, and it scatters just as much.

The trouble is in the far lags. 24.2’s estimate at lag ℓ\ell averages only N−ℓN-\ell products, so the far lags are noisy. The Blackman–Tukey method multiplies the estimate by a lag window w[ℓ]w[\ell], which is 1 at ℓ=0\ell=0, falls to 0 by a chosen largest lag and stays 0 beyond it. Then it takes the DTFT:

S^x(ejΩ)=∑ℓw[ℓ] R^x[ℓ] e−jΩℓ.\hat S_x(e^{j\Omega})=\sum_\ell w[\ell]\,\hat R_x[\ell]\,e^{-j\Omega\ell}.

A short lag window gives a smooth spectrum and a long one a sharp, noisy spectrum: the same trade. Bartlett’s method sits on the same road. By 25.1, its average is the DTFT of Rx[ℓ]R_x[\ell] tapered by the triangle 1−∣ℓ∣/L1-\lvert\ell\rvert/L, a lag window that reaches lag L−1L-1.

Multitaper estimates are a third relative: they average the periodograms of the whole record seen through several different, carefully chosen windows.

The maths behind it · correlation matrices

Write segment ii as a vector xi\mathbf{x}_i. Its periodogram is a squared coordinate of xi\mathbf{x}_i along a DFT row of The DFT as a matrix (13.5). Averaging over segments averages the rank-one matrices xixiH\mathbf{x}_i\mathbf{x}_i^{\mathsf H} into an estimate of the correlation matrix, before that projection.

Worked example

1. Bartlett. Cutting 4096 samples into segments of 4096, 1024, 256 and 64 gives K=4096/LK=4096/L = 1, 4, 16 and 64. The spread is 5.51 dB, 2.48 dB, 1.15 dB and 0.50 dB: each 4× in KK multiplies it by 0.45, 0.46 and 0.44 in turn.

2. Welch. Segments of 256 with hops of 256, 128 and 64 give

K=⌊3840/256⌋+1=16,K=⌊3840/128⌋+1=31,K=⌊3840/64⌋+1=61.\begin{aligned} K&=\lfloor3840/256\rfloor+1=16,\\ K&=\lfloor3840/128\rfloor+1=31,\\ K&=\lfloor3840/64\rfloor+1=61. \end{aligned}

The spread is 1.08 dB, 0.76 dB and 0.75 dB.

3. Choosing LL for this record. The true peak is 0.103 rad wide at half its height. At L=64L=64 one bin, 0.098 rad, is about as wide as the peak, and the peak loses 1.5 dB.

At L=256L=256 a bin is about a quarter of the peak’s width, the average peak is 21.7 dB, and Hann with 50 % overlap gives a spread of 0.76 dB. At fs=1f_s=1 kHz, L=256L=256 resolves about fs/L=3.9f_s/L=3.9 Hz.

Where you’ll meet this

Welch’s method is the usual PSD estimate in measurement software, vibration analysis and EEG, and Hann with 50 % overlap is the usual setting. SciPy’s welch uses it by default: a Hann window, segments of 256 and an overlap of half a segment.

Two of its defaults differ from this page. It removes each segment’s mean first. It also returns the positive frequencies only, doubled except at 0 and π\pi: the one-sided density of Reading a spectrum: scaling and units (15.4). Its fs argument turns the result into units per hertz.

No average beats the limit of about 2π/L2\pi/L. When that limit is too coarse for the record you have, other estimates help. Parametric models and linear prediction (25.3) fits a model to the data, and High-resolution frequency estimation (25.4) separates tones closer than one bin.

For more, see P. D. Welch, “The use of fast Fourier transform for the estimation of power spectra”, IEEE Trans. Audio and Electroacoustics (1967); J. G. Proakis and D. G. Manolakis, Digital Signal Processing, ch. 14; and P. Stoica and R. Moses, Spectral Analysis of Signals (2005), ch. 2.

Reference card

QuantityFormulaNotes
Bartlettaverage of KK periodograms of length LLspread roughly ∝1/K\propto 1/\sqrt K
Resolutionabout 2π/L2\pi/Lshorter segments blur; Hann’s main lobe is twice as wide
Welchwindowed, overlapping segmentsHann, 50 %: the usual choice
SegmentsK=⌊(N−L)/hop⌋+1K=\lfloor(N-L)/\text{hop}\rfloor+1hop=L(1−overlap)\text{hop}=L(1-\text{overlap})
Blackman–TukeyDTFT of w[ℓ]R^x[ℓ]w[\ell]\hat R_x[\ell]lag window

End of lesson 25.2

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look