Here are eight recordings of the same kind of noise, drawn one above the other. You can average across them at one instant, or along one recording. Watch the two averages: for plain noise both sit near 0, but once each recording gets a level of its own, the average along one recording follows its level.
Across the ensemble, along a realisation
8 of 200 realisations, 256 samples each (seed 242). A: white noise. B: the same noise plus a level drawn once per realisation.
Process A, white noise. Across all 200 realisations at n = 30 the average is −0.121; along realisation 8 it is 0.065: both near 0, the process's mean.
Describe this picture
Eight stacked strips, realisations 1 to 8 of 200, each a thin trace of 256 samples (seed 242), named “r = 1” to “r = 8”. They share the sample axis, 0 to 255, and each runs from −5 to 5. Realisation 8 is picked out, because the second readout follows it, and a dashed vertical line labelled “n = 30” marks the instant where the first readout looks down the ensemble. The readouts are the process, the average across 200 at and the average along realisation 8. There is no control. Process A, white noise, comes first. Dots light up where the dashed line crosses the strips, a line marks realisation 8’s own average, and a short bar beside each strip shows that realisation’s average, drawn 2.5 times taller than the strip; in process A all eight bars sit at 0. Across the ensemble the average is −0.121, and along realisation 8 it is 0.065: both near 0, the process’s mean. Then each realisation slides up or down by its own level, which turns it into process B. Each strip gets a dotted line at its level, with the band between 0 and the level shaded, and the bars now sit at different heights; a key names the lines “time average” and “level”. Across the ensemble the average is still near 0 (−0.082), but along realisation 8 it is −1.357, close to that realisation’s own level, −1.422. The end caption says: “One recording cannot reveal the ensemble mean: B is not ergodic.”
A whole signal at random
In Random variables for signals (24.1) each draw gave one number. Noise on a wire is more than one number, though. It is a whole signal, with a new value at every sample, and every recording of it comes out different.
So let’s draw a whole signal at once. A random process hands out whole signals at random. Each signal it hands out is a realisation, and I number them: is realisation . The collection of all the realisations it could hand out is the ensemble.
Picture 200 identical microphones in 200 identical quiet rooms, all recording at once. Each recording is a realisation. Line them up one above the other, and you are looking at a slice of the ensemble.
From 24.1 I assume what “Draws pile up into a density” and “A cloud that leans: correlation” teach: a density, the mean , the variance , estimates such as , independence and the correlation coefficient. I add one fact: averages add, so the mean of a sum is the sum of the means. Statistics proves these; this page only uses them.
Fix one instant and look down the ensemble. Every realisation has a value there, so is a random variable of 24.1, with its own density and its own mean . The mean carries the index because it may change with .
That gives two ways to average. Across the ensemble means many realisations at one instant. With 200 realisations it estimates :
Along a realisation means one realisation at many instants. That is the mean of “Four everyday sizes” in How big is a signal (1.3), the running average of “Average per sample” without the squaring:
When do the two agree? A process whose averages along one long realisation equal its averages across the ensemble is called ergodic. This matters, because in practice you have one recording, not 200.
The picture at the top of the page compares two processes. Process A is white noise of variance 1: every sample is a fresh Gaussian draw of 24.1, independent of all the others. The name gets its exact meaning later on this page.
Process B is the same noise plus a level: one more Gaussian draw of variance 1, made once per realisation and added to every sample of it. Within one realisation the level never changes; from one realisation to the next it does.
Across the ensemble, along a realisation
In process A, the picture’s two averages came out at −0.121 across all 200 realisations at , and 0.065 along realisation 8. Both are near 0, the process’s mean. In process B they are −0.082 across the ensemble, still near 0, but −1.357 along realisation 8, close to that realisation’s own level, −1.422. One recording cannot reveal the ensemble mean: B is not ergodic.
Let’s check what “near 0” means. At the 200 values are independent draws of variance 1. By “A longer average: less noise, more delay” in Simple smoothing filters (18.2), averaging 200 of them leaves a spread of .
So is 1.7 spreads from 0. Along realisation 8, 256 samples leave a spread of , and 0.065 is about one spread from 0. Both are as near to 0 as so few draws allow.
In B each sample is noise plus a level, two independent draws of variance 1. Their variances add (24.1), so each sample has variance 2, and an average of 200 has spread . B’s is well within it. It differs from A’s value only by the average of the 200 levels, 0.040.
Along realisation 8, the noise averages away as it did in A, but the level stays. B’s time average is A’s, 0.065, plus the level, , which gives . A longer recording would not help: the noise part would shrink towards 0, and the time average would close in on , not on 0.
The other realisations agree. Over all 200, A’s time averages scatter with a spread of 0.065, and B’s with a spread of 0.925, close to the levels’ own spread, 0.929. Each B recording reports its own level, not the process’s mean.
It is like the average height of everyone in a town today, against one person’s height averaged over a week. The first is the town’s; the second is that person’s, however long you measure.
Stationary, and ergodic
Neither process drifts as goes on. To say that properly, I need a second average across the ensemble: the mean of the product of the values at two instants, the correlation
When both values have mean 0 and the same variance, dividing by that variance gives 24.1’s correlation coefficient of and .
A process is wide-sense stationary, WSS for short, when two things hold. Its mean is the same at every , and depends only on the lag , how far apart the instants are. Then I write it . In words, its averages do not care when you look, only how far apart.
Process A has mean 0 at every . Two different samples are independent, so their product averages to 0; a sample times itself averages to its variance, 1. So is 1 at and 0 at every other lag.
Process B has mean at every . Its product at two instants is (noise + level) times (noise + level). Averages add, so take the four products one at a time. The noise parts are independent of each other and of the level, so only noise times noise at the same instant, and level times level, survive.
So for B, is 2 at and 1 at every other lag. The level is shared by every pair of instants, however far apart.
Both depend only on the lag, so both processes are WSS. Yet only A is ergodic. B’s correlation never dies away: samples 1000 apart still share their level. Each realisation never forgets its level, so it never sees the rest of the ensemble.
So stationary does not mean ergodic. In practice there is often only one recording, and analysing it alone means assuming ergodicity: its time averages are taken to stand for the ensemble’s. From here on I assume it too.
The maths behind it · stationary time series
This is what statistics calls a time series: a stationary series, its autocorrelation function (ACF), and the AR(1) model of the next section. For a process with mean 0, that track writes for and for .
White noise forgets, coloured noise remembers
For a WSS process, the autocorrelation
says how alike, on average, two samples apart are. At lag 0 it is the mean square, which is the variance when the mean is 0. It is symmetric, , because a pair read the other way round is the same pair.
If the process is ergodic, one long realisation of samples is enough. Average the products along it:
There are products, from to , and I divide by their number. The hat means “estimated”, as in 24.1. Some texts divide by instead; for the lags on this page that changes the values by less than 0.001.
White noise has at every : no two different samples are correlated. Process A is white. Power spectral density (24.4) shows why “white” fits.
I write white noise as , with variance . Dither (11.2) used the same letter for dither, which is white noise too. So is at and 0 elsewhere, which is with the unit impulse of Impulse, step and ramp (3.1).
Coloured noise is noise whose neighbouring samples resemble each other. A simple recipe feeds white noise into a loop: the leaky integrator of “With a loop and without one” in Difference equations (6.1). In “A 9-point average and its one-multiply twin”, 18.2 calls the same loop the EMA. With white noise as its input it reads
The output depends on its own last value, so this is an autoregressive process of order 1, AR(1) for short. The EMA of 18.2 is the same loop with the input scaled by .
How big is ? The value was built from noise up to only, so it is independent of the fresh . Their variances add, and multiplying by multiplies a variance by . In a stationary process and have the same variance, so
Now the likeness. Multiply both sides of the recursion by for some , and average. The term with averages to 0, because has mean 0 and is independent of the past. What is left is : each step of lag multiplies the likeness by . With the symmetry,
The instrument takes . With white noise of variance 1 as input, the output’s variance would be . So the input is scaled by , which scales its variance to 0.19 and brings the output’s variance to 1. Then , and white and coloured noise are compared at equal power.
Divided by , the autocorrelation is 24.1’s correlation coefficient of two samples apart. Next-door samples of this noise have a correlation coefficient of 0.9: in 24.1’s picture, a cloud of pairs leaning hard along a line. Think of coin tosses against the daily temperature: yesterday’s toss says nothing about today’s, but yesterday’s temperature says a lot.
White noise forgets, coloured noise remembers
Variance-1 white noise and AR(1) noise x[n] = 0.9x[n−1] + √0.19·v[n]; autocorrelation estimated from 4000 samples.
White noise: next-door samples are unrelated. Measured R_x[1] = 0.005, R_x[5] = −0.006: near zero apart from lag 0 (1.021).
Describe this picture
Two stacked panels for variance-1 white noise and AR(1) noise , with the autocorrelation estimated from 4000 samples. The first shows the first 64 samples as stems, from −3 to 3, and names the noise it shows, “white noise” or “AR(1) noise”. The second shows from −0.2 to 1.1 against lag ℓ from 0 to 10: the estimates are stems with square heads labelled “measured”, and the formula’s values open rings labelled “formula”. The readouts are the estimates and . There is no control. White noise comes first: next-door samples are unrelated, and and , near zero apart from lag 0 (1.021). Then the samples and stems change to the same white noise through the AR(1) loop, and the formula’s rings appear. Neighbours now resemble each other: and , beside the formula’s 0.9 and .
Watch the stems at lags 1 to 10: near 0 for white noise, fading slowly for coloured noise. Look at the white noise first. Its lag-0 value, 1.021, is the measured variance of these 4000 draws, close to the nominal 1. At lags 1 to 10 the estimates are not quite 0: the largest is 0.021.
That is the scatter you should expect. Each estimate averages about 4000 products of unrelated draws, and by 18.2’s rule that leaves a spread of about .
Now the coloured noise. Its estimates sit close to the formula: 0.914 against 0.9 at lag 1, and 0.593 against 0.590 at lag 5. At lag 10 the noise is still 0.366 alike, against the formula’s . The likeness fades step by step, and it first drops below a half at lag 7, where the formula gives 0.478.
The coloured estimates stray further from their formula than the white ones do, because neighbouring products are alike too, and alike errors do not average away as fast. The periodogram (25.1) comes back to how far an estimate can stray.
The maths behind it · Toeplitz matrices
The autocorrelations of a stationary process fill a Toeplitz matrix , one whose entries are constant along every diagonal: entry is . It is symmetric and positive semidefinite. It comes back in the Yule–Walker equations of Parametric models and linear prediction (25.3) and the Wiener–Hopf equations of The Wiener filter (26.2).
Stationary for a while
Real signals drift. Speech changes from one sound to the next, and a radio channel fades as a car drives along. So a recording is treated as stationary over a window short enough that its statistics do not drift.
For speech that window is about 20 to 30 ms. That is why “Reading a chirp” in Spectrograms & the STFT (15.5) used a 25 ms window for speech.
Worked example
1. Two averages. At the 200 realisations of process A average , and realisation 8 averages 0.065 along its 256 samples. Their expected spreads are and , so both are near the mean, 0. In process B, realisation 8 averages , its own level plus A’s average, while the ensemble still gives .
2. An estimate by hand. Take = 2, 1, −1, −2, so . At lag 0 there are four products, , so . At lag 1 there are three, , so . Next-door samples here are alike.
3. AR(1) with . With input variance 1 the output’s variance would be . Scaling the input by brings it to 1, and then . The 4000 seeded samples measure 0.593.
Where you’ll meet this
A noise model in a receiver or a sensor is a random process. Thermal noise in a resistor or an amplifier is modelled as white Gaussian noise, and its power sets the noise floor. A slowly wandering sensor offset is modelled as coloured noise, often AR(1).
The next pages build on this one. Correlation (24.3) compares two signals by the same products, Power spectral density (24.4) turns into a spectrum, and The periodogram (25.1) estimates that spectrum from one recording.
AR models describe speech in Parametric models and linear prediction (25.3). The Wiener filter of The Wiener filter (26.2) and the Kalman filter of RLS and the Kalman filter (26.4) start from models of the signal and the noise as random processes.
For more, see A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, appendix A; A. Papoulis and S. U. Pillai, Probability, Random Variables and Stochastic Processes (4th ed., 2002), ch. 9–11; and J. G. Proakis and D. G. Manolakis, Digital Signal Processing, ch. 12. In NumPy, np.dot(x[l:], x[:N-l]) / (N - l) gives this page’s estimate at lag l.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Ensemble mean | across realisations, at one | |
| Time average | along one realisation | |
| Autocorrelation | WSS: depends on only; | |
| Ergodic | time averages = ensemble averages | one recording stands for the process |
| White noise | no memory | |
| AR(1) | coloured; | |
| AR(1) autocorrelation | each lag step multiplies by | |
| Estimate | from one realisation |