Skip to content

Random variables for signals

Noise has no formula, but its values have a shape: a density, summed up by a mean and a variance; pairs add a correlation.

Before thisDither (11.2), Simple smoothing filters (18.2)

3 more before it

Kinds of signals (1.2), How big is a signal (1.3), Quantization & noise (11.1)

Before this11.2 · 18.2 · 3 more
Chapter 24 · Lesson 1 of 4

First, the picture

Noise has no formula, but draw many of its values, sort them into bins, and they pile up in the same shape every time. Watch the bars settle onto the curve as the draws grow from 10 to 10 000.

Draws pile up into a density

Gaussian draws from the site's generator (seed 241), bins 0.5 wide, bars scaled to area 1.

Ten draws: a ragged histogram. Their average is −0.005 and their spread (variance) 1.255.

draws N
10
mean
−0.005
variance
1.255
Density
0.00 / 16.00 s
Describe this picture

One panel of Gaussian draws from the site’s generator (seed 241), sorted into bins 0.5 wide, with the bars scaled to area 1: density from 0 to 0.7 against value from −4.5 to 4.5. The bars are outlined, the density is a thin solid curve labelled “density”, and the newest draws drop onto their bars as small dots. The readouts are the number of draws NN, their mean and their variance. The clip has no control. It opens on ten draws, a ragged histogram: mean −0.005 and variance 1.255. Draws rain in to 100, where the bars start to follow the curve (−0.051 and 0.898), then to 1000 (0.063 and 0.993), then to 10 000. There dotted lines appear at −1 and 1, labelled “68 %”, and the bars sit on the density: mean −0.002, variance 1.027, and 67.57 % of the draws within ±1. After the clip two buttons in a group named “Density”, “Gaussian” and “uniform”, switch the draws. “uniform” draws the flat density at 0.2887 between ±1.73, and the caption notes that the two end bins stick out past ±1.73.

Draws pile up into a density

In “Predictable, or not” of Kinds of signals (1.2), a steady tone sat beside noise. The tone has a formula and the noise has none. Yet noise is not lawless: draw many of its values, and they pile up in a shape that is the same every time. This page is about that shape.

A random variable is a number you get fresh each time you look. Each look is a draw. I write the random variable as xx, and its draws, numbered in the order they came, as x[0]x[0], x[1]x[1] and so on. A noise trace on this site is such a row of draws.

When the draws of a signal depend on each other, the signal is a random process, the subject of Random processes (24.2). Here every draw is made afresh, on its own. This page assumes no statistics: I define each term before I use it, and leave the proofs to statistics.

The probability of an event is the share of draws in which it happens, in the long run. I write it Pr⁡{⋅}\Pr\{\cdot\}. If xx is the roll of a fair die, Pr⁡{x=6}=1/6\Pr\{x=6\}=1/6: one roll in six, over many rolls.

To see the shape of many draws, sort them into bins of width 0.5 and count each bin. Then divide each count by N×0.5N\times0.5, where NN is the number of draws. A bar’s area is its height times 0.5, which is its count divided by NN: the share of the draws that landed in it. All the bars together have area 1.

As NN grows, the bars settle on a curve, the density px(α)p_x(\alpha). I use α\alpha for a value on the axis, to keep it apart from the random variable xx. The probability that a draw lands between α1\alpha_1 and α2\alpha_2 is the area under the density there:

Pr⁡{α1<x<α2}=∫α1α2px(α) dα.\Pr\{\alpha_1<x<\alpha_2\}=\int_{\alpha_1}^{\alpha_2}p_x(\alpha)\,d\alpha.

Every draw lands somewhere, so the total area is 1. A density is never negative, but it can be larger than 1: it is a probability per unit of width, not a probability.

The density met most often in signal work is the Gaussian, a bell. With mean 0 and variance 1 (both defined just below) it is

px(α)=12π e−α2/2,p_x(\alpha)=\frac1{\sqrt{2\pi}}\,e^{-\alpha^2/2},

which peaks at α=0\alpha=0 at a height of 1/2π=0.39891/\sqrt{2\pi}=0.3989. The site’s generator, gaussian in lib/random.ts, makes such draws from pairs of seeded uniform numbers, so every visit sees the same draws.

The picture at the top of the page has two readouts that follow the draws. The mean is their plain average, the xˉ\bar x of “Four everyday sizes” in How big is a signal (1.3). The variance is their average squared distance from that mean, a measure of how widely they spread.

Watch the variance readout: 1.255, 0.898, 0.993, 1.027. Ten draws can be far off, but ten thousand land within 3 % of the density’s 1. The tallest bar tells the same story. With ten draws it is 0.600, three draws in one bin; at 10 000 it is 0.380.

Even then the tallest bar stays a little below the curve’s peak, 0.3989. A bar shows the curve’s average over its width, and next to 0 that average is 0.3829.

Mean and variance

The expected value E{⋅}\mathbb{E}\{\cdot\} of anything that depends on xx is its average over the density: each value weighted by how likely it is. The mean μx\mu_x and the variance σx2\sigma_x^2 are two expected values, with the integrals running over every value:

μx=E{x}=∫α px(α) dα,σx2=E{(x−μx)2}=∫(α−μx)2 px(α) dα.\begin{aligned} \mu_x&=\mathbb{E}\{x\}=\int\alpha\,p_x(\alpha)\,d\alpha,\\ \sigma_x^2&=\mathbb{E}\{(x-\mu_x)^2\}\\ &=\int(\alpha-\mu_x)^2\,p_x(\alpha)\,d\alpha. \end{aligned}

The variance is the average power of 1.3, measured about the mean. For noise with mean 0 it is simply the average power E{x2}\mathbb{E}\{x^2\}. Its square root σx\sigma_x is the standard deviation, and for mean-0 noise that is the RMS. So the 0.2 kg RMS noise of “A longer average: less noise, more delay” in Simple smoothing filters (18.2) has a standard deviation of 0.2 kg.

The density is never known from a finite set of draws, only estimated. A hat means “estimated from draws”:

μ^x=1N∑n=0N−1x[n],σ^x2=1N∑n=0N−1(x[n]−μ^x)2.\begin{aligned} \hat\mu_x&=\frac1N\sum_{n=0}^{N-1}x[n],\\ \hat\sigma_x^2&=\frac1N\sum_{n=0}^{N-1}\big(x[n]-\hat\mu_x\big)^2. \end{aligned}

The first is 1.3’s mean xˉ\bar x of the draws. The estimates change with every new set of draws, while μx\mu_x and σx2\sigma_x^2 do not, and the first picture’s readouts are these two estimates. NumPy’s mean and var compute them; var divides by NN unless told otherwise.

A Gaussian can have any mean and any standard deviation. It is the bell above, moved to μx\mu_x and widened by σx\sigma_x:

px(α)=12π σx e−(α−μx)2/(2σx2).p_x(\alpha)=\frac{1}{\sqrt{2\pi}\,\sigma_x}\,e^{-(\alpha-\mu_x)^2/(2\sigma_x^2)}.

Its areas are worth knowing by heart. Within 1, 2 and 3 standard deviations of the mean lie 68.27 %, 95.45 % and 99.73 % of all draws. The 10 000 draws give 67.57 %, 95.44 % and 99.68 %.

A Gaussian has no edge, though. About 3 draws in 1000 land further out than 3 standard deviations, so noise has no fixed peak value.

The uniform density

The uniform density is flat. Every value in a width Δ\Delta is equally likely, and no value outside it ever comes. Its height is 1/Δ1/\Delta, so its area is 1, and its mean is the middle of the width.

“The error of a slow ramp” in Quantization & noise (11.1) found its variance: a value spread evenly over a width Δ\Delta has mean square Δ2/12\Delta^2/12 about its middle. That is the rounding error of 11.1’s noise model, a uniform draw over one step.

A width of 23=3.46412\sqrt3=3.4641 gives variance 1 and height 0.28870.2887. The generator’s uniform numbers uu run from 0 to 1, so 3 (2u−1)\sqrt3\,(2u-1) runs from −1.732-1.732 to 1.7321.732.

After the clip, press “uniform” under the first picture. It draws 10 000 values as 3 (2u−1)\sqrt3\,(2u-1), with uu from the same seed, and draws the flat density at 0.2887 between ±1.73. The mean and variance read −0.022 and 0.993: the width 232\sqrt3 was chosen for variance 1.

Look at the ends, though. The inner bars read 0.272 to 0.301, but the two end bars only 0.137 and 0.135. Their bins run from 1.5 to 2 and from −2 to −1.5, and the density stops at ±1.732, so less than half of each bin can catch draws.

Sums of draws

Add two draws that are made separately, with nothing linking them, and their variances add. This is “the powers add” of 18.2, now said of spread about the mean. The next section names this kind of pair: independent.

11.2’s triangular dither is the sum of two separate uniform values, each of width Δ\Delta. So its variance is 2×Δ2/12=Δ2/62\times\Delta^2/12=\Delta^2/6, twice that of rectangular dither. Add the rounding’s own Δ2/12\Delta^2/12, and you have the Δ2/4\Delta^2/4 that “Rectangular or triangular dither” in Dither (11.2) found.

The sum of two uniform values has a triangular density, which is where the name comes from. Add more of them and the shape rounds off further. Add many independent values, all with one density of finite variance, and the sum looks Gaussian, whatever that density was.

This is the central limit theorem, and statistics proves it. It is why thermal noise is Gaussian: it is the sum of countless tiny motions of electrons.

The maths behind it · sample mean and variance

Statistics writes a random variable as a capital X and its density fXf_X. This site’s DSP pages write a lowercase xx and pxp_x, because capitals are transforms here. The estimates μ^x\hat\mu_x and σ^x2\hat\sigma_x^2 are that track’s sample mean and (biased) sample variance.

A cloud that leans: correlation

Now take two values per draw, a pair (x,y)(x,y): say a sensor’s error and its neighbour’s. Pairs have a joint density, spread over a plane instead of along a line. The probability that a pair lands in a region is the volume under the joint density there.

Two random variables are independent when knowing one tells you nothing about the other. Then the joint density is the product of the two densities, and averages of products split: E{xy}=E{x} E{y}\mathbb{E}\{xy\}=\mathbb{E}\{x\}\,\mathbb{E}\{y\}.

Pairs that are not independent can hang together in many ways. The simplest measure of how they do is their correlation coefficient:

ρxy=E{(x−μx)(y−μy)}σxσy.\rho_{xy}=\frac{\mathbb{E}\{(x-\mu_x)(y-\mu_y)\}}{\sigma_x\sigma_y}.

It always lies between −1 and 1. It is 1 when the pairs lie on a rising straight line, −1 on a falling one, and 0 when there is no straight-line trend. Its estimate ρ^xy\hat\rho_{xy} puts estimates from the NN pairs in place of each expected value and each σ\sigma, and NumPy’s corrcoef computes it.

If xx and yy are independent, the top of that fraction splits into E{x−μx} E{y−μy}\mathbb{E}\{x-\mu_x\}\,\mathbb{E}\{y-\mu_y\}, which is 0×00\times0. So independent values are uncorrelated: ρxy=0\rho_{xy}=0. A figure after the instrument shows that the reverse fails.

To make pairs with a chosen correlation ρ\rho, take two independent Gaussian draws xx and zz, each with mean 0 and variance 1, and set

y=ρx+1−ρ2 z.y=\rho x+\sqrt{1-\rho^2}\,z.

The two parts of yy are independent, so their variances add: ρ2+(1−ρ2)=1\rho^2+(1-\rho^2)=1. And E{xy}=ρ E{x2}+1−ρ2 E{xz}=ρ\mathbb{E}\{xy\}=\rho\,\mathbb{E}\{x^2\}+\sqrt{1-\rho^2}\,\mathbb{E}\{xz\}=\rho, since E{xz}=0\mathbb{E}\{xz\}=0. With both variances 1, that makes ρxy=ρ\rho_{xy}=\rho.

A cloud that leans: correlation

500 pairs (x, y) with y = ρx + √(1 − ρ²) z; x and z are independent Gaussian draws (seed 241).

ρ = 0: a round cloud. Knowing x tells you nothing about y on average; the 500 pairs measure −0.080.

ρ, set
0.00
ρ, measured
−0.080
0.00 / 13.00 s
Describe this picture

One square panel of 500 pairs (x,y)(x,y) with y=ρx+1−ρ2 zy=\rho x+\sqrt{1-\rho^2}\,z, where xx and zz are independent Gaussian draws (seed 241): xx and yy each from −4 to 4. The pairs are small dots, and a dashed line labelled “y = ρx” marks the straight-line trend. The readouts are ρ as set, to two decimals, and ρ as measured, to three. The clip opens at ρ=0\rho=0, a round cloud that measures −0.080. Each point then slides vertically toward the line: at ρ=0.6\rho=0.6 the cloud leans along y=0.6xy=0.6x and measures 0.543. At ρ=0.95\rho=0.95 it is a thin cloud close to the line, measuring 0.944. After the clip a slider labelled “ρ”, named “Correlation ρ”, sets ρ from −0.99 to 0.99 in steps of 0.01. At other values the caption reads like “ρ = −0.50: measured −0.540.”

Watch the cloud lean toward the dashed line as ρ grows. That line marks the straight-line trend: for these pairs, yy is ρx\rho x on average.

Why does ρ=0\rho=0 measure −0.080? The 500 draws of xx and the 500 of zz are not perfectly unrelated over so few pairs: their own measured correlation is −0.080. Their measured means are 0.057 and 0.070, and their variances 0.960 and 1.025, not exactly 0 and 1.

After the clip, use the slider to set ρ yourself. Try 0.3 and 0.8: the pairs measure 0.219 and 0.772. Over the whole slider, 500 pairs miss the set value by at most 0.084, at ρ=0.16\rho=0.16. Near ±1 the miss shrinks: 0.006 at 0.95.

Uncorrelated, but not independent

Correlation measures only a straight-line relation. Here is a pair that is tied together completely and still has a correlation of 0.

Draw θ\theta uniformly from 0 to 2π2\pi, with the generator’s uniform numbers (seed 2410), and set x=cos⁡θx=\cos\theta and y=sin⁡θy=\sin\theta. Every point lies on a ring of radius 1, so once xx is known, yy is ±1−x2\pm\sqrt{1-x^2}: only its sign is left to chance.

−101−101xy
Fig. Points on a ring: y is fixed up to its sign once x is known, so the two are far from independent, yet their correlation is −0.039, near 0 (exactly 0 in the long run). Correlation sees only straight-line relations.

The ring is symmetric: a point above the xx-axis is as likely as its mirror image below it. So in the long run the products cancel: E{cos⁡θsin⁡θ}=12E{sin⁡2θ}=0\mathbb{E}\{\cos\theta\sin\theta\}=\tfrac12\mathbb{E}\{\sin2\theta\}=0. Both means are 0 too, so ρxy=0\rho_{xy}=0 exactly. The 500 drawn points measure −0.039.

So uncorrelated does not mean independent. For the instrument’s pairs the two do coincide: at ρ=0\rho=0 the recipe gives y=zy=z, which is independent of xx. Statistics shows that this holds for every pair of jointly Gaussian values.

The maths behind it · angles between vectors

Centred draws are vectors. ρ^xy\hat\rho_{xy} is the cosine of the angle between the vector of xx values and the vector of yy values, each minus its mean. It is ±1 when they are parallel, and 0 when they are orthogonal.

Worked example

1. Gaussian areas. Take the kitchen scale of 18.2: noise of standard deviation 0.2 kg on readings at 50 per second. Within ±0.2, ±0.4 and ±0.6 kg of the true weight lie 68.27 %, 95.45 % and 99.73 % of the readings.

So 0.27 % land further out than 0.6 kg, one reading in 370. At 50 readings a second, that happens about once every 7.4 s.

2. Uniform. In 11.1’s 3-bit quantizer the step is Δ=0.25\Delta=0.25, so the rounding error is uniform from −0.125 to 0.125. Its density is 1/0.25=41/0.25=4, its variance 0.252/12=0.0052080.25^2/12=0.005208, and its standard deviation 0.25/12=0.07220.25/\sqrt{12}=0.0722.

A width of 23=3.46412\sqrt3=3.4641 instead gives variance 12/12=112/12=1 and density 1/3.4641=0.28871/3.4641=0.2887.

3. Estimates from the seeded draws. With seed 241 the Gaussian draws estimate the variance as 1.255 from 10 draws and 1.027 from 10 000. The means are −0.005 and −0.002.

4. Making a correlation. For ρ=0.6\rho=0.6, 1−0.36=0.8\sqrt{1-0.36}=0.8, so y=0.6x+0.8zy=0.6x+0.8z. The 500 seeded pairs measure 0.543.

Where you’ll meet this

Every receiver and measuring instrument has a noise budget. Its noise sources are usually independent, so their variances add, as here. Two sensors whose errors have correlation ρxy\rho_{xy} and equal variance σx2\sigma_x^2 average to a variance of σx2(1+ρxy)/2\sigma_x^2(1+\rho_{xy})/2: half when the errors are uncorrelated, no gain at all when ρxy=1\rho_{xy}=1.

Quantization noise is modelled as uniform, as in 11.1, and the thermal noise of sensors and radios as Gaussian. How often Gaussian noise crosses a threshold, the area beyond a few standard deviations, is the question Matched filters and detection (26.1) answers.

Next, Random processes (24.2) lets draws depend on their neighbours, and Correlation (24.3) measures the correlation of a signal with a shifted copy of itself or of another signal.

NumPy computes this page’s estimates with histogram (its density=True scales the bars as here), mean, var and corrcoef. For more, see appendix A, “Random signals”, of A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing (3rd ed., 2010). See also A. Papoulis and S. U. Pillai, Probability, Random Variables and Stochastic Processes (4th ed., 2002), chapters 4 to 6.

Reference card

QuantityFormulaNotes
Probabilityarea under px(α)p_x(\alpha)total area 1
Mean, varianceμx=E{x}\mu_x=\mathbb{E}\{x\}, σx2=E{(x−μx)2}\sigma_x^2=\mathbb{E}\{(x-\mu_x)^2\}estimates μ^x\hat\mu_x, σ^x2\hat\sigma_x^2 from draws
Gaussian12π σxe−(α−μx)2/(2σx2)\frac{1}{\sqrt{2\pi}\,\sigma_x}e^{-(\alpha-\mu_x)^2/(2\sigma_x^2)}68.27 %, 95.45 %, 99.73 %
Uniform, width Δdensity 1/Δ1/\Delta, variance Δ2/12\Delta^2/12rounding error
Sum of independent drawsvariances addmany look Gaussian
Correlationρxy=E{(x−μx)(y−μy)}/(σxσy)\rho_{xy}=\mathbb{E}\{(x-\mu_x)(y-\mu_y)\}/(\sigma_x\sigma_y)straight-line relation only
Independentjoint density = product⇒ uncorrelated, not ⇐

End of lesson 24.1

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look