Skip to content

Finite word-length effects

Fixed-point numbers round a filter's coefficients and signals. Poles snap to a grid, sums overflow, and rounding inside a loop can stop it decaying.

Before this11.2 · 18.2 · 21.2 · 5 more
Chapter 21 · Lesson 3 of 4

First, the picture

On a small chip every number is stored with a fixed number of bits. Below, a filter’s two coefficients are rounded to fewer and fewer bits. Watch the rounded pole jump between the faint dots it is allowed to take, and the gain curve follow it.

Rounded coefficients put the poles on a grid

A biquad resonance at 111 Hz (f_s = 8 kHz, poles 0.980∠±5°), a₁ and a₂ rounded to B fractional bits.

B = 10 fractional bits: the grid is fine, and the pole lands 0.0004 from where it was wanted. Peak at 107.6 Hz.

fractional bits B
10
pole
0.9798∠4.98°
peak at
107.6 Hz
0.00 / 13.00 s
Describe this picture

Two panels for a biquad resonance at 111 Hz (fs=8f_s=8 kHz, poles 0.980∠±5°), with a1a_1 and a2a_2 rounded to BB fractional bits. The plane shows a small region near z=1z=1, the real part from 0.85 to 1.02 and the imaginary part from −0.02 to 0.19, on one scale both ways, with an arc of the unit circle. Every pole the rounded coefficients can make is a small faint dot. The wanted pole is a faint cross, and the rounded pole is a full-strength cross labelled with its value. The gain panel plots the gain from −20 to 20 dB against frequency from 0 to 400 Hz: the wanted response is a thin dashed curve and the rounded one solid. The readouts are the fractional bits BB, the pole and where the gain peaks.

The clip lasts 13 s, with a blank caption while the grid changes. At B=10B=10 the grid is fine: the pole lands 0.0004 from where it was wanted, at 0.9798∠4.98°, and the peak is at 107.6 Hz. Then the dots fade to a coarser grid and the pole jumps. At 5.75 s, B=7B=7: a much coarser grid near z=1z=1, but a grid point happens to sit close, 0.9803∠4.99°, with the peak at 107.9 Hz. Last the grid coarsens once more. At B=6B=6 the nearest grid values give two real poles, 0.953 and exactly 1, on the circle: the resonance is gone, the peak reads 0 Hz, and the filter is no longer stable. After the clip a slider named “Fractional bits B” runs from 4 to 10; the arrow keys move it by one bit, and Home and End jump to 4 and 10. At 10, 7 and 6 bits the caption is the clip’s; at other values it reads like “B = 8: pole 0.9803∠4.99°, peak at 107.9 Hz.” The choice is kept in the link, as grid.B.

Rounded coefficients put the poles on a grid

Most filters on this site run in floating point, where rounding is usually too small to see. On a small DSP chip or a microcontroller that changes. To save power and silicon, many of them store every number as a 16-bit integer and agree on where the binary point sits. That is fixed point.

The most common agreement reads the word as a fraction. The 16 bits hold a whole number from −32 768 to 32 767, and the value it stands for is that number divided by 215=32 7682^{15}=32\,768:

value=integer215,−32 768≤integer≤32 767.\begin{aligned} \text{value}&=\frac{\text{integer}}{2^{15}},\\ -32\,768&\le\text{integer}\le32\,767. \end{aligned}

So the values run from −1-1 to 1−2−15=0.999971-2^{-15}=0.99997, in steps of 2−15=3.05×10−52^{-15}=3.05\times10^{-5}. The first bit carries the sign, and the other 15 sit behind the binary point. I call those the fractional bits. The format’s name, Q1.15, counts them: one bit for the sign, 15 fractional bits.

Storing a value is the rounding of Quantization & noise (11.1), with a level at 0, the mid-tread placement. The step is Δ=2−15\Delta=2^{-15}, and 11.1’s Δ=2/2B\Delta=2/2^B gives the same with B=16B=16 bits in all. On this page BB counts only the fractional bits, so Δ=2−B\Delta=2^{-B}.

Floating point works differently. A 32-bit float keeps 24 significant bits and moves the binary point to suit each number, so its step is about 1.2×10−71.2\times10^{-7} of the number itself. Near 0.001 its step is 1.2×10−101.2\times10^{-10}, while Q1.15’s step there is still 3.05×10−53.05\times10^{-5}, 3 % of the value. Fixed point has one step everywhere, and that is where this page’s three problems start.

The first problem is the coefficients. A biquad’s bottom polynomial, from Transfer functions, poles & zeros (16.3), is A(z)=1+a1z−1+a2z−2A(z)=1+a_1z^{-1}+a_2z^{-2}. For poles at re±jθre^{\pm j\theta} it has a1=−2rcos⁡θa_1=-2r\cos\theta and a2=r2a_2=r^2. A stable biquad has a1a_1 between −2-2 and 22, inside the triangle of Stability and causality (16.4), so a1a_1 needs one more integer bit: Q2.14, from −2-2 to 2−2−142-2^{-14}.

Round a1a_1 and a2a_2 to BB fractional bits and they can only take values on a grid, 2−B2^{-B} apart. Each grid point (a1,a2)(a_1,a_2) gives one pair of poles:

r=a2,cos⁡θ=−a12r.r=\sqrt{a_2},\qquad\cos\theta=-\frac{a_1}{2r}.

So the poles can only sit where some grid point puts them. Second-order sections (21.2) saw rounding move poles; here we see where they are allowed to go. Think of a city map with streets only every kilometre: an address has to move to the nearest corner.

The picture at the top of the page rounds one low-frequency resonance at fs=8f_s=8 kHz. Its poles are 0.98 e±j5°0.98\,e^{\pm j5°}, and 5° of a full turn is 5/360⋅8000=111.15/360\cdot8000=111.1 Hz. So a1=−2⋅0.98cos⁡5°=−1.952542a_1=-2\cdot0.98\cos5°=-1.952542 and a2=0.9604a_2=0.9604.

The filter is H(z)=b0/A(z)H(z)=b_0/A(z) with b0=1+a1+a2=0.0078584b_0=1+a_1+a_2=0.0078584, which makes the gain 1 at 0 Hz. It rounds only a1a_1 and a2a_2 and keeps this b0b_0. The wanted gain peaks at 108.1 Hz, a little below the poles’ angle, as 16.3 found.

In that picture, ten and seven bits keep the resonance, and six bits destroy it.

Why did six bits destroy it? For a mirror pair, the height above the real axis is y=rsin⁡θy=r\sin\theta, and y2=r2−r2cos⁡2θ=a2−a12/4y^2=r^2-r^2\cos^2\theta=a_2-a_1^2/4. Our pole has y=0.0854y=0.0854, so y2=0.0073y^2=0.0073. One step of a2a_2 at six bits is 2−6=0.01562^{-6}=0.0156, more than twice as big.

The nearest grid values give a2−a12/4=−0.00055a_2-a_1^2/4=-0.00055. That is below 16.4’s parabola, so both poles are real. They also have 1+a1+a2=1−1.953125+0.953125=01+a_1+a_2=1-1.953125+0.953125=0, which by 16.4 puts one pole exactly at z=1z=1.

More bits did not always land closer. Ten bits left the pole 0.00038 from the wanted one, and seven bits only 0.00035. Each extra bit gives a finer grid, but where the nearest grid point falls is a matter of chance.

After the clip, set the number of fractional bits yourself with the slider, from 4 to 10. Eight and nine bits land on the same grid point as seven.

Now try 5. The poles jump to 0.9843 e±j10.18°0.9843\,e^{\pm j10.18°} and the peak to 225.4 Hz, about twice the wanted frequency. At 4 bits the poles are real again, 1 and 0.9375.

Why is the grid so sparse near z=1z=1? It is even in a1a_1 and a2a_2, not in the plane. Write a pole as x+jyx+jy; then a1=−2xa_1=-2x and a2=x2+y2a_2=x^2+y^2. Near the real axis yy is small, and y2y^2 smaller still, so one step of a2a_2 moves the pole a long way up.

In numbers, a patch of the plane holds grid poles in proportion to its height yy. They crowd near ±90°, where yy is largest, and thin out near z=±1z=\pm1. At six bits, the poles of radius 0.9 to 1 within 15° of z=1z=1 number 48. The same-sized wedge around 90° holds 382, eight times as many.

That is why low-frequency resonances suffer. Bass equalisers and slow sensors sampled fast put their poles close to z=1z=1, where the grid is thin. The fixes are more bits, or a different structure.

The coupled form stores rcos⁡θr\cos\theta and rsin⁡θr\sin\theta themselves, so its grid is even across the plane. Floating point avoids most of this, at more cost per operation.

The maths behind it · eigenvalue perturbation

Rounding the coefficients is a small perturbation of the 2×2 state matrix of 16.3, whose eigenvalues are the poles; how far they move depends on the matrix’s structure. Near a double eigenvalue, a small change can move them by about its square root: that is the square root in y=a2−a12/4y=\sqrt{a_2-a_1^2/4}. The coupled form’s matrix is a scaled rotation, an orthogonal matrix times rr. Its poles are spread evenly, and they move no further than its entries change.

Overflow wraps; saturation clips

The second problem is size. Q1.15 has no room above 0.99997, so what happens when a sum lands beyond it? Most chips use two’s complement integers. The 16-bit integer counts up to 32 767, and one more gives −32 768, the bottom of the range.

The value has gone all the way round, like a car’s odometer rolling from 99 999 to 00 000. This is wrap-around, and the stored value is off by 2, the width of the whole range. The other choice is saturation: the arithmetic holds the value at the limit, like an odometer that sticks at 99 999. Its error is only the excess above the limit.

Many DSP chips offer saturating additions in hardware for this reason. The instrument below stores the same sine both ways. Watch the tops of the sine once the amplitude passes the limit.

Overflow wraps; saturation clips

A sine stored as Q1.15 (−1 to 0.99997), its amplitude raised past the limit.

Amplitude 0.9: everything fits; both versions store the sine to within 0.00002.

amplitude
0.9
worst error, wrap
0.000
worst error, saturate
0.000
0.00 / 12.00 s
Describe this picture

Two stacked panels, “wrap-around” and “saturation”, for a sine stored as Q1.15 (−1 to 0.99997) with its amplitude raised past the limit. They share the sample axis, nn from 0 to 63. Each has dotted levels at ±1 labelled “limit” and the wanted sine as a thin dashed curve. The stored values are stems, with dot heads in the wrap-around panel and square heads in the saturation panel. The readouts are the amplitude and the worst error of each version. There is no control.

The clip lasts 12 s, and the stems follow the amplitude as it changes. At amplitude 0.9 everything fits, and both versions store the sine to within 0.00002. The amplitude rises past the limit. At 6 s, amplitude 1.2: wrap-around throws each top to the other side, to about −0.8, an error of 2.000, and saturation flattens the tops, an error of 0.200. At the end, amplitude 1.5: saturation’s error grows to 0.500, and wrap-around’s stays 2.000, the worst possible. The end caption advises scaling the signal down so the largest value fits, or saturating.

Look at a top in the wrap-around panel. At amplitude 1.2 the top sample is 39 322 steps, which does not fit. It comes back as 39 322−65 536=−26 21439\,322-65\,536=-26\,214 steps, which is −0.8, and its neighbours land at −0.82 and −0.89. At amplitude 1.5 the top comes back as −0.5.

So the further a value passes the limit, the further from −1 it lands, but every wrapped sample is wrong by 2. Saturation’s error grows with the excess, 0.2 and then 0.5, and never jumps.

The cure for both is scaling: multiply the signal by a constant below 1 before it can overflow, so the largest value fits. Halving the 1.5 sine gives 0.75, which fits with room to spare. The price is one bit: the step stays while the signal halves, so the signal-to-noise ratio falls by 6.02 dB (11.1).

In a filter, the values to watch are the stored ones inside the structure, not only the output. Filter structures (21.1) showed that direct form II’s inside signal w[n]w[n] can be much larger than its output. That is the value the scaling has to fit.

A filter that never goes quiet

The third problem hides inside a loop. Take the first-order loop of Difference equations (6.1), with no input: y[n]=a y[n−1]y[n]=a\,y[n-1]. Exactly, from y[0]=20y[0]=20, it decays as 20an20a^n. In fixed point the product a y[n−1]a\,y[n-1] has more bits than the word, so it is rounded before it is stored:

y[n]=Q{a y[n−1]}.y[n]=Q\{a\,y[n-1]\}.

Here I measure yy in steps, so QQ rounds to whole numbers. Engineers call one step one LSB, the value of the least significant bit. A value exactly halfway, like 4.5, rounds away from zero, to 5.

Each rounding adds an error of up to half a step. While ∣a y∣\lvert a\,y\rvert is large, that hardly matters. When it is small, the rounding can give back the same value it started from.

When does that happen for 0<∣a∣<10 < \lvert a\rvert < 1? Rounding gives back ∣y∣\lvert y\rvert when ∣a y∣\lvert a\,y\rvert is within half a step of ∣y∣\lvert y\rvert, that is when (1−∣a∣)∣y∣≤0.5(1-\lvert a\rvert)\lvert y\rvert\le0.5:

∣y∣≤0.51−∣a∣.\lvert y\rvert\le\frac{0.5}{1-\lvert a\rvert}.

That band around 0 is the dead band. For a=0.9a=0.9 it reaches 0.5/0.1=50.5/0.1=5 steps. Once the output is inside it, the rounded loop stops shrinking. An output that keeps repeating with no input, where the exact filter would decay, is a limit cycle.

Think of a ball that should come to rest but keeps bouncing on a step it cannot get over. The instrument below runs the loop for a=0.9a=0.9 and then a=−0.9a=-0.9.

A filter that never goes quiet

y[n] = Q{a·y[n−1]} from y[0] = 20, rounded to whole steps, no input.

Start at 20 steps, no input, and multiply by a each sample, rounding to whole steps.

a
not yet
rounded y[60]
not yet
exact y[60]
not yet
Loop gain
0.00 / 14.00 s
Describe this picture

One panel for y[n]=Q{a y[n−1]}y[n]=Q\{a\,y[n-1]\} from y[0]=20y[0]=20, rounded to whole steps, with no input: y[n]y[n] in steps, from −22 to 22, against the sample nn from 0 to 60. The exact values 20an20a^n are small open circles, and the rounded output is stems with dot heads. A hatched band from −5 to 5 is labelled “dead band”. The readouts are aa, the rounded y[60]y[60] and the exact y[60]y[60].

The clip lasts 14 s. First the axes and the band appear: start at 20 steps, no input, and multiply by aa each sample, rounding to whole steps. Then both sequences build from n=0n=0 to 60. At 6 s, a=0.9a=0.9: 20, 18, 16, 14, … then 5, 5, 5 for ever, because 0.9×5=4.50.9\times5=4.5 rounds back to 5, while the exact output has decayed to 0.0359. Then the picture cross-fades to a=−0.9a=-0.9, and the sequences build again. At the end the output flips between +5 and −5 for ever, a limit cycle: inside the dead band, ∣y∣≤0.5/(1−∣a∣)=5\lvert y\rvert\le0.5/(1-\lvert a\rvert)=5, rounding undoes the decay. After the clip two buttons in a group named “Loop gain” choose a=0.9a=0.9 or a=−0.9a=-0.9, each with its caption from the clip; the choice is kept in the link, as cycle.a.

Notice the exact circles: they cross into the band at n=14n=14 and keep shrinking towards 0. The rounded stems stop at its edge, 5, which is 0.5/(1−0.9)0.5/(1-0.9). With a=−0.9a=-0.9 the sign flips every sample, so the stuck value becomes a flip between +5 and −5.

After the clip, switch between the two loop gains with their buttons.

Round-off noise

With a busy input x[n]x[n], the loop is y[n]=Q{a y[n−1]}+x[n]y[n]=Q\{a\,y[n-1]\}+x[n], and the rounding errors behave like 11.1’s noise: each adds an error e[n]e[n] of mean square Δ2/12\Delta^2/12. So y[n]=a y[n−1]+e[n]+x[n]y[n]=a\,y[n-1]+e[n]+x[n], and e[n]e[n] enters as if it were a second input.

The loop filters that noise like everything else. Write hloop[n]h_\text{loop}[n] for the impulse response from the rounding point to the output; here it is ana^n. By Simple smoothing filters (18.2), noise power is multiplied by ∑nhloop[n]2\sum_nh_\text{loop}[n]^2:

noise power=Δ212∑n≥0a2n=Δ212⋅11−a2.\begin{aligned} \text{noise power}&=\frac{\Delta^2}{12}\sum_{n\ge0}a^{2n}\\ &=\frac{\Delta^2}{12}\cdot\frac{1}{1-a^2}. \end{aligned}

For a=0.9a=0.9 that factor is 1/(1−0.81)=5.261/(1-0.81)=5.26. Its square root, 2.29, is the noise gain of 18.2. Closer to the circle it grows fast: for a=0.99a=0.99 the factor is 50.25. I checked a=0.9a=0.9 by simulation with a busy input: 0.441 steps², against the predicted 5.26/12=0.4395.26/12=0.439.

In direct form II, the feedback products are rounded where w[n]w[n] is formed. So their noise passes through 1/A(z)1/A(z), the same part that made w[n]w[n] large in 21.1, and poles near the circle amplify both. The remedies are more bits, a different structure, or error feedback, which feeds each rounding error back into the loop to cancel part of it.

When the input goes quiet, the noise model breaks down. The error stops looking random and follows the signal, as it did for 11.1’s quiet sine, and that is exactly where limit cycles appear.

The maths behind it · variance propagation

Modelling round-off as uniform noise of variance Δ2/12\Delta^2/12 (11.1) and passing it through a filter is variance propagation. The output variance is the input variance times ∑nh[n]2\sum_nh[n]^2, the square of 18.2’s noise gain. The model fails exactly where limit cycles appear: there the error is no longer random.

Worked example

1. Q1.15 by hand. To store 0.3, multiply by 32 768 to get 9830.4, and round to 9830. That stands for 9830/32 768=0.2999889830/32\,768=0.299988, an error of 1.22×10−51.22\times10^{-5}, below half a step, 1.53×10−51.53\times10^{-5}. The largest value is 32 767/32 768=0.99996932\,767/32\,768=0.999969.

The whole span, 2, holds 2/2−15=65 5362/2^{-15}=65\,536 steps. As a level that is 20log⁡1065 536=96.320\log_{10}65\,536=96.3 dB, which is 6.02×166.02\times16: 11.1’s 6 dB per bit.

2. The resonance at five bits. With 25=322^5=32: a1⋅32=−62.48a_1\cdot32=-62.48 rounds to −62-62, so a1a_1 becomes −1.9375-1.9375. And a2⋅32=30.73a_2\cdot32=30.73 rounds to 31, so a2a_2 becomes 0.968750.96875.

Then r=0.96875=0.9843r=\sqrt{0.96875}=0.9843 and cos⁡θ=1.9375/(2⋅0.9843)=0.9843\cos\theta=1.9375/(2\cdot0.9843)=0.9843, so θ=10.18°\theta=10.18°. That angle is 10.18/360⋅8000=226.310.18/360\cdot8000=226.3 Hz, and the gain peaks a little lower, at 225.4 Hz. That is 2.08 times the wanted 108.1 Hz.

3. Overflow at amplitude 1.2. The top sample, 1.2, is 1.2⋅32 768=39 321.61.2\cdot32\,768=39\,321.6, rounded to 39 322 steps. Two’s complement wraps it to 39 322−65 536=−26 21439\,322-65\,536=-26\,214 steps, which is −0.79999-0.79999. Saturation stores 0.999970.99997 instead, an error of 0.200030.20003.

4. Dead bands. For a=0.5a=0.5 the dead band is 0.5/(1−0.5)=10.5/(1-0.5)=1 step. From 20 the rounded loop gives 20, 10, 5, 3, 2, 1, 1, and so on: 2.5 rounds to 3, and 0.5 rounds back to 1. For a=0.9a=0.9 the dead band is 5 steps, and the loop’s round-off noise power is 1/(1−0.81)=5.26321/(1-0.81)=5.2632 times Δ2/12\Delta^2/12.

Where you’ll meet this

Fixed point is common wherever power and cost matter. Audio codecs and hearing aids run their filters in Q15 and Q31, the names ARM’s library uses for Q1.15 and its 32-bit cousin Q1.31. Motor-control microcontrollers run their control loops the same way, and FPGA filters choose a custom word length for every signal.

ARM’s CMSIS-DSP library is a good example. Routines such as arm_biquad_cascade_df1_q15 work in Q15 and saturate by design. Their documentation also says when the input must be scaled down to avoid overflow. A real-time system then has to fit all this into its time budget, which is Real-time processing (21.4).

Rounding noise can also be pushed out of the band you care about, as converters do in Oversampling and noise shaping (11.3). In floating point, as in SciPy on a laptop, most of this page fades away. A high-order direct form can still move its poles a long way, though (21.2).

Reference card

QuantityFormulaNotes
Q1.15−1-1 to 1−2−151-2^{-15}, step 2−152^{-15}16-bit word: 1 sign bit, 15 fractional bits
Q2.14−2-2 to 2−2−142-2^{-14}a biquad’s a1a_1 lies in (−2,2)(-2,2)
StepΔ=2−B\Delta=2^{-B}BB counts fractional bits here
Biquad polesr=a2r=\sqrt{a_2}, cos⁡θ=−a1/(2r)\cos\theta=-a_1/(2r)from rounded a1a_1, a2a_2
Pole grideven in a1a_1, a2a_2sparse near z=±1z=\pm1
Wrap-arounderror 2, the whole rangetwo’s complement
Saturationerror = the excess above the limitprefer it, or scale down
Scalingmultiply so the largest stored value fitshalving costs 6.02 dB
Round-off in a loopΔ212∑nhloop[n]2\dfrac{\Delta^2}{12}\sum_nh_\text{loop}[n]^2factor 5.26 for a=0.9a=0.9
Dead band∣y∣≤0.5/(1−∣a∣)\lvert y\rvert\le0.5/(1-\lvert a\rvert) stepslimit cycles

End of lesson 21.3

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look