On a small chip every number is stored with a fixed number of bits. Below, a filter’s two coefficients are rounded to fewer and fewer bits. Watch the rounded pole jump between the faint dots it is allowed to take, and the gain curve follow it.
Rounded coefficients put the poles on a grid
A biquad resonance at 111 Hz (f_s = 8 kHz, poles 0.980∠±5°), a₁ and a₂ rounded to B fractional bits.
B = 10 fractional bits: the grid is fine, and the pole lands 0.0004 from where it was wanted. Peak at 107.6 Hz.
Describe this picture
Two panels for a biquad resonance at 111 Hz ( kHz, poles 0.980∠±5°), with and rounded to fractional bits. The plane shows a small region near , the real part from 0.85 to 1.02 and the imaginary part from −0.02 to 0.19, on one scale both ways, with an arc of the unit circle. Every pole the rounded coefficients can make is a small faint dot. The wanted pole is a faint cross, and the rounded pole is a full-strength cross labelled with its value. The gain panel plots the gain from −20 to 20 dB against frequency from 0 to 400 Hz: the wanted response is a thin dashed curve and the rounded one solid. The readouts are the fractional bits , the pole and where the gain peaks.
The clip lasts 13 s, with a blank caption while the grid changes. At the grid is fine: the pole lands 0.0004 from where it was wanted, at 0.9798∠4.98°, and the peak is at 107.6 Hz. Then the dots fade to a coarser grid and the pole jumps. At 5.75 s, : a much coarser grid near , but a grid point happens to sit close, 0.9803∠4.99°, with the peak at 107.9 Hz. Last the grid coarsens once more. At the nearest grid values give two real poles, 0.953 and exactly 1, on the circle: the resonance is gone, the peak reads 0 Hz, and the filter is no longer stable. After the clip a slider named “Fractional bits B” runs from 4 to 10; the arrow keys move it by one bit, and Home and End jump to 4 and 10. At 10, 7 and 6 bits the caption is the clip’s; at other values it reads like “B = 8: pole 0.9803∠4.99°, peak at 107.9 Hz.” The choice is kept in the link, as grid.B.
Rounded coefficients put the poles on a grid
Most filters on this site run in floating point, where rounding is usually too small to see. On a small DSP chip or a microcontroller that changes. To save power and silicon, many of them store every number as a 16-bit integer and agree on where the binary point sits. That is fixed point.
The most common agreement reads the word as a fraction. The 16 bits hold a whole number from −32 768 to 32 767, and the value it stands for is that number divided by :
So the values run from to , in steps of . The first bit carries the sign, and the other 15 sit behind the binary point. I call those the fractional bits. The format’s name, Q1.15, counts them: one bit for the sign, 15 fractional bits.
Storing a value is the rounding of Quantization & noise (11.1), with a level at 0, the mid-tread placement. The step is , and 11.1’s gives the same with bits in all. On this page counts only the fractional bits, so .
Floating point works differently. A 32-bit float keeps 24 significant bits and moves the binary point to suit each number, so its step is about of the number itself. Near 0.001 its step is , while Q1.15’s step there is still , 3 % of the value. Fixed point has one step everywhere, and that is where this page’s three problems start.
The first problem is the coefficients. A biquad’s bottom polynomial, from Transfer functions, poles & zeros (16.3), is . For poles at it has and . A stable biquad has between and , inside the triangle of Stability and causality (16.4), so needs one more integer bit: Q2.14, from to .
Round and to fractional bits and they can only take values on a grid, apart. Each grid point gives one pair of poles:
So the poles can only sit where some grid point puts them. Second-order sections (21.2) saw rounding move poles; here we see where they are allowed to go. Think of a city map with streets only every kilometre: an address has to move to the nearest corner.
The picture at the top of the page rounds one low-frequency resonance at kHz. Its poles are , and 5° of a full turn is Hz. So and .
The filter is with , which makes the gain 1 at 0 Hz. It rounds only and and keeps this . The wanted gain peaks at 108.1 Hz, a little below the poles’ angle, as 16.3 found.
In that picture, ten and seven bits keep the resonance, and six bits destroy it.
Why did six bits destroy it? For a mirror pair, the height above the real axis is , and . Our pole has , so . One step of at six bits is , more than twice as big.
The nearest grid values give . That is below 16.4’s parabola, so both poles are real. They also have , which by 16.4 puts one pole exactly at .
More bits did not always land closer. Ten bits left the pole 0.00038 from the wanted one, and seven bits only 0.00035. Each extra bit gives a finer grid, but where the nearest grid point falls is a matter of chance.
After the clip, set the number of fractional bits yourself with the slider, from 4 to 10. Eight and nine bits land on the same grid point as seven.
Now try 5. The poles jump to and the peak to 225.4 Hz, about twice the wanted frequency. At 4 bits the poles are real again, 1 and 0.9375.
Why is the grid so sparse near ? It is even in and , not in the plane. Write a pole as ; then and . Near the real axis is small, and smaller still, so one step of moves the pole a long way up.
In numbers, a patch of the plane holds grid poles in proportion to its height . They crowd near ±90°, where is largest, and thin out near . At six bits, the poles of radius 0.9 to 1 within 15° of number 48. The same-sized wedge around 90° holds 382, eight times as many.
That is why low-frequency resonances suffer. Bass equalisers and slow sensors sampled fast put their poles close to , where the grid is thin. The fixes are more bits, or a different structure.
The coupled form stores and themselves, so its grid is even across the plane. Floating point avoids most of this, at more cost per operation.
The maths behind it · eigenvalue perturbation
Rounding the coefficients is a small perturbation of the 2×2 state matrix of 16.3, whose eigenvalues are the poles; how far they move depends on the matrix’s structure. Near a double eigenvalue, a small change can move them by about its square root: that is the square root in . The coupled form’s matrix is a scaled rotation, an orthogonal matrix times . Its poles are spread evenly, and they move no further than its entries change.
Overflow wraps; saturation clips
The second problem is size. Q1.15 has no room above 0.99997, so what happens when a sum lands beyond it? Most chips use two’s complement integers. The 16-bit integer counts up to 32 767, and one more gives −32 768, the bottom of the range.
The value has gone all the way round, like a car’s odometer rolling from 99 999 to 00 000. This is wrap-around, and the stored value is off by 2, the width of the whole range. The other choice is saturation: the arithmetic holds the value at the limit, like an odometer that sticks at 99 999. Its error is only the excess above the limit.
Many DSP chips offer saturating additions in hardware for this reason. The instrument below stores the same sine both ways. Watch the tops of the sine once the amplitude passes the limit.
Overflow wraps; saturation clips
A sine stored as Q1.15 (−1 to 0.99997), its amplitude raised past the limit.
Amplitude 0.9: everything fits; both versions store the sine to within 0.00002.
Describe this picture
Two stacked panels, “wrap-around” and “saturation”, for a sine stored as Q1.15 (−1 to 0.99997) with its amplitude raised past the limit. They share the sample axis, from 0 to 63. Each has dotted levels at ±1 labelled “limit” and the wanted sine as a thin dashed curve. The stored values are stems, with dot heads in the wrap-around panel and square heads in the saturation panel. The readouts are the amplitude and the worst error of each version. There is no control.
The clip lasts 12 s, and the stems follow the amplitude as it changes. At amplitude 0.9 everything fits, and both versions store the sine to within 0.00002. The amplitude rises past the limit. At 6 s, amplitude 1.2: wrap-around throws each top to the other side, to about −0.8, an error of 2.000, and saturation flattens the tops, an error of 0.200. At the end, amplitude 1.5: saturation’s error grows to 0.500, and wrap-around’s stays 2.000, the worst possible. The end caption advises scaling the signal down so the largest value fits, or saturating.
Look at a top in the wrap-around panel. At amplitude 1.2 the top sample is 39 322 steps, which does not fit. It comes back as steps, which is −0.8, and its neighbours land at −0.82 and −0.89. At amplitude 1.5 the top comes back as −0.5.
So the further a value passes the limit, the further from −1 it lands, but every wrapped sample is wrong by 2. Saturation’s error grows with the excess, 0.2 and then 0.5, and never jumps.
The cure for both is scaling: multiply the signal by a constant below 1 before it can overflow, so the largest value fits. Halving the 1.5 sine gives 0.75, which fits with room to spare. The price is one bit: the step stays while the signal halves, so the signal-to-noise ratio falls by 6.02 dB (11.1).
In a filter, the values to watch are the stored ones inside the structure, not only the output. Filter structures (21.1) showed that direct form II’s inside signal can be much larger than its output. That is the value the scaling has to fit.
A filter that never goes quiet
The third problem hides inside a loop. Take the first-order loop of Difference equations (6.1), with no input: . Exactly, from , it decays as . In fixed point the product has more bits than the word, so it is rounded before it is stored:
Here I measure in steps, so rounds to whole numbers. Engineers call one step one LSB, the value of the least significant bit. A value exactly halfway, like 4.5, rounds away from zero, to 5.
Each rounding adds an error of up to half a step. While is large, that hardly matters. When it is small, the rounding can give back the same value it started from.
When does that happen for ? Rounding gives back when is within half a step of , that is when :
That band around 0 is the dead band. For it reaches steps. Once the output is inside it, the rounded loop stops shrinking. An output that keeps repeating with no input, where the exact filter would decay, is a limit cycle.
Think of a ball that should come to rest but keeps bouncing on a step it cannot get over. The instrument below runs the loop for and then .
A filter that never goes quiet
y[n] = Q{a·y[n−1]} from y[0] = 20, rounded to whole steps, no input.
Start at 20 steps, no input, and multiply by a each sample, rounding to whole steps.
Describe this picture
One panel for from , rounded to whole steps, with no input: in steps, from −22 to 22, against the sample from 0 to 60. The exact values are small open circles, and the rounded output is stems with dot heads. A hatched band from −5 to 5 is labelled “dead band”. The readouts are , the rounded and the exact .
The clip lasts 14 s. First the axes and the band appear: start at 20 steps, no input, and multiply by each sample, rounding to whole steps. Then both sequences build from to 60. At 6 s, : 20, 18, 16, 14, … then 5, 5, 5 for ever, because rounds back to 5, while the exact output has decayed to 0.0359. Then the picture cross-fades to , and the sequences build again. At the end the output flips between +5 and −5 for ever, a limit cycle: inside the dead band, , rounding undoes the decay. After the clip two buttons in a group named “Loop gain” choose or , each with its caption from the clip; the choice is kept in the link, as cycle.a.
Notice the exact circles: they cross into the band at and keep shrinking towards 0. The rounded stems stop at its edge, 5, which is . With the sign flips every sample, so the stuck value becomes a flip between +5 and −5.
After the clip, switch between the two loop gains with their buttons.
Round-off noise
With a busy input , the loop is , and the rounding errors behave like 11.1’s noise: each adds an error of mean square . So , and enters as if it were a second input.
The loop filters that noise like everything else. Write for the impulse response from the rounding point to the output; here it is . By Simple smoothing filters (18.2), noise power is multiplied by :
For that factor is . Its square root, 2.29, is the noise gain of 18.2. Closer to the circle it grows fast: for the factor is 50.25. I checked by simulation with a busy input: 0.441 steps², against the predicted .
In direct form II, the feedback products are rounded where is formed. So their noise passes through , the same part that made large in 21.1, and poles near the circle amplify both. The remedies are more bits, a different structure, or error feedback, which feeds each rounding error back into the loop to cancel part of it.
When the input goes quiet, the noise model breaks down. The error stops looking random and follows the signal, as it did for 11.1’s quiet sine, and that is exactly where limit cycles appear.
The maths behind it · variance propagation
Modelling round-off as uniform noise of variance (11.1) and passing it through a filter is variance propagation. The output variance is the input variance times , the square of 18.2’s noise gain. The model fails exactly where limit cycles appear: there the error is no longer random.
Worked example
1. Q1.15 by hand. To store 0.3, multiply by 32 768 to get 9830.4, and round to 9830. That stands for , an error of , below half a step, . The largest value is .
The whole span, 2, holds steps. As a level that is dB, which is : 11.1’s 6 dB per bit.
2. The resonance at five bits. With : rounds to , so becomes . And rounds to 31, so becomes .
Then and , so . That angle is Hz, and the gain peaks a little lower, at 225.4 Hz. That is 2.08 times the wanted 108.1 Hz.
3. Overflow at amplitude 1.2. The top sample, 1.2, is , rounded to 39 322 steps. Two’s complement wraps it to steps, which is . Saturation stores instead, an error of .
4. Dead bands. For the dead band is step. From 20 the rounded loop gives 20, 10, 5, 3, 2, 1, 1, and so on: 2.5 rounds to 3, and 0.5 rounds back to 1. For the dead band is 5 steps, and the loop’s round-off noise power is times .
Where you’ll meet this
Fixed point is common wherever power and cost matter. Audio codecs and hearing aids run their filters in Q15 and Q31, the names ARM’s library uses for Q1.15 and its 32-bit cousin Q1.31. Motor-control microcontrollers run their control loops the same way, and FPGA filters choose a custom word length for every signal.
ARM’s CMSIS-DSP library is a good example. Routines such as arm_biquad_cascade_df1_q15 work in Q15 and saturate by design. Their documentation also says when the input must be scaled down to avoid overflow. A real-time system then has to fit all this into its time budget, which is Real-time processing (21.4).
Rounding noise can also be pushed out of the band you care about, as converters do in Oversampling and noise shaping (11.3). In floating point, as in SciPy on a laptop, most of this page fades away. A high-order direct form can still move its poles a long way, though (21.2).
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Q1.15 | to , step | 16-bit word: 1 sign bit, 15 fractional bits |
| Q2.14 | to | a biquad’s lies in |
| Step | counts fractional bits here | |
| Biquad poles | , | from rounded , |
| Pole grid | even in , | sparse near |
| Wrap-around | error 2, the whole range | two’s complement |
| Saturation | error = the excess above the limit | prefer it, or scale down |
| Scaling | multiply so the largest stored value fits | halving costs 6.02 dB |
| Round-off in a loop | factor 5.26 for | |
| Dead band | steps | limit cycles |