What makes one filter the best? Below, two 27-tap designs face the same spec: one makes the average error smallest, the other the worst error. Watch which one stays out of the hatched zones.
Least squares or least worst
Two 27-tap designs for the running spec, each the best by its own measure. Stop-band errors count five times as much as pass-band ones.
The running spec as zones, 27 taps to meet it.
Describe this picture
One panel drawn like the spec drawings of 18.1, for two 27-tap designs for the running spec, each the best by its own measure; stop-band errors count five times as much as pass-band ones. It plots the gain from −80 to 5 dB against frequency from 0 to 4000 Hz, with the running spec’s forbidden zones hatched. The least-squares curve is dashed and the equiripple curve solid. The key adds “fails” or “passes” to each name, and a failing design gets a cross at its worst point. The readouts are the design, the stop band’s worst point, the stop band’s RMS (the square root of the mean square of the gain over the stop band, in dB) and the pass band’s worst point. Each shows “not yet” until a design is drawn.
The clip lasts 13 s. At the start only the zones are drawn. Then the least-squares design draws and holds: the smallest average error, −50.4 dB RMS in the stop band, but next to the edge it reaches −31.7 dB, and the pass band strays 0.076 from 1, so it fails. Next the equiripple design draws, and the least-squares curve fades to 40 % opacity. Every ripple is the same height, −41.9 dB in the stop band and 0.040 in the pass band: a worse average, −44.9 dB, but the lowest worst point, and it passes. After the clip two buttons in a group named “Design” choose which design is drawn at full strength; the other stays at 40 % opacity, and the caption shows the chosen design’s caption from the clip. The choice is kept in the link, as worst.d.
Least squares or least worst
In Window-method FIR design (19.1), Kaiser’s window met the running spec with 38 taps. In Frequency-sampling design (19.2), two well-chosen transition samples did it with 31. Every tap costs one multiplication per output sample, so I want to know the shortest filter that can do it. To find it, I first have to say what “the best filter” means.
Here is the running spec of Filter specifications (18.1) once more. The sample rate is 8 kHz. From 0 to 1 kHz the gain must stay within , and from 1.5 kHz to 4 kHz it must stay at or below 0.01, which is 40 dB down. In rad/sample the pass band ends at and the stop band starts at .
A filter with an odd number of symmetric taps has the real zero-phase response of Linear-phase systems (17.3):
and its gain is . For 27 taps the centre is . Only to are free, because the other 13 taps are their mirror copies. So a 27-tap design is a choice of 14 numbers.
What do we want to be? I call the target the desired response : 1 in the pass band and 0 in the stop band. In the transition band I ask for nothing, and errors there do not count. The error of a design is then in the two bands.
The spec is not equally strict in both bands. It allows 0.05 in the pass band but only 0.01 in the stop band, five times less. So I multiply the stop-band error by 5 before I judge it, and call this factor the weight : 1 in the pass band, 5 in the stop band. The weighted error is
With this weight, the whole spec becomes one condition: in both bands. In the pass band that says . In the stop band it says , which is a gain of at most 0.01.
There are two natural ways to make small. The first is to make its average size small. Least squares chooses the taps that make the mean square of over the two bands as small as it can be:
This is the measure of Convergence and the Gibbs phenomenon (7.3), where a partial sum was the best approximation in mean square. The second way is to make the largest value small. Minimax chooses the taps that make the worst point as low as it can be:
You have already met one least-squares design. The plain cut of 19.1 kept 27 terms of the ideal , and by 7.3 a partial sum is the best in mean square. So the plain cut is the least-squares answer with no transition band and . It was the best by that measure, and it still rang next to the edge.
Least squares is also easy to compute. 19.2 showed that the gain is linear in the design’s numbers, so is linear in the 14 taps. Its mean square is then a quadratic in the 14 taps, and its lowest point comes from 14 linear equations, solved once. SciPy’s firls does this.
Which measure should we use? Think of two ways to get to work. One usually takes 20 minutes but sometimes 60, and the other always takes 30. If you must never be late, you take the second.
A spec is like that: it does not look at the average, only at the worst point. The picture at the top of the page draws the best 27-tap design by each measure against the running spec. Both use the weight 5.
After the clip, press either design’s button to draw it at full strength.
Compare the two “stop band, RMS” readouts first. Least squares wins on average, −50.4 dB against −44.9 dB, and it still fails: right at the stop edge its gain climbs to −31.7 dB. The equiripple design spreads its error evenly, so no ripple sticks out. For comparison, Kaiser’s window at 27 taps strays 0.069 in the pass band and reaches −22.4 dB.
Equal ripples
The equiripple design is not a lucky guess: it is what minimax gives. Why should the best worst point come with ripples that are all the same height? Suppose one ripple were lower than the others. Its spare room would be wasted, because the filter could give some of it up to push the highest ripple down.
A theorem makes this exact. Call the largest size of the weighted error .
Take a type I filter (odd length, symmetric taps, 17.3) with free numbers. The alternation theorem says it has the smallest possible if and only if its weighted error reaches at least times, alternating in sign. Only the best filter does this. A design whose ripples all have one height is called equiripple.
For 27 taps, , so the best design touches at 15 frequencies. In terms of gain, the pass band strays at most from 1, and the stop band reaches at most , because the weight there is 5. I will not prove the theorem here, but the next instrument lets you count the 15 touches.
Remez exchange: move the points to the peaks
How do we find the equiripple filter? The Parks–McClellan program does it with the Remez exchange, a search that repeats four steps:
- Pick 15 trial frequencies in the two bands.
- Find the one filter whose weighted error is exactly at them.
- Look where its weighted error actually peaks.
- Move the trial points to those peaks, and repeat until no peak is higher than .
Step 2 is a set of linear equations. At each trial point, with counting the points,
is linear in the taps, so these are 15 linear equations for 15 unknowns: the 14 numbers to , and .
Think of levelling a wobbly table by always fixing its highest corner. Each step fixes the worst place, and the next worst place becomes the target. Here, of a trial filter can never be larger than the best possible worst error, and the largest error of that filter can never be smaller. So the answer is caught between the two numbers, and the search stops when they meet.
The instrument below runs the search for the running spec at 27 taps, and shows the weighted error , where equiripple means literally equal ripples.
Remez exchange: move the points to the peaks
27 taps for the running spec. The weighted error ε(Ω) is forced to ±δ at 15 trial points; then the points move to the error's peaks.
Iteration 1: 15 trial points spread evenly. The error is exactly ±0.0029 at each, but between them it swings off the scale, to 1.144.
Describe this picture
One panel for 27 taps and the running spec: the weighted error , from −0.35 to 0.35, against from 0 to rad/sample. The transition band, to , is left blank, with “transition” written in it. The error is a solid curve. Two dashed lines mark the levels and , and the 15 trial points are filled dots on the curve. Where the curve leaves the panel, an open triangle at the edge is labelled “off the scale”. The readouts are the iteration, the level and the largest error. There is no control.
The clip lasts 19 s and shows seven iterations. Between them the curve changes into the new error and the dots slide to the new points, with a blank caption. In iteration 1 the 15 trial points are spread evenly: the error is exactly ±0.0029 at each, but between them it swings off the scale, to 1.144. In iteration 2 the points have moved to the peaks: rises to 0.0201, and the largest error falls to 0.310. Iterations 3 to 6 name the two numbers as they close in. In iteration 7 every peak touches , 15 alternations, and the caption adds that no 27-tap filter has a smaller worst error: in the stop band that is , −41.9 dB.
The two numbers at each iteration are these.
| Iteration | level δ | largest error |
|---|---|---|
| 1 | 0.0029 | 1.1443 |
| 2 | 0.0201 | 0.3101 |
| 3 | 0.0336 | 0.1190 |
| 4 | 0.0391 | 0.0654 |
| 5 | 0.0400 | 0.0472 |
| 6 | 0.0400 | 0.0402 |
| 7 | 0.0400 | 0.0400 |
Watch the two readouts. The level grows at every step, and the largest error shrinks. In iteration 1 the dots are spread evenly, and between them the error is hundreds of times larger than . Once the dots sit on the peaks, the two numbers close in fast, and in iteration 6 they agree to three decimals.
In the last frame, count the dots. Five sit in the pass band and ten in the stop band, and their signs alternate: 15 touches, as the theorem demands. The stop-band ripples look as tall as the pass-band ones only because of the weight. In gain they are five times smaller.
The weight moves the ripple between the bands
Why the weight 5? Call the weight’s value in the stop band . The equiripple design makes the weighted error high in both bands, so the gain strays in the pass band and in the stop band. Their ratio is
So the weight decides how the error is shared between the bands, and the length decides how much error there is to share. It is like splitting a fixed budget between two teams: more for one means less for the other. The running spec allows , so that is the weight to use.
The instrument below shows three 27-tap equiripple designs with different weights.
The weight moves the ripple between the bands
27-tap equiripple designs for the running spec's edges, with stop-band weight 1, 25 and 5.
Weight 1: equal ripples, 0.022 in both bands. The stop band, −33.2 dB, fails.
Describe this picture
Two panels for 27-tap equiripple designs at the running spec’s edges, with stop-band weight 1, 25 and 5. The first, “pass band, close up”, plots the gain from 0.85 to 1.15 against frequency from 0 to 1000 Hz, with hatched zones above 1.05 and below 0.95. The second plots the gain in dB like the panel of the first instrument, with the running spec’s zones. One solid curve runs through both panels, and a failing point gets a cross. The readouts are the stop-band weight, the pass band’s ± deviation and the stop band’s level.
The clip lasts 13 s, with a blank caption while the curve changes. At weight 1 the ripples are equal, 0.022 in both bands, and the stop band, −33.2 dB, fails. At weight 25 the stop band sinks to −48.1 dB, but the pass band strays 0.098 from 1, and it fails the other way. At the end the weight is 5, the spec’s own ratio 0.05/0.01: 0.040 in the pass band and −41.9 dB in the stop band, and both pass. After the clip a slider named “Stop-band weight” sets any whole number from 1 to 50, with a value like “5: passes”. The arrow keys change the weight by 1, Page Up and Page Down by 5, and Home and End jump to 1 and 50. At other weights the caption reads like “Weight 10: pass band ±0.055, stop band −45.1 dB: the pass band fails.” or “Weight 7: pass band ±0.044, stop band −44.0 dB: both pass.” The weight is kept in the link, as weight.k.
After the clip, set the weight yourself with the slider, from 1 to 50. Move the weight up and down and watch the two readouts. The pass-band ripple and the stop-band level always move in opposite directions.
Only weights 4 to 8 pass: the pass band strays 0.038 to 0.048, and the stop band sits at −40.5 to −44.5 dB. At weight 3 the stop band fails at −38.8 dB, and at weight 9 the pass band fails at 0.051. At every weight the two ripples keep the ratio , within 1 %.
One spec, seven designs
Now we can answer the question from the start. For each method I increased the length until the design met the running spec. The delay of a linear-phase filter is samples (17.3), and the last column gives it at 8 kHz.
| Method | Shortest length that passes | Delay at 8 kHz |
|---|---|---|
| equiripple, weight 5 | 26 taps (27 if odd) | 1.56 ms |
| frequency sampling, two optimised samples | 31 taps | 1.88 ms |
| Kaiser window | 38 taps | 2.31 ms |
| least squares, weight 5 | 39 taps (odd lengths only) | 2.38 ms |
| Hann window | 50 taps | 3.06 ms |
| Hamming window | 50 taps | 3.06 ms |
| Blackman window | 66 taps | 4.06 ms |
The equiripple design is the shortest. SciPy’s remez also designs even lengths, which are linear phase too (17.3), and 26 taps are enough. Against Kaiser’s 38, that is a third fewer multiplications for every output sample.
Is there a formula for the length, as there was for the Kaiser window? Kaiser also gave an estimate for equiripple designs. With the transition width in hertz (18.1),
For the running spec it gives 22.9, against the true 26. Like Kaiser’s window formula in 19.1, it is an estimate to check.
Why not always use equiripple? It needs an iterative program, where a window design is one formula. And its stop band does not fall away.
The 27-tap design stays at −41.9 dB right up to 4 kHz. Kaiser’s 38 taps are at −41.7 dB next to the edge, but below −59.0 dB from 2500 Hz on and below −70.0 dB from 3500 Hz on. If most of the unwanted signal lies far from the edge, the window design removes more of it.
Worked example
Let’s design the running spec as an equiripple filter, step by step, and check the result.
The weight. The spec allows and , so .
A first length. In Kaiser’s estimate, , and dB. The transition is , so the denominator is . Then
Check and adjust. Start at 23 taps and add taps until the design passes. At 23 taps the pass band strays 0.078 and the stop band reaches −36.1 dB, so both fail. At 25 taps it is 0.056 and −39.1 dB, still failing. At 27 taps it is 0.040 and −41.9 dB, and the design passes.
With even lengths allowed, 26 taps pass with 0.044 and −41.1 dB.
The result at 27 taps. The Remez exchange takes 7 iterations from the even start and ends at . So and , which is −41.9 dB. The weighted error touches 15 times, alternating, so no 27-tap filter does better with this weight. The delay is samples, which is 1.625 ms at 8 kHz.
The alternatives at 27 taps. Least squares with the same weight fails at −31.7 dB, and it needs 39 taps to pass (43 with weight 1). Equiripple with weight 1, the same ripple in both bands, needs 35 odd taps, or 34 with an even length. The weight alone is worth 8 taps here.
Where you’ll meet this
scipy.signal.remez and MATLAB’s firpm are the standard tools for filters in fixed-point processors and hardware, where every tap costs. At 8 kHz, the 26-tap design needs 208 000 multiplications a second, and Kaiser’s 38 taps need 304 000. On a chip, fewer taps also means less memory and less power.
Least squares has its place too. When the unwanted signal is broadband noise, the noise power that leaks through depends on the mean square of the stop-band gain. That is the quantity least squares keeps small, and SciPy’s firls designs it in one step.
The same equal-ripple idea returns for recursive filters: elliptic filters are the equiripple designs among IIR filters, in Analog prototype filters (20.1). Choosing FIR or IIR (20.6) compares them for one spec. Long equiripple low-pass filters are the usual choice in Downsampling and decimation (22.1), and Finite word-length effects (21.3) shows what happens to their ripples when the taps are rounded.
The maths behind it · normal equations
Least squares is the normal equations . Here holds the cosines at the grid frequencies, has the weights on its diagonal, and holds the desired values. Minimax is Chebyshev approximation. The equal ripples are cousins of the Chebyshev polynomials, which swing between ±1 more often than any other polynomial of their degree.
The maths behind it · squared and worst-case loss
Squared loss against worst-case loss: regression with squared errors gives the best average fit and lets a few points sit far off. Minimax (Chebyshev) regression bounds every residual instead. The weights play the role of the inverse variances in weighted least squares: they say which errors matter more.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Weighted error | transition ignored | |
| Least squares | minimise over the bands | firls; best average |
| Minimax | minimise | remez; best worst point |
| Equiripple | , touching at least times, alternating | type I, |
| Weight | 5 for the running spec | |
| Remez step | ±δ at trial points → move them to the peaks | δ grows each step |
| Length estimate | check, then adjust |