Skip to content

Adaptive filters: LMS

A filter that learns its taps: LMS walks down the error bowl in noisy steps, and an adaptive canceller removes mains hum.

Before this25.4 · 26.2 · 4 more
Chapter 26 · Lesson 3 of 4

First, the picture

Here a filter with two taps learns to copy an unknown system, nudging its taps after every sample. Each pair of taps is a point on a map of the error, and the lowest error is at the cross. Watch the solid path wobble down towards it.

Downhill, one noisy step at a time

Two taps identify h_true = (0.8, −0.4) from a coloured input (seeds 263 and 2630); μ = 0.05.

Start at h = (0, 0), high on the bowl: J = 0.426.

step n
0
taps
(0.000, 0.000)
J
0.426
Step size
0.00 / 14.00 s
Describe this picture

One square panel: two taps identify htrue=(0.8,−0.4)\mathbf{h}_\text{true}=(0.8,-0.4) from a coloured input (seeds 263 and 2630), with μ = 0.05. The axes are h0h_0 from −0.2 to 1.4 and h1h_1 from −1.0 to 0.6, at equal scales. Faint ellipses mark the heights 0.02, 0.05 and 0.17 of the error bowl, and a cross labelled “bottom” marks htrue\mathbf{h}_\text{true}. LMS and steepest descent run side by side for 300 steps each. The steepest-descent path is dashed, with a small ring at its end, and the LMS path is a solid line with a dot at the current taps; a key names both. The readouts are the step nn, the taps h0h_0 and h1h_1, and JJ, the bowl’s height at the current taps. The dot starts at (0, 0), high on the bowl, at JJ = 0.426. After 30 steps LMS is at (0.483, −0.162) and has come down to JJ = 0.077, wobbling about the smooth steepest-descent path. After 300 steps it is at (0.812, −0.369), near the bottom, with JJ = 0.012 against the least possible 0.010, and keeps wandering a little around it. When the clip ends, three buttons in a group named “Step size”, “μ = 0.01”, “μ = 0.05” and “μ = 0.12”, redraw both paths with that step size, each with its own line in the key. With μ = 0.01, after 300 steps the taps are (0.609, −0.224) and JJ is 0.037; with μ = 0.12 they are (0.827, −0.365) and JJ is 0.013.

A filter that learns

The Wiener filter (26.2) found the best FIR filter from correlations. It needed the input’s autocorrelation matrix R\mathbf{R} and the cross-correlation vector r\mathbf{r}, known in advance. Often nobody knows them. A phone does not know the echo path of the room it is in, and that path changes when someone moves.

So let the filter learn them. It starts with some taps, watches its own error, and nudges the taps after every sample so that the error shrinks. A filter that does this is adaptive. This page builds the simplest and most widely used one, the LMS algorithm.

I assume that the input is stationary while the filter learns, as in “Stationary, and ergodic” of Random processes (24.2). If the statistics drift slowly, the filter keeps learning and follows them.

An unknown system to copy

Here is the task for the first half of the page. An unknown FIR system with two taps, htrue=(0.8,−0.4)\mathbf{h}_\text{true}=(0.8,-0.4), is driven by an input x[n]x[n]. I can measure its output, with noise v[n]v[n] of standard deviation 0.1 added. Finding the system from its input and output is called system identification.

The noisy output is what my filter tries to match, so it plays the part of 26.2’s desired signal:

ddes[n]=0.8 x[n]−0.4 x[n−1]+v[n].\begin{aligned} d_\text{des}[n]&=0.8\,x[n]-0.4\,{x[n-1]}\\ &\quad+v[n]. \end{aligned}

My filter has two taps too, h=(h0,h1)\mathbf{h}=(h_0,h_1). At sample nn it sees the two latest inputs, stacked as a column, newest first:

xn=(x[n], x[n−1])⊤.\mathbf{x}_n=\big(x[n],\,x[n-1]\big)^\top.

This xn\mathbf{x}_n is the tap-input vector. The snapshot of 25.4 ran forward in time; this one runs backward. The output is y[n]=h⊤xn=h0x[n]+h1x[n−1]y[n]=\mathbf{h}^\top\mathbf{x}_n=h_0x[n]+h_1{x[n-1]}, and the error is e[n]=ddes[n]−y[n]e[n]=d_\text{des}[n]-y[n], as in 26.2.

The input is coloured noise, the AR(1) process of “White noise forgets, coloured noise remembers” in 24.2. Each sample is 0.6 x[n−1]0.6\,{x[n-1]} plus white noise of variance 0.64. Its power is then 0.64/(1−0.62)=10.64/(1-0.6^2)=1, and Rx[ℓ]=0.6∣ℓ∣R_x[\ell]=0.6^{\lvert\ell\rvert}. With two taps, the matrix R\mathbf{R} of 26.2 has entries 1 and 0.6:

R=(10.60.61).\mathbf{R}=\begin{pmatrix}1&0.6\\0.6&1\end{pmatrix}.

The error is a bowl

How good is a choice of taps? In 26.2 the measure was the mean-square error J=E{e2[n]}J=\mathbb{E}\{e^2[n]\}. Square e[n]=ddes[n]−h⊤xne[n]=d_\text{des}[n]-\mathbf{h}^\top\mathbf{x}_n and average:

J(h)=σdes2−2 h⊤r+h⊤R h.J(\mathbf{h})=\sigma_\text{des}^2-2\,\mathbf{h}^\top\mathbf{r}+\mathbf{h}^\top\mathbf{R}\,\mathbf{h}.

Here σdes2\sigma_\text{des}^2 is the power of ddes[n]d_\text{des}[n], and r\mathbf{r} has the entries E{ddes[n] x[n−k]}\mathbb{E}\{d_\text{des}[n]\,x[n-k]\}.

The best taps solve 26.2’s Wiener–Hopf equations, Rhopt=r\mathbf{R}\mathbf{h}_\text{opt}=\mathbf{r}, and leave the least error Jmin=σdes2−r⊤hoptJ_\text{min}=\sigma_\text{des}^2-\mathbf{r}^\top\mathbf{h}_\text{opt}. Put r=Rhopt\mathbf{r}=\mathbf{R}\mathbf{h}_\text{opt} into JJ and collect the terms:

J(h)=Jmin+(h−hopt)⊤R (h−hopt).\begin{aligned} J(\mathbf{h})&=J_\text{min}\\ &\quad+(\mathbf{h}-\mathbf{h}_\text{opt})^\top\mathbf{R}\,(\mathbf{h}-\mathbf{h}_\text{opt}). \end{aligned}

The last term is the mean square of (h−hopt)⊤xn(\mathbf{h}-\mathbf{h}_\text{opt})^\top\mathbf{x}_n, so it is never negative. Over the two taps, JJ is a bowl with its bottom at hopt\mathbf{h}_\text{opt}, at height JminJ_\text{min}. This bowl is the error surface.

For our task the bottom is where you would hope. The noise is uncorrelated with the input, so r=Rhtrue\mathbf{r}=\mathbf{R}\mathbf{h}_\text{true} and hopt=htrue\mathbf{h}_\text{opt}=\mathbf{h}_\text{true}. What is left at the bottom is the noise: Jmin=0.12=0.01J_\text{min}=0.1^2=0.01. At h=(0,0)\mathbf{h}=(0,0) the filter outputs nothing, and JJ is the whole power of ddes[n]d_\text{des}[n], 0.426.

The shape of the bowl

Cut the bowl at a fixed height and the edge is an ellipse. Its axes are the eigenvectors of R\mathbf{R}, from “Directions that are only stretched” in High-resolution frequency estimation (25.4). Here you can check them by hand: R\mathbf{R} sends (1,1)(1,1) to (1.6,1.6)(1.6,1.6) and (1,−1)(1,-1) to (0.4,−0.4)(0.4,-0.4). So the eigenvalues are λmax=1.6\lambda_\text{max}=1.6 and λmin=0.4\lambda_\text{min}=0.4.

Step a distance ss from the bottom along a unit eigenvector, and JJ rises by λs2\lambda s^2. So the bowl is steep along (1,1)(1,1) and shallow along (1,−1)(1,-1). The ellipse’s long axis runs along (1,−1)(1,-1), the direction of the smaller eigenvalue.

At J=0.17J=0.17, for example, the ellipse reaches 0.632 from the bottom along (1,−1)(1,-1) and 0.316 along (1,1)(1,1). The ratio of the eigenvalues, 4 here, is the eigenvalue spread: the longer the ellipse, the larger it is.

Walking downhill

If you know the bowl, you can walk down it. The slope of JJ along each tap is its derivative with respect to that tap. Together the two derivatives form the gradient, which points uphill:

∂J∂h=2 R (h−hopt).\frac{\partial J}{\partial\mathbf{h}}=2\,\mathbf{R}\,(\mathbf{h}-\mathbf{h}_\text{opt}).

Steepest descent takes a small step against the gradient, again and again, each new value from the last as in Difference equations (6.1). I fold the 2 into the step:

hn+1=hn+μ R (hopt−hn).\mathbf{h}_{n+1}=\mathbf{h}_n+\mu\,\mathbf{R}\,(\mathbf{h}_\text{opt}-\mathbf{h}_n).

Here hn\mathbf{h}_n is the taps at step nn, and μ\mu is the step size, a small positive number. On this page and in 26.4, a bare μ\mu is a step size, not the fractional position of Resampling by any factor (22.3).

There is a catch. Steepest descent needs R\mathbf{R} and hopt\mathbf{h}_\text{opt}, and if I knew those I would not need to walk.

LMS: the slope from one sample

Look again at the step. Since Rhopt=r\mathbf{R}\mathbf{h}_\text{opt}=\mathbf{r}, it is r−Rhn\mathbf{r}-\mathbf{R}\mathbf{h}_n, the average of xn(ddes[n]−xn⊤hn)\mathbf{x}_n\big(d_\text{des}[n]-\mathbf{x}_n^\top\mathbf{h}_n\big). That is the average of e[n] xne[n]\,\mathbf{x}_n.

The LMS algorithm, for “least mean squares”, drops the average. It uses the error and the inputs of the current sample only:

e[n]=ddes[n]−hn⊤xn,hn+1=hn+μ e[n] xn.\begin{aligned} e[n]&=d_\text{des}[n]-\mathbf{h}_n^\top\mathbf{x}_n,\\ \mathbf{h}_{n+1}&=\mathbf{h}_n+\mu\,e[n]\,\mathbf{x}_n. \end{aligned}

Each tap costs about two multiplications per sample, one for the output and one for its update. No correlations are estimated at all. B. Widrow and M. E. Hoff published it in 1960.

Each step points downhill only on average. Think of walking downhill in fog: you judge the slope from the patch of ground under your feet. Each step is a little off, but on average you go down.

Downhill, one noisy step at a time

The picture at the top of the page runs LMS and steepest descent side by side on this bowl, for 300 steps each, with μ = 0.05. LMS starts at (0, 0), high on the bowl, at JJ = 0.426. After 30 steps it has come down to JJ = 0.077, wobbling about the smooth steepest-descent path. After 300 it is near the bottom, JJ = 0.012 against the least possible 0.010, and keeps wandering a little around it. When the clip ends, try the three step sizes.

Notice where the two paths end. After 300 steps at μ = 0.05, LMS is at (0.812, −0.369) and steepest descent at (0.799, −0.399). The smooth path has nearly reached the bottom; the noisy one is close and still moving.

Fast and slow directions

Why does the walk curve? Write the distance from the bottom as hn−hopt\mathbf{h}_n-\mathbf{h}_\text{opt}. Steepest descent multiplies it by I−μR\mathbf{I}-\mu\mathbf{R} at each step. Along an eigenvector, that is a plain number, 1−μλ1-\mu\lambda.

So the distance along each eigenvector shrinks by its own factor at every step. After 1/(μλ)1/(\mu\lambda) steps it has fallen to about 0.37 of its start, and that count is the time constant of the direction. At μ = 0.05 the steep direction has the time constant 1/(0.05⋅1.6)=12.51/(0.05\cdot1.6)=12.5 steps and the shallow one 50.0 steps.

That is the curve you saw. The path first drops quickly across the narrow width of the ellipse, then crawls along its long axis. A larger eigenvalue spread makes the crawl longer compared with the drop.

Speed against wander

Near the bottom the two walks behave differently. Steepest descent settles. LMS keeps taking noisy steps, so its taps keep wandering around the bottom.

A larger step arrives sooner and wanders more. On this record, the first step within 0.05 of the bottom is step 663 for μ = 0.01, 116 for μ = 0.05 and 45 for μ = 0.12. Over samples 1000 to 2000, the taps stray from htrue\mathbf{h}_\text{true} by RMS 0.013, 0.026 and 0.040.

The wander costs error. The bowl’s height at the wandering taps, averaged over samples 1000 to 2000, is 0.0101, 0.0106 and 0.0115 against Jmin=0.0100J_\text{min}=0.0100. The extra error as a fraction of JminJ_\text{min} is the misadjustment: 1.2 %, 5.5 % and 14.8 % here.

For small steps the misadjustment is about μNhRx[0]/2\mu N_hR_x[0]/2, with NhN_h taps of input power Rx[0]R_x[0]. That gives 1 %, 5 % and 12 %, close for the two smaller steps. So choosing μ is a trade: halve it, and you halve the extra error but wait twice as long.

Too big a step

Steepest descent converges when every factor 1−μλ1-\mu\lambda lies between −1 and 1. That needs

0<μ<2λmax,0<\mu<\frac{2}{\lambda_\text{max}},

which is 1.25 here. LMS needs a margin below that, because single noisy steps can overshoot. On this record, μ = 0.3 is about a quarter of that limit, yet at one point LMS throws the taps about 96 away from the bottom.

There is a second trap. The limit depends on the input’s level. Turn this record’s input up three times, and every entry of R\mathbf{R} grows nine times: λmax=14.4\lambda_\text{max}=14.4, and the limit falls to 0.139. On that louder record, LMS with μ = 0.05 blows up, its taps reaching about 3500.

NLMS, normalised LMS, removes that dependence. It divides the step by the energy of the current inputs, ∥xn∥2=x[n]2+x[n−1]2\lVert\mathbf{x}_n\rVert^2=x[n]^2+{x[n-1]}^2:

hn+1=hn+μ∥xn∥2 e[n] xn.\mathbf{h}_{n+1}=\mathbf{h}_n+\frac{\mu}{\lVert\mathbf{x}_n\rVert^2}\,e[n]\,\mathbf{x}_n.

Make the input three times louder and e[n] xne[n]\,\mathbf{x}_n grows nine times, but so does ∥xn∥2\lVert\mathbf{x}_n\rVert^2: the steps do not change. NLMS is stable for μ between 0 and 2. In practice a small constant is added to ∥xn∥2\lVert\mathbf{x}_n\rVert^2, so that a nearly silent moment does not make the step huge.

The hum fades as the weights settle

Now the second use, noise cancelling. A voice is recorded with 50 Hz mains hum on top, 0.8sin⁡(2π 50t+0.6)0.8\sin(2\pi\,50t+0.6). That recording is the primary input, and it plays the part of ddes[n]d_\text{des}[n]. The hum comes from the power line, so I can take a clean copy of it there: a 50 Hz sine and its cosine, the reference input.

The filter’s two weights rebuild the hum from the reference. Here xn\mathbf{x}_n is not a row of past samples: its two entries are the sine and the cosine at sample nn. At a sample rate of 1 kHz, 50 Hz is Ω=2π⋅50/1000=0.1π\Omega=2\pi\cdot50/1000=0.1\pi, so

xn=(sin⁡(0.1πn), cos⁡(0.1πn))⊤.\mathbf{x}_n=\big(\sin(0.1\pi n),\,\cos(0.1\pi n)\big)^\top.

The filter’s output y[n]y[n] is the rebuilt hum. The error e[n]e[n], primary minus rebuild, is the cleaned voice. LMS runs as before, with μ = 0.01.

Which weights rebuild the hum exactly? Expand the sine of a sum:

0.8sin⁡(0.1πn+0.6)=0.8cos⁡0.6 sin⁡(0.1πn)+0.8sin⁡0.6 cos⁡(0.1πn).\begin{aligned} &0.8\sin(0.1\pi n+0.6)\\ &\quad=0.8\cos0.6\,\sin(0.1\pi n)\\ &\qquad+0.8\sin0.6\,\cos(0.1\pi n). \end{aligned}

So the target weights are 0.8cos⁡0.6=0.6600.8\cos0.6=0.660 and 0.8sin⁡0.6=0.4520.8\sin0.6=0.452. The voice is unrelated to the reference, so on average it does not pull the weights anywhere.

The reference’s R\mathbf{R} has 12\tfrac12 on the diagonal, because a sine squared averages to 12\tfrac12. Its off-diagonal entries are 0, because a sine times its cosine averages to 0. Both eigenvalues are 0.5, so both weights learn at one rate. Its time constant is 1/(0.01⋅0.5)=2001/(0.01\cdot0.5)=200 samples, 0.20 s.

The voice here is a stand-in: seeded noise through a resonance at 200 Hz, scaled to an RMS of 0.3. There is no recording and no sound.

The hum’s RMS is 0.8/2=0.5660.8/\sqrt2=0.566. The next instrument measures the “hum left”: the RMS of the hum still in the output, over the last 0.1 s.

The hum fades as the weights settle

Voice stand-in (seed 2631, RMS 0.300) plus 50 Hz hum (RMS 0.566); reference: a clean 50 Hz sine and cosine; μ = 0.01, 1 kHz.

The primary input: a voice buried under hum of RMS 0.566. The weights start at 0.

time
0.000 s
hum left
0.566
weights
0.000, 0.000
0.00 / 15.00 s
Describe this picture

Two stacked panels against time from 0 to 2 s, for a voice stand-in (seed 2631, RMS 0.300) plus 50 Hz hum (RMS 0.566), with a clean 50 Hz sine and cosine as the reference, μ = 0.01, at 1 kHz; it runs all 2000 samples once. The first shows the signals from −2 to 2: the primary ddes[n]d_\text{des}[n] is a faint trace, drawn in full from the start, and the output e[n]e[n] a solid trace, drawn up to the current time; a key names both. The second draws the weights h0h_0 and h1h_1, from 0 to 0.8, up to the current time, with dotted levels at 0.660 and 0.452. The line of h0h_0 is solid and ends in a dot, the line of h1h_1 dashed and ends in a square, and each is named at its end. The readouts are the time, the hum left and the weights. There is no control. At the start the voice is buried under hum of RMS 0.566, and the weights are 0. Half a second in, the weights have reached 0.593 and 0.409, and the hum left is 0.069. At two seconds the hum left is 0.011, −34.1 dB relative to the start, and the weights are 0.665 and 0.435: the output is the voice.

Watch the solid output lose its hum as the two weights climb to their dotted targets. Half a second in, the hum left is 0.069; after two seconds it is 0.011, −34.1 dB relative to the start.

Notice the weights at the end: 0.665 and 0.435, close to the hum’s sine and cosine parts, 0.660 and 0.452. The output has become the voice.

They never sit still, though. The voice is in e[n]e[n] too, and every sample of it nudges the weights a little, as the noise did in the bowl. Over the last second the hum left moves between 0.005 and 0.017; at 1 s it read 0.005, lower than at the end. That is misadjustment again.

A notch that follows the hum

Once the weights have settled, the canceller is a fixed filter from the primary input to the output. Its transfer function, worked out by Glover in 1977, is a notch at θ=0.1π\theta=0.1\pi:

1−2cos⁡θ z−1+z−21−(2−μ)cos⁡θ z−1+(1−μ) z−2.\frac{1-2\cos\theta\,z^{-1}+z^{-2}}{1-(2-\mu)\cos\theta\,z^{-1}+(1-\mu)\,z^{-2}}.

Compare it with “Poles behind the zeros narrow the notch” in Resonators, notches and combs (17.4). The zeros sit on the unit circle at ±θ\pm\theta. The poles sit just inside, at radius 1−μ=0.995\sqrt{1-\mu}=0.995 and almost the same angle.

17.4’s rule, a width of 2(1−r)2(1-r) rad/sample, gives 1.60 Hz at 1 kHz. The measured −3 dB width is 1.58 Hz. Since rr is close to 1−μ/21-\mu/2, the width is about μ\mu rad/sample: a larger μ learns faster and cuts a wider notch.

17.4’s fixed notch for an ECG, with r=0.95r=0.95, was 8.1 Hz wide. That width leaves room for the mains frequency to drift. The adaptive notch can be much narrower, because it sits wherever the reference is: if the mains drifts, the notch drifts with it.

Worked example

1. Two LMS steps by hand. Start from h0=(0,0)\mathbf{h}_0=(0,0) with μ = 0.5. Feed one input sample of 1, so x0=(1,0)\mathbf{x}_0=(1,0) and x1=(0,1)\mathbf{x}_1=(0,1). The unknown system answers ddes[0]=0.8d_\text{des}[0]=0.8 and ddes[1]=−0.4d_\text{des}[1]=-0.4, without noise.

Step 0: e[0]=0.8−0=0.8e[0]=0.8-0=0.8, and h1=(0,0)+0.5⋅0.8⋅(1,0)=(0.4,0)\mathbf{h}_1=(0,0)+0.5\cdot0.8\cdot(1,0)=(0.4,0). Step 1: e[1]=−0.4−(0.4⋅0+0⋅1)=−0.4e[1]=-0.4-(0.4\cdot0+0\cdot1)=-0.4, and h2=(0.4,0)+0.5⋅(−0.4)⋅(0,1)=(0.4,−0.2)\mathbf{h}_2=(0.4,0)+0.5\cdot(-0.4)\cdot(0,1)=(0.4,-0.2).

Each tap went half-way to (0.8,−0.4)(0.8,-0.4). Here μ∥xn∥2=0.5\mu\lVert\mathbf{x}_n\rVert^2=0.5, so each step removes half of the error in the direction of xn\mathbf{x}_n.

2. Time constants. For the bowl, μ = 0.05 and the eigenvalues 1.6 and 0.4 give 1/(0.05⋅1.6)=12.51/(0.05\cdot1.6)=12.5 and 1/(0.05⋅0.4)=50.01/(0.05\cdot0.4)=50.0 samples. The step limit is 2/1.6=1.252/1.6=1.25.

3. The canceller’s targets. The hum 0.8sin⁡(0.1πn+0.6)0.8\sin(0.1\pi n+0.6) splits into 0.660sin⁡(0.1πn)+0.452cos⁡(0.1πn)0.660\sin(0.1\pi n)+0.452\cos(0.1\pi n). The reference’s eigenvalues are both 0.5, so at μ = 0.01 the time constant is 200 samples, 0.20 s. After 0.5 s the weights are 0.593 and 0.409, and after 2 s 0.665 and 0.435.

Where you’ll meet this

Echo cancellers are system identification. In a phone call or a video conference, the far-end voice comes out of your loudspeaker, travels round the room and enters your microphone. An adaptive filter fed with the far-end voice learns that echo path and subtracts its copy of the echo, often with NLMS. Network echo cancellers for telephone lines do the same, and ITU-T Recommendation G.168 sets what they must achieve.

Noise-cancelling headphones hear the noise with an outside microphone and play its opposite, using a variant of LMS. The 1975 paper by Widrow and his colleagues, “Adaptive noise cancelling: principles and applications”, used the canceller of this page on mains hum in an ECG. It also took the mother’s heartbeat out of a fetal ECG, with chest leads as the reference.

Receivers use adaptive filters as channel equalisers, which learn to undo the blur of a channel: Channels and equalisation (33.2). When LMS learns too slowly, for instance with a large eigenvalue spread, RLS and the Kalman filter (26.4) learns faster at a higher cost.

For more, see B. Widrow and S. D. Stearns, Adaptive Signal Processing (1985), and S. Haykin, Adaptive Filter Theory (5th ed., 2014), ch. 4 to 6.

The maths behind it · condition numbers

In the eigenvector coordinates of R\mathbf{R}, the bowl separates into independent one-dimensional bowls, each shrinking by 1−μλi1-\mu\lambda_i per step. The eigenvalue spread λmax/λmin\lambda_\text{max}/\lambda_\text{min} is the condition number of R\mathbf{R}, and it sets how slowly steepest descent finishes.

The maths behind it · stochastic gradient descent

LMS is stochastic gradient descent on a squared-error loss: one example at a time, a step against that example’s gradient. The same method trains linear regression and neural networks.

Reference card

QuantityFormulaNotes
Tap inputxn=(x[n],…,x[n−Nh+1])⊤\mathbf{x}_n=(x[n],\dots,x[n-N_h+1])^\topnewest first
Error bowlJ(h)=Jmin+(h−hopt)⊤R(h−hopt)J(\mathbf{h})=J_\text{min}+(\mathbf{h}-\mathbf{h}_\text{opt})^\top\mathbf{R}(\mathbf{h}-\mathbf{h}_\text{opt})an ellipse in two taps
Steepest descenthn+1=hn+μR(hopt−hn)\mathbf{h}_{n+1}=\mathbf{h}_n+\mu\mathbf{R}(\mathbf{h}_\text{opt}-\mathbf{h}_n)needs R\mathbf{R}
LMShn+1=hn+μ e[n] xn\mathbf{h}_{n+1}=\mathbf{h}_n+\mu\,e[n]\,\mathbf{x}_nno statistics
Time constants1/(μλi)1/(\mu\lambda_i)slow along the small eigenvalue
Step limit0<μ<2/λmax0<\mu<2/\lambda_\text{max}steepest descent; LMS needs a margin
Misadjustmentabout μNhRx[0]/2\mu N_hR_x[0]/2small μ: faster means more wander
NLMSstep μ/∥xn∥2\mu/\lVert\mathbf{x}_n\rVert^2, 0<μ<20<\mu<2safe when the power changes
Hum cancellernotch, poles at radius 1−μ\sqrt{1-\mu}width about μ\mu rad/sample

End of lesson 26.3

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look