Skip to content

Audio effects

See a delay effect as a comb, a compressor as a map from level to level, and a reverb as many decaying combs.

Before this21.4 · 5 more
Chapter 29 · Lesson 2 of 3

First, the picture

Most audio effects are filters you have met, and the simplest, a delayed copy added back, makes a comb. Watch the comb’s notches below as the delay grows: they crowd closer together, but get no deeper.

Every delay is a comb

y[n] = x[n] + 0.7 x[n − D] at 16 kHz.

1 ms (16 samples): 8 notches below 8 kHz, the first at 500.0 Hz.

delay
1 ms (16 samples)
first notch
500.0 Hz
notch spacing
1000.0 Hz
Delay
0.00 / 13.00 s
Describe this picture

The delay effect y[n]=x[n]+0.7 x[n−D]y[n]=x[n]+0.7\,x[n-D] at 16 kHz, in two panels, one above the other. The first is the impulse response, against time from 0 to 320 ms, with two stems, the dry one and the delayed one. The second is the gain, from −15 to 6 dB against frequency from 0 to 8 kHz. Where the notches are closer than 2 screen pixels, the curve is drawn as a band between −10.5 dB and 4.6 dB, labelled “teeth too fine to draw”. The readouts are the delay in ms and samples, the first notch and the notch spacing, in Hz with one decimal. The 13 s clip opens at 1 ms, 16 samples: 8 notches below 8 kHz, the first at 500.0 Hz, 1000.0 Hz apart. At 3.5 s it morphs to 5 ms, 80 samples: 40 notches, the first at 100.0 Hz, 200.0 Hz apart. From 8 s it morphs to 300 ms, 4800 samples: 2400 notches below 8 kHz, the first at 1.7 Hz and 3.3 Hz apart, far too close to draw, so the band shows; the delayed stem now stands well apart. When the clip ends, three buttons in a group named “Delay”, “1 ms”, “5 ms” and “300 ms”, choose the delay; the choice is kept in the link, as comb.d. Beside them, a button “Hear it” plays 3 s of sound through the comb at the chosen delay.

Every delay is a comb

In Pitch detection (29.1) I measured a sound. On this page I change one, with filters you have met already.

I take three families: delays, which turn out to be combs; reverberation, which is many combs at once; and compressors, which turn the sound down when it gets loud.

Every sound on this page is synthetic, made from a stated formula: there are no recordings.

A delayed copy

Start with the simplest effect. Delay a copy of the sound by DD samples, scale it by a gain aa, and add it back:

y[n]=x[n]+a x[n−D].y[n]=x[n]+a\,x[n-D].

This is the feed-forward comb at the end of “A delay in a loop grows teeth” in Resonators, notches and combs (17.4). In a real-time effect, the delayed copy comes from a buffer of DD slots, as in “A circular buffer: the write head goes round” of Real-time processing (21.4).

Its frequency response is H(ejΩ)=1+a e−jΩDH(e^{j\Omega})=1+a\,e^{-j\Omega D}. Where e−jΩD=1e^{-j\Omega D}=1 the copy adds to the dry sound, and where e−jΩD=−1e^{-j\Omega D}=-1 it works against it. So the gain swings between 1+a1+a and 1−a1-a.

I use a=0.7a=0.7 at fs=16f_s=16 kHz throughout. The gain then swings between 1.7 and 0.3, which is 4.6 dB and −10.5 dB, in the decibels of “Signal and noise, in decibels” in How big is a signal (1.3).

The notches sit where ΩD\Omega D is an odd number of half turns, ΩD=(2k+1)π\Omega D=(2k+1)\pi. With Ω=2πf/fs\Omega=2\pi f/f_s, that is at

f=(2k+1) fs2D,k=0,1,2,…f=\frac{(2k+1)\,f_s}{2D},\qquad k=0,1,2,\dots

The first notch is at fs/(2D)f_s/(2D), and after it they come every fs/Df_s/D. A longer delay packs the notches closer together. It does not make them deeper.

Colour or echo

One formula covers two very different sounds. With a short delay the notches are far apart, and you hear a hollow colour, like a voice in a tiled corridor. With a long one you hear the copy as a separate echo, like a shout across a valley.

The picture at the top of the page shows three delays: 1 ms, 5 ms and 300 ms. Its “Hear it” button plays 3 s of sound through the comb at the chosen delay. The sound is six bursts of noise, half a second apart. Burst kk is 0.5 e−(n−nk)/480 v[n]0.5\,e^{-(n-n_k)/480}\,v[n] for 150 ms from nk=8000kn_k=8000k, where v[n]v[n] is the seeded noise of the site with seed 292. Each burst fades by a factor ee every 30 ms, and the output is scaled so that its largest sample is 1.

Watch the readouts as the delay grows. The notch spacing falls from 1000.0 Hz to 200.0 Hz to 3.3 Hz, while the gain still swings between the same 4.6 dB and −10.5 dB.

Combs that move

Now let the delay change while the sound plays. If it sweeps slowly between 1 and 5 ms, the first notch slides between 500 Hz and 100 Hz, and the other teeth move with it. This is the flanger, which 17.4 called flanging.

A swept delay passes through values of DD between whole samples. The effect then needs a fractional delay, as in “Fractional delay” of All-pass systems (17.1).

A chorus uses somewhat longer delays that wander slowly and irregularly, often with several copies. A changing delay also shifts the pitch a little. So the copies drift against the dry sound, and one voice sounds like a few singing together.

A phaser keeps the idea but replaces the delay by a chain of all-passes. The first-order all-pass of 17.1 has a gain of 1 at every frequency. Its phase turns from 0 at 0 Hz to half a turn at fs/2f_s/2.

Four of them in a row turn the phase by up to two whole turns. Added to the dry sound, the chain’s output cancels it where its phase is an odd number of half turns. With each stage’s pole at 0.5 and fs=16f_s=16 kHz, that gives two notches, near 699 Hz and 3451 Hz. They are not evenly spaced, unlike a comb’s, and moving the poles slides them.

Loud notes turned down by a ratio

The second family changes the level, not the spectrum. A radio announcer leans into the microphone, then steps back, and the voice jumps by many decibels. A compressor measures the level as the sound goes and turns it down when it gets loud.

Levels in dBFS

I give levels in dBFS, as in “Signal and noise, in decibels” (1.3): LdBFS=20log⁡10(x/xFS)L_\text{dBFS}=20\log_{10}(x/x_\text{FS}), with the full scale xFS=1x_\text{FS}=1. On this page xx is the peak of the sound. So a note whose peak is 0.501 sits at −6.0 dBFS.

Measuring the level

First the compressor needs the level at each moment. It takes ∣x[n]∣\lvert x[n]\rvert and smooths it with the one-pole smoother of Simple smoothing filters (18.2). I call its constant cc, because aa is already the comb’s gain on this page:

xdet[n]=c xdet[n−1]+(1−c) ∣x[n]∣.\begin{aligned} x_\text{det}[n]&=c\,x_\text{det}[n-1]\\ &\quad+(1-c)\,\lvert x[n]\rvert. \end{aligned}

In “Choosing a” (18.2) the constant came from a time constant τ\tau, as e−Ts/τe^{-T_s/\tau}. The detector gets two time constants. While ∣x[n]∣\lvert x[n]\rvert is above xdet[n−1]x_\text{det}[n-1] it uses the attack time, and otherwise the release time.

With an attack of 1 ms and a release of 100 ms at 16 kHz, Ts=62.5T_s=62.5 µs. Then c=e−0.0625=0.9394c=e^{-0.0625}=0.9394 on the way up and c=e−0.000625=0.999375c=e^{-0.000625}=0.999375 on the way down. The detector follows a rise quickly and lets go slowly.

The gain rule

Then a rule turns the detector’s level, in dBFS, into a gain. Below the threshold LthrL_\text{thr}, nothing changes. Above it, each dB more going in gives only a fraction of a dB more coming out.

A ratio of 4:1 means 4 dB more in, 1 dB more out. Above the threshold, in dBFS,

out=Lthr+in−Lthr4.\text{out}=L_\text{thr}+\frac{\text{in}-L_\text{thr}}{4}.

Plotted as out against in, this is the compressor’s static curve: the diagonal below the threshold, and a line a quarter as steep above it. With Lthr=−20L_\text{thr}=-20 dBFS, a note 8 dB over the threshold comes out 2 dB over it, at −18.0 dBFS.

To get there, the compressor applies a gain of −(in−Lthr)(1−1/4)-(\text{in}-L_\text{thr})(1-1/4) dB to the sound, here −6.0 dB. This page adds no gain afterwards to bring the level back up, so loud notes only come down.

Let’s give it six notes at 440 Hz. Each rises over 10 ms, holds for 200 ms and falls over 10 ms, and a new one starts every 0.5 s.

Loud notes turned down by a ratio

Six synthetic 440 Hz notes at −6, −30, −12, −24, −6 and −18 dBFS; threshold −20 dBFS, ratio 4:1, attack 1 ms, release 100 ms.

Six notes, from −30 to −6 dBFS, and a 4:1 curve that bends at −20.

note
—
in
—
out
—
0.00 / 13.00 s
Describe this picture

Six synthetic 440 Hz notes at −6, −30, −12, −24, −6 and −18 dBFS through a compressor with threshold −20 dBFS, ratio 4:1, attack 1 ms and release 100 ms. Two panels. The first plots the level, from −40 to 0 dBFS, against time from 0 to 3 s; levels below −40 are drawn on the floor. The input level is dashed, the output level solid, and a dotted line marks the threshold at −20. The second, square, plots output against input, both from −40 to 0 dBFS: the static curve as a solid line, the diagonal dotted (“no change”), and a filled circle for each note. The readouts are the note, and its level in and out in dBFS. There is no control. The 13 s clip opens on the input level and the curve: six notes, from −30 to −6 dBFS, and a 4:1 curve that bends at −20. From 3 s the output level draws in, and a circle lights on the curve as each note passes. From 7 s it holds on the third note, picked out on both panels: the −12 dBFS note, 8 dB over the threshold, comes out at −17.8 dBFS, near the curve’s −18.0 dBFS. From 9 s two brackets on the curve panel mark how far apart the inputs and the outputs lie: notes below −20 pass unchanged, the loud ones come down to −16.3 dBFS, and the range shrinks from 24.0 dB to 13.7 dB.

Watch the loud notes come down while the quiet ones pass unchanged. The range between the loudest and the quietest note shrinks from 24.0 dB to 13.7 dB.

Notice the two −6 dBFS notes. The curve promises −16.5 dBFS, but they come out at −16.3. Why the 0.2 dB?

The detector smooths ∣x∣\lvert x\rvert, which falls to 0 twice in every cycle of the tone. So on a steady note it settles between 0.26 and 0.30 dB below the true peak, with a ripple under 0.04 dB. The compressor takes the note to be a little quieter than it is, and cuts three quarters of that, about 0.2 dB, too little.

The −30 and −24 dBFS notes are not touched at all. After a −6 dBFS note ends, the detector falls back below −20 dBFS within 0.16 s, well before the next note starts.

The attack shows at the start of each loud note. While the note is still rising, the detector lags behind it and the gain is too high. So the first 20 ms of each −6 dBFS note peak at −14.5 dBFS, 1.8 dB above where they settle.

Limiters and gates

A limiter is a compressor with a very high ratio. At 20:1, the static curve would bring the −6 dBFS note to −19.3 dBFS, only 0.7 dB over the threshold. Limiters guard the 0 dBFS ceiling of 1.3, to keep the peaks from being clipped.

A gate works the other way round. Below its threshold it turns the sound down hard, or off. That silences the hiss between notes or words.

Meters and units

A peak meter shows the largest ∣x∣\lvert x\rvert in dBFS, as the levels on this page do. An RMS meter shows the RMS over the last fraction of a second instead. For a steady sine the RMS reading is 3.0 dB below the peak, as 1.3 showed.

dBFS is a digital unit. In the air, sound is measured in dB SPL, 20log⁡10(prms/20 μPa)20\log_{10}(p_\text{rms}/20\,\mu\text{Pa}), where prmsp_\text{rms} is the RMS pressure. A pressure of 1 Pa RMS is 94.0 dB SPL.

Electrical levels are often given in dBm, 10log⁡10(P/1 mW)10\log_{10}(P/1\,\text{mW}), a power against one milliwatt. For example, 0 dBm in a 600 Ω600\,\Omega line is 0.775 V RMS.

A reverb is many decaying combs

Clap your hands in a large hall. You hear the clap, then a dense wash of reflections that fades over a second or two. That is reverberation. Its length is the reverberation time, T60T_{60} or RT60: the time the sound’s energy takes to fall by 60 dB.

Schroeder’s reverb

In 1962 Manfred Schroeder built a reverb from the parts on this page. Its core is the comb with a loop from 17.4, with the input delayed too:

y[n]=x[n−D]+g y[n−D].y[n]=x[n-D]+g\,y[n-D].

Each echo is gg times the one before, and they come every DD samples. Schroeder added up four such combs with different delays, so their echo trains interleave.

How large should gg be? Each trip round the loop takes D/fsD/f_s seconds and multiplies the echo by gg. In T60T_{60} seconds the sound goes round T60fs/DT_{60}f_s/D times, and by then it should be 10−310^{-3} of its start, 60 dB down:

g T60fs/D=10−3,g=10−3D/(T60fs).\begin{aligned} g^{\,T_{60}f_s/D}&=10^{-3},\\ g&=10^{-3D/(T_{60}f_s)}. \end{aligned}

A longer delay needs a smaller gg, because it makes fewer trips. For T60=1.5T_{60}=1.5 s at 16 kHz, the delays 475, 594, 658 and 699 samples, 29.7 to 43.7 ms, get the gains 0.872, 0.843, 0.827 and 0.818.

The four combs are averaged, then passed through two all-passes, from “Echoes with a flat gain” in 17.1. These multiply the echoes without colouring the sound. Both have gain 0.7, with delays of 80 and 27 samples, 5.0 and 1.7 ms.

Fig. Measured decay time (60 dB, from the −5 to −35 dB slope): 1.50 s; 741 of the first 1600 samples are already nonzero, 619 of them larger than 10⁻⁴.
Describe this picture

Two panels. The first is the impulse response h[n]h[n] of Schroeder’s reverb over its first 0.5 s: nothing for the first 475 samples, 29.7 ms, then echoes that pile up and fade. The second is its energy decay in dB, from −80 to 0, over 2 s: a nearly straight fall, through −5 dB at 0.150 s and −35 dB at 0.901 s, labelled “30 dB in 0.751 s”. The caption gives the measured decay time, 1.50 s, and says that 741 of the first 1600 samples are already nonzero, 619 of them larger than 10⁻⁴.

Look at the energy decay. The energy decay at sample nn is the energy still to come, the sum of h[m]2h[m]^2 for m≥nm\ge n, in dB against the total.

Its measured decay time is 1.50 s, from the slope between −5 and −35 dB. The curve passes −5 dB at 0.150 s and −35 dB at 0.901 s. That is 30 dB in 0.751 s, so 60 dB takes 1.50 s, as the gains were set for.

Nothing arrives for the first 475 samples, 29.7 ms, the shortest comb’s delay. After that the echoes pile up fast. Of the first 1600 samples, the first 100 ms, 741 are already nonzero, and 619 of them are larger than 10−410^{-4}.

Colour, echo and reverb

Now press “Hear it” on the first instrument. At 1 and 5 ms the bursts sound hollow: a colour, not a repeat. Sweeping the delay between those two would be a flanger. At 300 ms you hear each burst twice, as a separate echo.

Schroeder’s combs sit in between, at 30 to 44 ms. With feedback and all-passes, their echoes come too densely to hear one by one, and they blend into a tail.

Two later designs grew from this. A feedback delay network feeds several delay lines back into each other through a mixing matrix, so every echo spreads into all of them.

A convolution reverb skips the model. It convolves the dry sound with a room’s impulse response, measured or synthetic, as in “Naming what you already built” of Discrete convolution (5.2). A response of 1.5 s at 16 kHz is 24 000 samples long, so these reverbs use Fast convolution (14.3).

The maths behind it · orthogonal matrices

A feedback delay network is a vector of delay lines, mixed by an orthogonal matrix each time round. An orthogonal matrix keeps a vector’s length, so the mixing neither adds energy nor loses it. The decay comes only from the gains.

The maths behind it · survival functions

A dense reverb tail looks like noise with a decaying envelope. Its energy decay curve is a running sum from the end, built the same way as a survival function, which counts what is still to come.

Worked example

1. Notches of a 5 ms delay. At 16 kHz, 5 ms is D=80D=80 samples. The first notch is at 16 000/160=100.016\,000/160=100.0 Hz, and then they come every 16 000/80=200.016\,000/80=200.0 Hz. Below 8 kHz there are 40 of them.

2. The static curve. A note at −6 dBFS is 14 dB over the threshold of −20 dBFS. At 4:1 it comes out at −20+14/4=−16.5-20+14/4=-16.5 dBFS. The compressor’s gain is −14×(1−1/4)=−10.5-14\times(1-1/4)=-10.5 dB.

3. An attack constant. An attack of 1 ms at 16 kHz gives Ts/τ=0.0625T_s/\tau=0.0625, so c=e−0.0625=0.9394c=e^{-0.0625}=0.9394.

4. A reverb comb’s gain. For D=475D=475 samples and T60=1.5T_{60}=1.5 s at 16 kHz, 3D/(T60fs)=1425/24 000=0.05943D/(T_{60}f_s)=1425/24\,000=0.0594, so g=10−0.0594=0.872g=10^{-0.0594}=0.872. Each trip round the loop loses 1.19 dB, and in 1.5 s the echo makes 50.5 trips: 60 dB in all.

Where you’ll meet this

Digital audio workstations offer all of these effects: delays, flangers, choruses, phasers, reverbs and compressors. A guitar pedal usually holds one or two of them. Equalisers, the other everyday effect, are the biquads of Audio equalisers and biquads (20.5).

Broadcasters compress and limit their programmes to keep the loudness even. Hearing aids use fast compressors to fit loud and quiet sounds into a smaller range of hearing. Many modern reverbs use a feedback delay network in place of Schroeder’s four combs.

A loud sound can hide a quieter one close to it in frequency, and audio coders use that. Perceptual audio coding (29.3) takes it up.

I left out the exact shapes of the slow sweeps in flangers and choruses, look-ahead limiting, compressors that split the sound into bands, and loudness units (LUFS). For more, see U. Zölzer (ed.), DAFX (2nd ed., 2011), chapters 2, 4 and 5; J. D. Reiss and A. McPherson, Audio Effects (2015); and M. R. Schroeder, “Natural sounding artificial reverberation” (J. Audio Eng. Soc., 1962).

Reference card

QuantityFormulaNotes
Delay effecty[n]=x[n]+a x[n−D]y[n]=x[n]+a\,x[n-D]notches at (2k+1)fs/(2D)(2k+1)f_s/(2D), every fs/Df_s/D
Gain range1−a1-a to 1+a1+aa=0.7a=0.7: −10.5 dB to 4.6 dB
Flanger, chorusa swept or wandering DDcombs that move
Phaserdry sound plus a chain of all-passesnotches where the phase is an odd number of half turns
Reverb comby[n]=x[n−D]+g y[n−D]y[n]=x[n-D]+g\,y[n-D]Schroeder: four combs, then two all-passes
Comb gaing=10−3D/(T60fs)g=10^{-3D/(T_{60}f_s)}T60T_{60}: 60 dB of energy decay
Level detectorxdet[n]=c xdet[n−1]+(1−c)∣x[n]∣x_\text{det}[n]=c\,x_\text{det}[n-1]+(1-c)\lvert x[n]\rvertc=e−Ts/τc=e^{-T_s/\tau}, attack or release
Compressorout =Lthr+(in−Lthr)/ratio=L_\text{thr}+(\text{in}-L_\text{thr})/\text{ratio} above LthrL_\text{thr}levels in dBFS; limiter: a very high ratio
dBFS20log⁡10(xpeak/xFS)20\log_{10}(x_\text{peak}/x_\text{FS})0 dBFS is full scale
dB SPL20log⁡10(prms/20 μPa)20\log_{10}(p_\text{rms}/20\,\mu\text{Pa})1 Pa is 94.0 dB SPL
dBm10log⁡10(P/1 mW)10\log_{10}(P/1\,\text{mW})0 dBm is 1 mW

End of lesson 29.2

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look