Hann windows placed one every hop, and their sum. Watch the dashed sum as the windows slide closer: at some hops it goes flat.
Windows that add up to a constant
Hann windows of N = 64 samples, one every hop. Their sum is what overlap-add multiplies the signal by.
Hop 64, no overlap: the windows meet only at their zero ends. Their sum falls to 0 between frames, so overlap-add would silence the signal there.
Describe this picture
Hann windows of samples, one every hop; their sum is what overlap-add multiplies the signal by. One panel, with sample from 0 to 320 and value from 0 to 2.5. Each shifted window is a thin line, neighbours alternating between solid and dotted, and their sum is a thick dashed line labelled “sum of windows”. When the sum is not flat, the gaps between it and its maximum are hatched. The readouts are the hop and the sum.
The clip plays once, in 15 s, and holds on its last frame; the caption is blank while the windows move. At hop 64, no overlap (“64 samples, 0 % overlap”, “0.000 to 1.000”): “Hop 64, no overlap: the windows meet only at their zero ends. Their sum falls to 0 between frames, so overlap-add would silence the signal there.” At hop 48 (“48 samples, 25 % overlap”, “0.293 to 1.000”): “Hop 48: the sum ripples between 0.293 and 1.000. Overlap-add would leave a wobble every 48 samples.” At hop 32 (“32 samples, 50 % overlap”, “flat at 1.000”): “Hop 32, half overlap: the sum is exactly 1 everywhere, so adding the frames back gives x itself. This is the COLA condition: constant overlap-add.” The clip ends at hop 16 (“16 samples, 75 % overlap”, “flat at 2.000”): “Hop 16: flat again, at 2; divide by 2 and x comes back. A Hann window is COLA whenever N/hop is a whole number of 2 or more.”
When the clip has finished, a control named “Hop”, with value text such as “32 samples, 50 % overlap”, sets the hop from 8 to 64 samples; the hint reads “Drag across the plot, or use the arrow keys, to set the hop (8 to 64).” At hop 24 the caption reads “Hop 24: the sum ripples between 1.293 and 1.383: not COLA.” At hop 8 it reads “Hop 8: flat at 4.000: COLA.”
Adding the frames back
Spectrograms & the STFT (15.5) cut a signal into overlapping frames, multiplied each by a window , and took the DFT of each one. The result is a grid of cells , one for frame and bin , with the frames samples apart. This lesson goes the other way. Resynthesis is the making of a signal from those cells.
The first step is the inverse DFT of The DFT (13.2). Applied to frame it returns the windowed piece that was cut out, for . Frame is centred on sample , so it starts samples before it.
The second step is overlap-add of the frames. I write every piece back at the place it came from and add where pieces overlap, as the blocks of Fast convolution (14.3) were added. A sample of lies in several pieces, and each piece holds it multiplied by the window’s value at that spot. So the sum is
Here is the window placed on frame , and it is 0 outside its samples. I leave out the shift of from a frame’s start to its centre; it moves the whole sum along and does not change its shape. The signal comes back multiplied by the sum of the shifted windows, and everything below rests on that sum. Every window on this page is the periodic Hann window of Windowing & spectral leakage (15.1).
Windows that add up to a constant
The picture at the top of this page adds Hann windows of samples, one every hop. At hop 64, no overlap, the windows meet only at their zero ends: their sum falls to 0 between frames, so overlap-add would silence the signal there. At hop 48 the sum ripples between 0.293 and 1.000, and overlap-add would leave a wobble every 48 samples. At hop 32, half overlap, the sum is exactly 1 everywhere, so adding the frames back gives itself. At hop 16 it is flat again, at 2: divide by 2 and comes back.
When the clip has finished, drag across the plot to set the hop yourself, from 8 to 64. At hop 24 the sum ripples between 1.293 and 1.383, and at hop 8 it is flat at 4.000.
The condition is COLA, constant overlap-add: the sum is the same constant at every . Then , and dividing by gives back. For Hann the constant comes from a short sum. Put and suppose it is a whole number. The windows covering a sample are apart, and their sum is
The cosines are arrows spaced equally round a circle, and for they add to 0, as the roots of unity did in Complex numbers for signals (3.3). So : 1 at hop 32, 2 at hop 16 and 4 at hop 8, as the picture shows. At hop 64, and the cosine survives, so the sum is , which falls to 0 and rises to 1. Hops 48 and 24 give of and , which are not whole, and the arrows do not close.
SciPy asks the same question of a window: check_COLA('hann', 64, 64 - hop) is true for hops 32, 16 and 8 and false for 64, 48 and 24. The window must be periodic, as here. A symmetric Hann, with in the cosine as in NumPy’s np.hanning, is not quite COLA at hop 32: its sum ripples from 0.975 to 0.999.
Weighted overlap-add
SciPy’s istft does a little more than the sum above. It multiplies each frame by the window a second time, adds the frames, and divides by . This is weighted overlap-add. For frames that were not edited, the numerator is , so comes back whenever the divisor is not 0. That weaker requirement is NOLA, nonzero overlap-add: the sum of squared windows is never 0. Hann at hops 48 and 24 passes NOLA but not COLA.
Take a signal of 1 s at 8 kHz that holds a 440 Hz cosine of amplitude 0.5, and a 2500 Hz cosine of amplitude 0.3 from 0.3 s to 0.7 s. It is the signal of the next picture. I take its STFT with Hann and invert it with istft. At hop 128 the largest error is , and at hop 64 it is : rounding only. At hop 256 the windows are 0 where they meet, NOLA fails, and the error is 0.8. That is the signal itself, at sample 3200, where it peaks at 0.8 and the output is 0.
Erase a whistle, keep the tone
Once a spectrogram can be turned back into sound, its cells can be edited in between. A mask is a number from 0 to 1 for each cell, and spectral masking multiplies the cell by it before resynthesis:
Erasing is the simplest mask: inside a box and 1 everywhere else. What the erased cells held is gone, and every other cell is as before. It is like painting over one word on a page.
Watch the whistle’s readout as the box slides up over its cells, while the tone’s stays at 0.0 dB.
Erase a whistle, keep the tone
1 s at 8 kHz: a 440 Hz tone (amplitude 0.5) and a 2500 Hz whistle (0.3) from 0.3 to 0.7 s. Hann N = 256, hop 64. Cells inside the box are set to 0, then the frames are added back.
The eraser sits on empty cells, between the tone and the whistle. Nothing changes: whistle 0.0 dB, tone 0.0 dB.
Describe this picture
1 s at 8 kHz: a 440 Hz tone (amplitude 0.5) and a 2500 Hz whistle (0.3) from 0.3 to 0.7 s, with Hann and hop 64; cells inside the box are set to 0, then the frames are added back. Two stacked panels. The first is a spectrogram, with time from 0 to 1 s and frequency from 0 to 4000 Hz; the level runs from 0 to −60 dB re amplitude 1 (15.4 §1), and a scale bar names it. The eraser is a hatched rectangle labelled “eraser”, from 0.25 to 0.75 s and 400 Hz tall, and the cells it has erased are drawn as paper. The second panel is the resynthesised waveform, from −1 to 1, against time. The readouts are the whistle left, its amplitude in the output in dB re 0.3, and the tone kept, in dB re 0.5, both to 1 decimal; “gone” means below −60 dB. Two buttons, “Hear before” and “Hear after”, play the input and the current output.
The clip plays once, in 13 s, and holds on its last frame. The box starts centred at 1500 Hz: “The eraser sits on empty cells, between the tone and the whistle. Nothing changes: whistle 0.0 dB, tone 0.0 dB.” It then slides up in frequency while the cells under it blank, and the waveform and readouts change when it stops. At the middle of the clip (−15.6 dB, 0.0 dB): “Box from 2100 to 2500 Hz: it covers the lower half of the whistle’s cells. The whistle drops by 15.6 dB; what is left comes from the cells just above 2500 Hz.” At the end (“gone”, 0.0 dB): “Box from 2300 to 2700 Hz, over the whole whistle: it is gone, and the tone is untouched. The frames add back to the tone alone.”
When the clip has finished, a control named “Eraser band”, with value text “2300 to 2700 Hz”, moves the box; dragging it up or down, or the arrow keys, set its centre from 200 to 3800 Hz in steps of 50 Hz, and the hint reads “Drag the eraser, or use the arrow keys, to choose the band it removes.” Resynthesis runs when the drag settles. At 500 Hz the caption reads “Box 300 to 700 Hz: whistle 0.0 dB, tone gone.”
The box starts on empty cells, between the tone and the whistle, and nothing changes. Over 2100 to 2500 Hz it covers the lower half of the whistle’s cells, and the whistle drops by 15.6 dB: what is left comes from the cells just above 2500 Hz. Over 2300 to 2700 Hz, the whole whistle, the whistle is gone and the tone is untouched. The frames add back to the tone alone. Press “Hear after” at each stop to hear it.
The box covers the 13 bins within 200 Hz of its centre, 31.25 Hz apart, and the 62 frames whose centres lie from 0.25 s to 0.75 s, which is 806 cells. The whistle sits exactly on bin 80, at 2500 Hz. A Hann window spreads such a tone over bins 79, 80 and 81 (15.1). The box from 2100 to 2500 Hz ends at bin 80, so bin 81 survives. What it holds is 15.6 dB below the whistle, and not 0.
When the clip has finished, drag the box to choose the band it removes. Over the tone, from 300 to 700 Hz, the whistle stays and the tone is gone: at a box centre of 600 Hz the tone is 80.2 dB down.
An edit also has a limit. After it, the cells are no longer the STFT of any signal, because neighbouring frames share samples and now disagree about them. Weighted overlap-add still returns a signal, and the Linear Algebra note below says which one.
The maths behind it · the pseudo-inverse
The STFT is a tall matrix (15.5), and weighted overlap-add is its pseudo-inverse . Editing cells gives a set of numbers that is no exact STFT. The pseudo-inverse returns the signal whose STFT is nearest, in least squares (Griffin and Lim, 1984).
Keep the loud cells
A mask can also be chosen by the data. Noise spreads thinly over every cell, and a tone gathers in a few. So keep the cells that are loud, and set the rest to 0: threshold denoising.
Take a 440 Hz cosine of amplitude 0.5 in white noise of RMS 0.2, 1 s at 8 kHz, with the same Hann and hop 64 as before, and 129 bins by 126 frames, 16 254 cells. I measure the SNR against the clean tone, as of the tone’s energy over the energy of the error. The noisy signal has 4.9 dB. The median cell sits at −31.7 dB re amplitude 1. I keep only the cells more than 10 dB above that median, which is above −21.7 dB.
Of the 397 cells that pass the 10 dB threshold, 380 are in the four bins of the tone, and 17 are noise. For noise, a cell 10 dB above the median is rare, about 1 in 1024, so about 15 of the 15 750 cells outside the tone’s bins pass by chance. The kept cells hold the tone almost whole, and the SNR rises from 4.9 dB to 21.0 dB.
The cost has two forms, and the threshold sets the balance. Too low and noise cells pass: at 6 dB the page keeps 8.0 % of the cells and the SNR is only 13.3 dB. Too high and the tone loses its cells: at 20 dB only 225 of the 504 cells in the tone’s four bins survive, and the SNR falls to 12.9 dB. The surviving isolated noise cells, each a short burst, make the musical noise of the caption.
The maths behind it · shrinkage estimators
Threshold denoising is a hard shrinkage estimator: a coefficient is kept whole or set to 0. It is the same idea as wavelet thresholding in Wavelets (23.2) and as setting small regression coefficients to 0. The Wiener filter of The Wiener filter (26.2) is its soft, optimal version.
Twice as long, same pitch
The last edit is not a mask. Write the frames further apart than they were read, and the sound lasts longer. Read frames every samples, the analysis hop, and write them every , the synthesis hop. That stretches time by .
Frames moved apart do not join up. A tone of rad/sample turns between neighbouring input frames, but in the longer output it must turn between neighbouring output frames. A 440 Hz tone at 8 kHz turns rad between input frames, and needs 44.23 rad between output frames. A frame written with its old phase restarts the tone at the wrong point of its cycle.
The fix needs each bin’s frequency, and the phases carry it. Bin has its own frequency . Between frames and the bin’s phase should turn . Whatever it turns beyond that is the extra turn
brought into by adding or subtracting whole turns of , as in the unwrapping of The DTFT (12.2 §3). The bin’s true frequency is , and the output phase advances by that frequency times :
The magnitudes stay as they are. This is the phase vocoder. The wrap works while the tone is within rad/sample of the bin, which is Hz here, two bins.
For the 440 Hz tone, the nearest bin is 14, at 437.5 Hz, with rad. The measured turn is 22.12 rad, so rad. The extra frequency is rad/sample, which is Hz, so the true frequency is Hz.
Watch the middle strip swell and shrink where each frame restarts the tone, and the last strip stay smooth.
Twice as long, same pitch
0.5 s of a 440 Hz tone at 8 kHz, Hann N = 256. Frames read every 64 samples and written every 128.
Describe this picture
0.5 s of a 440 Hz tone at 8 kHz, Hann ; frames are read every 64 samples and written every 128. Three stacked strips. The first, “input”, shows time from 100 to 140 ms. The second, “frames moved apart, phases unchanged”, and the third, “phase vocoder”, show time from 200 to 280 ms, with tick marks at the frame joins every 16 ms. The readouts are the output length and the two strongest spectral peaks of the strip just drawn. Three buttons, “Hear the input”, “Hear frames moved” and “Hear phase vocoder”, play the strips; there is no other control.
The clip plays once, in 14 s, and holds its last frame. The input draws first, and holds with the caption “The input: a 440 Hz tone, one cycle every 2.27 ms.” (0.50 s, 440.0 Hz). Then the middle strip draws, with the caption blank, and holds with “Frames written twice as far apart, phases unchanged: each frame restarts the tone at the wrong point of its cycle, so neighbours partly cancel. The level wobbles by 2.6 dB, and the tone splits into 470 and 407.5 Hz.” (0.96 s, “470.0 and 407.5 Hz”). Then the last strip draws, and ends with “Phase vocoder: each bin’s phase is moved on by its frequency times the new hop, so every frame continues the last. Smooth (0.2 dB wobble), nearly twice as long, and still 440 Hz.” (0.96 s, “440.0 Hz”).
The input is a 440 Hz tone, one cycle every 2.27 ms. Written twice as far apart with their phases unchanged, the frames each restart the tone at the wrong point of its cycle, so neighbours partly cancel. The level wobbles by 2.6 dB, and the tone splits into 470 and 407.5 Hz: press “Hear frames moved” to hear what the wobble sounds like. The phase vocoder moves each bin’s phase on by its frequency times the new hop, so every frame continues the last. Its output is smooth, with a 0.2 dB wobble, lasts 0.96 s, nearly twice as long, and is still 440 Hz.
Two numbers are worth working out. The output has 7680 samples, 0.96 s, not 1 s: 59 frames fit the 4000 input samples, and the last one stops at sample 3968. Writing them 128 apart gives . And the split is no accident. Per input hop the tone turns 3.52 times, and per output hop it should turn 7.04. The unadjusted frames start each at the phase they had in the input, so each start is turns off, which is 0.48 turn once whole turns are dropped. Neighbouring frames are nearly opposite, and partly cancel. The output frame rate is Hz, and the leftover turn per frame is a shift of that rate: Hz and Hz.
Pitch by stretching
A stretch gives a pitch shift. Stretch by 2, then play the result at twice the rate: low-pass it and keep every second sample, as in Downsampling and decimation (22.1). In SciPy that is resample_poly(y, 1, 2). The result lasts 0.48 s, close to the original again, and its tone is 880 Hz, an octave up. In general, stretch by and then resample by : the length stays about the same and the pitch is multiplied by .
A pure tone is the easy case. Real sounds with attacks need more care: a sharp onset smears across the frame, and phase locking keeps the phases of a spectral peak’s neighbouring bins together. Griffin and Lim (1984) go further and rebuild phases when only the magnitudes are known. I leave all three out.
Worked example
- Hann sums, , steady part. Hop 64: 0 to 1. Hop 48: 0.293 to 1.000. Hop 32: 1. Hop 24: 1.293 to 1.383. Hop 16: 2. Hop 8: 4. The constants 1, 2 and 4 are : , and .
- Round trip, Hann 256. Hop 128: largest error . Hop 64: . Hop 256: 0.8, because NOLA fails.
- Eraser, box centre and what is left of the whistle. 1500 Hz: 0.0 dB. 2300 Hz: dB. 2400 and 2500 Hz: gone, below dB. The tone stays at 0.0 dB in all of these, and a box at 600 Hz takes it down by 80.2 dB.
- Threshold denoising, noise RMS 0.2. The input SNR is 4.9 dB. A threshold 6 dB above the median cell keeps 8.0 % of the cells and gives 13.3 dB. At 10 dB: 2.4 % and 21.0 dB. At 15 dB: 2.3 % and 21.2 dB. At 20 dB: 1.4 % and 12.9 dB.
- Phase vocoder, 440 Hz at 8 kHz, , hop 64 to 128. The phase turn is 22.12 rad per input hop and 44.23 rad per output hop. For bin 14 at 437.5 Hz, rad. Without the fix the 10 ms RMS runs from 0.254 to 0.343, a wobble of 2.6 dB, and the peaks are 470.0 and 407.5 Hz. With it, 0.350 to 0.357, 0.2 dB, and 440.0 Hz.
- Pitch shift, by stretching 2 and decimating by 2. The result lasts 0.48 s, and its tone is 880 Hz.
Where you’ll meet this
Every audio plug-in that stretches, retunes or cleans a recording does some form of this: an STFT, an edit of the cells, and a weighted overlap-add. The spectral denoiser is the simplest. Its thresholds become soft gains in The Wiener filter (26.2), and hard thresholds on wavelet coefficients appear in Wavelets (23.2). Keeping every second sample, in the pitch shift above, is Downsampling and decimation (22.1); Resampling by any factor (22.3) covers ratios that are not whole.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Overlap-add | placed at | gives |
| COLA | for every | Hann: whole, at least 2; |
| Weighted OLA | window again, divide by | SciPy istft; needs NOLA |
| Mask | , | erase: in a box |
| Denoise | where exceeds a threshold | musical noise |
| Vocoder phase | into | true frequency |
| Output phase | advance by | stretch |
| Pitch shift | stretch by , then resample by | same length, pitch |