Skip to content

Speech features

Pool a speech frame into mel bands, take logs and a DCT, and see why these MFCCs follow the vowel more than the pitch.

Before this13.6 · 23.1 · 30.1 · 2 more
Chapter 30 · Lesson 2 of 2

First, the picture

A speech recogniser wants a dozen numbers per slice of sound that follow the vowel and mostly ignore the pitch. Watch one slice below pooled into bands that widen as they climb, then squeezed to 13 numbers.

Bands as wide as hearing's

One 50 ms frame of 30.1's synthetic /a/ at 125 Hz, 8 kHz; 26 mel triangles.

The frame's spectrum: 125 Hz harmonics under the /a/ envelope.

stage
spectrum
bands
—
coefficients kept
—
0.00 / 13.00 s
Describe this picture

One 50 ms frame of 30.1’s synthetic /a/ at 125 Hz, 8 kHz, and 26 mel triangles, in three stacked panels. The first shows the frame’s power spectrum as a solid line, in dB relative to its peak, from −60 to 5 dB against frequency from 0 to 4000 Hz; values below −60 dB are drawn on the floor. The second, on the same frequency axis, draws the 26 triangles as thin lines, alternately solid and dashed, from 0 to 1.05. The third has two parts: ln⁡Pm\ln P_m as bars, from −10 to 8, for bands mm from 1 to 26, and the 13 DCT numbers as stems, ckc_k from −16 to 18, for kk from 0 to 12. The readouts are the stage, the bands and the coefficients kept. There is no control. The 13 s clip opens on the frame’s spectrum: 125 Hz harmonics under the /a/ envelope. At 3 s the triangles fan in: 26 triangles, equally spaced in mel, 106.0 Hz wide at the bottom and 618.3 Hz at the top. From 7 s the band powers and their logs grow as bars, and from 8.5 s the DCT stems appear: 13 numbers for the frame, the first few for the envelope.

Bands as wide as hearing’s

A speech recogniser looks at the sound a few tens of milliseconds at a time. On this page a frame is 50 ms: at 8 kHz that is 400 samples, and its spectrum has 257 numbers. That is far more than the recogniser wants.

In The source–filter model of speech (30.1), in “Same pitch, different vowels”, a voiced sound was harmonics under an envelope. The harmonics carry the pitch, and the envelope’s peaks, the formants, carry the vowel. To know which word was said, I mostly need the envelope.

So the goal of this page is a short list of numbers, about a dozen per frame, that describes the envelope the way hearing does and mostly ignores the harmonics. Every signal on this page is synthetic: the vowels are 30.1’s, made from an impulse train and three formant resonators, never recorded.

The mel scale

Play a tone at 200 Hz, then one at 300 Hz. The step sounds large. Now play 3200 Hz, then 3300 Hz: the same 100 Hz sounds much smaller. Hearing spreads its attention unevenly across frequency.

The mel scale measures frequency the way listeners judge pitch distance. It comes from listening tests in which people judged the pitch distances between tones. The version I use is the one in the HTK speech toolkit:

mel(f)=2595log⁡10(1+f700),\mathrm{mel}(f)=2595\log_{10}\Bigl(1+\frac{f}{700}\Bigr),

with ff in hertz. It is built so that 1000 Hz is very nearly 1000 mel. Other toolkits use slightly different formulas; the idea is the same.

Below about 1 kHz, mel and hertz stay fairly close: 500 Hz is 607.4 mel. Above, it grows more and more slowly. The 1000 Hz from 1 to 2 kHz add 521 mel, but the 2000 Hz from 2 to 4 kHz add only 625.

To turn a mel value back into hertz, undo the formula: divide by 2595, raise 10 to that power, subtract 1, and multiply by 700.

Triangles equally spaced in mel

Think of a graphic equaliser whose sliders are spaced the way the ear hears: close together at low frequencies, far apart at high ones. A mel filter bank is that.

In Filter banks (23.1), in “One prototype, eight channels, a perfect sum”, every channel had the same width. Here I want widths that follow the mel scale, and I use the simplest band shape: a triangle.

The top of my range is 4000 Hz, half the sampling rate, which is mel(4000)=2146.1\mathrm{mel}(4000)=2146.1. I mark 28 points equally spaced in mel from 0 to 2146.1, which is 27 steps of 79.48 mel. Then I turn each point back into hertz.

Triangle mm, for m=1m=1 to 26, rises from point m−1m-1 to a peak of 1 at point mm, and falls to 0 at point m+1m+1. So each triangle’s peak sits on the feet of its two neighbours, and the triangles’ centres are 79.48 mel apart.

In hertz the first centre is 51.2 Hz and the last 3679.9 Hz. The first triangle is 106.0 Hz wide at its base, and the last is 618.3 Hz wide, 5.8 times as wide.

One frame, 26 band powers

Now the frame. I take N=400N=400 samples, 50 ms, of 30.1’s /a/ at a pitch of 125 Hz: samples 1800 to 2199, from 225 ms to 275 ms into the half-second vowel.

Three steps prepare it. First comes 30.1’s pre-emphasis, 1−0.97z−11-0.97z^{-1}, from “The tract, recovered from the sound”: each sample minus 0.97 times the one before. I apply it to the whole vowel and then cut the frame, so every sample of the frame has been through it.

Second, a Hamming window, from Window functions compared (15.2), “Five windows, one trade”. I use its symmetric form, 0.54−0.46cos⁡(2πn/(N−1))0.54-0.46\cos\bigl(2\pi n/(N-1)\bigr) for n=0n=0 to N−1N-1, with N−1N-1 where 15.2 has NN. Third, a 512-point FFT, with the 400 samples padded by zeros. Its bins are 8000/512=15.6258000/512=15.625 Hz apart, 257 of them from 0 to 4000 Hz.

This is one column of a spectrogram, as in Spectrograms & the STFT (15.5), “A row of short spectra”. The frame is long enough to show the harmonics. A Hamming window’s peak is 4 bins of the 400-point DFT wide at its base, 4fs/N=804f_s/N=80 Hz. That is less than the 125 Hz between harmonics.

The power spectrum is ∣X[k]∣2\lvert X[k]\rvert^2. Each triangle pools it into one number, the band power PmP_m: every bin’s power, weighted by the triangle’s height at that bin, added up.

Pm=∑k=0256(height of triangle m at bin k) ∣X[k]∣2\begin{aligned} P_m=\sum_{k=0}^{256}&\bigl(\text{height of triangle } m\\ &\ \text{at bin } k\bigr)\,\lvert X[k]\rvert^2 \end{aligned}

The first triangle touches only 6 bins, and the last 39.

Two more steps finish the job, and the picture at the top of the page shows them too. I take the natural log of each band power, ln⁡Pm\ln P_m, since loudness is heard on a log scale. Then I take a DCT of the 26 logs and keep the first 13 numbers. The next section looks at both steps closely.

Notice how the triangles widen as they climb: the first is 106.0 Hz wide, the last 618.3 Hz. Notice also the tallest bar, band 10, centred at 717 Hz, beside /a/’s first formant at 730 Hz.

Vowels far apart, pitches close together

Now the last two steps, and why they leave the vowel and drop most of the pitch.

Logs, then a DCT

Why a log? A band twice as loud should count as one step louder, whatever its level, and the log does that. It is the decibel idea of How big is a signal (1.3), from “Signal and noise, in decibels”, in other units. A step of 1 in ln⁡Pm\ln P_m is 10/ln⁡10=4.3410/\ln 10=4.34 dB.

The 26 logs ln⁡Pm\ln P_m make the log-mel spectrum of the frame. Next comes the orthonormal DCT-II of The discrete cosine transform (13.6), on 26 points instead of 8:

ck=wk∑m=126ln⁡Pm cos⁡πk(2m−1)52,c_k=w_k\sum_{m=1}^{26}\ln P_m\,\cos\frac{\pi k(2m-1)}{52},

with w0=1/26w_0=\sqrt{1/26} and wk=2/26w_k=\sqrt{2/26} for k≥1k\ge1. The bands are counted from 1 here, hence 2m−12m-1 where 13.6 had 2n+12n+1.

As in 13.6’s “Keep a few numbers, rebuild”, I keep the first few: k=0k=0 to 12, 13 numbers. These are the MFCCs, the mel-frequency cepstral coefficients. Taking a transform of a log spectrum gives what is called a cepstrum, a name made by turning round the first letters of “spectrum”.

A warning about the name: these ckc_k are not the Fourier series coefficients ckc_k of Fourier series coefficients (7.2). On this page ckc_k is always a cepstral coefficient.

For the /a/ frame of the first instrument, c1c_1 to c4c_4 are 5.57, −12.49, −3.93 and −0.85.

What each coefficient keeps

Each ckc_k measures how much of one cosine shape runs across the 26 bands. c1c_1 is half a cosine period across the bands: it compares the low bands with the high ones. c12c_{12} is twelve half periods, a ripple about every four bands. So the first few describe the envelope’s broad shape.

c0c_0 is different. Its cosine is flat, so c0c_0 is 26\sqrt{26} times the average of the ln⁡Pm\ln P_m: it follows how loud the frame is. A louder speaker shifts every ln⁡Pm\ln P_m by the same amount, which changes c0c_0 alone. So I compare vowels with c1c_1 to c12c_{12} only.

Where does the pitch go? The harmonics are 125 Hz apart. The last triangle runs from 3381.7 to 4000 Hz and holds four of them, 3500 to 3875 Hz, so its power changes little when they move.

The lowest triangles are a different story. The first three are narrower than 125 Hz at the base, and the first, from 0 to 106.0 Hz, holds no harmonic at all. So its power depends on where the nearest harmonic sits.

Lower the pitch to 100 Hz and the first harmonic falls inside it: band 1’s log power rises by 2.1 to 3.6 for the three vowels. Raise the pitch to 160 Hz and the first harmonic moves away: band 1’s log power falls, by 0.6 to 4.0. So pitch does not vanish from the MFCCs; it mostly falls away.

Distance between two vowels

To compare two frames, I treat c1c_1 to c12c_{12} as a point in twelve dimensions. Their distance is the straight-line one: the square root of the sum of the squared differences, coefficient by coefficient.

The test uses 30.1’s /a/, /i/ and /u/ at three pitches, 100, 125 and 160 Hz, each with the same frame and recipe as above. If MFCCs follow the vowel, a vowel at another pitch should stay close to itself, and different vowels at one pitch should be far apart. A recogniser needs that to hear the same word from a child and from an adult.

Vowels far apart, pitches close together

30.1's synthetic /a/, /i/, /u/ at 100, 125 and 160 Hz; distance between c_1 to c_{12}.

Three vowels at 125 Hz: three different shapes.

pair
—
distance
—
0.00 / 12.00 s
Describe this picture

30.1’s synthetic /a/, /i/ and /u/ at 100, 125 and 160 Hz, compared by the distance between their c1c_1 to c12c_{12}. Two panels. The first draws the nine coefficient lists as lines, ckc_k from −16 to 18 against kk from 1 to 12: /a/ solid, /i/ dashed and /u/ dotted, each labelled, with the 125 Hz versions standing out from the other pitches. The second is a 3 × 3 grid of the distances between the vowels at 125 Hz, with the numbers printed, and beside it each vowel’s largest pitch shift. The readouts are the pair and its distance. There is no control. The 12 s clip opens on the three vowels at 125 Hz: three different shapes. At 3 s the versions at 100 and 160 Hz join, and each vowel’s lines stay together; the largest shift is 4.20, for /a/. From 8 s the grid fills: /a/–/i/ 21.38, /a/–/u/ 13.77, /i/–/u/ 20.71, all larger than any pitch shift, and a nearest-vowel rule gets all 6 shifted versions right.

Watch each vowel’s lines as the other pitches join: the three versions of a vowel stay together, while the three vowels stay apart.

Notice the two scales. The largest pitch shift is 4.20, and the closest two vowels are 13.77 apart. The other shifts are smaller still: /i/ moves at most 3.25 and /u/ 3.45. /a/ moves most at 160 Hz, /i/ and /u/ at 100 Hz.

The “nearest-vowel rule” names each shifted version after the 125 Hz vowel it is closest to. It names all six correctly. It is a very simple recogniser, and it works here because the pitch moves the points less than the vowel does.

Speaker identification uses the same features the other way round. Instead of the vowel, it looks for the fine differences between speakers’ envelopes, which depend on the length and shape of each speaker’s vocal tract.

The maths behind it · matrix products

The filter bank is a 26 × 257 matrix applied to the power spectrum, and the DCT an orthogonal 26 × 26 one. MFCCs are two matrix products with a log in between.

The maths behind it · nearest-neighbour classification

Recognisers model MFCC vectors with mixtures of Gaussians (older systems) or feed log-mel frames to neural networks. The nearest-vowel rule above is nearest-neighbour classification, a basic method of statistical learning.

Worked example

Every number below was recomputed with NumPy and SciPy.

  1. Hertz to mel. For 2000 Hz, 1+2000/700=3.8571+2000/700=3.857, and log⁡103.857=0.5863\log_{10}3.857=0.5863. Times 2595, that is 1521.4 mel.
  2. The first triangle. The mel step is 2146.1/27=79.482146.1/27=79.48. Back in hertz, 79.48 mel is 51.2 Hz and 158.97 mel is 106.0 Hz. So triangle 1 runs from 0 to 106.0 Hz with its peak at 51.2 Hz, and it covers bins 1 to 6, 15.6 to 93.8 Hz.
  3. A louder voice. Double the amplitude and every PmP_m grows 4 times, so every ln⁡Pm\ln P_m grows by ln⁡4=1.386\ln 4=1.386. Only c0c_0 changes, by 26×1.386=7.07\sqrt{26}\times1.386=7.07.
  4. Nearest vowel. /u/ at 160 Hz is 1.88 from /u/ at 125 Hz, 13.68 from /a/ and 20.74 from /i/. The nearest is /u/, so the rule reads it correctly.

Where you’ll meet this

Speech recognisers start from these features. Older ones, built on hidden Markov models, read MFCCs frame by frame. Modern neural recognisers read the log band powers directly. Stacked frame after frame, as in 15.5, they make the log-mel spectrogram.

Recognisers usually take 20 to 30 ms frames every 10 ms. The longer 50 ms frame here makes the harmonics easy to see. Speaker verification, which checks that a voice belongs to the person it claims to be, compares the same features against a stored model of that person. Music software uses them too, to sort recordings by genre or to name an instrument.

I left out the delta features, which add each coefficient’s change from frame to frame, cepstral mean normalisation, which removes a fixed channel’s colouring, and features learned by neural networks. The next chapter turns to rhythms in the body, in EEG and rhythms (31.2).

For more, see S. B. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition” (IEEE Trans. ASSP, 1980); Rabiner and Schafer, Theory and Applications of Digital Speech Processing (2011), chapter 8; and Jurafsky and Martin, Speech and Language Processing (3rd ed. draft), chapter 16.

Reference card

QuantityFormulaNotes
Mel2595log⁡10(1+f/700)2595\log_{10}(1+f/700)HTK form; 1000 Hz ≈ 1000 mel
Filter banktriangles equally spaced in mel26 up to 4 kHz here, 79.48 mel apart
Band powerPmP_m: triangle-weighted sum of ∣X[k]∣2\lvert X[k]\rvert^2one number per band
Log-mel spectrumln⁡Pm\ln P_m, m=1m=1 to 26stacked over frames: log-mel spectrogram
MFCCckc_k = orthonormal DCT-II of ln⁡Pm\ln P_m13 kept, k=0k=0 to 12
Loudnessc0=26×c_0=\sqrt{26}\times mean of ln⁡Pm\ln P_mleft out of distances
What they keepthe envelopepitch mostly falls away
Distancestraight-line distance between c1c_1 to c12c_{12}nearest vowel: 6 of 6 here

End of lesson 30.2

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look