Question

How does speech recognition work?

Vault Verified
Curated Intelligence
Definitive Source
Answer

By converting sound into a numerical representation, then using a model trained on enormous quantities of speech to produce the most probable sequence of words — and "most probable" is the key, because the system is not identifying sounds so much as guessing sentences.

The pipeline:

Capture and preprocess. Audio is sampled, and noise reduction, echo cancellation and voice isolation are applied. Multiple microphones allow beamforming, focusing on a direction.

Feature extraction. The waveform is converted into a compact representation of how energy is distributed across frequencies over short overlapping windows — a spectrogram-like form, historically mel-frequency coefficients.

Acoustic modelling, mapping those features to speech units.

Language modelling, which is where most of the accuracy comes from. The system knows which word sequences are plausible, which is how it distinguishes "recognise speech" from "wreck a nice beach" — acoustically similar, and vastly different in likelihood.

Decoding, searching for the sequence that best fits both the audio and the language model.

What changed recently. Modern systems are end-to-end neural networks trained directly from audio to text, trained on hundreds of thousands of hours, which removed the hand-built pronunciation dictionaries and separate components of earlier designs and improved robustness enormously.

Why it still fails:

Accents and dialects underrepresented in training data, where measured error rates are documented as substantially higher — a well-established fairness problem.

Overlapping speakers, which remains hard.

Background noise and reverberation.

Domain-specific vocabulary — names, medical and legal terms — which the language model considers improbable and therefore corrects away.

Children's speech, which differs acoustically from the adult speech that dominates training data.

Why wake words work differently. A small always-listening model runs locally, detecting only one phrase. Only after it triggers is audio usually sent onward, which is the technical answer to whether the device is recording everything — though false triggers do send audio, which is how recordings occasionally surface.

Speaker identification is a separate task from recognising the words.

Related Questions