Speech recognition (ASR)
The technology that turns spoken language into written text. It powers voice assistants, live captions, meeting transcripts, and call-centre tools — and is a particular strength of several UK companies.
Speech recognition, often called automatic speech recognition or simply speech-to-text, is the task of converting the sound of a human voice into words on a page. It is deceptively hard: people speak quickly, run words together, use different accents, and talk over background noise. Modern systems handle this with deep learning, trained on huge quantities of recorded speech paired with accurate transcripts.
The results are now good enough to be woven into everyday life — dictation, live subtitles, the transcript of a meeting, the assistant that takes a spoken request. Getting it to work reliably across many accents, languages, and noisy real-world conditions is where much of the engineering effort goes, and where the better systems distinguish themselves.
It is an area of notable UK strength. Speechmatics, based in Cambridge, is known for speech recognition that aims to work fairly across a wide range of accents and voices, while companies such as PolyAI build on the technology to run natural telephone conversations for businesses.