Models & technology

Speech Recognition

Speech recognition converts spoken words into text. It underlies transcription, voice input, and the listening half of voice conversation with AI.

From speech to text

Speech recognition has a computer listen to speech and convert it into text — also called speech-to-text, or STT. Voice input on a phone, automatic meeting transcription, and speaking to a smart speaker have made it an everyday technology.

Deep learning made it usable

Earlier systems were inaccurate enough that you had to speak slowly and clearly for them to be useful at all. Deep learning improved accuracy dramatically, to the point of transcribing natural conversational speed reliably. Capable models such as OpenAI’s Whisper substantially advanced multilingual support and robustness to noise.

Everyday uses

  • Automatic transcription of meetings, interviews, and video, and automatic captions
  • Dictating messages on a phone, often faster than typing
  • The listening half of voice conversation with an assistant
  • Turning call center recordings into text for analysis

Transcription in particular is the standard example of AI efficiency. An hour of meeting takes minutes to transcribe, and summarizing it can be delegated as well.

Paired with speech synthesis

If speech recognition is the ear, text-to-speech is the mouth. Add an LLM as the mind and you have voice conversation with AI. As multimodal capability advances, listening, thinking, and speaking happen together, and talking to AI keeps moving closer to talking to a person.

Accuracy on specialist terminology and regional speech has improved enough for practical meeting notes. In a workplace with a lot of meetings, transcription is the AI deployment most likely to show value immediately.

Try voice input on your phone once. For longer messages and notes, speaking is faster more often than people expect.

Related terms

Sources and review information

Last reviewed July 17, 2026

Back to the AI Glossary

Search this site