AI Glossary

Speech recognition

Speech recognition converts spoken audio into text or word sequences. Also called automatic speech recognition (ASR), it estimates what was said; understanding the intent or verifying the claims is a separate task.

Also known as: automatic speech recognition

· Updated · Chain of Thought

A meeting assistant first transcribes microphone audio, then sends the text to search or summarization. A recognizer can use separate acoustic and language components or an end-to-end neural model. Either approach must handle the audio conditions of the application.

For example, confusing “fifteen” with “fifty” can change an action item even when the rest of a transcript looks fluent. Evaluate representative accents, background noise and specialist vocabulary. Word error rate counts substitutions, deletions and insertions relative to a reference transcript; it does not measure whether a resulting summary preserves the most important fact.

This matters because later language processing can amplify an earlier transcription error. Keep the original audio available where appropriate for review, and check consequential names and numbers. Recognizing words also differs from identifying which speaker said them: speaker diarization adds another task. A useful evaluation should reflect the complete workflow as well as transcription fidelity.

Sources

Go deeper