A microphone captures a sound wave. A model has to turn that into words it has never technically “heard” before, and it still gets thrown off by the same things that trip up a person on a bad phone line: accents, background noise, and two people talking at once.
Key takeaways
Before a model can do anything with speech, the audio is sampled and converted into a numerical representation that captures pitch and timing over very short slices of time — the audio equivalent of the tokens a language model processes when it reads text. The model was trained on many hours of audio paired with its correct transcript, and it uses that training to predict, for each slice, which sequence of words most likely produced the sound pattern it is seeing.
That framing — predicting the most likely words, rather than hearing in any human sense — explains a detail that trips people up: the system does not process sound the way an ear does. It converts a continuous waveform into discrete slices, and its output for each slice is a best guess informed entirely by statistical patterns in the training audio, with no independent sense of the room, the speaker, or the topic beyond what those patterns encode.
This is genuinely deployed technology today, not a lab demonstration: Canada's Cyber Centre lists it among current business applications, noting that AI can understand voice and text to analyze, respond and carry out tasks. Call centres and website chatbots use generative AI to analyze initial requests and offer information to try and solve common questions without needing human interaction. In each of those settings, the system is doing the same underlying job — matching an audio pattern to the words most likely to have produced it, based on everything it saw during training.
Prediction, rather than direct perception, is also why a speech system can produce a confident, well-formed sentence that is simply the wrong words — it has not misheard in the way a distracted person mishears; it has matched the sound pattern to the most statistically likely words given its training, and occasionally the most likely match is not the correct one.
Because the system is matching new audio to patterns it has seen before, anything under-represented in that training audio produces more errors: an accent or dialect the training data barely covered, background noise, two people talking over each other, or a business's own specialized vocabulary — product names, file numbers, acronyms — that never appeared often enough in general training data to be learned reliably. This is exactly the property the NIST framework calls robustness: the ability of a system to maintain its level of performance under a variety of circumstances. A speech system that performs well in a quiet, single-speaker test and badly on a noisy call centre floor has a robustness gap in precisely that sense — the same system, a different circumstance.
Specialized vocabulary is a genuinely common, easily overlooked version of this problem for a Canadian business specifically: a client's name, a file or policy number, a product SKU, or a French-language term used in an otherwise English call are all exactly the kind of low-frequency input a general-purpose training set is least likely to have covered well, which is why these are disproportionately the words that come back wrong in an otherwise fluent transcript.
Picture two calls handled by the same speech-to-text system on the same day. The first is a single speaker, on a good line, in plain English, asking a routine question — conditions close to what the training data represented well, and the transcript comes back essentially clean. The second is two people talking over each other on a poor mobile connection, one of them giving a client's name and a file number the system has never encountered before. The underlying mechanism has not changed between the two calls; what changed is how far each one sits from the conditions the training audio represented. The same tool can be genuinely reliable and genuinely unreliable in the same afternoon, depending entirely on which of those two calls it is handling — which is why a single, blended accuracy claim for a speech tool is rarely the useful number.
Statistics Canada's most recent survey puts real numbers on adoption: speech or voice recognition using AI was reported by 20.6% of AI-using businesses in the second quarter of 2026, essentially flat against 20.0% the year before. This describes how widely the capability has been taken up, not how accurately any particular tool performs — no Canadian source publishes an accuracy or error-rate figure for speech recognition in business use, and none should be assumed.
That the figure has barely moved between the two years is itself informative: it suggests a technology that has settled into steady, specific use cases — largely call-centre and voice-interface applications, per the Cyber Centre's own examples above — rather than one still spreading rapidly into new corners of business operations the way large language model adoption has over the same period.
A transcription error reads exactly as fluently as a correct passage — a wrong name or number sits in the sentence with the same confidence as the right one would have. That is the same discipline covered in can AI cite its sources reliably applied to a different kind of output: before a transcript is relied on for anything consequential, a person needs to confirm it, because fluency is not a signal of accuracy in a transcript any more than it is in a written answer.
See also how AI image generation works and can AI cite its sources reliably.
Only if it is deliberately set up to adapt to that voice or vocabulary. A generic tool run cold on each new call does not retain anything from a previous one unless that memory has been built in on purpose.
It is one of several, and not necessarily the largest. Accent and dialect representation in the training data, and unfamiliar or specialized vocabulary, can matter just as much, per the same robustness gap described above.
No. A person should confirm anything from a transcript that will be relied on for a decision, a file, or a record — the same review discipline that applies to any AI-generated output applies here too.
Better audio capture reduces noise-driven errors, but it does nothing for a gap caused by under-represented training data — an accent, dialect or specialized term the model rarely saw in training will still be misheard on a perfectly clean recording, because the limitation sits in what the model learned, not in the audio quality it received.
Feeding a transcript into the rest of a business's systems — a CRM, a ticketing tool, a case file — is an integration question. For that, see connecting AI to your CRM without breaking what already works.