Glossary · Stentor

What is word error rate (WER)?

What is word error rate (WER)? A clear explanation for Azerbaijani business — and how Stentor applies it.

Understanding Word Error Rate (WER)

Word Error Rate (WER) is the industry-standard metric used to evaluate the accuracy of speech-to-text (automatic speech recognition) systems. It is calculated by comparing a machine-generated transcript against a gold-standard reference transcript, counting substitutions, deletions, and insertions at the word level, and dividing that sum by the total number of words in the reference. A lower WER indicates a more precise system with fewer transcription mistakes, which is critical because the reliability of all downstream processes—such as automated quality assurance scoring, compliance monitoring, and sentiment analysis—depends entirely on the accuracy of the initial text conversion. For businesses operating in Azerbaijan, WER is a pivotal factor in the viability of call analytics. Most general-purpose transcription engines are not trained on the specific nuances of the Azerbaijani language or the common practice of mixing Azerbaijani and Russian within a single conversation. This lack of specialization leads to high error rates that can render automated insights useless. By focusing on reducing WER through purpose-built models, organizations can transform raw audio into a dependable data asset, ensuring that compliance risks and customer complaints are detected based on factual evidence rather than transcription noise.

Capabilities

Why Low WER is Critical for Your Business

Ensures every word in a customer conversation is captured correctly, making automated scoring and compliance checks trustworthy and precise.

Enables complaint detection and negative sentiment analysis to trigger based on real linguistic signals rather than transcription errors.

Allows for the safe automation of quality assurance across 100% of calls, eliminating the need to rely on small, potentially biased statistical samples.

Significantly reduces the operational cost of QA programs by minimizing the time staff must spend manually correcting inaccurate transcripts.

Maintains high accuracy even during code-switching conversations (mixed AZ/RU), which are frequent in local Azerbaijani contact centers.

Creates a consistent, low-error auditable record that provides a reliable foundation for regulatory compliance and dispute resolution.

How Stentor Optimizes Transcription Accuracy

Purpose-Built Azerbaijani Speech Recognition

Stentor's speech-to-text engine is specifically trained on Azerbaijani and mixed Azerbaijani-Russian audio, the language reality of most local contact centres. This targeted training keeps WER low where general-purpose engines typically struggle most.

100% Call Coverage, No Sampling

Because transcription accuracy is high enough to trust, Stentor analyses every single call rather than a statistical sample. Decisions about agent performance and compliance risk are based on the full picture, not an estimate.

Speaker Diarisation

Stentor separates agent and customer speech before transcribing, so errors are not compounded by mixed-speaker confusion. Each speaker's words are attributed correctly, which is essential for accurate sentiment and compliance scoring.

Hybrid QA Scoring

Transcripts feed a three-layer scoring system combining rule-based checks, semantic AI analysis, and human override. Even where residual transcription errors exist, the semantic layer can still interpret meaning, and human reviewers can correct edge cases.

Private Single-Tenant Cloud

Audio and transcripts never leave your dedicated environment. There is no data egress to shared infrastructure, which matters for industries with strict data-handling obligations.

Complaint and Risk Detection

Low WER enables reliable detection of complaints, negative sentiment, and compliance risk phrases. When the underlying transcript is accurate, these signals can be acted on with confidence rather than treated as indicative.

From Raw Audio to Actionable Insight

1Every call is ingested into Stentor's single-tenant private cloud environment, ensuring the audio never leaves your controlled infrastructure.
2The purpose-built Azerbaijani speech-to-text engine transcribes the audio, handling mixed AZ/RU speech to keep WER as low as possible for your language context.
3Speaker diarisation separates agent and customer turns, attributing each word to the correct speaker before any analysis begins.
4The transcript is scored using a hybrid QA model: rule-based criteria check for specific phrases or omissions, semantic AI interprets meaning and tone, and human reviewers can override any automated decision.
5Complaints, negative sentiment, and compliance risk signals are flagged automatically across 100% of calls, giving supervisors a prioritised queue rather than a random sample.
6Results are surfaced in dashboards and reports, providing a complete, auditable record of every conversation and its quality score.

WER and Stentor: Frequently Asked Questions

What is a good Word Error Rate for contact centre use?

While there is no universal threshold, the WER must be low enough that automated scoring decisions are reliable. The acceptable rate depends on the language complexity and the specific use case. Purpose-built models for Azerbaijani typically achieve significantly lower WER than general-purpose alternatives, making them viable for compliance and QA.

Why do general-purpose transcription tools struggle with Azerbaijani calls?

Most global systems are trained on high-resource languages. Azerbaijani is a lower-resource language, and the frequent switching between Azerbaijani and Russian mid-conversation adds complexity. Without targeted training data, these systems produce high WER, making the resulting transcripts unreliable for automated analysis.

Does Stentor analyse all calls or only a sample?

Stentor transcribes, diarises, and scores 100% of conversations with no sampling. This comprehensive coverage is possible because the transcription accuracy is high enough to trust automated scoring at scale.

How does the hybrid QA scoring model handle residual transcription errors?

Stentor uses a three-layer approach: rule-based checks, semantic AI, and human override. The semantic layer can often interpret the intended meaning even if individual words are imperfectly transcribed, while human reviewers provide a final safety net to correct edge cases.

Is our call data kept private during this process?

Yes. Stentor utilizes a single-tenant private cloud architecture. This ensures that your audio and transcripts are processed and stored in a dedicated environment with no data egress to shared infrastructure.

Experience Accurate Azerbaijani Call Transcription

If your contact centre relies on sampled reviews or struggles with inaccurate transcripts, Stentor's purpose-built speech recognition and hybrid QA scoring can help you move to full call coverage with confidence. Reach out to the Allmaz team to learn how Stentor handles your specific language environment.

Request a demo