Multilingual call QA: testing scoring in three languages
In Baku, Azerbaijani and Russian often mix within the same call, and product and tariff names are said in English. An evaluation system that works well in one language can make mistakes in another, and that can systematically lower the scores of agents who work in it. This article covers five differences between languages, an illustrative test set, why a translated transcript is not evidence, language-independent criteria, a terminology glossary and fair comparison across language groups.
Short answer
In a multilingual contact center, agent scoring has to be tested for each language separately. A system that works well in one language may not do so in another: recognition quality, speaker separation and the interpretation of criteria all differ between languages. In Baku, everyday reality is more complex still — Azerbaijani and Russian mix within one call, and product names are said in English.
So the question is not "does the system know Azerbaijani?". The question is: on your calls, with your terminology and your criteria, how well does the result match human evaluation for each language and each language mix? Only a separate test set shows that.
What differs between languages
- Recognition qualitySpeech-to-text accuracy varies by language. A number or name recognised well in one language may be written wrongly in another.
- Dialect and accentBaku, Ganja or Nakhchivan pronunciation, the accent of an Azerbaijani speaking Russian — all of these affect recognition.
- Code-switchingSentences that switch between Azerbaijani and Russian mid-phrase are normal in a multilingual region, but hard for automatic recognition.
- Domain termsTariff names, product codes, banking terms. If the recogniser does not know them, it substitutes an approximate word.
- Politeness and styleA phrase that is polite in one language can sound cold in another. Style criteria are interpreted differently depending on the language.
Illustrative test set
The design below is an illustrative example. The goal is to see results for each language and language mix separately; adjust the numbers to your own call flow.
- Language groupsAzerbaijani-only, Russian-only, English-only and mixed-language calls — each group separately.
- Variety within each groupDifferent agents, customer profiles and call types; a few noisy recordings.
- TerminologyEach group should have several calls with product names, prices and tariff names — that is where recognition errors show most.
- Human referenceTwo people fluent in that language score each call independently. Mixed-language calls need a reviewer who knows both languages.
- Results tableAgreement between AI and the reference is recorded for each language group and each criterion separately. No overall average is drawn.
A translated transcript is not evidence
Reading the evaluation of a Russian-language call in Azerbaijani is convenient. But translation adds another layer of error: the translating system can soften a phrase, harden it or change its meaning.
- The line in the original language is always used as evidence; the translation is only for understanding
- On a disputed critical criterion, a person who knows the original language has the final word
- Criteria such as politeness, promises and prohibited phrases are the most distorted by translation
- If a comment written in the report language disagrees with the original line, the original wins
Writing criteria that do not depend on language
When a criterion is tied to a specific phrase, a multilingual setting causes problems. A criterion that requires one particular Azerbaijani greeting will never be met on a call held in Russian.
- Intent, not wording"The agent greeted the customer and introduced themselves" — whatever the language.
- Mandatory text per languageWhere exact wording is required, such as a legal disclosure, an approved version is written for each language.
- A language-choice criterionIf it is company policy: "the agent switched to the language the customer was speaking". That is a separate, clear criterion.
- Examples in every languageThe criterion's explanation includes one "yes" and one "no" example for each language.
Fair comparison across language groups
If the system makes more recognition errors in Russian than in Azerbaijani, agents who serve Russian-speaking customers may systematically score lower. That is a measurement error, not the agent's quality.
- When comparing agents, show the language split of their calls
- If the test set found a difference between languages, apply more human review in that language
- Track the share of "unable to assess" on mixed-language calls separately
- Do not merge language groups into one ranking until the difference has been removed
A terminology glossary
In multilingual evaluation, the most useful small tool is a terminology glossary. It helps both to spot recognition errors and to write criteria. Each entry holds:
- The term and its variantsThe official name of a tariff, product or service, and the variants said on calls — in Azerbaijani, Russian and English.
- Typical recognition errorHow the system got the term wrong in the test set. It gives the specialist reading the transcript a quick answer to "what is this word really?"
- Link to a criterionWhich criterion the term matters for — for example, stating the correct tariff name.
- Update dateThe glossary is updated with each new product and campaign; old terms are archived, not deleted.
It does not need to be large: the 30–50 most often misrecognised terms usually cover most of the problem.
Typical mistakes
- Treating a vendor's "we support Azerbaijani" as proof of accuracy
- Building the test set in one language and applying the result to all languages
- Leaving mixed-language calls out of the test — even though real traffic has many of them
- Showing an agent a translated transcript as evidence
- Tying criteria to phrases in a specific language
Ongoing monitoring
A test set is not a one-off job. New product names, a new campaign, a new agent group or a change in the customer mix all affect per-language results. A practical rhythm: each month, people re-score a few calls from each language group and compare them with the AI result. If agreement drops in one language, look for the cause: a new term, a new dialect, recording quality. The result is written up as a short monthly note — how many calls were checked in each language, where a gap was found and what was fixed. The note also helps when an agent disputes a result: if a Russian-language call is contested, you already know how agreement looked in Russian that month.
Limits
- Accuracy measured in one language does not carry over to another language or to a language mix
- On mixed-language calls, both recognition and speaker separation may make more errors
- Translation can lose meaning and tone; the original line is what counts
- No equal accuracy should be claimed for any language without proof
- An agent's accent or language is no basis for conclusions about their personality or quality
What Vexvon Audio Analyzer offers
- The spoken language of every call is detected automatically and shown with the call
- The transcript and summary can be translated into the languages you choose; the original line is kept separately
- The language the summary and evaluation are written in — the report language — is set per company, so the evaluation of a Russian-language call can be read in Azerbaijani
- Each step of the call standard shows the original transcript lines as evidence
More: Vexvon Audio Analyzer. How to build a test set in general is covered in the speech analytics pilot plan, and speaker separation in speaker diarization.
First step
Pull the language split of last month's calls: how many were in Azerbaijani, Russian, English and how many mixed. Then take 10 calls from each group and test them with the set above. This article belongs to the audio analysis implementation and reliability section. To try it on your own recordings, get in touch.
Further reading on this topic: speech analytics accuracy testing.