Latest Blog
Multilingual AI Interviewing: What to Test Before You Roll Out Globally
Researcher
•
5 min read
Share this post
Multilingual AI interviewing means running a structured interview in the candidate's own language instead of the employer's, by phone, video or text. Vendors compete on language count. Whether a rollout survives depends on something else, which is how accurately the speech model transcribes the varieties your candidates speak. A transcription error doesn't stay in the transcript. It becomes the evidence a scoring model reads, and then the evidence a recruiter reads.
The distance between those two things has been measured. Testing five commercial speech recognition systems from Amazon, Apple, Google, IBM and Microsoft, Koenecke et al. found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers. That was one language, one country, clean recordings. Over 20% of samples from Black speakers hit a word error rate of 0.5 or worse, against fewer than 2% of samples from white speakers. If one country's range of accents produces a spread that wide, then a deployment across forty languages is not simply a larger version of the same problem. The problem changes shape, and you have to measure it language by language.
What does "supports 40+ languages" mean on a vendor spec sheet?
It usually means the speech model was trained on data tagged with those language codes and produces intelligible output on a benchmark recording. It rarely means the vendor has measured accuracy on the varieties your candidates speak. Language coverage is a claim about a training corpus, and what you're buying is a claim about your applicant pool. Verified performance across the six languages you hire in beats a spec sheet advertising eighty. That difference shows up in three places.
Dialect spread inside a "supported" language
Arabic is the clearest case. One effort to build Arabic speech recognition covering the dialects spanned 17 Arabic-speaking countries and at least 11 varieties, which is roughly the distance between "supports Arabic" and "works in Jeddah and Cairo and Casablanca." Gulf, Egyptian, Levantine and Maghrebi Arabic diverge enough in phonology and lexicon that a model weighted toward Modern Standard Arabic can be reliable on a formal answer and unreliable on a candidate describing a shift handover. Modern Standard is the register of broadcast news, which is where a lot of training data comes from. The same pattern separates Quebec French from metropolitan French, Chilean Spanish from Castilian, and Swiss German from standard German, if less dramatically. A vendor who has never been asked which variety they benchmarked on has probably not benchmarked.
Code-switching
In Philippine BPO hiring, candidates answer in Taglish as a matter of course. A sentence starts in Tagalog, the technical noun arrives in English, and the sentence closes in Tagalog. Educated urban speech in Manila runs that way as a matter of course, and a candidate doing it is not failing to follow your instructions. The same holds for Hinglish in Indian metros and Spanglish along the US-Mexico border. The data needed to evaluate this arrived late. CS-FLEURS, released in September 2025, covers 113 code-switched language pairs across 52 languages, and its authors frame it as a step toward code-switched research "beyond high-resourced languages." That benchmark is less than a year old, so if a vendor tells you code-switching is handled, ask what they measured it on and when.
Role vocabulary
Certifications, equipment names, safety terms and shift jargon are low-frequency tokens, which is what speech models get wrong most often. They're also the tokens carrying the hiring signal. A warehouse interview in Polish turns on whether the candidate holds a uprawnienia UDT forklift licence. A CDL interview turns on endorsement codes. If the model mis-transcribes that certification, then the recruiter reads an answer with the decisive fact missing from it, and scores the candidate as though they never said it.
Book a Language Coverage Review. Bring the six languages you hire in most and we will go through what to ask any vendor about the varieties inside them. No product demo.
Has speech recognition bias improved enough to stop testing for it?
Partly, and not enough. Current multilingual models are considerably better than the 2019-era systems Koenecke tested. But the disparity hasn't been engineered out. A 2025 sociophonetic analysis of four major commercial systems using the Pacific Northwest English Corpus introduced a Phonetic Error Rate metric to trace individual recognition errors back to specific linguistic variables, and concluded that "acoustic modeling of dialectal phonetic variation, rather than lexical or syntactic factors, remains a primary source of bias in commercial ASR systems." Vowel-quality variation correlated most strongly with errors, with the most pronounced effects for African American speakers across every system tested.
That has direct procurement consequences. A primary source of the bias sits deep in the stack, in acoustic modelling, which means a better question set or rubric can't correct for it later. So change what you ask. "How accurate is your speech recognition?" is usually the wrong question. "Whose model, which version, and what did you measure on our language pairs?" is the right one. Ask it before the demo, not after.
What should global employers test before rolling out multilingual AI interviews?
Test six things, in this order, with the pass threshold set before anyone looks at results. A failure at step one invalidates everything measured after it.
1. Measurement equivalence across languages. Run the same role and rubric in two languages and compare score distributions: mean, variance, shape. A question set that survives translation word-for-word but not difficulty-for-difficulty shows up as a shifted distribution long before anyone complains.
Set the threshold using the rule regulators already apply. 29 CFR 1607.4(D) treats a selection rate below four-fifths of the highest group's rate as evidence of adverse impact. Don't file the result as an internal metric, because language sits inside protected-class analysis. Under 29 CFR 1606.7, "the primary language of an individual is often an essential national origin characteristic," and §1606.6 brings selection procedures expressly within the national-origin guidelines. A Portuguese pass rate sitting at 61% of your English one is therefore potentially direct adverse-impact evidence on a protected characteristic. Where the mechanism producing that difference is speech recognition performing worse on accented speech, it is close to a textbook national-origin fact pattern. Run this test first.
2. Speech performance on messy audio, per market. Collect your own recordings instead of vendor samples: mobile handsets, background noise, hands-free, weak connections. Score word error rate on a held-out set per variety. Where you can't get audio before purchase, ask for the per-variety figures. A vendor who can't produce them probably hasn't measured them.
3. Translation transparency. Confirm that a recruiter reviewing a Vietnamese candidate can reach the Vietnamese transcript and the original audio, not just the English summary. A summary of a translation of a transcript of a noisy phone call sits four lossy steps from what the candidate said, and the chain should be navigable inside a minute, without a support ticket.
4. Candidate completion, disaggregated. Aggregate completion rates hide everything interesting. Split by language, country, device class and time of day. A drop confined to Android in one market is usually technical. A drop across devices in one language usually means the instructions read as machine-translated. Don't assume the direction of the effect before you measure it. It doesn't always run the way you'd expect.
5. Recruiter adoption. Watch behaviour, not training-session feedback. Do recruiters open transcripts? Do they change the recommendations they're given? If nobody ever changes one, that more likely means nobody is reading them than that everybody trusts them.
6. Governance and audit trail. Establish that you can reconstruct, six months later, which version of which question set a candidate saw, who changed it, and when. Regional teams will ask for local variations by week three, and whether those take a settings change or a services engagement decides a great deal about year two.
Book a Test Plan Session. Ninety minutes turning the six tests above into a scored evaluation sheet you can run against whichever vendors you are already talking to.
Which laws apply to multilingual AI interviews in 2026?
Four regimes matter for a global rollout, and one of them moved this summer. A fifth, Colorado, is worth reading as a cautionary tale about planning around dates.
European Union. AI used for recruitment and candidate evaluation is high-risk under Annex III, point 4(a) of the EU AI Act. The deadline for stand-alone Annex III systems is no longer 2 August 2026. The AI Omnibus entered into force on 27 July 2026, and the European Commission now states that for high-risk AI systems in Annex III, "rules apply starting 2 December 2027". A good deal of published guidance still says August 2026, so check the date on whatever you're relying on.
Article 50 was not deferred. Those transparency duties took effect on 2 August 2026 and are in force now, with a narrow grace period to 2 December 2026 for the Article 50(2) watermarking obligation on systems already marketed. The deferral gives you sixteen extra months for conformity work and no extra time at all for telling candidates they're speaking to a machine.
New York City. Local Law 144 requires an independent bias audit within the preceding twelve months, a published summary, and at least ten business days' notice before an automated employment decision tool is used on a candidate. Enforcement has been thin. Across 391 employers, Wright et al. found published audit reports for 18 and transparency notices for 13, a picture the authors attribute largely to employers self-determining that the law doesn't reach them. Thin enforcement is a risk you're choosing, not a plan.
Illinois. The Artificial Intelligence Video Interview Act (820 ILCS 42) requires pre-interview notice, an explanation of how the AI works and what characteristics it evaluates, affirmative consent, and deletion of the video within 30 days of a candidate's request, including copies held by third parties. HB 3773 separately amended the Illinois Human Rights Act effective 1 January 2026 to prohibit discriminatory AI use in employment decisions and to bar zip code as a proxy for a protected class.
Federal United States. This one gets misread. The EEOC removed its 2023 Title VII technical assistance on AI in January 2025, and as employment counsel noted at the time, withdrawing guidance did not withdraw the statute. Title VII, the Uniform Guidelines, and the national-origin guidelines at 29 CFR 1606 all remain in force. Less federal guidance means less clarity about how to comply, not less exposure if you don't.
Colorado shows why to build around evidence rather than dates. Its 2024 AI Act was pushed from 1 February to 30 June 2026, then repealed and replaced by SB 26-189, signed 14 May 2026 and effective 1 January 2027. The replacement drops the impact-assessment and risk-programme duties, but it adds a duty, on request and where commercially reasonable, to provide meaningful human review and reconsideration of adverse outcomes from automated decisions, employment included. That last one is test #5 written into statute. Maryland and Texas add requirements of their own.
How does accessibility apply when the interview is a phone call?
It applies directly, and in a way that catches most voice-first designs. WCAG 2.2 Success Criterion 2.2.1, Timing Adjustable is a Level A requirement. Where content imposes a time limit, the user must be able to turn it off, adjust it to at least ten times the default, or extend it after a warning. A 90-second cap on a spoken answer is a time limit. So is a countdown before the next question fires.
This bites hardest on the candidates a multilingual programme exists to reach. Somebody interviewing in a second language needs formulation time, and so does a candidate with a stammer or a cognitive disability. The awkward part is that response caps usually exist to deter cheating, so loosening them for accessibility weakens a control added for a different good reason. You cannot have both at once. The workable version is an accommodation path that's easy to find, easy to request, and doesn't oblige a candidate to disclose a diagnosis to a recruiter, plus a non-voice route for anyone who can't use voice at all. Ask for the vendor's VPAT or EN 301 549 conformance statement and read what it claims, at what level, against which criteria.
Book an Accommodation Path Review. We will walk your current voice flow against WCAG 2.2 timing and trace the accommodation route a candidate would have to find. You keep the notes whether or not you buy anything.
What tends to break after launch?
The failures cluster, and none of them are visible in a demo, because a demo has one language, clean audio and a cooperative participant.
Signal to watch | What it usually means | Where to measure it |
|---|---|---|
Completion 10+ points below your best market | Instructions read as machine-translated, or the flow assumes a device profile the market lacks | Funnel analytics, split by language and device |
Pass rates diverging between two languages, same role | The localised question set drifted in difficulty | Score distributions per language |
Recruiters never opening transcripts | Low trust, or review costs more time than it returns | Product usage logs |
Recommendations changed almost never, or almost always | Rubber-stamping, or a recommendation carrying no information | Decision logs |
Accommodation requests rising and unanswered | The alternate path exists but isn't findable | Support queue, tagged |
None of these require a vendor's cooperation to measure, and all of them are worth instrumenting before go-live. If you don't record a baseline before launch, then when one of these numbers moves in month six you will have nothing to compare it against.
What can you not test before you buy?
Three things stay untestable before you sign, and any test plan promising certainty about them is selling you something.
Word error rate is blunt. It weights every token equally, so fumbling "the" and fumbling "forklift certification" score the same, when only one of them changes a hiring decision. Measure error rates on job-relevant terms separately where you can.
Standardisation and localisation pull against each other. Ask identical questions everywhere and you get comparability across markets, with prompts that read oddly in half of them. Localise properly and you get the reverse. No configuration delivers both. Most global teams standardise the competency model and localise the prompts, which trades some comparability for prompts that make sense in each market.
Pre-purchase testing doesn't predict post-launch drift. Upstream models get updated, sometimes without notice, and a version change can move your score distributions without anyone on your side touching a setting. Short of version pinning and scheduled re-testing, there isn't much defence against this one.
Where Tenzo fits
One result in Tenzo's own data runs against the warning in test #4. Candidates who answer in a language other than English rate the experience higher than English-speaking candidates do, and they decline to take an AI interview at all less often. Our read is that they are comparing it against something different. For a candidate whose first language isn't English, the usual alternative is a screening call conducted in their second language, on someone else's schedule, with an interviewer who may or may not be patient about it. Measured against that, a structured interview in their own language is a relief. We can't tell you whether it holds outside our own book of business, so treat it as something to check on your own candidates rather than as a settled finding.
Here are three numbers from our own data, with their denominators where we have them. Of candidates who apply to a role and are invited to interview, 80% go on to complete an interview. The denominator is candidates we invited, not everyone who applied, and it is an aggregate, so split it on your own data. For blue-collar roles, 10 to 15% of interviews are conducted in a language other than English, and Spanish is the most common by a distance. Average candidate satisfaction is 4.6 out of 5.
Tenzo runs structured interviews across phone, video and text, sitting alongside your ATS rather than replacing it. Recruiters review full transcripts, interview design is configurable, accommodation paths are documented, and there's version history for audit. Humans review throughout and make all final decisions. Underneath, Tenzo runs multiple models in parallel, for redundancy.
On test #2, we hold ourselves to the same standard we just asked you to demand. In any evaluation we measure per-variety word error rate on your own audio, in the session, and show you the failures alongside the averages. An average hides the varieties that fail, and those are the ones your candidates are speaking.
On test #5, we measure the recruiter transcript-open rate with you during a pilot, on your own data, week by week.
To run the six tests against our platform with your own audio and your own roles, book a 30-minute working session. Bring recordings from your worst market.
FAQ
What is multilingual AI interviewing? Multilingual AI interviewing is a structured interview conducted by an AI system in the candidate's own language, usually by phone, video or text, with responses transcribed, evaluated against a defined rubric, and passed to a human recruiter for review. It differs from translated interviewing in that the candidate speaks their own language throughout.
Is AI interviewing legal in the EU? Yes, subject to conditions. AI used for recruitment and candidate evaluation is classified as high-risk under Annex III, point 4(a) of the EU AI Act, which brings obligations covering risk management, data governance, human oversight, logging and technical documentation. Those obligations now apply from 2 December 2027 rather than 2 August 2026, following the AI Omnibus that entered into force on 27 July 2026. The Article 50 transparency duties were not deferred and have applied since 2 August 2026, so candidates must be told they are interacting with an AI system today.
How accurate is AI speech recognition across accents? Accuracy varies substantially, and the variation is not random. Published research on five commercial systems found average word error rates of 0.35 for Black speakers against 0.19 for white speakers of American English, and a 2025 analysis of four commercial systems found the disparity persisting, tracing it to acoustic modelling of dialectal phonetic variation rather than to vocabulary or grammar. Performance on second-language accents, regional dialects and code-switched speech is generally worse than on the standard variety of the same language.
Does code-switching break AI interviews? Code-switching degrades them. Benchmark data covering code-switched speech beyond a handful of high-resource pairs only became available in 2025, so vendor confidence in this area is newer than it sounds. Where candidates habitually mix languages, as with Taglish, Hinglish or Spanglish, test on code-switched samples rather than monolingual ones.
Is a difference in pass rates between languages a legal problem? Potentially yes, in the United States. Under 29 CFR 1606, an individual's primary language is often treated as an essential national-origin characteristic, and national origin is protected under Title VII. A difference in selection rates between language groups can therefore support an adverse-impact claim, particularly where the mechanism is speech recognition performing worse on accented speech. The four-fifths rule at 29 CFR 1607.4(D) is the usual screening threshold.
Do we need a bias audit for AI interviews? In New York City, yes. Local Law 144 requires an independent bias audit within the preceding twelve months, a public summary, and ten business days' notice to candidates. Elsewhere in the US no statute mandates one, but Title VII and the Uniform Guidelines on Employee Selection Procedures still apply, so an adverse impact analysis remains the practical way to show a selection procedure is job-related.
What accessibility standard applies to a voice interview? WCAG 2.2 at Level AA is the usual contractual benchmark, with EN 301 549 the European equivalent. The criterion most often missed in voice flows is 2.2.1 Timing Adjustable, because capped response windows and countdowns are time limits under the standard.
How many languages does an AI interview platform need to support? Fewer than most buyers assume. Count the languages you hired in over the past twelve months, weight by volume, and evaluate depth on the top five. Verified performance across six live languages beats a spec sheet advertising eighty.



