Beyond English: Semantic Matching Across Languages
RESOURCES · EXPERTINI ATS

Beyond English: Semantic Matching Across Languages

Keyword matching fails at a border. Reading for meaning does not — and that difference decides who reaches a shortlist in a multilingual market.

6 min read · Updated July 2026 · Expertini Editorial

Recruitment technology has spent a decade moving from string matching to semantic matching: from asking whether a CV contains the word in the job description to asking whether it demonstrates the thing the job description is asking for. That shift is usually justified with English examples — 'led a team' against 'team leadership', 'Postgres' against 'PostgreSQL'. The examples are real, but they understate the case.

The stronger argument for semantic matching is not synonymy. It is that a person and a job are frequently written down in different languages, and no amount of string matching survives that. 'Teamleitung' and 'team leadership' share no characters at all. A keyword system scores that candidate zero on leadership, and does so silently — there is no error, no flag, just a shortlist that quietly excluded someone qualified.

This page is about that failure and what it takes to fix it properly, including where the honest limits are.

0characters shared by 'Teamleitung' and 'team leadership'
2languages held at once — the CV's and the job description's
1published formula, applied identically whatever the language
75interface languages; cross-lingual reading is not limited to them

01The failure is invisible, which is what makes it expensive

A recruiter who receives no German applicants for a Berlin role does not conclude that their screening is language-bound. They conclude there are no German applicants. The absence of a signal looks identical to the absence of candidates, and nothing in a conventional pipeline distinguishes them. This is why the problem persists in organisations that would fix it immediately if they could see it.

It compounds in the direction you would least want. The candidates most likely to write a CV in their own language are the ones least embedded in an English-speaking professional network — which is to say, precisely the people a company opening a new market is trying to reach. The screening artefact filters hardest against the population the expansion exists to hire.

02What reading across languages actually requires

Two things have to be true, and they are independent. The system must be able to *read* a document in one language against a requirement written in another, and it must *report* what it found in the language the reader works in. Systems that translate a CV into English before scoring it satisfy neither cleanly: translation loses the evidence trail, and a recruiter reading translated text cannot tell which words the candidate actually used.

Expertini's screening layer reads both documents in their original languages and extracts evidence per competency, citing what the source document said. A German CV assessed against an English job description yields English evidence that references the German original. The candidate is never asked to write in the employer's language; the employer is never asked to read in the candidate's.

The direction reverses cleanly too. An English CV against a German job description produces German evidence for the German-speaking hiring team, with the same competencies recognised. This is not translation in either direction — it is extraction against a requirement, performed in a model that holds both languages at once.

03Why the score is deliberately not the clever part

It would be straightforward to ask a language model for a match percentage and present it. Expertini does not, and the reason matters more in a multilingual setting than a monolingual one. The Candidate Match Score is computed by a fixed published formula from the extracted evidence — AI reads, arithmetic scores. Same evidence, same number, every time.

In one language that property buys reproducibility. Across languages it buys something stronger: it makes the language question separable. If a German candidate scores lower than an English one, that difference is traceable to which evidence was found, not to how a model happened to feel about the two documents. You can inspect the dimensions, see what was and was not evidenced, and argue with it. A single generated number gives you nothing to argue with, and gives a regulator nothing to audit — see auditable hiring reports.

It also bounds the blast radius of a translation problem. If cross-lingual reading misses something, it shows up as missing evidence on a named dimension, which is visible and correctable. It does not silently redistribute weight across the whole score.

04The limits, stated plainly

Cross-lingual extraction is a model capability, not a guarantee. It is strongest for languages with abundant training data and weaker for those without, and Expertini's 75 supported languages are emphatically not equal in this respect. A CV in a widely-written language will be read more reliably than one in a language with a thin digital corpus. We would rather name that asymmetry than let a language count imply uniform quality.

Domain vocabulary is the sharpest edge. Professional and legal titles are often untranslatable rather than merely untranslated — a German *Meister*, a French *cadre*, a Japanese 主任 carry structural meaning about seniority and qualification that has no clean English equivalent. A system that maps them to the nearest English word loses exactly the information a hiring manager needed. Where the evidence is genuinely ambiguous, the honest output is an unevidenced dimension a human can look at, not a confident guess.

And none of this removes the human. Cross-lingual screening widens the pool a hiring team can see; it does not decide who to hire, and our own published methodology says so explicitly.

05What this means for hiring in more than one market

The practical consequence is that language stops being a filter applied before assessment. A company hiring in six countries can run one requirement against candidates who wrote in six languages, and read the results in one. The alternative — a separate localised pipeline per market, or an implicit English-only requirement — either multiplies the operational cost or narrows the field to people comfortable in a second language.

That second option is worth naming for what it is. Requiring a CV in English is a proficiency filter, applied at the earliest stage, to a skill that many roles do not need. Sometimes it is a legitimate requirement. Frequently it is an artefact of the tooling, and it should not be mistaken for a hiring standard.

Frequently asked questions

Does the CV get translated before scoring?
No. Both documents are read in their original language and evidence is extracted against the job's competencies. Translating first would lose the evidence trail — you could no longer see which words the candidate actually used, only what a translator made of them.
Is a candidate disadvantaged by not writing in English?
Not by the mechanism. Evidence is extracted from the CV in whatever language it is written, and the score is computed from that evidence by a fixed formula. The honest caveat is that extraction quality varies with how well-resourced a language is, so this is a real engineering property rather than a guarantee of perfect parity.
Which language does the screening report come back in?
The hiring team's. They are the ones reading it. The CV is treated as evidence to be understood, not as a voice to be echoed back — so a German CV against an English role produces an English report citing the German source.
Can it handle a CV that mixes languages?
Usually — mixed-language CVs are common and the extraction reads the document as it is rather than assuming one language throughout. As with everything else here, what you get is evidence per dimension, so a passage that was not understood shows up as an unevidenced dimension rather than as a silently lower score.
Is semantic matching just a better keyword search?
No, and the distinction is easiest to see across languages. A better keyword search finds more strings. Semantic matching asks whether the document demonstrates the requirement — which is why it can credit 'Teamleitung' for a team-leadership requirement, something no string search can do at any level of sophistication.
Does this mean the AI decides who gets hired?
No. The AI reads and extracts; a published formula scores; a human decides. Our own methodology papers state that a match score is a screening aid and never a hiring decision, and cross-lingual reading does not change that.

At a glance

  • Keyword matching scores 'Teamleitung' as zero evidence of team leadership
  • The failure is silent — it looks like an absence of candidates
  • Both documents are read in their original language; neither is translated first
  • AI reads, a published formula scores — so the language question stays separable
  • Extraction quality varies by language, and we say which way
  • Requiring an English CV is a proficiency filter, not a hiring standard

See beyond english: semantic matching across languages on your own hiring.

Bring a real job description to a 30-minute demo — free trial included.

Book a demo
Expertini AI
Online now
Hi! I'm Expertini's AI Product Expert. Ask me anything about our solutions, get guidance on any of our Hiring Tools, or just tell me what you're trying to do — I'll point you in the right direction. For account-specific issues, email support@expertini.com.