● LIVEECB RATE 2.4%·EUR/USD 1.1643·EUR/GBP 0.8572·BRUSSELS 19°C — MAINLY CLEAR·EU AI RULES MUST GUARANTEE CHATBOT PERFORMANCE IN ALL OFFICIAL LANGUAGES·EU FACES CRITICISM OVER FOCUS ON REFORMS INSTEAD OF SANCTIONS AS ISRAELI ELECTION CANDIDATES REJECT PALESTINIAN STATE·ICELAND TO HOLD REFERENDUM ON RESTARTING EU ACCESSION TALKS·RECORD LOW RIVER LEVELS EXPOSE HISTORIC HUNGER STONES AND THREATEN EUROPEAN SUPPLY CHAINS·PUTIN WARNED OF POSSIBLE ESCALATION AFTER CIA VISIT, ANALYSTS WARN OF NUCLEAR BRINKMANSHIP·MOLDOVA'S INDEPENDENCE PARADE DEMANDS CONCRETE SECURITY GUARANTEES FROM EUROPE·WEEK 35 · VOL. XV · N°241·● LIVEECB RATE 2.4%·EUR/USD 1.1643·EUR/GBP 0.8572·BRUSSELS 19°C — MAINLY CLEAR·EU AI RULES MUST GUARANTEE CHATBOT PERFORMANCE IN ALL OFFICIAL LANGUAGES·EU FACES CRITICISM OVER FOCUS ON REFORMS INSTEAD OF SANCTIONS AS ISRAELI ELECTION CANDIDATES REJECT PALESTINIAN STATE·ICELAND TO HOLD REFERENDUM ON RESTARTING EU ACCESSION TALKS·RECORD LOW RIVER LEVELS EXPOSE HISTORIC HUNGER STONES AND THREATEN EUROPEAN SUPPLY CHAINS·PUTIN WARNED OF POSSIBLE ESCALATION AFTER CIA VISIT, ANALYSTS WARN OF NUCLEAR BRINKMANSHIP·MOLDOVA'S INDEPENDENCE PARADE DEMANDS CONCRETE SECURITY GUARANTEES FROM EUROPE·WEEK 35 · VOL. XV · N°241·
Brussels · Est. 2012
Independent European journalism

Unionpress

Vol. XV · N°241
Saturday, 29 August 2026
Home/Opinion/ai-chatbots-need-to-work-in-all-eu-languages-not-just-english
Opinion29 August 2026

EU AI rules must guarantee chatbot performance in all official languages

New enforcement powers for the EU AI Office raise the risk that multilingual public services will be tested only in English, leaving citizens of smaller language groups with inferior access.

EU AI rules must guarantee chatbot performance in all official languages

EU regulators have just received expanded powers to enforce the bloc's artificial‑intelligence rules, but the framework still leaves a crucial gap: there is no requirement that high‑risk systems be tested in each of the Union's 24 official languages. For citizens who rely on public‑service chatbots, a missing clause could mean incomplete or even misleading answers, effectively creating a two‑tier system of digital rights.

New enforcement powers, old language blind spot

On 2 August the European Commission activated a second phase of its AI governance regime. The AI Office and national supervisory bodies can now impose fines and demand corrective measures when providers breach the rules that will become fully applicable in December 2027 for most high‑risk applications and in August 2028 for the most sensitive uses.

Article 15 of the AI Act obliges the Commission to promote benchmarks that assess accuracy, robustness and cybersecurity. Article 10 adds that training, validation and testing data must reflect the geographical, contextual, behavioural and functional setting in which the system will operate. While the text does not spell out a blanket duty to test every model in every language, it does set a clear principle: evidence of performance must match the real‑world context of use.

Language is part of that context. A chatbot that correctly interprets a French eligibility rule but omits a crucial condition when queried in Romanian is not merely a translation error, it is a denial of equal access to a public service. The problem is not abstract multilingualism; it is the concrete risk that a system's compliance score will hide language‑specific failures.

Evidence from recent benchmarks

Recent independent studies confirm the worry. The MuBench project, which evaluated a range of large language models across 61 languages, found that many systems performed well enough to be deemed "fluent" in low‑resource languages but fell short on factual accuracy and controllability. A 2026 benchmark focusing on European versus Brazilian Portuguese (P3B3) showed that most models favoured the Brazilian variety, with noticeable drops in precision when prompted in European Portuguese.

These gaps matter most in domains where legal nuance is essential, recruitment, credit scoring, education and public‑service advice. A model that answers a question about social‑security eligibility correctly in English may omit a qualifying condition when the same query is posed in Polish, leading to misinformed decisions for thousands of citizens.

Why a narrow benchmark would be hard to reverse

Standards, procurement templates and testing protocols are being drafted now, ahead of the December 2027 deadline. Once a narrow benchmark, for example, an English‑only test suite, is embedded in conformity‑assessment procedures, changing it later will be costly and politically fraught. The EU therefore faces a choice: embed multilingual testing from the start or risk a future scramble to retrofit the system.

From a practical standpoint, Europe does not need 24 completely separate regulatory regimes, nor does every AI product have to be evaluated in every language. The proportionality principle should apply: a system intended for cross‑border banking, a pan‑EU public chatbot or a recruitment platform used in several member states must be assessed in the languages and dialects it is likely to encounter.

Four steps toward a multilingual evaluation commons

1. Funding native, domain‑specific test modules. Rather than relying on translated English queries, the EU should invest in test suites created by national regulators, universities and language experts. Translation can preserve vocabulary but often loses legal concepts, administrative idioms and cultural references that are vital for correct outcomes.

2. Reproducible, versioned results. A performance score without details on model version, prompt template, test date and number of repetitions offers little evidential value. Updates to a model should trigger targeted retesting in the languages where the system has material impact.

3. Language‑specific disclosures in public procurement. Suppliers should be required to attach a "language‑performance card" to conformity documentation, indicating which languages and varieties were tested, the size of the test set, known failure modes and the version of the model. Accuracy, safety and domain knowledge should be reported separately from fluency.

4. Incident reporting that records language. If a chatbot provides a correct answer in German but an incomplete one in Greek, the incident log must capture the language variable. Otherwise, authorities will see isolated errors rather than a pattern of unequal performance.

Implications for workers and citizens

For public‑sector employees, especially those in contact centres and citizen‑service desks, reliable AI assistance can reduce workload and free up time for complex cases. If the assistance works only in the majority language, staff in regions with smaller language communities will face higher pressure and may have to double‑check AI outputs manually, eroding the promised efficiency gains.

Consumers in multilingual regions, from Catalonia to the French‑speaking parts of Belgium, could see their rights diluted. A recruitment algorithm that misclassifies qualifications in Slovene but not in Italian could affect job prospects and reinforce existing labour‑market inequalities.

From a broader European perspective, the issue touches on the Union's commitment to equal treatment across its linguistic landscape. The AI Act is meant to protect citizens from harmful outcomes, yet a testing regime that privileges English risks creating a de‑facto second class of users whose interactions are measured indirectly.

Industry response and the road ahead

Technology providers argue that developing and maintaining test suites for 24 languages would be costly and could slow innovation. Some vendors claim that high‑quality multilingual models already exist and that English‑centric benchmarks are a pragmatic starting point.

Trade unions and consumer groups counter that the cost of unequal AI performance will be borne by citizens and workers, not by the firms that profit from the technology. They call for a transparent, language‑aware compliance regime that forces providers to demonstrate real‑world reliability, not just linguistic fluency.

European Parliament committees are expected to debate the language‑specific provisions in the coming months. If they adopt the four‑step approach outlined above, the EU could set a global standard for inclusive AI governance, ensuring that digital public services truly serve all 447 million citizens, regardless of the language they speak.

Until then, the risk remains that a regulation published in 24 languages will be enforced on evidence gathered almost exclusively in English, cementing a digital divide that runs parallel to the Union's cultural diversity.

■ ENDOpinion© UnionPress 2026