Europe's main AI safety law gained teeth on 2 August, and the standards that count as evidence have to cover all 24 EU languages before high risk rules apply in late 2027.
An EU public-service chatbot answers an eligibility question in French and drops a condition in Romanian. A citizen with the same facts gets unequal access to a public service, in a system that officially speaks 24 languages. That gap is the test the EU's new AI rules have to pass.
On 2 August 2026, the EU's main AI safety law, the AI Act, entered a new phase. The AI Office and national authorities gained enforcement powers over the general-purpose and prohibited-practices provisions, as the European Commission outlined in its enforcement framework. The high-risk provisions apply later, in December 2027 and August 2028. Those are the rules that govern hiring, credit, education, and public services, which is where the language question gets teeth.
Europe is a political community of 24 official languages. Citizens may address EU institutions in any of them and receive a reply in the same language. That right has not been wired into the AI systems those institutions buy.
The law's text is what makes that wiring mandatory. Article 15 of the AI Act requires an "appropriate level of accuracy" and robustness for high-risk systems, and tells the Commission to encourage benchmarks and measurement methods for assessing those qualities. Article 10 requires that the data used to train, validate and test those systems reflect "the geographical, contextual, behavioural and functional setting" where the system will run. When the system runs in 24 member states, language is part of that setting, because a system that misreads legal categories, instructions, or safeguards in a given language is not the same system the rule was written for.
"Multilingual support" is a supplier claim, not a measured compliance result. A model that posts a high score on an English benchmark can drop accuracy sharply when asked the same question in Polish, Portuguese, or Greek, especially on local law, recruitment, credit, education, or public services. MuBench, a multilingual evaluation published in the ACL 2026 Findings, tests 61 languages and has begun to expose those gaps. P3B3, a multi-turn benchmark on European and Brazilian Portuguese, goes further: it asks whether a model treats two varieties of the same language as the same jurisdiction, which is exactly the question a recruitment or benefits chatbot has to answer.
There is at least one European precedent. Bürokratt, operated by the Estonian Information System Authority, is a multilingual virtual assistant network that connects Estonian citizens to public services in their own language. Estonia is one country, not 24, but it shows the wiring is possible.
The hard part is the window. Harmonised standards, the conformity-assessment templates a vendor runs to prove a system meets the AI Act, and the procurement criteria public buyers use are all being designed now, before December 2027. If those documents accept an English-first aggregate score as evidence of "appropriate accuracy," that English-first test is what the French-versus-Romanian gap has to beat in court. Article 10 does not create a general duty to test every AI system in every European language, but it sets the principle that evidence must reflect the context of use. That principle has to be the place where disaggregated-by-language evidence gets required, not just encouraged.
Three asks belong on the standards bodies' table before the window closes. Per-language accuracy and robustness scores should be a default output of any benchmark cited in a conformity assessment, not an optional add-on. Procurement templates public buyers use to hire AI vendors should require evidence of performance in the languages those systems will actually run in, signed off by the buyer, not the vendor. Civil-society and academic multilingual evaluations, of the kind MuBench and P3B3 represent, should be listed in the AI Office's measurement methods, so that a French-only system cannot point to an English-aggregate score and call it compliance.
The August 2026 enforcement phase is the first sign that the AI Act has teeth. The high-risk provisions are where those teeth meet the daily work of public services, hiring, and credit. If the standards being written now treat language as a footnote, the 24-language promise becomes a press release. If they treat it as a measurement requirement, the chatbot that drops a condition in Romanian has a chance of failing the conformity test, which is the only place a citizen can meet it on equal terms.