When Latvian is 0.09% of Common Crawl: what the EU Institutional LLM is for

When Latvian is 0.09% of Common Crawl: what the EU Institutional LLM is for

BackerLeader 3 15 133
calendar_today agoschedule5 min read
— Originally published at apogeewatcher.hashnode.dev

DG Translation is training an EU Institutional LLM on Euramis data and EuroHPC supercomputers, with all 24 official EU languages in scope. What that means for multilingual sites, benchmarks, and sovereign AI tooling.

A client site ships in six EU languages. Marketing runs every string through a proprietary LLM for tone checks. Latvian copy comes back fluent and slightly wrong: formal where the brand is plain, and vague on product terms the model rarely saw during training. Nobody lied about multilingual support. The training data did. DG Translation’s EU Institutional LLM programme names the skew plainly: in Common Crawl, Latvian is about 0.09% of the dataset, Irish about 0.07%, Maltese about 0.03%. The least-represented half of EU official languages add up to roughly 2.4%. Global models are not neutral on language; they mirror what was crawled.

The Commission’s response is institutional, not consumer-facing hype. Engineers are continuing pre-training of European open models (Mistral Mixtral 8x7B and 8x22B) on Euramis, the multilingual corpus from EU institutions, using EuroHPC supercomputers including MeluXina, Leonardo, and MareNostrum 5. The goal is an EU Institutional LLM with better coverage of all 24 official languages, tuned for public-sector use cases, with models available to EU-based legal entities through the European Language Data Space. That is a different bet from “pick the biggest US model and hope fine-tuning fixes Lithuanian.”

Why Common Crawl under-represents EU languages

Common Crawl is convenient because it is huge and open. It is not balanced. Web publishing volume, historical crawl bias, and English as a lingua franca concentrate tokens in a handful of languages. For low-resource EU languages, the model’s prior is thin even when your site is impeccable. Fine prompts help at the margin. They do not replace missing institutional vocabulary in legal, procurement, and public-health contexts.

DG Translation also calls out quality, copyright safety, transparency, and bias as reasons proprietary crawls are risky for public administrations. Agencies serving EU clients hit the same wall on a smaller scale: you can localise UI strings and hreflang tags correctly and still lose nuance in generated summaries, chat widgets, and internal knowledge tools trained on English-heavy corpora. The failure mode is quiet; the page looks translated, but procurement language, liability clauses, and product names drift toward English defaults the model saw more often.

What Euramis and EuroHPC change in the training story

Euramis is not a marketing blog scrape. It is a large, curated multilingual archive of EU institutional text, aligned with quality standards and, per the Commission, free of copyright infringements in the training pipeline described publicly. EuroHPC supplies compute at a scale most teams will never run themselves. The published results are directional but striking: on an EU institutional benchmark, the adapted model beat its Mistral base across tested languages, with Irish nearly quadrupling its score, Estonian up almost 80%, Greek nearly doubling, and Latvian and Lithuanian gaining on the order of 70–75%.

That does not mean your WordPress plugin should download v1 tomorrow and replace human reviewers. It does mean European public-sector and NGO buyers now have a documented path to models trained with EU language balance as a stated requirement, not an afterthought. The instruct-tuned variant is meant for administration workflows; eSummary already runs on the model for multilingual summarisation inside the Commission’s tooling.

EU MMLU and why English-only benchmarks mislead

Building the model is half the problem. Measuring it fairly is the other. Most multilingual benchmarks rely heavily on machine translation of English exams. DG Translation released EU MMLU, adapted from MMLU with human translation and revision through the European Master’s in Translation network, covering subjects from law and economics to public affairs. Sixteen languages are available now, with more planned toward full EU coverage.

For web teams, the lesson is operational: if you evaluate a model on English-only QA sets, you will ship confident dashboards and embarrassed Latvian FAQs. When you compare vendors or open weights for a multilingual client, ask which benchmark languages match the client’s markets, not which leaderboard screenshot looked best on Twitter. If EU MMLU (or an equivalent human-translated set) does not cover those markets yet, treat vendor “multilingual” claims as provisional until you spot-check real pages.

What agencies should do with multilingual sites while models catch up

Institutional LLMs complement, not replace, your publishing stack. You still need correct hreflang, stable URLs, crawlable HTML, and performance that keeps bots willing to fetch every locale. AI Overviews and other answer engines already change how clicks arrive; low-quality localised pages lose twice. Our Watcher guide on AI Overviews and CTR covers the traffic side. Pair that with locale-level monitoring: a fast English homepage does not excuse a Slovak product page that times out.

For AI visibility measurement across markets, start with crawl and index signals per language rather than one aggregate score. Are we visible in ChatGPT? What agencies can measure first outlines logs, Search Console slices, and performance checks you can run before you trust a vendor’s “multilingual AI” badge. A single portfolio average hides a slow /lv/ product page behind a healthy English homepage.

Apogee Watcher schedules PageSpeed tests per URL. Treat each locale’s priority templates as scheduled URLs in their own right, not as alternate links you check once after launch. Institutional models may improve summarisation and translation assist inside EU workflows; they do not fix broken hreflang or missing lang attributes on the open web.

How the EU Institutional LLM fits a wider sovereignty map

Compute on EuroHPC, weights in the Language Data Space, benchmarks like EU MMLU, and production tools like eSummary are pieces of the same story: Europe building AI capacity that matches its linguistic reality. We map the wider trend lines for developers and agencies in European tech sovereignty in 2026 (Hashnode #33). Read that companion when you need a single slide for a client QBR; read this one when they ask why their Baltic microsite still sounds generic in ChatGPT.

Layer, do not rip and replace: keep your CMS, your analytics, and your monitoring. Add sovereign models where procurement requires EU-hosted options. Keep human reviewers on anything customer-facing until EU MMLU scores exist for your client’s languages and you have spot-checked real pages.

What to check on a multilingual client site this week

List priority URLs per locale (home, pricing, contact, top three articles). For each, verify lang, hreflang reciprocals, and a manual read of the H1 (not machine-translated title tags alone). Run PageSpeed or scheduled monitoring per URL, not only on /en/. If the client uses AI assist for copy, log which model and which training claims the vendor makes; ask for per-language evals, not a single English MMLU screenshot.

When DG Translation publishes updated model weights, check whether your client’s sector can access them under Language Data Space rules. Public sites still live on the open web; institutional models do not remove the need for crawlable, fast, well-structured HTML. Keep the publishing checklist even if the summariser improves: bots and answer engines still fetch what you ship.

On Monday, pick one low-resource EU locale the client cares about and compare a human-edited paragraph with the same paragraph passed through their default LLM. The gap you see is the digital language skew problem in miniature. The EU Institutional LLM is one institutional answer. Your monitoring and publishing discipline is the part agencies still own.

Originally published on Hashnode.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

TypeScript Complexity Has Finally Reached the Point of Total Absurdity

Karol Modelskiverified - Apr 23

I’m a Senior Dev and I’ve Forgotten How to Think Without a Prompt

Karol Modelskiverified - Mar 19

Your Tech Stack Isn’t Your Ceiling. Your Story Is

Karol Modelskiverified - Apr 9

Snyk: Enterprises Can See a Third of Their Own AI Attack Surface. The Other Two-Thirds Is Where th

Tom Smithverified - Aug 4

The Audit Trail of Things: Using Hashgraph as a Digital Caliper for Provenance

Ken W. Algerverified - Apr 28
chevron_left
6.9k Points151 Badges
123Posts
31Comments
73Connections
We bring quality & creative intelligence since 2002. We design, develop and operate custom informati... Show more

Related Jobs

View all jobs →

Commenters (This Week)

4 comments
2 comments
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!