EuroBERT: Scaling Multilingual Encoders for European Languages

Boizard, Nicolas; Gisserot-Boukhlef, Hippolyte; Alves, Duarte M.; Martins, André; Hammal, Ayoub; Corro, Caio; Hudelot, Céline; Malherbe, Emmanuel; Malaboeuf, Etienne; Jourdan, Fanny; Hautreux, Gabriel; Alves, João; El-Haddad, Kevin; Faysse, Manuel; Peyrard, Maxime; Guerreiro, Nuno M.; Fernandes, Patrick; Rei, Ricardo; Colombo, Pierre

Computer Science > Computation and Language

arXiv:2503.05500 (cs)

[Submitted on 7 Mar 2025 (v1), last revised 26 Mar 2025 (this version, v2)]

Title:EuroBERT: Scaling Multilingual Encoders for European Languages

Authors:Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo

View PDF

Abstract:General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in generative decoder-only models. However, many innovations driving this progress are not inherently tied to decoders. In this paper, we revisit the development of multilingual encoders through the lens of these advances, and introduce EuroBERT, a family of multilingual encoders covering European and widely spoken global languages. Our models outperform existing alternatives across a diverse range of tasks, spanning multilingual capabilities, mathematics, and coding, and natively supporting sequences of up to 8,192 tokens. We also examine the design decisions behind EuroBERT, offering insights into our dataset composition and training pipeline. We publicly release the EuroBERT models, including intermediate training checkpoints, together with our training framework.

Comments:	28 pages, 8 figures, 13 tables
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2503.05500 [cs.CL]
	(or arXiv:2503.05500v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2503.05500

Submission history

From: Nicolas Boizard [view email]
[v1] Fri, 7 Mar 2025 15:13:58 UTC (2,349 KB)
[v2] Wed, 26 Mar 2025 18:43:59 UTC (1,543 KB)

Computer Science > Computation and Language

Title:EuroBERT: Scaling Multilingual Encoders for European Languages

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:EuroBERT: Scaling Multilingual Encoders for European Languages

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators