Over 90% of foundation-model training data is in English. Over 90% of the planet speaks another language.
Over 90% of foundation-model training data is in English. Over 90% of the planet speaks another language.
Over 90% of foundation-model training data is in English. Over 90% of the planet speaks another language.

The numbers
By most estimates, over 90 percent of foundation-model training data is in English, and most of it is American. Under 10 percent is left for every other legal system to share. Models default to the dominant framing in their training data. So for Croatia, Poland, the Philippines or Thailand, the information may simply not be there.
Why legal content makes it worse
Law is language. Legal meaning is jurisdiction specific, shaped by doctrine, case law, institutions, history and procedure. A model trained overwhelmingly on American legal text does not just miss vocabulary. It imports assumptions about remedies, enforceability and procedure that quietly reshape its answers about other systems. Fluency in the output conceals the gap in the input.
Where the approximation breaks down
TransLegal’s research maps where that approximation breaks down, jurisdiction by jurisdiction. Concepts are compared on purpose, scope, conditions of application and legal effect, then graded: functionally analogous, partially overlapping or structurally distinct.
A gap that scale won’t close
The missing information was never written down, so more training data of the same kind does not solve the problem. TransLegal builds the necessary knowledge directly: comparative research done once, in advance, with more than twenty university law faculties and hundreds of lawyer-linguists, across projects in 65+ countries, and licensed as a data layer beneath any model.
Founded 1989 · 20+ university law faculties · hundreds of lawyer-linguists · 65+ countries · 65+ jurisdictional datasets
See where the gap is mapped
Our data demo illustrates our term-by-term comparative layer. The article series explains the reasoning behind what we’re doing.
See the data demo · Read the article series
Can your AI explain why two legal concepts aren’t equivalent?
The numbers
By most estimates, over 90 percent of foundation-model training data is in English, and most of it is American. Under 10 percent is left for every other legal system to share. Models default to the dominant framing in their training data. So for Croatia, Poland, the Philippines or Thailand, the information may simply not be there.
Why legal content makes it worse
Law is language. Legal meaning is jurisdiction specific, shaped by doctrine, case law, institutions, history and procedure. A model trained overwhelmingly on American legal text does not just miss vocabulary. It imports assumptions about remedies, enforceability and procedure that quietly reshape its answers about other systems. Fluency in the output conceals the gap in the input.
Where the approximation breaks down
TransLegal’s research maps where that approximation breaks down, jurisdiction by jurisdiction. Concepts are compared on purpose, scope, conditions of application and legal effect, then graded: functionally analogous, partially overlapping or structurally distinct.
A gap that scale won’t close
The missing information was never written down, so more training data of the same kind does not solve the problem. TransLegal builds the necessary knowledge directly: comparative research done once, in advance, with more than twenty university law faculties and hundreds of lawyer-linguists, across projects in 65+ countries, and licensed as a data layer beneath any model.
Founded 1989 · 20+ university law faculties · hundreds of lawyer-linguists · 65+ countries · 65+ jurisdictional datasets
See where the gap is mapped
Our data demo illustrates our term-by-term comparative layer. The article series explains the reasoning behind what we’re doing.
See the data demo · Read the article series
Can your AI explain why two legal concepts aren’t equivalent?


