Insight
•
Multilingual legal AI requires data, not just better models
Multilingual legal AI requires data, not just better models
Multilingual legal AI requires data, not just better models

When legal AI struggles with cross-border meaning, the instinctive response is often to look for a better model. Better models will undoubtedly solve some problems, but legal meaning remains jurisdiction-specific and develops through doctrine, legislation, case law, procedure, institutions, history and legal culture. Two systems may address similar problems while constructing the relevant concepts differently.
Large language models can infer a great deal from text. However, some of the distinctions that matter most in comparative law are subtle, inconsistently documented or apparent only when legal rules are considered in context. Comparative-law analysis consequently involves more than finding similar wording. It asks what a concept does, what conditions govern it, what legal consequences follow and where an apparently similar concept in another system begins to diverge.
This information does not automatically emerge as a structured resource simply because a model has been exposed to large volumes of legal text. It has to be identified, organised and, in many cases, deliberately created. A system supporting cross-border legal work therefore benefits from structured representations of concepts, jurisdiction-specific context and mappings between legal systems in addition to access to source documents.
Prompting can improve how a question is framed, retrieval can supply relevant sources and better models can improve reasoning and synthesis. However, where the relevant comparative information is absent, these approaches have limited means of identifying it. If a system has no information showing that two concepts overlap only partially, for example, it has little basis on which to explain the boundary between them.
The practical implication is that organisations deploying legal AI internationally should ask questions about the information surrounding the model as well as the model itself. Where does jurisdiction-specific knowledge come from? How are comparative relationships represented? Who made those judgements? Can uncertainty and non-equivalence be expressed?
At TransLegal, we are addressing these questions by building structured, human-curated and AI-assisted legal data around legal concepts and their relationships across jurisdictions. As model capabilities continue to improve, the quality of the legal information available to them will remain an important part of cross-border reliability.
When legal AI struggles with cross-border meaning, the instinctive response is often to look for a better model. Better models will undoubtedly solve some problems, but legal meaning remains jurisdiction-specific and develops through doctrine, legislation, case law, procedure, institutions, history and legal culture. Two systems may address similar problems while constructing the relevant concepts differently.
Large language models can infer a great deal from text. However, some of the distinctions that matter most in comparative law are subtle, inconsistently documented or apparent only when legal rules are considered in context. Comparative-law analysis consequently involves more than finding similar wording. It asks what a concept does, what conditions govern it, what legal consequences follow and where an apparently similar concept in another system begins to diverge.
This information does not automatically emerge as a structured resource simply because a model has been exposed to large volumes of legal text. It has to be identified, organised and, in many cases, deliberately created. A system supporting cross-border legal work therefore benefits from structured representations of concepts, jurisdiction-specific context and mappings between legal systems in addition to access to source documents.
Prompting can improve how a question is framed, retrieval can supply relevant sources and better models can improve reasoning and synthesis. However, where the relevant comparative information is absent, these approaches have limited means of identifying it. If a system has no information showing that two concepts overlap only partially, for example, it has little basis on which to explain the boundary between them.
The practical implication is that organisations deploying legal AI internationally should ask questions about the information surrounding the model as well as the model itself. Where does jurisdiction-specific knowledge come from? How are comparative relationships represented? Who made those judgements? Can uncertainty and non-equivalence be expressed?
At TransLegal, we are addressing these questions by building structured, human-curated and AI-assisted legal data around legal concepts and their relationships across jurisdictions. As model capabilities continue to improve, the quality of the legal information available to them will remain an important part of cross-border reliability.


