Better models still need better legal data.
Better models still need better legal data.
Better models still need better legal data.

The fair case for waiting
Waiting for better models costs nothing and requires no decision. Each generation is cheaper and more capable than the last, multilingual performance is genuinely improving, and much of 2024’s legal AI tooling has simply been absorbed into base models. For many capabilities, waiting is the correct call.
Why the gap stays where it is
Multilingual legal AI does not fail because models lack power. It fails because they lack data showing how legal concepts function across systems. Legal meaning is not universal: it is the product of culture, history, doctrine, procedure and institutional development. And the comparison between legal systems was never written down systematically, because comparative lawyers resolved it in their heads and in prose. It is in no corpus. Scale cannot create data that was never written.
Why tooling can’t repair it downstream
Prompting, retrieval and interface design improve presentation and relevance, but they cannot supply what is absent. If a system holds no information on how two concepts differ in scope or effect, it has no basis on which to flag the difference. Ontologies, knowledge graphs and term banks are containers: they give you somewhere to put a comparative judgement, they do not make the judgement, and retrieval returns only what the bank records. Fine tuning on parallel legal corpora teaches the conventional translation, which is the trap, and hardens it into the weights. No degree of downstream optimisation can reconstruct distinctions that were never encoded in the first place.
The data problem, solved as a data problem
TransLegal’s approach: do the comparative research once, in advance, with lawyer-linguists and university law faculties, and license it as a data layer beneath any model. Purpose, scope, conditions of application and legal effect are compared term by term. Non-equivalence is recorded instead of smoothed away. The core is human curated, with AI generated layers built with experts in the loop. It is complementary to every model and competitive with none.
65+ jurisdictional datasets · 20+ university law faculties · hundreds of lawyer-linguists
See the data your model is missing
Our data demo shows the comparative layer directly: graded equivalence, comparative notes, non-equivalence flags.
See the data demo · Talk to us
The gap is not a capability problem. It is a data problem. Waiting does not close it.
The fair case for waiting
Waiting for better models costs nothing and requires no decision. Each generation is cheaper and more capable than the last, multilingual performance is genuinely improving, and much of 2024’s legal AI tooling has simply been absorbed into base models. For many capabilities, waiting is the correct call.
Why the gap stays where it is
Multilingual legal AI does not fail because models lack power. It fails because they lack data showing how legal concepts function across systems. Legal meaning is not universal: it is the product of culture, history, doctrine, procedure and institutional development. And the comparison between legal systems was never written down systematically, because comparative lawyers resolved it in their heads and in prose. It is in no corpus. Scale cannot create data that was never written.
Why tooling can’t repair it downstream
Prompting, retrieval and interface design improve presentation and relevance, but they cannot supply what is absent. If a system holds no information on how two concepts differ in scope or effect, it has no basis on which to flag the difference. Ontologies, knowledge graphs and term banks are containers: they give you somewhere to put a comparative judgement, they do not make the judgement, and retrieval returns only what the bank records. Fine tuning on parallel legal corpora teaches the conventional translation, which is the trap, and hardens it into the weights. No degree of downstream optimisation can reconstruct distinctions that were never encoded in the first place.
The data problem, solved as a data problem
TransLegal’s approach: do the comparative research once, in advance, with lawyer-linguists and university law faculties, and license it as a data layer beneath any model. Purpose, scope, conditions of application and legal effect are compared term by term. Non-equivalence is recorded instead of smoothed away. The core is human curated, with AI generated layers built with experts in the loop. It is complementary to every model and competitive with none.
65+ jurisdictional datasets · 20+ university law faculties · hundreds of lawyer-linguists
See the data your model is missing
Our data demo shows the comparative layer directly: graded equivalence, comparative notes, non-equivalence flags.
See the data demo · Talk to us
The gap is not a capability problem. It is a data problem. Waiting does not close it.


