Cross-border legal AI often doesn't fail by hallucinating. It fails by being almost right.
Cross-border legal AI often doesn't fail by hallucinating. It fails by being almost right.
Cross-border legal AI often doesn't fail by hallucinating. It fails by being almost right.

Why almost right is worse than wrong
A legal term carries embedded assumptions about enforceability, remedies, procedure and interpretation. When a familiar term appears, those assumptions fire automatically, and readers fill the gaps without realising it. So the wording can be correct and still invite the reader to apply the wrong legal framework. The text is credible, precise and patterned the way lawyers are trained to recognise. That is exactly what makes the failure hard to detect.
Only a lawyer qualified in both legal systems catches the error: the scarce, expensive person the tool was meant to replace. Nobody knows the error rate. Including the vendor.
A worked example: the penalty clause
In common law, liquidated damages are enforceable only if they are a genuine pre-estimate of loss. Punitive clauses are struck down. In many civil law systems, a contractual penalty is enforceable even beyond compensatory loss, subject to judicial adjustment. A translation based system maps one term to the other without hesitation. It reads cleanly, but is still wrong in substance.
A similar pattern can be seen for terms like consideration, good faith, fiduciary duties, termination and bankruptcy, and in the divergence between authentic EU language versions.
Why no error rate captures it
Evaluating a cross-border answer needs a ground truth about cross-border meaning. That ground truth was never written down: comparative lawyers resolved it in their heads and in prose, so it sits in no corpus and no benchmark. If a system holds no information about how two concepts differ in scope or effect, it has no basis on which to flag the difference.
Making the difference visible
TransLegal’s World Law Dictionary records how legal concepts correspond across jurisdictions, and where they do not. Terms are mapped as functionally analogous, partially overlapping or structurally distinct, with non-equivalence explicitly identified and highlighted. Equivalence is graded, not binary: an equivalency score (for example 4.25 of 5) built from structured comparison of function, scope, requirements, effects and procedural context.
The result: “these two concepts do not correspond” becomes a machine readable fact your system can act on, before the assumption breaks downstream.
Founded 1989 · built with 20+ university law faculties · hundreds of lawyer-linguists · projects in 65+ countries
See the problem in the data
Our data demo flags non-equivalence term by term, with comparative notes and equivalency grading.
See the data demo · Talk to us
What happens when your legal AI is almost right? Who is accountable when it is wrong?
Why almost right is worse than wrong
A legal term carries embedded assumptions about enforceability, remedies, procedure and interpretation. When a familiar term appears, those assumptions fire automatically, and readers fill the gaps without realising it. So the wording can be correct and still invite the reader to apply the wrong legal framework. The text is credible, precise and patterned the way lawyers are trained to recognise. That is exactly what makes the failure hard to detect.
Only a lawyer qualified in both legal systems catches the error: the scarce, expensive person the tool was meant to replace. Nobody knows the error rate. Including the vendor.
A worked example: the penalty clause
In common law, liquidated damages are enforceable only if they are a genuine pre-estimate of loss. Punitive clauses are struck down. In many civil law systems, a contractual penalty is enforceable even beyond compensatory loss, subject to judicial adjustment. A translation based system maps one term to the other without hesitation. It reads cleanly, but is still wrong in substance.
A similar pattern can be seen for terms like consideration, good faith, fiduciary duties, termination and bankruptcy, and in the divergence between authentic EU language versions.
Why no error rate captures it
Evaluating a cross-border answer needs a ground truth about cross-border meaning. That ground truth was never written down: comparative lawyers resolved it in their heads and in prose, so it sits in no corpus and no benchmark. If a system holds no information about how two concepts differ in scope or effect, it has no basis on which to flag the difference.
Making the difference visible
TransLegal’s World Law Dictionary records how legal concepts correspond across jurisdictions, and where they do not. Terms are mapped as functionally analogous, partially overlapping or structurally distinct, with non-equivalence explicitly identified and highlighted. Equivalence is graded, not binary: an equivalency score (for example 4.25 of 5) built from structured comparison of function, scope, requirements, effects and procedural context.
The result: “these two concepts do not correspond” becomes a machine readable fact your system can act on, before the assumption breaks downstream.
Founded 1989 · built with 20+ university law faculties · hundreds of lawyer-linguists · projects in 65+ countries
See the problem in the data
Our data demo flags non-equivalence term by term, with comparative notes and equivalency grading.
See the data demo · Talk to us
What happens when your legal AI is almost right? Who is accountable when it is wrong?


