Insight
•
Where comparative mapping determines reliability
Where comparative mapping determines reliability
Where comparative mapping determines reliability

Training data in legal AI is not a neutral collection of legal text. It reflects choices about jurisdictional scope, categories, annotation and the relationships between concepts. Once those choices are embedded upstream, they influence what a system can reliably preserve downstream.
Consider the term consideration. In common law contract doctrine, consideration has a particular role in determining whether promises are enforceable. Civil law systems can recognise binding contractual obligations without employing a directly corresponding doctrine.
A cross-border dataset therefore has several options. It could identify another concept as the equivalent of consideration, describe concepts as functionally analogous while preserving their structural differences or record that there is no direct equivalent and explain how the relevant legal function is performed instead. Each approach creates a different comparative mapping.
If distinct legal constructions are collapsed into a shared category too early, the model inherits that compression and may produce a fluent comparison while obscuring the fact that the underlying systems reach similar outcomes through different doctrinal structures. Comparative mapping can therefore become part of a reliability problem.
This also helps explain why simply adding more legal text does not necessarily resolve cross-border accuracy. Greater volume increases the information available to a model, but does not automatically provide an explicit comparative structure explaining how concepts from different systems relate. That structure has to be developed separately.
At TransLegal, our approach is to treat comparative relationships as data in their own right. Definitions, contextual information, equivalence assessments and comparative explanations can then form part of the information available to downstream applications. This does not mean that every legal relationship can be reduced to a simple score or classification, as comparative judgement remains necessary. Rather, the aim is to preserve that judgement in a form AI systems can use.
Two legal AI systems may appear similarly capable at surface level. An important difference between them may lie much further upstream, in how carefully the legal world was mapped before the model was asked to reason about it.
Training data in legal AI is not a neutral collection of legal text. It reflects choices about jurisdictional scope, categories, annotation and the relationships between concepts. Once those choices are embedded upstream, they influence what a system can reliably preserve downstream.
Consider the term consideration. In common law contract doctrine, consideration has a particular role in determining whether promises are enforceable. Civil law systems can recognise binding contractual obligations without employing a directly corresponding doctrine.
A cross-border dataset therefore has several options. It could identify another concept as the equivalent of consideration, describe concepts as functionally analogous while preserving their structural differences or record that there is no direct equivalent and explain how the relevant legal function is performed instead. Each approach creates a different comparative mapping.
If distinct legal constructions are collapsed into a shared category too early, the model inherits that compression and may produce a fluent comparison while obscuring the fact that the underlying systems reach similar outcomes through different doctrinal structures. Comparative mapping can therefore become part of a reliability problem.
This also helps explain why simply adding more legal text does not necessarily resolve cross-border accuracy. Greater volume increases the information available to a model, but does not automatically provide an explicit comparative structure explaining how concepts from different systems relate. That structure has to be developed separately.
At TransLegal, our approach is to treat comparative relationships as data in their own right. Definitions, contextual information, equivalence assessments and comparative explanations can then form part of the information available to downstream applications. This does not mean that every legal relationship can be reduced to a simple score or classification, as comparative judgement remains necessary. Rather, the aim is to preserve that judgement in a form AI systems can use.
Two legal AI systems may appear similarly capable at surface level. An important difference between them may lie much further upstream, in how carefully the legal world was mapped before the model was asked to reason about it.


