A second legal system: what breaks when a legal pipeline changes country
In short: taking a legal pipeline built for Canadian law to a second legal system is less about code that crashes than about assumptions: that a document is in English unless told otherwise, that a…
- published
- read time
- 5 min
- words
- 1,044
- lang
- en
- filed under
- Engineering
In short: taking a legal pipeline built for Canadian law to a second legal system is less about code that crashes than about assumptions: that a document is in English unless told otherwise, that a jurisdiction is a two-letter province code, that a court is a dataset code, and that the assistant describes itself as Canadian. The fix is to turn the jurisdiction into data instead of code.
What the Canadian pipeline looked like
The Canadian version was big and, by the end, boring in the good way. It covered case law and legislation from many public sources, federal and provincial. Every source followed the same three-part pattern:
- a scraper that discovers URLs and yields text one document at a time,
- a small adapter that turns each document into one standard record of text, metadata and the raw original,
- and a command-line script that can skip what already exists, do a dry run, wait between requests and stop after a set number of documents.
Behind them, one shared ingestion step did the chunking, the embedding, the storage of vectors and the metadata rows in PostgreSQL. When HTML parsing was not enough, an LLM filled in structured metadata. PDFs went through a text extractor. And when a source site blocked the scraper, a fallback could fetch the same law by its two-letter province code.
That shape is the reason a second legal system was possible at all. It is also where Canada was hiding.
Where Canada was hiding
Most of what breaks is not code that throws an exception. It is code that runs fine and is quietly Canadian. These are the places I would check first in any legal pipeline before it crosses a border:
| Assumption | Where it lived | What it has to become |
|---|---|---|
| Documents are English unless a flag says otherwise | A default language on the ingestion command, with bilingual sources handled one by one | A language field on every document and every chunk, set from the source, checked against the jurisdiction |
| A jurisdiction is a province | Two-letter province codes passed to the fallback scraper | A jurisdiction key plus its own levels of government |
| A court is a dataset code | Short court abbreviations used as dataset names | A court hierarchy declared per jurisdiction, not read off a string |
| Metadata looks like Canadian metadata | The LLM enrichment step's idea of which fields exist | Fields named by the jurisdiction profile, so the model is asked for what that system actually records |
| There is a universal fallback | One fallback source assumed to cover everything | No fallback assumed. Each source is an adapter, or it is not covered |
| The assistant is Canadian | The prose that describes the service, and its prompts | A scope sentence that comes from configuration |
The last row is my favourite, because it is so ordinary. A scope sentence in a description or a prompt can name a country long after the code has moved on. Nothing fails because of that line. But it is a perfect sample of the problem: the country was written into prose, prompts and defaults, not into anything a test would ever notice.
Canada had already taught me about two languages
Canada is not monolingual either. New Brunswick publishes its laws in English and French, and Quebec's legislation needed a source of its own. In the Canadian pipeline that was handled source by source, with a flag. That works when one or two sources are the exception. It stops working when more than one language is the normal case, because now every source is the exception.
The same goes for retrieval. A question in one language has to find a passage that may be stored in another. Before ingesting a large corpus in a new language, check what the embedder does with it on a handful of known pairs. Re-embedding a whole corpus afterwards is the expensive way to find out.
The fix: the jurisdiction is data
The pattern is small. A jurisdiction is a frozen object. Every adapter receives it, and every document leaves the adapter stamped with it, rather than inheriting it from whatever the script's defaults happened to be.
from dataclasses import dataclass
@dataclass(frozen=True)
class Jurisdiction:
key: str # "ca", "ca-qc", ...
languages: tuple[str, ...] # every language the law is published in
levels: tuple[str, ...] # e.g. ("federal", "provincial")
scope: str # the sentence the assistant uses about itself
def to_record(doc, jurisdiction: Jurisdiction) -> dict:
"""Every adapter returns this shape. The jurisdiction is stamped, never assumed."""
language = doc.language or jurisdiction.languages[0]
if language not in jurisdiction.languages:
raise ValueError(f"{doc.url}: {language!r} is not a language of {jurisdiction.key}")
return {
"text": doc.text,
"metadata": {
"jurisdiction": jurisdiction.key,
"language": language,
"level": doc.level,
"source_url": doc.url,
},
"raw": doc.raw,
}
The raise is the part that earns its keep. A document in a language the jurisdiction does not publish in is almost always a scraper picking up the wrong page, and it is far better to stop on the first one than to find a few thousand of them in the index later.
What does not need to change
Plenty of the Canadian design carries over as it is, and it is worth saying which parts, because they are the parts I would build the same way again:
- The three-part source pattern. A new source is still a scraper, an adapter and a script.
- A dry run and a document cap on every script. Fetching five documents and printing stats is still the first thing I run against any new site.
- Retry with exponential backoff on embedding rate limits, and ingesting one source at a time to stay inside them.
- PostgreSQL as the record of what was ingested, so a run that died halfway can resume by skipping what is already there.
Before your pipeline crosses a border
Search your codebase, prompts and READMEs for the name of your country, its courts and its language codes. Every hit is an assumption. Move each one into a single jurisdiction object, make every adapter take it, and make the pipeline refuse a document that does not match it. Then run each new source as a dry run on five documents before you let it anywhere near the index.
related