Writing

A second legal system: what breaks when a legal pipeline changes country

In short: taking a legal pipeline built for Canadian law to a second legal system is less about code that crashes than about assumptions: that a document is in English unless told otherwise, that a…

published
read time
5 min
words
1,044
lang
en
filed under
Engineering

In short: taking a legal pipeline built for Canadian law to a second legal system is less about code that crashes than about assumptions: that a document is in English unless told otherwise, that a jurisdiction is a two-letter province code, that a court is a dataset code, and that the assistant describes itself as Canadian. The fix is to turn the jurisdiction into data instead of code.

What the Canadian pipeline looked like

The Canadian version was big and, by the end, boring in the good way. It covered case law and legislation from many public sources, federal and provincial. Every source followed the same three-part pattern:

  • a scraper that discovers URLs and yields text one document at a time,
  • a small adapter that turns each document into one standard record of text, metadata and the raw original,
  • and a command-line script that can skip what already exists, do a dry run, wait between requests and stop after a set number of documents.

Behind them, one shared ingestion step did the chunking, the embedding, the storage of vectors and the metadata rows in PostgreSQL. When HTML parsing was not enough, an LLM filled in structured metadata. PDFs went through a text extractor. And when a source site blocked the scraper, a fallback could fetch the same law by its two-letter province code.

That shape is the reason a second legal system was possible at all. It is also where Canada was hiding.

Where Canada was hiding

Most of what breaks is not code that throws an exception. It is code that runs fine and is quietly Canadian. These are the places I would check first in any legal pipeline before it crosses a border:

AssumptionWhere it livedWhat it has to become
Documents are English unless a flag says otherwiseA default language on the ingestion command, with bilingual sources handled one by oneA language field on every document and every chunk, set from the source, checked against the jurisdiction
A jurisdiction is a provinceTwo-letter province codes passed to the fallback scraperA jurisdiction key plus its own levels of government
A court is a dataset codeShort court abbreviations used as dataset namesA court hierarchy declared per jurisdiction, not read off a string
Metadata looks like Canadian metadataThe LLM enrichment step's idea of which fields existFields named by the jurisdiction profile, so the model is asked for what that system actually records
There is a universal fallbackOne fallback source assumed to cover everythingNo fallback assumed. Each source is an adapter, or it is not covered
The assistant is CanadianThe prose that describes the service, and its promptsA scope sentence that comes from configuration

The last row is my favourite, because it is so ordinary. A scope sentence in a description or a prompt can name a country long after the code has moved on. Nothing fails because of that line. But it is a perfect sample of the problem: the country was written into prose, prompts and defaults, not into anything a test would ever notice.

Canada had already taught me about two languages

Canada is not monolingual either. New Brunswick publishes its laws in English and French, and Quebec's legislation needed a source of its own. In the Canadian pipeline that was handled source by source, with a flag. That works when one or two sources are the exception. It stops working when more than one language is the normal case, because now every source is the exception.

The same goes for retrieval. A question in one language has to find a passage that may be stored in another. Before ingesting a large corpus in a new language, check what the embedder does with it on a handful of known pairs. Re-embedding a whole corpus afterwards is the expensive way to find out.

The fix: the jurisdiction is data

The pattern is small. A jurisdiction is a frozen object. Every adapter receives it, and every document leaves the adapter stamped with it, rather than inheriting it from whatever the script's defaults happened to be.

jurisdiction data, not code source A source B source C shared pipeline vectors + rows
The adapters and the pipeline keep their shape. What moves is everything they used to assume.
from dataclasses import dataclass


@dataclass(frozen=True)
class Jurisdiction:
    key: str                    # "ca", "ca-qc", ...
    languages: tuple[str, ...]  # every language the law is published in
    levels: tuple[str, ...]     # e.g. ("federal", "provincial")
    scope: str                  # the sentence the assistant uses about itself


def to_record(doc, jurisdiction: Jurisdiction) -> dict:
    """Every adapter returns this shape. The jurisdiction is stamped, never assumed."""
    language = doc.language or jurisdiction.languages[0]
    if language not in jurisdiction.languages:
        raise ValueError(f"{doc.url}: {language!r} is not a language of {jurisdiction.key}")
    return {
        "text": doc.text,
        "metadata": {
            "jurisdiction": jurisdiction.key,
            "language": language,
            "level": doc.level,
            "source_url": doc.url,
        },
        "raw": doc.raw,
    }

The raise is the part that earns its keep. A document in a language the jurisdiction does not publish in is almost always a scraper picking up the wrong page, and it is far better to stop on the first one than to find a few thousand of them in the index later.

What does not need to change

Plenty of the Canadian design carries over as it is, and it is worth saying which parts, because they are the parts I would build the same way again:

  • The three-part source pattern. A new source is still a scraper, an adapter and a script.
  • A dry run and a document cap on every script. Fetching five documents and printing stats is still the first thing I run against any new site.
  • Retry with exponential backoff on embedding rate limits, and ingesting one source at a time to stay inside them.
  • PostgreSQL as the record of what was ingested, so a run that died halfway can resume by skipping what is already there.

Before your pipeline crosses a border

Search your codebase, prompts and READMEs for the name of your country, its courts and its language codes. Every hit is an assumption. Move each one into a single jurisdiction object, make every adapter take it, and make the pipeline refuse a document that does not match it. Then run each new source as a dry run on five documents before you let it anywhere near the index.

related

Keep reading