Writing

Ingesting a country's law when every province has its own format

In short: Canada's legislation lives on a different website for nearly every province and territory, as PDFs, Word files, an XML API and plain web pages. When a product depends on that text, the…

published
read time
5 min
words
962
lang
en
filed under
Engineering

In short: Canada's legislation lives on a different website for nearly every province and territory, as PDFs, Word files, an XML API and plain web pages. When a product depends on that text, the scraper for each one is not plumbing. It is the product. What made it manageable was giving every source the same three-part shape, so nothing after the scraper ever knows which website a law came from.

I was working on a product that depends on public legal text, and the first real job was getting that text in. Not a sample. All of it: federal, provincial and territorial legislation, and on top of that a large body of case law. Federal legislation alone runs to several thousand statutes and regulations. The case law is bigger again: one provincial superior court alone accounts for tens of thousands of decisions.

Every source, its own headache

Each jurisdiction publishes its own laws, its own way. Here is what the legislation side looks like.

JurisdictionWhere it comes fromWhat you get
FederalJustice Laws websitePDF
Ontarioe-Laws.doc files
QuebecLégis QuébecHTML
British ColumbiaBC LawsXML API
AlbertaA public legal aggregatorHTML
SaskatchewanA public legal aggregatorHTML
ManitobaA public legal aggregatorHTML
New BrunswickProvincial laws siteWeb pages, English and French
Nova ScotiaLegislature and government sitesWeb pages
Prince Edward IslandProvincial government siteWeb pages on a CMS
Newfoundland and LabradorHouse of Assembly siteWeb pages
YukonTerritorial laws sitePDF behind a CMS
Northwest TerritoriesDepartment of Justice sitePDF
NunavutTerritorial legislation sitePDF embedded in pages

Even where the format column repeats, nothing else does. Two "web pages" sources have different URL schemes, different ways to list every act, different ways to mark a regulation versus a statute, and different ideas about where the title is. One format per source is not an exaggeration. If anything, it is an undercount.

One shape for every source

The rule that saved me: every source gets the same three pieces, and only those pieces know anything about the website.

  • A scraper that discovers URLs and yields documents one at a time from a generator.
  • An adapter that turns one raw document into a standard record: text, metadata and the raw original.
  • A command-line script that wires them to the shared pipeline, with the same options every time.

Here is the shape, stripped down:

import time

# the scraper: knows the website
def iter_documents(session, delay=0.5):
    for url in discover_urls(session):
        yield {"url": url, "body": fetch(session, url)}
        time.sleep(delay)

# the adapter: knows the website's markup
class ExampleAdapter:
    jurisdiction = "NB"

    def to_record(self, raw):
        text = extract_text(raw["body"])
        return {
            "text": text,
            "metadata": {
                "jurisdiction": self.jurisdiction,
                "source_url": raw["url"],
                "doc_type": guess_doc_type(raw),
                "language": detect_language(text),
            },
            "raw": raw,
        }

Everything after the adapter is shared. One pipeline does chunking, embedding, storing the vectors and writing the metadata to Postgres. Every script can do a dry run, cap the number of documents, skip what already exists and wait between requests. When a source breaks, and they break, I fix one folder and nothing else moves.

PDF .doc XML HTML CMS scraper + adapter one per source pipeline chunk, embed vectors postgres knows the website never does
All the knowledge about a website stays in the shaded box. The pipeline after it is the same for every source.

What each format costs you

PDFs

The federal and territorial laws come as PDFs, and PDFs are slow and full of surprises. For the federal set I split the work in two passes. The first pass reads only the first few pages and the last page of each file and writes a metadata CSV, so I can see titles, dates and oddities before spending anything on embeddings. The second pass extracts the full text, chunks it at roughly a thousand characters with some overlap, and ingests. Text extraction uses pdfplumber with pypdf behind it.

Word files

Ontario gives you .doc downloads. They are easier to read than PDFs, but they are files you download, not pages you parse, so the scraper is really a download manager with a converter on the end.

The XML API

British Columbia is the easy one. Structure comes for free: titles, sections and types are in the markup. If every jurisdiction did this, this post would be one paragraph.

Web pages

The HTML sources are where layouts change without warning and metadata hides in odd places. When the parser cannot find a field reliably, a language model call reads the page and returns the structured metadata instead. It is a fallback, not the default.

CarefulA broken scraper does not crash. It ingests empty pages, error pages, CAPTCHA pages and garbled PDF columns, and all of it embeds perfectly well. Then it shows up in retrieval as nonsense with a real citation attached. Run every new source as a dry run on five documents and read the text before anything touches the database.

Before you scrape a jurisdiction

If you are about to do this for your own country, or even one province, here is the checklist I wish I had started with:

  1. Make the table above for your sources first. Format, entry point, how to list everything, which language. It is your project plan.
  2. Fix the adapter output shape before writing a single scraper. Every source must hand over the same dictionary.
  3. Put jurisdiction, source URL, document type and language on every chunk. You cannot filter or cite what you did not store.
  4. Be polite. Use a delay, run one source at a time, and read each site's terms. Embedding APIs have rate limits too, so build in retries with backoff.
  5. Make skipping what already exists work early. You will rerun everything more times than you think.
  6. Start with the source that has an API. Get one jurisdiction all the way into search before you start the second.

related

Keep reading