Ingesting a country's law when every province has its own format
In short: Canada's legislation lives on a different website for nearly every province and territory, as PDFs, Word files, an XML API and plain web pages. When a product depends on that text, the…
- published
- read time
- 5 min
- words
- 962
- lang
- en
- filed under
- Engineering
In short: Canada's legislation lives on a different website for nearly every province and territory, as PDFs, Word files, an XML API and plain web pages. When a product depends on that text, the scraper for each one is not plumbing. It is the product. What made it manageable was giving every source the same three-part shape, so nothing after the scraper ever knows which website a law came from.
I was working on a product that depends on public legal text, and the first real job was getting that text in. Not a sample. All of it: federal, provincial and territorial legislation, and on top of that a large body of case law. Federal legislation alone runs to several thousand statutes and regulations. The case law is bigger again: one provincial superior court alone accounts for tens of thousands of decisions.
Every source, its own headache
Each jurisdiction publishes its own laws, its own way. Here is what the legislation side looks like.
| Jurisdiction | Where it comes from | What you get |
|---|---|---|
| Federal | Justice Laws website | |
| Ontario | e-Laws | .doc files |
| Quebec | Légis Québec | HTML |
| British Columbia | BC Laws | XML API |
| Alberta | A public legal aggregator | HTML |
| Saskatchewan | A public legal aggregator | HTML |
| Manitoba | A public legal aggregator | HTML |
| New Brunswick | Provincial laws site | Web pages, English and French |
| Nova Scotia | Legislature and government sites | Web pages |
| Prince Edward Island | Provincial government site | Web pages on a CMS |
| Newfoundland and Labrador | House of Assembly site | Web pages |
| Yukon | Territorial laws site | PDF behind a CMS |
| Northwest Territories | Department of Justice site | |
| Nunavut | Territorial legislation site | PDF embedded in pages |
Even where the format column repeats, nothing else does. Two "web pages" sources have different URL schemes, different ways to list every act, different ways to mark a regulation versus a statute, and different ideas about where the title is. One format per source is not an exaggeration. If anything, it is an undercount.
One shape for every source
The rule that saved me: every source gets the same three pieces, and only those pieces know anything about the website.
- A scraper that discovers URLs and yields documents one at a time from a generator.
- An adapter that turns one raw document into a standard record: text, metadata and the raw original.
- A command-line script that wires them to the shared pipeline, with the same options every time.
Here is the shape, stripped down:
import time
# the scraper: knows the website
def iter_documents(session, delay=0.5):
for url in discover_urls(session):
yield {"url": url, "body": fetch(session, url)}
time.sleep(delay)
# the adapter: knows the website's markup
class ExampleAdapter:
jurisdiction = "NB"
def to_record(self, raw):
text = extract_text(raw["body"])
return {
"text": text,
"metadata": {
"jurisdiction": self.jurisdiction,
"source_url": raw["url"],
"doc_type": guess_doc_type(raw),
"language": detect_language(text),
},
"raw": raw,
}
Everything after the adapter is shared. One pipeline does chunking, embedding, storing the vectors and writing the metadata to Postgres. Every script can do a dry run, cap the number of documents, skip what already exists and wait between requests. When a source breaks, and they break, I fix one folder and nothing else moves.
What each format costs you
PDFs
The federal and territorial laws come as PDFs, and PDFs are slow and full of surprises. For the federal set I split the work in two passes. The first pass reads only the first few pages and the last page of each file and writes a metadata CSV, so I can see titles, dates and oddities before spending anything on embeddings. The second pass extracts the full text, chunks it at roughly a thousand characters with some overlap, and ingests. Text extraction uses pdfplumber with pypdf behind it.
Word files
Ontario gives you .doc downloads. They are easier to read than PDFs, but they are files you download, not pages you parse, so the scraper is really a download manager with a converter on the end.
The XML API
British Columbia is the easy one. Structure comes for free: titles, sections and types are in the markup. If every jurisdiction did this, this post would be one paragraph.
Web pages
The HTML sources are where layouts change without warning and metadata hides in odd places. When the parser cannot find a field reliably, a language model call reads the page and returns the structured metadata instead. It is a fallback, not the default.
Before you scrape a jurisdiction
If you are about to do this for your own country, or even one province, here is the checklist I wish I had started with:
- Make the table above for your sources first. Format, entry point, how to list everything, which language. It is your project plan.
- Fix the adapter output shape before writing a single scraper. Every source must hand over the same dictionary.
- Put jurisdiction, source URL, document type and language on every chunk. You cannot filter or cite what you did not store.
- Be polite. Use a delay, run one source at a time, and read each site's terms. Embedding APIs have rate limits too, so build in retries with backoff.
- Make skipping what already exists work early. You will rerun everything more times than you think.
- Start with the source that has an API. Get one jurisdiction all the way into search before you start the second.
related