Classical ML or an LLM? The decision before touching code
In short: if the input is numbers or a signal and you have labelled examples, I start with a classical model. If the input is language and the job is to read or write it, I start with an LLM. Most of…
- published
- read time
- 5 min
- words
- 908
- lang
- en
- filed under
- Engineering
In short: if the input is numbers or a signal and you have labelled examples, I start with a classical model. If the input is language and the job is to read or write it, I start with an LLM. Most of the real work is being honest about which of those two problems you actually have.
Every few weeks someone asks me whether they should "use AI" for a problem, and lately that means an LLM. Sometimes the answer is yes. Often the answer is a gradient boosted tree that trains in a minute and nobody finds exciting. The choice should happen before anyone opens an editor, because it decides the data you collect, the budget, the evaluation and what you'll be able to explain later.
A note on words. By classical I mean a model trained on your own labelled data for one task. That covers logistic regression and random forests, and for this post also small task-specific neural networks. By LLM I mean a general language model you prompt, with or without retrieval, and with or without fine-tuning.
Start with three questions
- What goes in? A row of numbers, an audio clip, an image, or text that a person wrote.
- What comes out? A label, a number, or text that a person will read.
- What does a wrong answer cost? And who has to explain it when it happens.
Most of the time those three answers make the decision for you. In health, where I've spent most of my work, the third question carries the most weight. Being wrong has a cost, and someone will ask why the model said what it said.
Two defaults
What each one gives you
Classical
- Needs labelled data
- Fast and cheap per prediction
- Same input, same output
- Feature-level explanations with SHAP or LIME
- Runs anywhere, including on-site
- Weak on free text
LLM
- Works from a prompt and a few examples
- Slower, paid per call
- Output can vary between runs
- Explains itself in words, which isn't the same as a reason
- Data often leaves your servers
- Strong on language, weak on arithmetic
Neither column is better. They're good at different shapes of problem.
The decision rules I use
| Question | Leans classical | Leans LLM |
|---|---|---|
| What is the input? | Numbers, signals, images | Free text, documents, conversation |
| Do you have labels? | Hundreds or more, and trusted | Few or none |
| Must the answer be repeatable? | Yes, audited or regulated | Some variation is acceptable |
| What explanation is needed? | Which inputs drove the decision | A quoted source or a readable summary |
| Volume and latency | Many calls, milliseconds | Modest volume, seconds are fine |
| Can the data leave your servers? | No | Yes, or you can host a model |
| How will you measure it? | Held-out set and one metric | A rubric and human review |
When the rows disagree, the first two usually win. A model that matches the shape of your input and the labels you actually have will beat a better-sounding model fighting against both.
How the decision plays out
Voice screening: classical side
Take a typical voice-screening task, where a model listens to a recording and flags a condition. It is audio in and a label out. There are labelled recordings, the decision has to be repeatable, and a clinician will want to know what the model heard. That's a trained model on acoustic features or spectrograms. A language model has nothing to add to the core prediction.
Clinical decision support: classical side
At a hospital network I built decision support with federated learning, where the data stayed at each site. Tabular clinical features, a label, and a hard rule that patient data doesn't travel. Sending records to a hosted LLM wasn't an option, and feature-level explanations, the kind SHAP gives you on a tree model, are something a clinician can actually read and argue with.
A companion that talks: LLM side
When I built a voice companion for patients with dementia, the whole product was conversation. Nobody has a labelled dataset of good replies to every sentence an older adult might say. That was an LLM from the first day, and the engineering went into speech recognition, voice and privacy around it.
The hybrid that keeps winning
The most useful pattern sits in the middle. Use an LLM to turn messy text into structured fields, then hand those fields to something simple and testable. I did a version of this for literature reviews: the model reads a title and abstract and answers include or exclude against written criteria, and the result is a plain list you can check against a hand-labelled sample. The LLM does the reading. The bookkeeping and the measurement stay in plain code.
Signs you picked wrong
- You're writing longer and longer prompts to make an LLM produce a number. That's a regression problem. Train a regressor.
- You're hand-crafting dozens of text features for a classical model and it still misses the point of the sentence. Let a language model read it.
- You can't say how you'd evaluate it. That's not a model choice problem yet. Go back to the three questions.
Next time a project lands on your desk, write the three answers on one line before you write any code: what goes in, what comes out, what wrong costs. Then go through the table row by row. If you still can't decide, build the classical baseline first. It takes an afternoon, and it gives the LLM a real number to beat.
related