Audio preprocessing that matters more than the model
In short: on medical audio, the steps before the model decide most of what the model can learn. Resampling, trimming, filtering, fixed-length windows and a split by patient each remove a shortcut the…
- published
- read time
- 5 min
- words
- 952
- lang
- en
- filed under
- Engineering
In short: on medical audio, the steps before the model decide most of what the model can learn. Resampling, trimming, filtering, fixed-length windows and a split by patient each remove a shortcut the model would otherwise take. Set up the ablation so you can switch each one off.
Lately my work has moved on to a different audio problem. Before that, I built a pipeline that classifies respiratory sounds from the ICBHI 2017 dataset into normal and abnormal, and into broader groups like chronic respiratory disease and respiratory infection. Every kind of audio is a different signal, but the pipeline problems are the same. This post is about those problems, and about how to find out which step is actually doing the work.
Why the front of the pipeline matters so much
A medical audio dataset is never one clean recording setup. It is several devices, several rooms, several people holding the microphone differently. Some recordings are long, some are short. Some start with a few seconds of rustling. Some are clipped because the gain was too high.
Every one of those differences can line up with the label by accident. If most of the sick patients were recorded on one device, a model will happily learn the device. It will look accurate. It will fail the first week in production. Preprocessing is the part of the system whose job is to make those accidents invisible, so the only thing left to learn is the signal you care about.
The steps, and what each one is for
In that pipeline the training workflow runs the same order every time: filtering and sampling, feature extraction, a split into train, validation and test, SMOTE oversampling after the split, then hyperparameter search with Optuna. Every run is logged to MLflow. The script takes the feature type as an argument (MFCC, log-mel, or MFCC with augmented features) and the task as another (binary or multiclass), so one command can run every combination.
That structure is what makes an ablation cheap. You don't rewrite code to test a step. You flip an argument and compare runs in MLflow.
| Step | What it removes | Shortcut the model takes without it |
|---|---|---|
| Resample to one rate | Different devices and apps | Learns the sample rate, which often tracks the recording site |
| Trim silence and handling noise | Rustle at the start, dead air at the end | Learns how long the operator waited before speaking |
| Band-pass filter | Mains hum, low rumble, hiss above the useful band | Learns the room and the cable |
| Loudness normalisation | Gain differences between devices | Learns microphone distance instead of the voice |
| Fixed-length windows | Recordings of very different lengths | Learns recording length, which often differs by group |
| Split by patient | The same person in train and test | Recognises the speaker, not the condition |
| SMOTE on training data only | Class imbalance | With SMOTE before the split, synthetic copies of test patients leak into training |
Read the table as a list of questions to test, not as results. Which rows matter most depends on how your data was collected. On a dataset recorded with one device in one room, resampling buys nothing. On a dataset pulled from several clinics, it can be the difference between a real model and a device detector.
How to run the ablation
The rule is simple: one step off at a time, everything else fixed, same split, same seed. Log every run. Then look at two things, not one.
- The headline metric on a patient-grouped test set. If turning a step off makes the score go up, be suspicious, not happy. It usually means the step was hiding a shortcut and the model just found it again.
- The metric per device or per site. A step that barely moves the overall number can still close a big gap between subgroups. That gap is what you'll see in production.
The steps also interact. Loudness normalisation before trimming gives a different result from trimming first, because the silence drags the average down. Fix the order, write it down, and only ever ablate within that order.
Here is the core of a preprocessing function in that order. Each step can be switched off by a flag, which is all an ablation needs.
import numpy as np
import librosa
from scipy.signal import butter, sosfiltfilt
def preprocess(path, sr=16000, seconds=5.0, band=(60, 4000),
trim=True, filt=True, norm=True):
y, _ = librosa.load(path, sr=sr, mono=True) # resample + mono
if trim:
y, _ = librosa.effects.trim(y, top_db=30) # drop leading/trailing quiet
if filt:
sos = butter(4, band, btype="band", fs=sr, output="sos")
y = sosfiltfilt(sos, y)
if norm:
rms = np.sqrt(np.mean(y ** 2)) + 1e-8
y = y * (0.1 / rms) # common loudness
n = int(sr * seconds)
y = np.pad(y, (0, max(0, n - len(y))))[:n] # fixed length
return librosa.feature.mfcc(y=y, sr=sr, n_mfcc=20)
The band edges are an example for voice. For lung sounds you'd pick a different band. The point is that the value is an argument, logged with the run, and not a constant buried in a notebook cell.
Where the model choice fits
I'm not saying the model doesn't matter. It does, once the inputs are honest. But I've learned to spend the first week on the front of the pipeline and only then on architectures. A bigger network on leaky inputs just learns the leak faster.
If you want a place to start, take any public respiratory sound dataset, run your pipeline end to end on random data first, then run the full grid of feature types and compare the runs in MLflow. Then add one flag of your own for a step you suspect, and turn it off.
related