Writing

Merging two sensor streams that tick at different rates

In short: to put a 100 Hz signal and irregular detections on one 10 Hz clock, summarise the dense stream per bin, count and carry forward the sparse one, and make sure no bin ever sees a sample from…

published
read time
5 min
words
980
lang
en
filed under
Engineering

In short: to put a 100 Hz signal and irregular detections on one 10 Hz clock, summarise the dense stream per bin, count and carry forward the sparse one, and make sure no bin ever sees a sample from its future. Every strategy throws something away. Pick the one whose loss your model can live with.

This problem shows up any time you build on wearables or smart-home sensors. One stream is a steady signal, say an accelerometer at 100 samples a second. The other is a stream of events that arrive when they arrive: a detector that fires on a tap, a classifier that reports a label every so often, a button press. The model or the dashboard downstream wants one row every 100 milliseconds. Somebody has to decide what goes in each row.

100 Hz signal detections 10 Hz clock one row = 10 samples + 0..n events
Each 10 Hz row has to stand in for ten signal samples and any number of detections, including none.

The dense stream: ten samples into one

Going from 100 Hz to 10 Hz means each output row covers ten input samples. The tempting move is to keep every tenth sample. Do not. Anything faster than 5 Hz in the original signal folds back into the slow one as a fake pattern. That is aliasing, and it looks like real data.

The safer choices, from least to most information kept:

  • Mean per bin. A crude low-pass filter. Keeps the slow trend, flattens short spikes.
  • Proper decimation. A real anti-aliasing filter, then downsample. Better frequency behaviour than the mean, same single value per bin.
  • Several statistics per bin. Mean, standard deviation and max. Three columns instead of one, but a sharp tap that the mean would hide still shows up in the max and the spread.

For movement data I usually start with the third option. Models care about "something sharp happened in this 100 ms" far more often than about the exact average.

The sparse stream: zero, one or many events per bin

Detections are the opposite problem. Most bins have none. A few have one. Now and then a bin has three. Each way of turning that into a column answers a different question:

  • Count per bin. How many events happened. An empty bin is a real zero.
  • Max score per bin. How confident the strongest one was. Empty bins are missing, not zero.
  • Last value carried forward, with its age. What the detector said most recently, and how stale that is. The age column matters: a reading from two seconds ago is not the same as one from 50 ms ago, and the model cannot tell unless you say so.

What I avoid for events is interpolation. Drawing a line between two detections invents values that were never measured, and it needs the next detection to exist, which brings us to the real trap.

The trap: bins that see the future

If the merged data will ever be used in real time, each row must only contain what was known at the end of its bin. Three common ways to break that without noticing:

  1. Interpolation uses the next sample, which has not arrived yet.
  2. Centred rolling windows average half a window of the future.
  3. Bin labels. If a bin covering 0 to 100 ms is labelled 0 ms, and you join it with something else at 0 ms, you have handed 100 ms of the future to that row.

In pandas, close and label each bin on the right, and join the sparse stream with merge_asof looking backward only.

import pandas as pd

# imu: 100 Hz, columns ax, ay, az.  det: irregular, column score.
# Both have a sorted DatetimeIndex named "t".
grid = dict(rule="100ms", closed="right", label="right")

dense = imu.resample(**grid).agg(["mean", "std", "max"])
dense.columns = ["_".join(c) for c in dense.columns]

sparse = det["score"].resample(**grid).agg(["count", "max"])
sparse.columns = ["det_count", "det_max"]

# Last detection known at the end of each bin, and its age. Never looks forward.
seen = det["score"].rename("det_last").reset_index().rename(columns={"t": "t_det"})
merged = pd.merge_asof(
    dense.reset_index(), seen,
    left_on="t", right_on="t_det",
    direction="backward", tolerance=pd.Timedelta("2s"),
)
merged["det_age_s"] = (merged["t"] - merged["t_det"]).dt.total_seconds()
merged = merged.set_index("t").drop(columns="t_det").join(sparse)

The tolerance is a decision, not a detail. Two seconds here means "after two seconds without a detection, I no longer know". Set it from what the detector actually does.

What each strategy throws away

StrategyStreamKeepsThrows awaySees the future?
Every tenth sampledenseExact values at grid pointsEverything between; adds aliasingNo
Mean per bindenseSlow trendPeaks and short eventsNo
Mean, std, max per bindenseTrend, spread, peaksShape and order inside the binNo
Count per binsparseHow many eventsScores and exact timesNo
Last value plus agesparseLatest reading and how stale it isOlder events in the same binNo
Linear interpolationeitherA smooth lineThe fact that data was missingYes
Centred rolling meaneitherA smooth lineSharp edgesYes

Checks before you trust the merged table

CarefulTwo devices never share a clock. A phone and an earbud can drift apart by more than a whole bin over a long recording. Line them up on a shared event first, like a tap both sensors can see, before any resampling.
  • Plot one raw minute of both streams over the merged rows. Your eyes catch an offset faster than any test.
  • Count empty bins in the dense stream. They are gaps in recording, and they should stay missing, not become zeros.
  • Check for duplicate timestamps. Bluetooth batching loves to deliver several samples with the same time.
  • Confirm every timestamp is in the same timezone. One stream in UTC and one in local time is a classic.

Next time you merge streams, write one line per column in the merged table saying which strategy produced it and what it lost. If you cannot fill in the "lost" part, that column is the one to look at first when the model behaves strangely.

related

Keep reading