Writing

What an earbud's motion sensor can tell you

In short: a six-axis motion sensor in an earbud sees enough to tell sitting from walking from nodding. Before any of that is learnable, though, you have to remove the differences between people and…

published
read time
4 min
words
826
lang
en
filed under
Engineering

In short: a six-axis motion sensor in an earbud sees enough to tell sitting from walking from nodding. Before any of that is learnable, though, you have to remove the differences between people and how they wear the earbud, and that normalisation is most of the work.

An earbud sits on your head, which turns out to be a good place for a motion sensor. It moves with every step, every nod, every turn to look at someone. Many earbuds already carry an inertial measurement unit, an IMU, for things like tap gestures and head tracking. Six numbers, many times a second: acceleration on three axes and rotation on three axes.

I've been exploring EarSet, a public dataset of ear-worn sensor recordings released for research, and building a small tool to look at it before I train anything. The tool is a public repo on my GitHub, earbudIMU. It loads the recordings per user and per activity, normalises them, and plots them either as a static figure or as an animation, in a little desktop window made with CustomTkinter and Matplotlib.

Look before you classify

The temptation with sensor data is to cut it into windows and hand it to a model straight away. I wanted to see it first. Plotting each activity for each person answers questions a confusion matrix can't. Is the signal there at all? Does one person look completely different from everyone else? Is one axis just noise?

Once you look, the activities have recognisably different shapes.

still
nod
walk
Sketches of the typical shapes, not recorded traces: flat when still, slow swings for a nod, a sharp repeating beat for each step.

Walking has a rhythm. Nodding is slow and large on one rotation axis. Sitting still is almost flat. With shapes this different, telling them apart for one person isn't hard. Telling them apart across many people is the real problem.

Why normalisation is most of the work

Raw IMU data from two people doing the same thing can look nothing alike. A few reasons:

  • Fit. Everyone wears an earbud at a slightly different angle. Gravity is a constant pull on the accelerometer, so a different angle shifts every axis by a different fixed amount.
  • Bodies. A tall person with a long stride and a short person with a quick one walk with different rhythms and different force.
  • Units. Acceleration and rotation are measured in different units and live on very different scales. Without rescaling, the larger one drowns out the other.
  • Drift. Gyroscopes have a small bias that differs between devices and sessions.

A model trained on raw data learns all of this along with the activity. It gets very good at recognising who is wearing the earbud and how, and worse at what they are doing. That is why the tool centres every recording and rescales each axis to a z-score before plotting.

raw: user A above, user B below z-scored per user
Same activity, same axis. The offset between the two people is fit and gravity, not movement. After normalising, the movement lines up.

The code that does it

The core is short. Compute a few orientation-free signals first, then normalise every axis within each person's recording, so one person's offset never leaks into another's.

import numpy as np
import pandas as pd

ACC = ["ax", "ay", "az"]
GYR = ["gx", "gy", "gz"]

def add_magnitudes(df):
    # The length of the vector doesn't care how the earbud is rotated.
    df["acc_mag"] = np.sqrt((df[ACC] ** 2).sum(axis=1))
    df["gyr_mag"] = np.sqrt((df[GYR] ** 2).sum(axis=1))
    return df

def normalise_per_user(df, cols):
    g = df.groupby("user")[cols]
    df[cols] = (df[cols] - g.transform("mean")) / g.transform("std")
    return df

def windows(x, size, step):
    for start in range(0, len(x) - size + 1, step):
        yield x[start:start + size]

df = pd.read_csv("user01_walk.csv")      # one file per user and activity
df["user"] = "user01"
cols = ACC + GYR + ["acc_mag", "gyr_mag"]
df = normalise_per_user(add_magnitudes(df), cols)

The column names here are mine for the example. Two things in it matter. The magnitudes are computed before normalising, because they're the one signal that is the same whichever way the earbud sits. And the statistics come from each user alone. If you compute them over everyone, including your test users, your evaluation quietly gets to peek.

What I'd try next, and what you can try now

With clean, normalised windows, a small classifier on simple features per window, such as the mean, the spread and the dominant frequency of each axis, is the obvious next step. The important part is how you split the data. Hold out whole people, not random windows. A random split puts windows from the same walk in both training and test, and the score tells you the model can recognise a person it has already met.

If you have any IMU data, from an earbud, a watch or a phone in a pocket, plot one activity for three different people on the same axes before you train anything. If the traces sit at different heights, you've found the work that comes before the model.

related

Keep reading