№ 69 · control & signals

Trusting two liars

You cannot read the number you want — a satellite's altitude, a car's speed on a wet road. You have a noisy sensor, and a model of how the number should behave. Neither is good enough alone. The Kalman filter is the arithmetic for combining them.

Two ways to be wrong

Write the thing you want as a state: the numbers describing where the system is now. You never observe it directly — you observe a measurement, the state plus noise. You also have a model: the previous state carried forward through the dynamics, plus noise of its own. The filter assumes the two noises are independent, white and normally distributed, and the step from one state to the next is linear. Each noise has a size, a variance, and that size decides who is believed.

The whole idea is one weighted average

The filter carries a guess and a number for how wrong that guess is. It predicts the reading the guess implies, then subtracts: measurement minus prediction is the residual, and a residual of zero means the two agree completely. What is left is multiplied by a gain. The gain is built from the two uncertainties, so the estimate lands nearer whichever source is more certain. Drive the measurement variance towards zero and the gain weights the residual more heavily: the reading is trusted more, the prediction less. Drive the filter's own variance towards zero and the reverse happens. Two imperfect opinions, averaged by their relative uncertainty.

Interactive Watch the estimate follow a value that drifts and occasionally jumps. Push sensor noise up and the filter starts ignoring the readings and smoothing hard; push model noise up and it leans on them instead. Step one measurement at a time, or start again to watch the uncertainty collapse.

One number, tracked. The true value swings on a slow drift and takes occasional jumps of about a tenth of a volt; the sensor adds noise whose standard deviation is the first slider. The filter's model is the weakest one it could have — the value simply persists from step to step — so the two variances are all it knows. The ribbon is plus or minus one standard deviation of the filter's own uncertainty, drawn at least two pixels thick so it stays visible when that uncertainty gets tiny: it opens at ± 1 V, a deliberately rough guess, and collapses within the first handful of measurements as the readings accumulate. That collapse is the filter learning. The bar underneath shows how the gain splits the two opinions; the readout prints the arithmetic behind it. The model slider changes only what the filter is told about how much the value wanders — the value's own drift does not move — which is the source's point about process noise: you cannot watch it, so you choose a number and tune it. Both sliders set a variance; the labels show the standard deviations. Worked example with our own numbers, and no matrix algebra anywhere, because in one dimension the gain reduces to P⁻ / (P⁻ + R). The plot covers ± 0.6 V, and readings past that edge are drawn at the edge.

Predict, correct, repeat

Two steps run on every measurement. The time update projects the current estimate and its uncertainty forward to the next step; the measurement update folds in the new reading, then starts over from the corrected estimate. The filter never stores the history: the previous estimate already carries what the earlier readings said, so the next needs only it and the fresh one. That is the point: as the source notes, one designed to work on all the data directly for each estimate, as a Wiener filter is, is far less feasible. R. E. Kalman published this recursive solution to the discrete-data linear filtering problem that year.

What it actually claims

The estimate is the mean of the distribution the filter believes the state has; the uncertainty it carries is its variance. The gain minimises the mean of the squared error. That guarantee is narrow — our reading, not the source's: it holds for the model the filter was handed and the noise statistics it was told. Change the model and the guarantee goes with it. The filter also supports estimates of past and future states, not just the present, the paper notes.

Where it breaks

Four assumptions carry the whole thing: the two noises are independent of each other, white, and Gaussian, and both the dynamics and the measurement relationship must be linear. Their sizes must be known too. Measurement variance is usually measurable off-line, from readings you were taking anyway. Process variance is the hard one: the process is exactly what you cannot watch. So tuning is normal practice: the source allows that even a poor model gives acceptable results with enough uncertainty injected through it. Feed in a nonlinear system and the extended Kalman filter linearises about the current estimate, approximating the optimum. And if the readings never pin the state down, the estimate diverges.

In short

Predict with the model, correct with the reading, split the difference by their two variances. That one weighted average, repeated, tracks a hidden number through noise — for as long as the model is linear, the noise is Gaussian, and its size is roughly right.

Where this comes from

  1. An Introduction to the Kalman Filter, UNC-Chapel Hill TR 95-041 linked only, not reproduced
    G. Welch & G. Bishop · 2006
    web.archive.org/web/20260522234924/www.cs.unc.edu/%7Ewelch/media/pdf/kalman_intro.pdf