Somewhere between a microphone and a pair of ears, a stream of numbers is being reshaped. Frequencies nobody wants are pushed down, the ones that matter are left alone, and it all happens tens of thousands of times a second. That reshaping is the work of a digital filter, and it is one of the friendlier corners of signal processing: a little theory goes a long way.
You do not need to derive a z-transform to use a filter well. You need to know what the two main families do, what each costs you in latency and processor time, and when each is the sensible choice. That is what follows.
Audio arrives as a list of numbers. A signal sampled at 48 kHz gives you 48,000 measurements every second for each channel, and a digital filter is simply an equation that produces a new list from the old one. Most of the time it reads like this: take the current input, add a fraction of the previous inputs, add a fraction of the previous outputs, write the result out. Those fractions are the coefficients, and each one is called a tap.
If you have ever averaged the last three readings from a noisy sensor, you have built a filter: three equal taps, each a third, making a crude low-pass. Audio filters use the same arithmetic with more taps and more carefully chosen numbers. Hardware does not care — multiplication and addition are cheap. The design work is all in choosing the coefficients.
The impulse response is the filter's fingerprint. Feed it a single spike and the output tells you almost everything about how it behaves. Run a Fourier transform on that response and you get the frequency response: the same filter described as gain against frequency. Two views, one object.
A filter does not remove noise. It decides which parts of a signal are allowed to stay loud.
A finite impulse response filter uses inputs only. No feedback, no memory of its own output. That one restriction buys a lot: an FIR filter cannot become unstable, its response to an impulse always ends, and if you make the coefficients symmetric you get perfectly linear phase. Linear phase means every frequency is delayed by the same amount, so the shape of a waveform survives filtering intact.
The price is arithmetic and delay. A sharp low-pass at 100 Hz with a 48 kHz sample rate may need several hundred taps, and a linear-phase design delays the signal by roughly half its length. A 512-tap filter adds about 5.3 ms of latency. That is invisible in a recording chain and very audible in a live two-way conversation.
An infinite impulse response filter feeds part of its output back into its input. It rings, it decays gradually, and it does far more work per coefficient. Where an FIR filter needs hundreds of taps to carve out one steep edge, a few cascaded second-order sections — usually called biquads — will manage it.
The classic analogue designs carry over almost intact. Butterworth is flat in the passband with no ripple. Chebyshev permits a little ripple and gets a steeper roll-off in return. Bessel trades steepness for a gentler phase response. Elliptic is the steepest of the four, at the cost of ripple in both passband and stopband. Most practical equalisers, crossovers and tone controls are cascaded biquads, and the standard biquad coefficient equations used in audio work are a dependable starting point.
IIR filters demand respect. Get the coefficients wrong and the output can grow instead of decaying. Run a low-frequency high-pass in single-precision fixed point and coefficient rounding may shift the response or destabilise it. They are also not linear phase, so transients are smeared rather than preserved.
Before an analogue-to-digital converter samples at 48 kHz, energy above 24 kHz has to be removed. Otherwise it folds back down the spectrum and appears as a whistle nobody can trace to its source. The low-pass that prevents this is the anti-aliasing filter; its counterpart after the digital-to-analogue converter is the reconstruction filter.
A high-pass at 80 to 100 Hz removes rumble, desk knocks and plosive thumps. A narrow notch at 50 Hz and its harmonics deals with mains hum. A gentle low-pass around 7 kHz makes a voice channel sound like a voice channel. These small, unglamorous filters are what make a call intelligible.
Every band on a parametric equaliser is a filter. A telephone channel is a deliberate band-pass from roughly 300 Hz to 3400 Hz, chosen to keep speech readable while using as little bandwidth as possible. In data links, pulse-shaping filters such as the raised-cosine limit bandwidth without letting one symbol interfere with the next.
Photo: Egor Komarov / Pexels