BEEP BOOP
Public beta

Sample-accurate timing on an iPhone

A drum machine has one job before all others: every step on time. Beep Boop places each step on an exact sample of sound, and was measured doing it on an iPhone. This is how it works, how it was measured, and what the measurement does not cover.

Fig. 1
The worst error of a step in two runs on an iPhone 14 Pro Max at 48,000 samples a second. Steps counted in samples: 0 samples over 5,083 steps. Steps fired by a main-thread timer: worst 619 samples, 12.9 milliseconds, with 99 in 100 under 240 samples. One audio buffer is 120 samples.0 SAMPLES0 MS1603.33206.748010.064013.3COUNTED IN SAMPLESWORST 0 OF 5,083 STEPSFIRED BY A TIMERWORST 619 · 99 IN 100 UNDER 240ONE BUFFER, 120 SAMPLES
How far a step landed from its place, at worst, in two runs on the same phone. The black dot is the engine. The orange line is the same beat fired by a timer.

Two problems, not one

"Tight in time" and "instant to the touch" are different problems. Tight is a scheduling problem: each step of a pattern starts where the grid puts it. Lateness does not matter there. A pattern that is 10 milliseconds late everywhere is still perfectly in time. Instant is a latency problem: a finger lands and a sound must follow, and nothing can be planned ahead because nobody knew the finger was coming.

This page is about the first. The second is at its foot.

Count samples, never time

An iPhone plays sound as a stream of samples, 48,000 a second on its own speaker. iOS asks the app for them a buffer at a time. Beep Boop's buffer is 120 samples, about 2.5 milliseconds.

The engine is Apple's AVAudioEngine with a single AVAudioSourceNode, whose render block is called once for every buffer. The sequencer lives inside that block. It keeps a count of every sample it has made since PLAY, and it knows the sample on which the next step falls. When a buffer contains that sample, the step's sound starts at that place in the buffer, not at the buffer's edge.

No timer is involved. A timer asks the system to be woken at a time, and the system wakes it when it can. A count of samples cannot be early or late, because the samples are the sound.

Never add, always multiply

At 127 BPM a sixteenth note is 5,669.29 samples long. A step cannot start on a fraction of a sample, so each is rounded to the nearest whole one. If the engine added 5,669 each time, it would lose 0.29 of a sample a step and drift. Instead every step's place is worked out from the first: the step's number times the exact length, then rounded. The error never exceeds half a sample and never builds up.

Swing is the same sum with a delay added to every second step.

How it was measured

A claim like this has to be able to fail, so the measurement was designed before the engine was written.

  1. The app plays a test signal: a short 1.7 kHz click on every sixteenth note at 127 BPM. The tempo is chosen so a step is not a whole number of samples and the rounding is exercised.
  2. A tap on the engine's own output records the samples to a file on the phone.
  3. The file is copied to a Mac and read by an analyser that shares no code with the engine. It finds each click and compares its place with the place the grid gives it.
  4. A run passes only if every step is on its exact sample. An error of one sample fails.

The recording alone is not trusted. A dropout happens after the samples are made, and a recording of them would still look perfect. So the render block also keeps witnesses: a buffer whose sample time does not follow the last one, and a buffer that arrives more than half a buffer late by the system's clock. A run is called tight only when the grid is exact and every witness reads zero.

The result

On an iPhone 14 Pro Max, on 3 October 2026:

Four timing runs on an iPhone 14 Pro Max: the load, the steps played, and the worst error.
RunStepsWorst error, samples
60 s At rest.5140
600 s Every processor core busy.5,0830
300 s The screen's thread held for 12 ms of every frame.2,5480
240 s Sent to the background after 67 s.2,0360

The zero is not rounding. The engine and the analyser place steps by the same rule and every click is the same waveform, so an error of one sample moves the detected click by exactly one sample.

The fault planted to prove the test

A test that has never failed proves nothing. So the app carries the very fault the engine avoids: a mode that fires each step from a timer on the main thread. Each click then starts at the edge of whichever buffer happens to be next.

The same analyser reads that run as not tight: a worst error of 619 samples, 12.9 milliseconds, with 99 steps in 100 inside 240 samples. That is the orange line in Fig. 1, on the same phone, the same day.

A second planted fault holds the audio thread for twice a buffer, once every thousand buffers. In a 120 second run the witnesses counted 48 breaks, one for each. On this hardware a dropout cannot hide.

What is not measured

Instant to the touch

When a finger plays a key, the note is a message the audio thread picks up at the start of its next buffer. Measured over about 400 taps, that wait has a median of 1.2 milliseconds and never exceeded about 2.5, one buffer. iOS then reports 17.9 milliseconds for its own path to the built-in speaker. The engine's share is the small part.

Why the audio thread can answer that fast, every time: the audio thread never waits.

Sources

More sheets

  1. A beat in 40 bytesTempo, swing, eight voices and a check byte.
  2. The audio thread never waitsNo allocation, no lock, one queue.
  3. Ghost notes and accents on drumsLoud and quiet strokes in one bar.
  4. A kick drum is twelve numbersHow a drum machine's kit is built.
  5. PatternsDrum patterns drawn on sixteen steps.
  6. Tempo by genreBPM for twelve genres, on one scale.