Reflex Trainer
All articles

Why Online Reaction Tests Disagree (and How We Measure Ours)

· 6 min read

Take three different online reaction tests in a row and you'll probably get three different averages, sometimes 50 ms or more apart. Your reflexes didn't change in those few minutes. What changed is the test: how it draws the stimulus, when it starts and stops the clock, and how much delay your device adds along the way.

This post explains where that error comes from, and the choices we made in Reflex Trainer to keep it as small as we can.

Where the extra milliseconds come from

Screen refresh

A screen doesn't change the instant code asks it to. A typical 60 Hz display draws a new frame roughly every 16.7 ms, so a stimulus requested just after a frame started waits for the next one. Some displays then add their own processing delay before the image actually lights up.

If a test starts its clock when the code requeststhe colour change, rather than when the frame is actually shown, it charges you for time you couldn't possibly have reacted in. That alone can add anywhere from nothing to a full frame, at random, to every rep.

Input latency

Your input takes time to reach the software, too:

  • Touchscreens scan for contact at a fixed rate and process the touch before reporting it.
  • Wireless and Bluetooth mice and keyboards add radio and polling delay.
  • Busy devices may take a moment to deliver the event to the page or app.

When the clock stops

Some tests react to a click, which only fires when you release the button, adding the time your finger spends pressed down. Others read the current time inside their event handler, which might run a little after the input actually happened if the device was busy.

Browser throttling

Browsers deliberately slow down timers and animations in background tabs, on battery saver and on some low-power devices. A test that relies on precise timers can drift when that happens.

How we measure reaction time

We designed Reflex Trainer around two rules that apply on every platform:

  1. Start the clock when the stimulus is actually presented. In the iOS app we record the time the stimulus frame is scheduled to appear on the display, not when our code asked for it.
  2. Stop the clock when the input actually happened. We use the timestamp the system attaches to the touch-down, pointer-down or key-down event itself, not the moment our code got round to handling it, and never the release.

Both times come from the same steady, monotonic clock, so the reaction time is simply the input time minus the stimulus present time. Changes to the device's wall-clock time can't creep in.

Punches instead of taps

In the mobile apps' combat-sports mode, the “input” is a punch. A phone's accelerometer, or an Apple Watch's motion sensors, sample movement 100 times a second, and each sample carries its own timestamp. We time the punch from the sensor sample where it's detected, not from when the detection code ran.

With a watch there's an extra problem: the phone shows the stimulus, but the punch is detected on your wrist, and the two devices have separate clocks. Before a test, the watch measures the offset between its clock and the phone's by exchanging timed messages and keeping the most reliable measurement. Punches are then converted onto the phone's timeline. The watch also keeps a short buffer of recent punches, so a fast punch that lands before the “stimulus shown” message reaches the watch is still timed correctly.

Making the stimulus unpredictable

If you can predict when the stimulus will appear, you can start moving before it does. Plenty of tests use a random wait spread evenly between, say, 2 and 5 seconds. The problem is that the longer you've waited, the more certain it becomes that the stimulus is about to appear, so a patient player can “feel” it coming.

We draw each wait (between 1.5 and 5 seconds) from a distribution where the chance of the stimulus appearing in the next moment stays roughly constant however long you've already waited. Counting in your head doesn't help. Live duels use the same formula, with the server scheduling the stimulus for both players.

False starts and misses

  • Under 100 msis treated as a false start. That's faster than people can genuinely react to a visual signal, so it was almost certainly a guess.
  • No reaction within 2 seconds counts as a miss, and the rep is repeated rather than dragging down your average.
  • Reacting before the stimulus is always a false start. On Easy it repeats the rep; on Medium and Hard it resets the set.

Why one best score misleads

Even with careful timing, a single rep is noisy. A lucky guess that scrapes past the false-start line, or a rep that happens to line up perfectly with a screen refresh, can produce a best time that you'll rarely match. An average over five or ten reps, alongside your best and worst, gives a much truer picture. That's why our results show all three, and why the leaderboard ranks each set length separately.

What we can't fix

No consumer device is a laboratory instrument. Touchscreen scanning, display processing and Bluetooth delays happen before any app or website sees the input, and they vary between devices. In a browser, we also can't see exactly when a frame reaches the screen the way a native app can.

So the fairest comparison is always you against yourself, on the same device. Comparing a laptop with a wireless mouse to a new phone tells you as much about the hardware as about your reflexes. For a sense of what typical scores look like, see what is a good reaction time?

Try it yourself

Take the test in your browser and do a set of five or ten reps on the device you use most. Then compare your average on the leaderboard, or check how your results hold up over time in does reaction time slow with age?

Test your reaction time

Free in your browser. Get the app to time real punches with your Apple Watch or phone.

Download on the App StoreGet it on Google Play