How it performs

In one paragraph, before the table. This app tries to hear gunfire on ordinary phones. It detects roughly nine real gunshots in ten on the recordings it has been tested against, and it raises far more false alerts than it is supposed to — enough that alerting is still treated as unproven rather than as a product. No alert has ever fired from a real event. Two different false-alert figures appear below rather than one, because they answer two slightly different questions and both are honest answers; whenever a figure from this project is quoted anywhere, it should say which of the two it is. If somebody quotes you one without saying, ask.

Measured figures for the current build. The app is given away unproven rather than sold, so anyone deciding whether to install it should see these first. They change as the app changes, and this page changes with them.

How it performs today

Gunshots it detects (one phone, the whole app)
~89%, against a 90% target
False alerts, per school year (target: one)
9 to 27
The same, if the phone asks a person first (simulated only)
under 1, on both site sizes tested
How far it can warn other phones at the school
Bluetooth range
How far it can forward an alert to a paired guardian
Anywhere
That forwarding measured over mobile networks, or across a school day
not yet
Detections forwarded to a guardian
~4 in 5
Can it call emergency services
no
Can it tell you where the sound came from
only with four or more phones
How accurately, on real phones
not measured yet
Can it count shots, or tell you the weapon
no
How quickly an alert reaches a phone that is waiting for one
about 3 seconds
How quickly one reaches a police force
it does not; no site can invite an agency in this version
Real gunshots it has alerted on so far
none

Two notes on the headline rows, because a single number would mislead. The false-alert figure is a range because it rests on an assumption nobody can settle from recordings: whether two phones hearing the same wrong sound make the same mistake. Assuming they always do gives about 27 a school year; using the error correlation actually measured between devices gives about 9. Both are correct answers to slightly different questions, and both are above the design target of one. That gap is why this app is a research prototype and not a product.

The detection figure is measured on held-out recordings, through the whole app: the always-listening stage that decides a sound is worth examining, and then the classifier. It is not a figure for a real building — no real gunshot has ever raised an alert from this app.

The detection row moved again on 2026-09-14, and this time nothing about the app changed. What changed is which version of the app the figure describes. The ~95% printed here until today came from a test that pauses the always-listening stage while it waits out the tail of a sound it has already decided to examine. A phone does not pause it. The difference is real, and it runs in both directions at once. A sound loud enough to be worth examining deliberately does not raise the phone's running idea of how noisy the room is, so that the shots after it are not missed. The echo of that same sound is loud but not sudden, so it is not treated as worth examining, and it does raise that idea. The phone therefore stops to examine fewer stretches of audio than the test believed. Measured the phone's way, on the same recordings, detection is about 89% rather than about 95%. The rate at which it fires on non-gunshot recordings falls at the same time, to about four fifths of what was published, although this many recordings can show that direction without settling its size. Fewer examinations cost real gunshots and buy back false alarms, so quoting only the drop would be half the story. The design target is 90%, and the measured range runs from about 87% to just over 90%. It neither meets that target nor misses it, and neither word is available here. The older figure is not withdrawn and was not a lie: it is a real measurement of a real cascade, reproduced exactly today, and it is simply not the cascade a handset runs. The false-alert rows in the table are counted a different way, which this finding does not reach, and they are unchanged.

Both headline rows moved on 2026-09-13, in the same direction, and that is worth explaining because it is unusual. Detection rose to about 95%, as detection was measured then, and false alerts fell to 9–27, both at once. This page said about 63% and 89–170 before that day, and neither of those old numbers is a fair "before" for the new one. On detection, 63% was an older measurement taken against a setting the app had already stopped using, and it was too low; measured the way that 95% was measured, the app was at about 85% the day before the change. On false alerts, the recordings the figure is computed over are defined as the ones the always-listening stage wakes for — so making that stage more willing to wake enlarged the test set at the same time, and 89–170 and 9–27 were counted over different sets of recordings. The right way to read them is as two separate measurements, not as one falling. The model was not retrained and not changed in any way; what changed is the two settings either side of it. The always-listening stage was made more willing to wake the classifier, and the classifier was made stricter about what it calls a gunshot. Those two are not the same trade-off, so moving both together improved both numbers — where moving either one alone had always cost the other. The cost is battery, because the classifier now runs about twice as often. Measured on real phones, that came to about six minutes of extra processor time per hour on a cheap handset, which was too much — so the arithmetic inside the classifier was rewritten to do the same work in half the time, and checked to produce bit-for-bit identical results. A cheap phone now spends about as much power on this as it did before any of these changes, while detecting more gunshots than it did on the day before the change. It still works harder than a recent phone does, by roughly seven times, and reducing that further is ongoing work.

The row before that said ~96%, until 2026-09-11, and that figure was wrong twice over: it was measured on the classifier alone, skipping the always-listening stage, and on a simulated campus of forty phones spread over 150 metres, which is the arrangement most favourable to two phones agreeing. What two phones agreeing adds to the current figure has not been measured, so no number for it is published here.

The third row is the one to read most carefully, because it is not a measurement of this app in use. The app can ask whoever is nearest whether they heard gunfire, and count the answers before it alerts anybody. In simulation, on a large site with many phones, that removes every false alert in the test — both figures go to zero. On a two-phone family the same simulation now also lands below one a year on both figures, with the whole confidence interval below one; it was unresolved here until 2026-09-13, and what settled it was running fifty times as many simulated events, not a change to the app. Three things stop any of that being a claim about your school, and they apply to both site shapes equally. It is simulated, not observed. The prompt has never run on a real event. And it assumes a rate at which a frightened person answers wrongly, which nobody has measured — so the row rests on one number that has never been checked against a real person.

“Forwarded to a guardian” is the same measurement read at the higher confidence the app requires before it will send an alert off the school site: of the shots it detects at all, about four in five clear that bar. It said ~3 in 5 until 2026-09-11 and the basis for that could not be found.

In an emergency, always call your local emergency number.

What a phone has to be able to do

There is no list of supported handsets here, and there is not going to be one. A list is out of date the day a phone launches, and the thing that matters is not which badge is on the back — it is whether that particular phone can finish one check of the sound before the next one arrives. So the app measures the phone it is actually on, at setup, by running the real detector nine times and timing it.

Two limits come out of work already published on this page, and neither is a round number because neither was chosen:

That produces one of three answers, and the app tells you which:

Separately, the app needs Android 8.0 or newer and a microphone it can open. Every phone that has run this cascade so far has been inside the hard limit; the slowest fell in the middle band, over the battery budget by about a fifth.

What this does not cover. Speed is not the only way a phone can let you down. Some makers shut background apps down aggressively whatever their settings say, and that is a different problem with a different fix — the app detects those makers and points at the setting that stops it. A fast phone with an aggressive battery manager and a slow phone with a permissive one are two different faults, and we do not report one as the other.