How these numbers are made

This page is the argument behind every figure on this site. If you think a number here is unfair, this is the page to attack.

The six outcomes

We take the operator's own published timetable, and for every scheduled trip we assign exactly one outcome. There is no “other” bucket, and nothing is quietly dropped. (This list was five until 22 July 2026; the sixth was added by amendment G3, described below, after a real feed outage showed the five could not honestly absorb one.)

COMPLETED
A vehicle was seen running this trip, and nothing in the evidence says it did not finish — whether because it was tracked most of the way, because it was still reporting as the trip's scheduled end approached, or because what little evidence we have was simply too thin to say otherwise.
VANISHED
A vehicle was seen assigned to this trip — running it, or merely keyed to it in the few minutes before departure, wherever it actually was — and then stopped reporting before the trip was over.
UNTRACKED
No vehicle was ever seen on this trip while we were watching.
CANCELLED
The operator's own real-time feed told us the trip was cancelled. We take their word for it.
EXCLUDED
We were not watching for enough of the trip to judge it. Our problem, not theirs.
EXCLUDED (FEED)
The national feed's own vehicle-position stream was degraded for this operator while the trip ran, so silence about this bus proves nothing. Not the operator's failure, not ours — the feed's. Amendment G3, below.

These six are exhaustive and mutually exclusive for every trip we do classify: excluded plus excluded (feed) plus cancelled plus completed plus vanished plus untracked equals the scheduled count, for every route and every day. An automated check enforces that arithmetic before anything is published — catching a corrupted or unrecognised outcome label — and if it fails we publish nothing at all rather than publish something that does not add up. That check is about internal consistency, not completeness: it confirms the trips we did classify add up correctly, not that every trip in the operator's timetable was classified in the first place.

UNTRACKED means we could not see it. It does not mean it did not run.

This is the most important sentence on the site.

A bus with a broken transmitter, a vehicle swapped in at short notice without the right equipment, a gap in the operator's real-time feed, an outage at our end that we failed to detect — every one of those looks exactly like a bus that never came. We cannot tell them apart, and we do not pretend we can. An untracked trip is reported as untracked and as nothing else.

If you want the number that points at real trip failures, that is the vanished rate: a vehicle was there, and then it was not. It is our strongest signal, not a certainty — see below for the ways it can still be wrong — and it is the number we rank on.

Why there are two rates and never one

It would be easy to add the vanished rate and the untracked rate together and call the total “ghost trips”. It would also be dishonest, because it would assert precisely the thing we just said we cannot know. So we never add them. They are computed separately, published in separate columns of the dataset, displayed in separate columns of the table, and no code in this project sums them — there is a test whose only job is to fail if that ever changes.

The untracked rate is worth showing anyway. It tells you how much of a route's service is invisible to us, which is useful context and is sometimes itself a story about an operator's equipment. But it never moves a route up or down the table, because ranking on it would rank operators by how good their telematics is, not by whether your bus came.

Untracked trips do still count in the shared denominator behind both rates, and that has a direction worth stating plainly: a route with a lot of untracked trips has its vanished rate diluted by a bigger denominator. That pulls a telematics-blind route's vanished rate, and its rank, down — never up. Cancelled trips stay in that denominator too, for the same reason: a route with a lot of cancellations gets its vanished rate diluted the same way, never inflated by them.

EXCLUDED is our downtime, and it never counts against the operator

When our tracker is down, restarting, or otherwise not watching, we cannot judge the trips that were running at the time. Those trips are marked EXCLUDED and removed from the denominator of both rates. A route is never punished for the minutes we were not looking.

That protection has a threshold, and we state it plainly rather than leave it implicit: a trip is EXCLUDED only when our tracker's own uptime over that trip's own window falls below 90%. Up to 10% of a trip's window can pass with us not watching and the trip is still judged as if we had watched all of it. A gap in our own coverage that stays under that threshold does not, by itself, keep a trip out of the ranked numbers — see the third case under "How we decide a vehicle finished the trip" below, the one adverse case on this page that is caused by us, not by the operator's equipment.

Our own uptime CSVs are published in full from the first day of tracking; the strip on the front page shows only the last 30 days of it. Both exist precisely so you can see how often that happens and hold us to it.

How we decide a vehicle finished the trip

Nobody tells us a trip completed. We infer it from two signals, and either is enough. We track how far through the trip's own stop sequence any of the vehicle's reports reached — combining the operator's own feed position with our own geographic match, crediting a report to the nearest scheduled stop it comes within a set distance of (250 metres today, chosen before launch and open to revision once burn-in data suggests a better one), and keeping whichever goes further. If that reaches 90% of the way through the trip's stop sequence, we call it completed. Separately, if the vehicle's last report — timed by the vehicle's own clock, not by when we downloaded it; see the staleness section below — falls in the last 10 minutes of the trip or later, we also call it completed. A trip is only called vanished if neither of those holds and the vehicle's reports stopped below 75% of the way through, with more than 15 minutes still left to run. Everything else — most of the way there, or last heard from just outside that final 10 minutes — is completed too, on the same benefit of the doubt as everywhere else on this page.

The geographic half of that has a known error, and it runs in one direction only: because a geographic match is merged in by keeping the higher of it and the feed's own progress, it can only raise a trip's credited progress, never lower it. A report that happens to fall near a stop further along the route than the feed claims can push a trip from vanished to completed; a report that falls short of the feed's own claim can never pull a completed trip back down to vanished. So a generous geographic match can hide a ghost. It can never invent one — for that step.

That guarantee is about the geographic step, not the classifier as a whole, and we would rather name where it does not hold than imply it never happens. Three cases run the other way. A vehicle that reports just once, tied to this trip, in the few minutes before it was due to leave, and is never heard from again, is classified vanished even though we have no evidence the trip ever ran. More consequentially: if a vehicle's transmitter fails partway through a trip that actually finished, we record that as vanished too — a dead transmitter looks exactly like a dead trip, which is the same problem the untracked rate exists to name, just arriving mid-route instead of never at all. Losing the signal late does not have this problem: a trip still counts as completed if the vehicle got 90% of the way, or if we last heard from it in the trip's final 10 minutes, so this failure mode is confined to signal loss earlier than that. The third case is different in kind from the other two, because it is ours, not the operator's: EXCLUDED only fires when our own tracker's uptime over a trip's window drops below 90%, so up to 10% of a trip's window can pass with us not watching and the trip is still judged as if nothing had happened. If our own outage lands where the vehicle would otherwise have shown late progress, a bus that genuinely finished can read as vanished — and unlike the first two cases, no equipment failure on the operator's side is involved at all; it is systematic, not a rare edge, because every sub-threshold gap in our own coverage carries this risk for every trip running inside it. All three cases are genuinely ambiguous from the evidence alone, and all three are resolved against the operator, not for them — even though the third is a failure of ours. We do not call the vanished rate a floor: it is our best available evidence, not a guarantee, and it can be pushed up by exactly the kind of failure the untracked rate exists to describe, including our own.

Benefit of the doubt

Where the geographic test cannot resolve a trip either way, we resolve it in the operator's favour. Outside the two cases above, if we cannot show that a trip failed, we do not say it failed. We would rather understate a problem than manufacture one — a league table of bus routes is an accusation, and an accusation has to clear a higher bar than a hunch.

When the feed itself fails: excluded (feed) — amendment G3, 22 July 2026

EXCLUDED exists because a gap in our own polling is indistinguishable, trip by trip, from a gap in a bus's telematics — so we refuse to grade what we could not see, and we count that against ourselves. On 21 July 2026 we met the same situation one level up: the national feed's vehicle-position stream partially collapsed for about forty minutes. Two operators' reporting fell by roughly four fifths at the same moment while a third barely moved, and our own tracker was healthy throughout — so this was not buses breaking down, and it was not our downtime either. Trip by trip, every bus caught in it looked exactly like a vehicle vanishing mid-route. Left alone, that one interval would have put hundreds of false accusations into the vanished column, the only column that ranks routes.

So the classifier now watches for that signature: for each operator, in each ten-minute interval, the share of its scheduled, in-progress trips that produced at least one position report. The expected share comes from the operator's own timetable, so quiet Saturdays and bank holidays are compared with their own schedules and not with a busier day's. When that share collapses below half across a sustained stretch — with enough scheduled trips running that the number means something, and only while our own tracker was watching — trips overlapping the stretch that would have been called vanished or untracked are instead set aside as excluded (feed): not the operator's failure, not ours, the feed's. They leave the denominator of both rates exactly as tracker downtime does. Completed and cancelled verdicts are never touched — evidence that exists still counts — and a stretch our own tracker missed can never be blamed on the feed; it stays ours.

Three honest costs of this rule, named rather than buried. First, a genuine mass failure that somehow also silenced position reporting would be set aside rather than counted — when we cannot tell the difference, we refuse to accuse, and the set-aside count itself is published so you can see how often that happens. Second, the tracker-watching requirement cuts the other way in one narrow band: if our own coverage briefly dips inside a real feed outage, the dipped stretch is not evaluated, which can break the outage into pieces too short to arm the rule — so a trip that deserved to be set aside can still be called vanished. That is the one place this rule favours the accusation rather than the operator, and we state it rather than smooth it over. Third, the day that motivated this rule, 21 July 2026, is not reclassified but withdrawn entirely: its verdicts were contaminated before the rule existed, and a caveat under a ranking table protects nobody. Withdrawn days are listed, with reasons, on the about-the-data page, and they do not count toward the fourteen-day baseline.

Feed staleness: acted on from 22 July 2026 (amendment G2)

Real-time positions arrive with a timestamp, and that timestamp is sometimes older than the moment we receive it. An earlier version of this page said we measured that gap but did not act on it, and promised that when we did, the change and its reasoning would be written down here before it affected a single published figure. This is that write-up. What the measurement showed: the feed re-serves old positions routinely — on every full day of our baseline, roughly one position in forty arrived already older than the ten-minute completion rule it could have satisfied, most often late at night, and unevenly between operators. Until this amendment we timed every report by when we received it, so a re-served position could keep a bus "alive" on our page minutes after it last actually reported. That error ran in the operator's favour, and it is the exact kind of quiet flattery this project exists to refuse.

The fix sets no threshold, because the evidence did not support one: the exposure declines smoothly as you move any proposed cutoff, with no natural break, so every candidate number would have been an arbitrary line we would then have had to defend. Instead we now time each report by the vehicle's own report time — the moment the vehicle itself says it reported, which the feed carries on effectively every position and which, across more than a million positions in our baseline, never once claimed a time later than our download. A re-served position no longer moves a trip's "last heard from" forward, and a position reported before a trip's scheduled start no longer earns geographic progress credit, however late the feed re-serves it. (If the feed itself ever attached explicit stop-sequence progress to a report — on this feed it never has — that channel is not time-gated; it is the operator's own claim about its own vehicle, and it can only flatter them.) Two things deliberately did not change: a report with no vehicle timestamp is timed by our download clock exactly as before, and whether a vehicle was seen at all — the untracked test — still uses our download clock, so this reinterpretation can only remove credit a re-served position was wrongly earning; it can never, by itself, put a trip into the accusatory classes for want of a timestamp. The one residual risk runs the other way and we name it: a vehicle whose own clock freezes while its position keeps updating would look silent to us and could read as vanished; our burn-in monitoring watches for that signature, and if it ever occurs at material rates this section will change again — with the reasoning, before it affects a published figure.

How to read a confidence interval

A route with 30 trips and 2 vanished shows 6.7%. But 30 trips is a small sample. Watch the same route for another month and you might see 1, or 5, without anything about the service having changed. The headline percentage alone invites you to over-read it.

So every rate we publish carries a 95% confidence interval: the range the true rate plausibly sits in, given how much data we actually have. We use the Wilson score interval rather than the textbook one, because the rates here sit close to zero on small samples, and the textbook formula produces nonsense there — negative lower bounds, and intervals of zero width when nothing has gone wrong yet.

Two rules for reading them:

Why the table is ordered the way it is

We rank by the lower bound of the vanished interval, worst first — not by the headline rate. That ordering is conservative, not a claim that any two neighbouring routes are statistically distinguishable — routes next to each other in the table will often have overlapping intervals. What it does guarantee is that an unlucky handful of trips on a quiet route cannot leap it to the top: outranking another route takes a worse lower bound, which folds in both the rate and the sample size behind it. The table is conservative by construction: to be ranked badly here, a route has to have both a bad rate and enough trips to be sure of it.

When we stay quiet

Check it yourself

The site you are reading is built from the published CSV files and nothing else — it has no access to our database, so a number here cannot differ from the number in the data you can download. Take the files, recompute, and tell us if you get something different.

Download the data · Back to the scoreboard