howbadisthetrain

NYC · methodology

How the scores work

The full picture of what each line's score means, what data feeds it, and where the line between signal and snark lives.

What the score is

Every line gets a number from 0 to 10, one decimal. Higher is better. The number is a deliberate vibescore, not a literal statistical model — it's built to communicate "should I get on this train right now or grab an Uber?" in a single glance. A 7.4 means what riders expect, not "the literal mathematical answer."

How the inputs blend

Each score is a weighted blend of five inputs. The weights aren't fixed at a magic number; they tune against rider behavior and the MTA's own on-time stats so a line that the city says is bad is also bad here.

1. Service status (the largest input)

Is the line running normally, delayed, suspended, or rerouted? A suspended line is a 0.0 — there is no train to score. A line in a planned reroute is whatever the reroute normally scores (often fine, sometimes awful, depends on the weekend). Real-time delay minutes feed in linearly: a five-minute delay costs roughly the same as five minutes of commute expectation.

2. Crowding

How full is the next train, on average? Crowding is the line's normal rush-hour state plus any current spike. We don't have per-car crowding data (the MTA doesn't publish it), so this input is a baseline set against rider reports and historical patterns: the L is busy, the J is usually OK, the 6 at 8:45am is chaos. Updates when the live feed wires up.

3. Planned work

Weekend reroutes, late-night service changes, and track maintenance get factored in. Planned work reduces the score based on severity: a single-station skip costs less than a full-line shutdown. This is where the weekend "is the F running normal today?" question gets answered.

4. Open alerts

Open MTA alerts for the line in the last hour, weighted by severity. Signal problems hit harder than a debris report. An "all clear" doesn't move the score; a stuck-train alert drops the line a third of a point or so. Multiple alerts stack but plateau fast — a line with five open alerts isn't five times worse than a line with one.

5. Historical baseline

Lines that are chronically bad start a few points lower. Lines with a reputation for reliability start higher. This is intentional and disclosed so the numbers stay recognizable across days: a 7.4 on the 7 today means what a 7.4 on the 7 meant last week. The baseline updates annually based on the MTA's own service-delivered stats.

What the labels say

The one-line blurbs under each score are written by a human (us), not generated. The tone is dry, occasionally snarky, and firmly on the side of public transit — if you're reading a line we said is "bumper-to-bumper, too many humans in a tin can, you know this," that's a real L train observation, not a made-up complaint.

We try not to editorialize on ridership in ways that read as class-coded. The F is not a "scary" line; the L is not a "broken" line. They're transit lines with capacity and reliability issues, and we describe them as such.

Driving comparison numbers

The "vs. driving" cost and time estimates on each line page are rough, not measured. Uber and parking costs vary wildly by borough, time of day, and whether you're a regular or a first-timer. Treat them as an order-of-magnitude nudge, not a quote. We pull baseline figures from NYC TLC averages and apply a rough mid-borough surcharge — not a precise number for any specific trip.

Current status: stub data

What you're seeing right now is stub data, not the live MTA feed.We've shipped the scoring math, the per-line pages, and the layout against a hand-curated set of scores (one per line) so the design could be reviewed against the real shape of the data. The MTA GTFS-realtime integration is the next planned build.

Until that ships:

  • Scores don't change when you refresh the page — they're baked in at build time.
  • The "NYC · ranked" tag on each page refers to the static ranking, not a live feed. We intentionally do not use the word "live" anywhere on the site until the feature actually ships.
  • Alerts, planned-work flags, and labels are static. Don't make a real-life decision off them.

What changes when the real feed wires up

  • ISR revalidation drops from hours to ~60 seconds, and the home list reorders itself as live data shifts.
  • Each line's score recomputes every time the MTA feed changes, blended with the same historical baseline so the numbers stay recognizable across days.
  • The label under each score updates to reflect the current state (delays, planned work, normal service) instead of being a fixed one-liner.
  • An "as of" timestamp appears on each line page so riders can see how stale the data is at a glance.

How we handle MTA quirks

The MTA's feeds are well-documented but occasionally weird: a line marked "delays" with no estimated resolution time, an alert that's been open for three weeks, a planned-work flag that contradicts the schedule PDF. We treat each quirk the same way — the score reflects the rider experience, not the data purity. If the MTA says "delays" but trains are running on time, we don't tank the score; if the MTA says "good service" but Twitter is on fire, we dock it.

Refresh cadence

When the live feed ships, scores update roughly every minute during peak hours and every 5 minutes overnight. The home page revalidates in the background so you almost never see a loading state. Page-level ISR keeps the per-line pages snappy without hammering the MTA feed.

What this site does not do

  • No account creation. No login. No cookies set by us.
  • No third-party tracking scripts beyond what Google AdSense sets (and AdSense only loads when this site has ads enabled).
  • No selling or sharing of any data we collect — and we collect almost none. See the Privacy page for the full list.

Questions or corrections? Get in touch— we read every email.