cityflow

NYC trip records, queried in the browser

Dates
to
Service
Day
Borough
Reading source provenance from mart_source_freshness.

What does a normal week look like, and when did that change?

The tiles and the two charts read the same window from the filter row above. Until the engine finishes loading they run on a small precomputed summary of the default window, which is why a number can move once without a filter changing. Engine still loading.

Table view: the daily values behind the tiles and their sparklines
One row per day inside the selected window.
DateTotal tripsRevenueMean durationTip rateAirport share
No rows match the current filters.

0 rows

Daily trip volume against its trend

Observed trips per day with the STL trend component, and every level shift the changepoint search kept.

  • Observed trips
  • STL trend
  • Detected changepoint

No changepoint falls inside this window for the selected services. The search runs PELT on the STL trend component, period 7, minimum segment 14 days, q 0.05. Running it on the observed series returned eighty five candidates, because the weekly cycle is larger than most level shifts and PELT was finding the cycle. A detected shift is a change in level, not a cause: the search knows nothing about fare policy, weather or a source that changed shape. The trend layer is joined on date and service only, because agg_daily_decomposition carries no day type, so the day type filter moves the observed line and not the trend.

The week, hour by hour

One hundred and sixty eight cells, one per hour of the week.

Trips in the cell, quietest to busiest00

Every cell is a count, measured at the grain agg_hour_of_week publishes: month, service, borough, day of week and hour. All four filters reach this chart, and the 0 cells drawn here carry 0 trips, the same total the tiles above report for the same window. Trips whose pickup zone was never recorded are counted here, because this is a question about time rather than about place, and dropping them would put a different denominator under this chart than under the rest of the panel.

Where do trips start and end, and how concentrated is that?

Geography is the one place a trip with no recorded location cannot be shown, so it is the one place the page has to say what it is leaving out. Every chart here states its exclusion and the share it costs.

Pickups by zone

Every taxi zone the source can place, shaded by the trips that started there.

Trips starting in the zone, cut at sextiles0up to 0
Unknown locations are not on this map. Zone ids 264 and 265 are the TLC's own codes for a trip whose location was never recorded. They have no geometry, so they cannot be drawn, but they carry real volume: share loading of trips in this window. They stay in every total on this page and are excluded only from geography.

The style has no external sources: there is no tile server, no basemap and no font fetch, so the page works from a static host with no network beyond its own origin. Geometry comes from zones.geojson, which holds 260 shapes. Those zones keep their trips in the table view and in every total.

Which pairs of zones move the most people?

The 150 heaviest zone pairs, drawn as arcs between zone centroids and weighted by volume.

Trips on the pairfewer trips0
The arcs carry 0 trips between them. An arc bows to the left of its own direction, so a pair that runs both ways draws two curves rather than one line.

A centroid is not where a trip started, it is the middle of the zone, so an arc is a claim about which pair of zones moved people, not about a route. agg_od_flow holds no day type, so that filter does not reach this chart; the date, service and borough filters do, with borough applied to the origin. Trips whose origin or destination was never recorded are absent from this mart entirely, which is why its totals sit below the citywide totals in panel one.

How concentrated are pickups?

Zones ordered by volume, with the running total.

  • Share of trips in the zone
  • Cumulative share

The first 40 zones are drawn; the table has all 0. Both marks are shares of the same total, so they sit on one axis: a Pareto with a second axis is two charts pretending to be one, and the reader has no way to know which gridline belongs to which mark. Unknown locations are excluded here as well, for the same reason they are excluded from the map, so the denominator is trips with a known pickup zone.

Does any zone tip differently from its service?

Zone tipping rates against the service rate, each with a 95 percent Wilson interval.

0 zones are compared against the For hire vehicle rate, and every one of those comparisons is a chance to find a difference that is not there. The q values are Benjamini Hochberg adjusted and the threshold is 0.05: without that correction, testing 0 zones at five percent would be expected to flag about 0 zones by chance alone. No zone survives the correction, so nothing is marked. Tipping is measured only where a tip is observable, which is card payments: a cash trip records no tip, and counting those as zero tips would invent a difference between zones that differ in how people pay. The date and day type filters do not reach this chart, because zone_comparisons is computed once over the whole source window; the service and borough filters do.

How do trips differ by hour, and what does that cost the rider?

Three of the four marts behind this panel are keyed by service and hour rather than by date, so the date and day type filters reach less of it than they do elsewhere. Each chart says which filters it answers to.

How long does a trip take, hour by hour?

One ridge per hour of the day. Every ridge is a distribution over the same denominator and they share one height scale, so a narrower ridge means a tighter spread of trip times.

Hour of day00:0023:00

Height is the share of that hour's trips falling in a one minute bucket, not volume, so the ridges compare shape rather than size. The last bucket is everything at sixty minutes or longer and is labelled 60+, which is why it stands up: it is a tail folded into one column, not a spike at exactly an hour. agg_duration_dist is keyed by service and hour only, so the date and day type filters do not reach this chart; the service filter does.

What does distance buy, and where does the meter stop mattering?

Yellow taxi: hex bins of mean fare against mean distance, shaded by the trips inside the bin, with a volume weighted cubic through them.

Trips in the bin, log scaled1 trip0

The horizontal band at $70 is the JFK flat fare, and it is where the curve stops being a curve: 0 of the 0 flat rate trips in this window sit inside one dollar of it, at every distance from eight miles to twenty. Newark, rate code three, does not draw a band, because it is not a flat fare: it is the meter plus a fixed surcharge, so its trips scatter along the same curve as everything else and only the intercept moves. The cubic is weighted by the trips in each bin and drawn over the distance range that holds ninety percent of the bins; beyond that the bins hold single trips and a fit through them would be a drawing, not a model. agg_fare_distance is keyed by service, so the date and day type filters do not reach this chart, and one service is drawn at a time: the mart holds a mean per cell, not the component sums, so pooling services here would mean averaging averages.

Does the tip rate move with the hour, and does it differ by borough?

Tip as a share of fare, over the hours of the day, one panel per borough, ordered by volume.

The rate is tip over fare across the trips where a tip can be seen at all, which is card payments: a cash fare records a tip of zero whether or not one was handed over, and folding those in would turn a payment mix difference into a generosity difference. No interval is shown, and the catalog says why: the catalog could not be read. The observable trip count is in the tooltip and the table so a point built on a few hundred trips is not read as firmly as one built on a hundred thousand. All four filters reach this chart; borough is applied to the pickup zone.

How have yellow, green and for hire vehicles traded share?

For hire vehicles run roughly seven times the volume of yellow and three hundred times that of green, so the first two charts answer the question twice: once as relative movement against a common base, once as a share of the month. Neither is a second axis on the other.

Which service grew, relative to where it started?

Monthly trips for each service, indexed to a common base.

Each line is that service against itself, not against the others: an index of 120 means twenty percent more trips than the service ran in the base month, and says nothing about whether it carries more people than another service. Changing the date range moves the base, so the lines rebase and the shape changes; that is the index working, not a bug. Absolute volumes are in the table, and the stacked area below shows the levels as shares.

How is the month split between the three services?

Share of all trips in the month, stacked to one hundred percent.

A share is only readable against a denominator, and the denominator here is the services the filter keeps: dropping one from the filter row does not shrink the stack, it redistributes it. The band at the bottom is the only one whose height can be read against the axis directly; the two above it are read as thicknesses, which is what a stacked area is for and the reason the indexed chart above exists alongside it.

Does a trip take longer on one service than another?

Quartiles and the median interval for trip duration, one box per service.

The quantiles are read off agg_duration_dist, which is a histogram of one minute buckets, so every quartile and every median on this chart is accurate to a minute and no finer: a median printed as twelve means the twelfth minute bucket is where the running count crosses half, not that the median is 12.0 minutes. The waist is the usual 1.58 times the interquartile range over the square root of the count, the approximation that makes two boxes comparable at a glance: where two waists do not overlap, the medians differ. The upper whisker sits at the 60+ bucket wherever any trips fall in it, because that bucket is the top of the scale rather than a value.

Where does a borough send its trips?

Origin borough on the left, destination borough on the right, band thickness by trips.

No borough to borough flow survives the current filters.

0 trips are on this diagram. A borough appears twice, once as an origin and once as a destination, and the two are different quantities: the left hand Manhattan node is trips that started there, the right hand one is trips that ended there. Band colour is the origin, fixed by borough rather than by size, so filtering does not repaint the diagram. agg_od_flow carries no day type, so that filter does not reach this chart, and the borough filter applies to the origin side only, which is why selecting one borough leaves a single band on the left and several on the right.

Can you trust these numbers?

The three marts behind this panel are the pipeline auditing itself: what it read, what it removed and what was missing from the files it read. They are keyed by source period and service, so the date filter applies at month resolution and the day type and borough filters do not reach them at all.

Reading source provenance from mart_source_freshness.

What did the ingest throw away, and why?

Rows removed by each quarantine rule, as a share of the rows the source held.

The share is the comparable number: a rule that catches two thousand rows in a busy month and two hundred in a quiet one has not changed behaviour. The interval is a Wilson score on rows caught out of rows read, which is what makes a rule with a handful of catches distinguishable from one with none. Quarantined rows are not in any figure on the other panels: they were removed before the fact table, which is why the totals here and the totals there differ by exactly this much.

Did every month arrive, and how much of it survived?

Rows kept per source period, one frame per service because their volumes differ by a factor of three hundred. Backend reported as unknown.

No source period matches the current filters.

Each frame has its own axis and none of them is a second axis on another: for hire vehicles run about 1.9M rows a month against 298k for yellow and 5.5k for green, and on one linear scale the smaller two are a flat line on the baseline. The frames share an x domain, so a month is in the same place in all three. No month in this window follows a gap, so every line is continuous. If one did, the line would stop and restart rather than slope across the hole: a segment drawn over a missing month asserts a number nobody read. The 0 schema vintages in the table are what the loader matched each file against, which is how a column that appears partway through the history is handled without a null rate alarm.

Which columns were missing, and when?

Null share per watched column and source period, one panel per service.

Null share0 percent null100 percent null

A null share of exactly 1.0 followed by a drop is a column being introduced, not a data quality problem. congestion_surcharge appears in 2019, airport_fee in 2022, cbd_congestion_fee in 2025, and the high volume for hire vehicle schema has no passenger_count or ratecode_id at all, which is why those two rows are solid for that service. No such transition falls inside this window. What should raise an alarm is a share that moves in the middle of a service's history, because that is a source that changed shape without announcing it.

What does a panel query cost?

Published with the build, in manifest.json.

No benchmark was published with this build.

Where does each number come from?

Every aggregate on this page is spliced out of the catalog below at runtime rather than written again in the dashboard, so what is printed here is the definition the page actually ran, not a description of it.

The metric catalog

Search by name, by definition, or by the column a metric is built from.

Loading metric_catalog.json.

Sixteen metrics, one owner, one definition each. The two SQL blocks are the same definition compiled for two engines: warehouse_sql runs over the fact table in dbt, browser_sql runs over whatever aggregate a panel puts in a CTE named agg. The page lifts the projection out of the second one and splices it into its queries, which is why a definition change in dbt reaches this dashboard without a TypeScript edit.

The dbt graph

Source on the left, the panels that display the metric on the right. The highlighted chain is the selected metric's.

No dbt manifest was published with this build, so the lineage graph has nothing to draw.

Pick a metric to highlight its chain. Hover a node for its description.

Read left to right: a source table, a staging model that types and renames it, an intermediate model that resolves geography, the fact table, and the panel that puts a number on screen. An exposure is a dbt object, not a label added here, so the rightmost column is the warehouse's own record of which dashboard depends on which model. The layout is a fixed layered assignment rather than a force simulation: the columns are the dbt layers and nothing should be free to drift out of its own.