How the data is cleaned and how the heatmap is computed

The raw data cannot be plotted as it arrives. This page covers every step in between and the problem each step fixes. All figures are measured, not estimated.

1. One row is one party, not one accident

This is the easiest step to get wrong. In the raw data an accident produces one row per party involved. A car hitting a motorcycle is two rows with identical date, time and coordinates, differing only in the party sequence number.

2,403,922 raw rows de-duplicated on date, time and coordinates give 1,042,189 accidents, an average of 2.26 parties each. Counting rows instead would inflate the accident count by more than a factor of two.

2. Removing geocoder fallback coordinates

Some coordinates are not where the accident happened. They are the value the geocoder fell back to when it could not resolve the address, usually the county government building.

The worst single point is 121.481, 25.009, which is the New Taipei City Government. Left in, that one point carries more than two thousand accidents whose addresses span 20 districts, and the largest accident hotspot in Taiwan appears at the city hall's front door. It is an artefact.

Detecting them is simple: if the accidents sharing one coordinate have addresses in three or more districts, the coordinate is fake. A real intersection is not in Sanxia and Bali at the same time. 139 such points were found, carrying 19,542 accidents or 1.84% of the total. All are removed.

3. Aggregating into cells rather than plotting raw points

Handing 1,042,189 points to a browser will kill a phone. Accidents are aggregated into cells instead: the database holds one grid of roughly 12 metre cells with a month dimension, and queries roll those up with integer division to whatever coarseness the current zoom needs. Each tile is always divided into 96 cells.

So every bright spot is a total for a cell, not a single accident. At national zoom a cell is 3.3 km across; at street level it is 6 metres.

4. How the colour is decided

Each cell's weight is on a logarithmic scale: the natural log of its count divided by the log of the largest count currently on screen.

This part was rewritten twice and both earlier versions failed, which is worth recording. The first used the maximum on a linear scale. Nationally the distribution runs from a median of 38 to a maximum of 15,601, a factor of four hundred, so everything except that one point rounded to zero and the map was solid blue. The second used the 97th percentile as full scale, and was solid blue for the same reason: 97% of cells are still far below that divisor. On the current log scale the median cell gets 0.38, a cell with five hundred accidents gets 0.64 and the largest gets 1, which is what opens up the middle of the range.

5. What "fatal only" counts

It keeps only cells where someone died and switches the weight to the number of deaths. The panel then reports the number of fatal accidents, not the total accidents in those cells.

This was wrong in the first version: it showed the total accidents in cells that happened to contain a fatality, a number two orders of magnitude larger than the death count, which made it look as though there had been more than ten thousand fatal accidents. It is fixed. Injuries are hidden in this mode, because most injuries in those cells come from other, non-fatal accidents and mixing them in misleads.

Back to the map

Other tools on this site

All run on one machine at home; the data tools use Taiwan government open data and state their limits.