← 網誌 · Turning the police's daily lost-property notices into a searchable map: the item name is buried in a sentence, the location is just text

Turning the police's daily lost-property notices into a searchable map: the item name is buried in a sentence, the location is just text

When you lose something, the first thought is "did anyone find it, and where did it go?" The National Police Agency actually posts found items every day and lets you download them — it is just that they are spread across 34 receiving agencies, are text only, and even "what was found" has to be picked out of a sentence of officialese. This writes up the process of turning that into a search tool.

2026-09-09 作者 William Hsu

這篇講什麼
  1. The data is public — nobody had just organised it
  2. "What was found" isn't a field
  3. The location is text only, with no coordinates
  4. Personal data: the source already redacts it, and we don't undo that
  5. Why we don't connect to the official search site
  6. Some sources: being able to scrape doesn't mean you should
  7. After launch: one dataset actually held two files
  8. Deliberately left out
  9. Being clear about what it can't do
Lost-and-found across Taiwan, blue dots accurate to the door, red to the road, grey to the district
The lost property the police post every day, put on a map: blue dots are accurate to a street number or lane, red to the road, grey only to the district's average position

The data is public — nobody had just organised it

The National Police Agency publishes the "Police lost-and-found online notices digital batch files" (dataset 7317) on the government open-data platform: a single CSV, downloadable directly with no API key, about ninety thousand records, kept on a rolling last-six-months basis, under the Open Government Data License, and usable commercially.

It covers not just the 22 city and county police departments, but the specialist police agencies too:

The data may be public, but it is nearly impossible for an ordinary person to search for their own lost item — it is scattered, text only, and has no usable front door. Merging these 34 agencies, six months and ninety thousand records into one search box and putting them on a map is what this tool does.

"What was found" isn't a field

You would assume there is a column called "item name." There isn't. What was found is wrapped inside a sentence of officialese:

"The finder handed in: one black wallet containing some cash. The owner should come to claim it within the six-month notice period, bringing their own seal and photo ID."

The only useful part is the stretch between "The finder handed in:" and ", the owner should." A single rule pulls it out — cleanly for 99.3% of the ninety thousand records; the under-1% that don't follow the format (bulk inventories, truncated sentences) are put in a separate category rather than guessed at.

Only once it is extracted can you search at all. When you type "wallet," "EasyCard" or "phone," you are matching against that item-name string, not the whole official sentence.

The location is text only, with no coordinates

A found-property notice has only the "where found" text, no latitude and longitude. Of the ninety thousand records:

I brought in another open dataset — each county's street-number coordinates — matched it offline, and turned the addresses that have a road name into coordinates, so in the end about 95% of the records can go on the map. For those that can't even be placed to a county, I work back from the name of the receiving police unit (the Aviation Police go to Taoyuan; a given railway police precinct goes to where its station is).

The map draws dots only, not radius circles. The data has a single location and no extent, so forcing a fixed-radius circle onto it would just be faking precision. It is the same as with the flood and campsite tools: don't pretend, on screen, to have something the data doesn't.

Personal data: the source already redacts it, and we don't undo that

480 notices carry a redacted name in their text (like "Chang ○-ling"). The source de-identified these to begin with, so I show them exactly as they are — I don't restore them and don't cross-reference them. The whole dataset has no contact details for any owner — claiming has always required going to the receiving unit in person with ID, and that is written plainly on the data page.

Why we don't connect to the official search site

The National Police Agency has its own search site, but that route is a dead end: a search requires three mandatory fields (meaning you have to know the answer before you can search), and since July 2025 it also has an image captcha. That design isn't meant to be connected to by a program. So I don't touch the web version at all, and use only the CSV it outputs every day — which is the whole point of open data: the agency dumps out the entire batch, and how to present it is left to everyone else.

Some sources: being able to scrape doesn't mean you should

I looked into whether lost property from other transport hubs could be pulled in too: Taipei Metro, Taichung Metro, Kaohsiung Metro, Taoyuan Metro, Taoyuan Airport and the freeway service areas all have public, captcha-free interfaces, and could be included.

But there are also ones I won't touch under any circumstances. Taiwan High Speed Rail's lost-property front end spits out a big blob of JSON and is easy to grab, yet its terms flatly forbid reproduction and distribution; and one department store's privacy policy forbids crawlers and AI agents word for word. Things like these I don't do, even where I technically could. Whether you can scrape something and whether you should are two different questions.

After launch: one dataset actually held two files

I got this one wrong at first, so I'll write it up honestly.

One day I noticed the newest item on the site was stuck at 3 September, with nothing new for days after. My first thought was "the Police Agency has paused updates" — government data running a few days late is common, and I even checked whether there was an official announcement. There wasn't.

It turned out that wasn't it. Under this dataset on data.gov.tw there was actually more than one download file, two of them identical in format: the one we had been fixed on fetching had stopped at 9/3 and gone no further; the other, under the same dataset, was updating right up to that day.

Our update program has a safeguard: if it finds that the "newest date" hasn't moved forward, it won't take that possibly-incomplete batch and overwrite the existing data. So the site didn't break and didn't show anything wrong — it just stayed at 9/3. The safeguard did its job, but what it was safeguarding was data we had pointed at the wrong source ourselves.

Once I knew the cause, I didn't just swap the URL for the other file and call it done — who knows what day that one might stop too. Instead I made the program ask data.gov.tw every time: which files does this dataset have right now? It fetches them all, compares the dates, and automatically uses the newest. From now on, whether it changes an ID, retires one file, or adds another, it will switch to the newest on its own and never quietly stall again.

While I was at it I also blocked a small edge case: the data occasionally contains fake dates, like 211360621530 (year 2113, month 60 of the ROC calendar). If you just compare magnitudes, a fake future date like that would be treated as "the newest," so before picking, it now validates dates and filters those out.

Deliberately left out

Being clear about what it can't do

The tool itself is at lazetool.com/lost-found.

這篇文章寫的是本站實際的實作與量測。工具本身在老照片動起來與變老變年輕,都可以直接試。

← 上一篇同一個地點多出三十隻走失貓狗,深入調查之後,答案讓人意外

看其他文章