What this gives you: for any patch of ground in greater Jakarta, and any month since January 2023, how far away the nearest air sensor that actually sent readings was, how much of that month it sent, and what it read.
What it does not give you: a reading for your own street. There is no measurement between the sensors and none is invented here. It is also not a health warning and not a forecast you could plan a school day around — the forecast this case built is real, and further down you can see the two things it cannot do.
The register says 24. The record says one.
A public air-quality register is a list of registrations, not a list of working instruments, and the gap between those two things is this whole story. Greater Jakarta has 24 registrations. They sit at 16 distinct addresses — 6 of them are one rooftop in South Jakarta, listed 6 times under names that differ by a word, and not one of those 6 has ever returned an hour.
Of the 24, 11 instruments ever sent a reading and 13 never did. The record peaked in July 2024 at 9 instruments across 8 addresses. The last month with even three addresses reporting was December 2025. In August 2026 there is one: a Clarity unit registered as “Jakarta”, which returned 44 of 703 hours, and 7% of the last three months.
The register marks 2 of them as live. That is not a contradiction: 1 of the 2 has never returned a single hour to this record, so the number of instruments actually sending readings is one.
| Greater Jakarta's air sensors | Count |
|---|---|
| Registrations on the public register | 24 |
| Distinct addresses among them | 16 |
| Instruments that ever returned an hour | 11 |
| Registrations that never returned an hour | 13 |
| Addresses reporting in July 2024, the fullest month | 8 |
| Addresses reporting in August 2026 | 1 |
| Sensors meeting the standard this project set itself | 0 |
That last row is the one to hold on to. Before any of this was built, the project wrote down what a usable network would have to be: at least 3 sensors returning four fifths of their hours over three months. 0 sensors meet it. The standard was not lowered afterwards, and the result is published in red on the technical dashboard rather than softened.
What the sensors are measuring. The reading is PM2.5: soot and dust in particles under two and a half thousandths of a millimetre across, small enough to pass out of the lungs and into the blood. It is counted in micrograms per cubic metre of air — the numbers on this page are all in those units. The World Health Organization's guideline for a day's average is 15. Nothing about the number is visible; on a bad day in Jakarta the sky looks much the same as on a good one.
How far is far? For half this region, 55 kilometres.
Put the map on August 2026 and half of the study box is more than 55 km from a sensor that reported, with the worst corner 113.1 km away. Even if every registration on the register worked perfectly — every dot bright, including the ones that never switched on — half the box would still be 24.1 km from the nearest, and only 17% of it would be within 10 kilometres. The instruments that ever reported occupy a strip 25 km wide inside a region 100 km across.
The box is the box the pipeline pulls its weather and its sensors for, and it includes the Java Sea along the top, so treat these distances as a description of the sensor pattern rather than of where people live. The pattern is the point: a north–south line through Jakarta and Depok down to Bogor, and nothing much either side of it.
What the air was, where anyone was listening
Where there were readings, they were not reassuring. Averaged over every hour any sensor returned, the air held 44.71 micrograms per cubic metre — about three times the guideline. On 94.3% of the 3779 station-days complete enough to count, the day's average was over the guideline. On 27.4% it was over 55.5, the level at which health agencies tell everyone, not just the vulnerable, to change their plans.
The day has a shape. Worst at 06:00 local time, at 52.65 — the morning rush under a lid of cool night air that has not yet lifted. Best in the middle of the afternoon, 15:00 at 34.83, once the ground has warmed and the air has room to mix upward. The year has a shape too: January is the cleanest month at 28.2 and June the dirtiest at 54.99, though which sensors were alive changes month to month, so read that as a rough seasonal swing rather than a measured trend.
And the instruments disagree with each other. The two reference-grade monitors in the record average 38. The consumer-grade ones average 52.8 — 39% higher, which is a bigger gap than the distance between "fine" and "watch out". Two registrations 43 metres apart agreed to within 6.85 micrograms across 7297 shared hours, and then, in their last month together, one read 9.9 while the other read 63 over the same 21 hours. The one reading 9.9 is the instrument that is still reporting today, and its neighbour has gone silent, so there is now nothing left in greater Jakarta to check it against.
Is this normal for a big city?
There are two WHO lines, not one: 15 micrograms for a single day, which is the one this page has been using, and 5 as a yearly average. Against the yearly line, the 44.71 in this record is about 9 times over. An independent assessment of the same metro area for 2024, by the Centre for Research on Energy and Clean Air, puts Jabodetabek at 30–55 micrograms, 6 to 11 times that yearly line. This case's own stations run 28.4 to 67.4, which brackets it — two different networks, read two different ways, landing in the same place.
What this page will not give you is Jakarta's rank against other cities. The WHO's own city database — 8,023 settlements across 127 countries, the largest compilation there is — states in its documentation that it must not be used to rank the most polluted cities, because monitoring capacity and data completeness vary too much between them. The disagreement above is what that means in practice. Jakarta's official network is 10 reference-grade sensors and 110 low-cost ones, and on this record those two kinds of instrument differ by 39%. A table that mixes them can move a city further than the distance between cities.
The forecast works, and cannot make the call you need
This case does contain a real fitted forecast: given the weather, the sensor's recent history and the fires burning upwind, predict the air a day ahead. Against the crudest possible rule — assume the air tomorrow is whatever it is right now — it wins. Its typical miss at 24 hours ahead is 16.04 micrograms against 19.49, an improvement of 17.7%. That was the bar set in advance and it clears it.
Then someone tried a slightly less crude rule: carry the trailing daily mean forward. That rule misses by 18.15, and against it the forecast is only 11.6% better — somewhere between 9.4% and 13.7% once you allow for luck, an interval lying entirely below the 15% the project demanded. The original result stands as recorded, because thresholds are not moved after the fact. The fairer comparison is published beside it.
Two failures matter more than either number. The first: the forecast draws a band around itself and says the truth will land inside it 80 times in a hundred. It actually lands inside 63 times in a hundred, so the band is too narrow and should not be used for planning until it is fixed.
The second is worse, and it is the one a school or a hospital would care about.
| Hours that turned out to be… | Happened | Warned of | Caught, per 100 |
|---|---|---|---|
| over 35.5Unhealthy for sensitive groups | 14303 | 16081 | 85 |
| over 55.5Unhealthy | 6084 | 5582 | 49 |
| over 125.5Very unhealthy | 247 | 0 | 0 |
| over 225.5Hazardous | 20 | 0 | 0 |
Read the bottom two rows. 247 hours in the stretch of record the model never saw reached the level at which everyone is told to stay indoors, and 20 reached the level above that. The forecast predicted 0 of either. Its highest day-ahead prediction anywhere in the record is 99.1, against an observed maximum of 338. It has a ceiling, and the hours that hurt people live above it.
That figure could not be loaded. Everything it shows is stated in the table above.
The six checks, in plain words
Six questions were written down before any model was fitted, so that the answers could not be chosen afterwards. Here they are with their verdicts.
| Asked in advance | Answer |
|---|---|
| Does the day-ahead forecast beat simply repeating the last reading? | yes |
| Does it beat that at every lead time, from one hour to three days? | yes |
| Does it warn about more than half of the hours that turn out unhealthy? | no |
| Are at least three sensors returning four fifths of their hours, over three months? | no |
| Does the band the forecast draws around itself hold as many readings as it claims? | no |
| Is the forecast driven by the weather, rather than by the sensor's own recent history? | no |
2 of 6. Nothing was retuned to turn a red into a green, and the reason to publish all six is that the four failures tell a client more than the two passes do.
What this cannot do
Everything above rests on a record that thins out to a single instrument, and a reader should know exactly where that bites.
- It cannot tell you the air on your street. Not approximately, not with a caveat. Where the map is dark there is no measurement, and nothing here fills that in. A sensor 55 kilometres away is not a reading for your house.
- The level itself is uncertain by more than the difference between good and bad. Consumer instruments read 39% above reference ones here. Two instruments 43 metres apart ended up disagreeing by a factor of six. Any single number quoted from this record carries that spread whether or not it is printed next to it.
- "Complete" can be counted two ways, and only one of them is honest. You can divide the hours a sensor sent by the hours in the month, or by the hours the sensor happened to be switched on. The second flatters badly: in March 2026 one sensor returned 36 of 744 hours — that is 5% of the month — and against its own switched-on hours it scores as nearly complete. Every completeness figure on this page divides by the calendar.
- The forecast was tested on a panel that dissolves. Of the 10 addresses that ever reported, 9 have since gone quiet, and they are what the test was made of. One station (Bogor Selatan) had no training history at all and scored 5.6% against 21.3% for stations the model had seen — so the honest expectation at a brand-new sensor is the lower number, not the headline.
- Weather is not what drives it. The check asked whether mixing and wind were among the forecast's strongest inputs. Neither is. The eight strongest are led by which sensor it is, the sensor's newest reading and the same hour yesterday. A model leaning that hard on which sensor it is, is a model describing a sparse network as much as it is describing air.
What you must not conclude from this page. It is not a health assessment and cannot become one. Do not read a dark patch as clean air — it means nobody measured, and the measured parts of this region were over the guideline on 94.3% of complete days, so the honest prior for an unmeasured patch is "probably also bad". Do not use a dot's colour to decide whether a child can play outside, whether a ward needs a clinic, or whether one neighbourhood is safer than another: a monthly average from one instrument, of uncertain calibration, 55 kilometres from most of the map, is not a diagnosis. And do not treat the forecast as a warning system. It misses about half the unhealthy hours and has never once predicted a very unhealthy one. Nobody's asthma plan should depend on it.
What a government or an NGO can actually do with this
Three things follow, and the first is cheap.
- Fix the register before buying anything. 13 of 24 registrations have never returned an hour and 6 of them are one rooftop listed 6 times. A register that cannot distinguish an instrument from an entry makes any coverage claim built on it meaningless — including optimistic ones.
- Keeping sensors alive is worth more than adding them. This network did not fail for want of instruments; it had 9 working at once and lost them one at a time to nothing more dramatic than time. Maintenance, not procurement, is the binding constraint, and it is the cheaper of the two.
- Put two instruments where you put one. Everything anyone knows about how far these readings can be trusted comes from the single pair that happened to share an address. When its partner went silent, the survivor became unverifiable. Pairing is how a network stays checkable.
Could this run in daily operations?
BLOCKED BY INPUTS — none, and the model is not the reason
Every pipeline here is a batch backfill: it fetches a season or an archive, not a live feed. Anything operational means rebuilding the ingest for near-real-time arrival, whatever the verdict below says.
What stops it. Four of six checks failed, but the one that decides it is data coverage — it asks for a handful of stations reporting near-continuously for three months. Almost none of the registered public sensors still report, and two last published nearly a decade ago.
What it would need. A working sensor network. That is a procurement and maintenance problem rather than a modelling one — a production system on this feed would have degraded silently.
If you want the detail
The methods, the checks and how each came out, the comparison against published forecasts elsewhere, and everything still uncertain are on the technical article, and the technical dashboard carries the live forecast, the airshed and the fire season behind it. Every figure here is read from the same record as those, so they cannot drift apart.
Words on this page, in plain language 4 terms
- ERA5
- A reconstructed record of past weather — a model reconciled with observations to give a consistent history of wind, rain and temperature.
- hold-out
- Data deliberately kept away from the model while it learns, then used to test it. Without one, a model is graded on the answers it was shown.
- calibration
- Adjusting a model's output so it lines up with a trusted measurement. It can genuinely fix a scale error — or quietly force agreement and teach you nothing, which is why the before-and-after is always published here.
- FIRMS
- NASA's Fire Information for Resource Management System, a near-real-time feed of hotspots detected from orbit. It records where a satellite saw heat — not what was burning, and not who lit it.
Ground readings from OpenAQ and their originating providers; weather from the Copernicus ERA5 record; fires from NASA FIRMS. Sensor record 2023-01-16 to 2026-08-30; the forecast is scored only on the period after 2025-04-07, which the model never saw. Squares on the map are 2.8 km on a side. Written to be read without a background in air chemistry or statistics; nothing was rounded to make a point.
Found something wrong on this page? Report a correction — it opens a pre-filled issue — or email taufik.adi@openstudy.id. Corrections are credited by name in the errata, and one that changes a finding says so on the page.