How bad will
the air be
tomorrow?
Greater Jakarta gets its air quality the way it gets its weather: after the fact. This case turns three open feeds — ERA5 meteorology, NASA fire detections and the handful of ground sensors still reporting — into a PM2.5 forecast with a horizon and an error bar, graded against every trivial rule that could stand in for it: carrying today forward, carrying yesterday's daily mean forward, and the long-run average for this hour and season.
Jakarta's air is not occasionally bad. It is usually bad.
Of the 3,779 complete days on record at the city's ground monitors, 94.3% exceeded the WHO 24-hour guideline of 15 µg/m³, at a mean of 44.7 µg/m³. The 95th percentile hour reaches 94.4.
Read that mean carefully. It pools instrument classes: across the 16,021 hours from the two reference-grade regulatory monitors the mean is 38.0 µg/m³, against 52.8 across the 40,690 hours from consumer-grade units. The gap is instrument, not geography, and the technical page quantifies it.
Below, everything the most recently reporting sensor has actually sent in the last three weeks — not a smoothed line, the raw hours, gaps and all. It is a thin record, and that thinness is the operating condition this whole case is built around. Under it, the shape of the problem across the full archive: the daily and seasonal cycles that a forecast has to get right before it earns any attention at all.
Observed PM2.5 · last 21 days · station
Hover for the hour. Gaps are real: the sensor's own uptime, not smoothing.
The daily rhythm · every observed hour, all stations
So what: a dial showing the current number is worthless to anyone who has to schedule something. The value is in the next 72 hours.
Beating “tomorrow looks like today”
At the 24-hour horizon the model's mean absolute error is 16.0 µg/m³ against persistence's 19.5 — a 17.7% improvement on the held-out future, and 11.6% against the strongest trivial rule available at that lead.
Which baseline matters. Persistence is not the hardest trivial rule at a 24-hour lead — a last 24 h mean reaches 18.1 µg/m³ where persistence reaches 19.5, because at a 24-hour lag persistence is handed the entire daily cycle for free. Measured against that stronger rule the improvement is 11.6% [9.4%, 13.7%], which is below the 15% this case registered in advance. The skill check below is reported against persistence, as it was written; on the fairer ruler it would not pass, and the technical page works through why.
Every number here comes from a single time-based split: the model never sees an hour later than the cut, and is then asked about the months that follow. A random split would leak tomorrow into today's training set and inflate every figure on this page. Two baselines are drawn below — persistence (carry the last reading forward) and diurnal climatology (the average for this hour of day) — and the technical page adds four more.
Error by horizon · held-out period
So what: skill is not a single number, and it is not only a property of the model. The peak at 12 hours is where persistence is at its worst, not where the forecast is at its best — at a 12-hour lead the baseline is comparing night to afternoon. Read the curve against the line it is drawn on.
What the model actually leans on
Feature attribution pending.
Permutation importance at the 24-hour horizon: how much worse the model gets when each feature is shuffled. Features are grouped by what they represent, because the interesting question is not which column wins but which physics does — the air's own memory, the depth of the layer it mixes into, the wind that ventilates it, the fires upwind, or the clock.
Feature importance · 24 h model · permutation, 5 repeats
Δ MAE ON LOG SCALESo what: if this chart were pure autoregression — lagged PM2.5 and nothing else — the model would be a clock, not an air-quality forecast, and it would fail the moment conditions changed. One of the published checks exists to test exactly that.
Where the bad air comes from
Jakarta does not breathe alone. The rose below splits every measured hour by the direction the wind was blowing from: petal length is how often, petal colour is the mean PM2.5 that arrived on it. Switch to fires and the same eight sectors show NASA's VIIRS hotspot counts across the airshed — Sumatra's peatlands, the length of Java, the southern edge of Kalimantan — split into three distance rings.
Wind rose · mean PM2.5 by arrival direction
Hover a petal for the numbers. Fires are attributed to a three-sector arc around the bearing from Jakarta and lagged one day, so an afternoon satellite overpass never informs that morning's forecast. Volcanoes, gas flares and other permanently hot ground are excluded — see below.
Fire hotspots detected across the airshed, by month
VIIRS S-NPP · 95–119°E · 9°S–6°NNot every hot pixel is a fire. Indonesia has around 130 active volcanoes and a great deal of gas flaring, and VIIRS sees all of it. The satellite does label those detections — but only in the standard-processing archive; the near-real-time product that covers the last few months omits the label entirely. Filtering on the label alone would therefore clean the history and leave the recent tail full of volcanoes, putting a step change in this chart exactly where the two products meet — and fire counts feed the forecast, so the step would propagate into the model. Instead the archive is swept across its full record for everything it ever flagged as a volcano, flare or offshore source; those locations are gridded to 0.01° cells with a one-cell buffer for pointing jitter, and the resulting mask (3,908 cells, 708 of them flagged directly) is applied to every row of both products. It removes 5.9% of otherwise-qualifying near-real-time detections against 3.6% of archive ones — the archive needs less help because its own labels already caught most of them, and that asymmetry is exactly what used to land at the seam. In April 2026, the one month the seam falls inside, both products now read as roughly a third static (34% archive, 36% near-real-time); before the fix the near-real-time side read as zero.
So what: the same fire is worth a lot or nothing depending on where the wind is. That is why the model sees upwind fire counts, not a regional total.
A metropolis of 32 million is watched by 1 working public sensor.
The open record lists 24 PM2.5 stations inside Jabodetabek. Only 11 have ever returned an hour of data. 2 reported in the last fortnight — and 1 of those has delivered no hourly measurements at all.
The denominator is softer than it looks. Those 24 registrations resolve to 16 distinct addresses — 6 of them share one coordinate under six different names, none of which ever reported — and 13 of the 24 have never delivered an hour of data, including legacy diplomatic-post entries that duplicate feeds registered separately. Told honestly it is 11 instruments that ever worked, which is a smaller number and a worse one.
Each line below is one station's life. The solid bar is the hours this build actually holds; the faint rule behind it is the lifetime the provider's metadata claims — the two diverge, and the gap is itself worth knowing before anyone budgets for an analysis. Read left to right and the network does not grow: it flickers. This is not an access problem a bigger API key would fix. The instruments are not there, and the ones that are there are mostly low-cost units run by embassies, NGOs and volunteers.
And a network this thin loses something a bigger one takes for granted: the ability to catch itself. Two of these instruments sat 43 metres apart and tracked each other to 6.9 µg/m³ across 7,297 shared hours — until, in the last months both were alive, they diverged to 9.9 against 63.0 µg/m³. Then one of them went dark. Nothing in Jabodetabek can now say which was right.
Station lifetimes · every PM2.5 station in the bbox
Where they are
2-D CANVAS · NO WEBGL · HOVER OR TAB FOR DETAILSo what: a dense city-wide nowcast grid cannot be honestly validated here, so this case does not claim one — and the deeper cost is not coverage but corroboration. A single sensor cannot be shown to be wrong. The station-coverage check is published failed rather than quietly relaxed, and on the review's reading it is the most important output this pipeline has.
The held-out months, unedited
The 24-hour forecast plotted against what actually happened, for every hour after the training cut. The shaded band is the model's own 80% prediction interval — quantile regression, so it is allowed to be lopsided, which for pollution it always is. Toggle persistence to see the baseline the model has to beat.
24-hour forecast vs observation · held-out period
Gates · thresholds fixed before the first model run
2/6 PASS- pass G-E124 h forecast beats persistence by >= 15% MAE MAE 16.04 vs persistence 19.49 µg/m³ on the held-out future (+17.7%). On RMSE the same comparison is +20.8%.
- pass G-E2Beats persistence at every horizon weakest horizon is h=1 at +4.4%; all horizons positive
- fail G-E3Episode recall at 24 h (PM2.5 ≥ 55.5 µg/m³) 6084 episode hours in the held-out window. Model recall 0.489 at precision 0.534; persistence recall 0.476 at precision 0.475. The threshold is not moved to meet a near miss.
- fail G-E43+ stations at 80%+ hourly completeness over 90 days 24 PM2.5 stations exist in the bbox; 2 have reported in the last 14 days; 0 clear 80% completeness over the trailing 90 days
- fail G-E580% prediction interval is calibrated (72-88% coverage) empirical coverage 62.8% at h=24
- fail G-E6Meteorology drives the model, not autoregression alone mixing term present: False; wind term present: False
One failure the checks above do not capture. The episode check is written at a single threshold. Above it the US index has two more categories, and at those the forecast is not merely weak — it is silent. Across the held-out window 247 hours reached the “very unhealthy” level of 125.5 µg/m³ and the model called 0 of them. Its highest 24-hour prediction anywhere in the record is 99.1 µg/m³ against an observed maximum of 338.0. Training on log1p and scoring on absolute error pulls predictions toward the middle, and the result is a hard ceiling below the categories a city would act on. No school advisory or outdoor-work suspension should be triggered by this forecast in its current form.
So what: a failed gate published with its diagnosis is worth more to a client than six green ticks they cannot audit — and a failure the gates were never written to catch is worth publishing too.
Audited adversarially, by its own author.
A second read of the same data against the published literature, written to attack the first. It was done by the person who built the instrument — no external reviewer has checked it, and one person reviews everything here — if you would like to be the first, breaking a gate here is worth more than another instrument. It re-derives the headline against five baselines this page did not try and finds the skill check fails on the strongest of them; it shows the skill curve's peak at 12 hours is the shape of the pollutant's autocorrelation, not of the model; it dates the month the city's last two co-located instruments stopped agreeing; and it argues the thin network is not this case's limitation but its result.
Read the technical article →Instrument
Method, models, evaluation and instrument · Taufik Adi Nugraha
Estimator choice, model construction, baselines and validation design. Implementation was AI-assisted (Claude), under the builder's direction. Thresholds were fixed before any result was seen, and every gate is published with its outcome including the failures.
Domain interpretation · open
This instrument has no domain author yet. Someone who knows Indonesian air-quality monitoring. Only 2 of 24 registered public sensors still report; two US diplomatic stations last published 3,580 days ago. What is missing is why the network collapsed, what BMKG's own record would add, and what a 24-hour PM2.5 forecast is actually useful for here.
Write the interpretation and take first author; the instrument's author above becomes methods co-author for the analysis and evaluation already built here.
Found something wrong on this page? Report a correction — it opens a pre-filled issue — or email taufik.adi@openstudy.id. Corrections are credited by name in the errata, and one that changes a finding says so on the page.