Indonesia's official poverty rate is published for 514 regencies and cities, once a year, from a household survey. Below that there is nothing official. So a tempting idea has taken hold, here and in many countries: look down from orbit — at roofs, at night-time lights, at how much ground is built on — and fill in the gaps. Governments and aid agencies are already doing it. This is a test of whether it works.
What this gives you: an honest measurement of how well it works, and a warning. Predicting a province the model has never seen, it accounts for about 40% of the differences between regencies. That is real signal, and it is far less than the same model appears to have if you test it carelessly.
What it does not give you: a way to decide who gets help. Within a single province — which is where the choosing actually happens — it accounts for about 8% of the differences between neighbouring regencies. On cities it is worse than simply assuming the national average. Do not use this to allocate anything.
Both halves are the finding. The number is not zero, and the number is not good enough. Publishing only the first half is how a map like this ends up deciding something it cannot support.
The same model, tested three ways
Nothing about the model changes between these three rows. Only the way it is marked does — and it is the same exam paper, graded honestly or leniently.
| How it was tested | Differences explained | What that test means |
|---|---|---|
| Split the map at random | 65% | the quick way, and the way most results are reported |
| Keep neighbours apart, 200 km blocks | 58% | so a regency cannot be predicted from the one next door |
| Hold out a whole province | 40% | predicting somewhere the model has never seen — the honest number |
The first row is the number a hurried write-up would report, and it is 26 points too generous. The reason is simple: poor places sit next to poor places. If you scatter the test across the whole country, almost every regency you are trying to predict has a near neighbour in the training set, and the model is half remembering rather than inferring.
The fairest test, and the one least often run
Take the satellite away completely. Give every regency its own province's poverty rate — one number repeated across all of it — and score that. Then score the satellite model on the same 514 regencies.
| Way of estimating a regency's rate | Differences explained | Typical error |
|---|---|---|
| Just use the province's rate — no satellite at all | 69% | 2.7 pts |
| The satellite model | 73% | 2.5 pts |
The satellite is closer to the truth on 54% of regencies — better than a coin toss, and not by much — and it improves the typical error by 0.2 of a percentage point. That is the honest answer to the question in the headline. Most of what looks like seeing poverty from orbit is knowing which province you are looking at, and the household survey already told you that.
Is that a good score, or a bad one?
The bar this case sets itself is 0.50, and the model reaches 0.40, so it fails. But a bar is only meaningful next to how the score was earned. Studies that hold whole areas out — the same hard way this model is judged — report 0.67, 0.56 and 0.45. This model's 0.40 is below all of them. Tested the easy way instead, by splitting the map at random, the very same model reads 0.65 — which would sit comfortably among them. The gap between those two numbers is the whole argument of this page.
As for the 0.50 itself: it comes from a single study of 38 regencies in one Indonesian province, and that study measured its score in-sample — on the same places it learned from, which is easier again than either test above. So the bar this case fails was set by a study working under gentler conditions than the case holds itself to. That does not rescue the score. It does mean the honest comparison is with the numbers above, not with the bar.
Why it fails exactly where it would be used
A national figure hides the thing that matters. Three checks show where the apparent skill actually comes from.
- Most of it is knowing which province you are in. About 72% of Indonesia's variation in poverty is between provinces rather than within them — and the official survey already tells you that. Strip it out and the satellite explains about 8% of what is left, across 37 provinces.
- More than half the error is a single wrong number per province. 57% of the total error is the model being consistently too high or too low across an entire province at once. It is partly re-learning each province's own offset, not seeing poverty.
- On cities it is worse than a coin toss dressed as a guess. Across the 98 cities it performs worse than someone who simply assumed the national average for every one, while on the 416 rural regencies it explains about 37%. Roofs and lights mean something different in a dense city than in a village.
One more trap, and it is an easy one to fall into
Test it by training on earlier years and predicting 2025 and it looks superb — about 95% of the differences explained. Almost all of that is an illusion: a regency's poverty rate barely moves from one year to the next, so the model is largely repeating what it was told about the same place a year earlier. Remove that and the same test gives about 33%.
Both numbers are published, because the first is the one most write-ups would print.
What a government or an NGO can actually do with this
Three things follow, and the first is the only safe use.
- Use it to decide where to survey next, not who receives a benefit. On the fairest split — a whole province held out — the model explains 40% of the variation with a typical error of 5.4 percentage points of poverty rate. That is accurate enough to rank where enumerators should go next, and nowhere near accurate enough to decide a household transfer. Targeting fieldwork is cheap and reversible; targeting money is neither.
- Read the interval before the estimate, and notice which split produced it. The same model scores 65% when neighbouring districts help train it and 40% when a province is held out whole. The gap between those two numbers is the honest measure of how far it travels, and the second is the one that matches how it would actually be used. Two districts whose intervals overlap cannot be ranked against each other however different their central estimates look.
- Benchmark against the official survey rather than replacing it. The bar set before the analysis ran was to account for 50% of the variation between districts; the honest split reached 40%, so the check is published as a failure rather than quietly relaxed. BPS remains the authority on the rate. What this adds is coverage between survey rounds — presenting it as a rival number invites a fight it would deserve to lose.
What you must not conclude. Not that satellites are useless for this — they carry real information, and this model does find some. And not that the official survey is wrong; it is the thing being measured against. What you must not do is take a map built this way and use it to decide which neighbourhood, village or family receives assistance. At the level where that choice is made, the evidence here says the map does not know.
Could this run in daily operations?
FEASIBILITY ONLY — none — the model surface is deliberately not published
Every pipeline here is a batch backfill: it fetches a season or an archive, not a live feed. Anything operational means rebuilding the ingest for near-real-time arrival, whatever the verdict below says.
What stops it. The blocking check failed — skill measured by holding out one whole province at a time. A random split of the same rows scores far higher, and that gap is the flattery a non-spatial split buys rather than skill.
What it would need. Out-of-sample skill measured the way it would be used, predicting a province the model has never seen. Until then the page ships the official figures and no surface of its own.
If you want the detail
The features, every fold design, the failed checks and the things still uncertain are on the technical article, and you can move the controls yourself on the dashboard. Every figure above is read from the same record, so they cannot drift apart.
Words on this page, in plain language 6 terms
- a look
- One pass of the satellite over a field. More looks means the crop's cycle is sampled more often; too few and a short stage — flooding, heading — can happen entirely between two looks and never be seen.
- fold
- One slice of the data in a repeated test. How the slices are drawn can matter more to the score than the model does.
- orbit (ascending/descending)
- Whether the satellite passed heading north or south. It looks from a different angle each way, so measurements from the two are never mixed.
- BPS
- Badan Pusat Statistik, Indonesia's national statistics agency. Its published figures are the benchmark every case here is scored against.
- regency
- The English name for a kabupaten: the administrative level below a province, and the level Indonesia publishes most official statistics at. A regency is not a city — cities are counted separately here, and on several of these cases they behave differently enough to be reported on their own.
- kecamatan
- An Indonesian sub-district, below a regency. Official poverty figures are not published at this level, which is why estimating them there is both useful and risky.
Data vintage 2026-08-30. Covers 514 regencies and cities, 2016–2025. Written to be read without a background in statistics; nothing was rounded to make a point.
Found something wrong on this page? Report a correction — it opens a pre-filled issue — or email taufik.adi@openstudy.id. Corrections are credited by name in the errata, and one that changes a finding says so on the page.