← The instrument REVIEW ARTICLE · DATA VINTAGE 2026-08-30
METHODS & VALIDATION · POVERTY MAPPING FROM SPACE

Can a satellite count Indonesia's poor?

This case publishes a failed out-of-sample check and blames the target: a headcount against a nominal, region-specific poverty line, it argues, is not the asset index the satellite-poverty literature reports high numbers for. We put that defence to the case's own data, four ways. It does not survive any of them — and the real explanation is better news than the one being offered.

Abstract

Question. A gradient-boosted model on roofs, night lights, population and land cover predicts the BPS poverty headcount for 514 Indonesian regencies at leave-one-province-out R² 0.395, against a published threshold of 0.50. The case attributes the shortfall to its target: a monetary headcount rather than an asset index. Is that diagnosis correct?

Method. We recompute every published skill statistic from the case's own cross-validated predictions, then run five pre-specified tests on the identical folds: a feature-family horse race; the official poverty line added as a diagnostic oracle; the same features re-targeted onto the poverty line itself; the shipped disaggregation procedure re-run one administrative level up where ground truth exists; and the operational baseline of simply reusing an older published rate.

Findings. Every published number reproduces exactly: the largest discrepancy across 12 recomputed statistics is 0.0000. The diagnosis does not. One: handing the model the official poverty line it says it cannot see moves R² from 0.395 to 0.409 — +0.014, against the 0.105 it needs. Two: the same features predict that poverty line in a province they have never seen at R² 0.378 and ρ 0.62 — better in rank than they predict the headcount — so the price level is not a blind spot. Three: the province-level offsets the case attributes to the line are explained by the line at R² 0.014, and by ordinary regression-to-the-mean compression at R² 0.357. Four: the thresholds said to come from an asset-wealth literature come, on the case's own configuration file, from Putri et al. (2022), a study of the BPS poverty rate in one Indonesian province — where there is only one poverty line. Held out on that same province, this model reaches ρ 0.62 against their 0.77, at RMSE 3.38 pp against their 3.18 pp. Separately, 22 of the 31 features are not paying their way: night lights and population alone score R² 0.412, above the published 0.395, and the 110-million-footprint building layer adds rank skill (ρ 0.52 against 0.42) while costing level skill.

Conclusion. The case's conclusion — that a monetary headcount is harder than an asset index — is right and well supported: every study that has measured both on one dataset agrees. Its stated mechanism is wrong. What actually separates this case from the three studies that set its thresholds is that none of them held space out, and a random fold on these same rows scores R² 0.653. Measured against the one published satellite model of a monetary headcount that does hold space out (Engstrom et al. 2022, R² 0.45 linear / 0.564 random forest), 0.395 over an archipelago is unremarkable rather than damning. The product that survives is narrow and genuinely defensible: benchmarked disaggregation beats the flat parental rate on 54.1% of regencies and cuts mean absolute error from 2.69 to 2.49 pp. It must never be used at a level BPS already publishes, where last year's official number scores R² 0.978, nor for household targeting, where the best-known satellite wealth index still excludes 32.8% of Indonesia's eligible poor.

1 The claim under test

BPS publishes a poverty headcount for each of 514 kabupaten and kota, once a year, from the March Susenas. Below that line nothing official exists. Indonesia has 7,069 kecamatan and no published poverty rate for any of them. That gap is the case's reason to exist, and it is a real one: the social registry that targets assistance, the village-fund formula and every provincial prioritisation exercise all operate below the level at which the statistic is measured.

The case fits satellite features to the regency rate, holds out one province at a time, and publishes the result: R² 0.395, Spearman ρ 0.55, RMSE 5.4 pp, against thresholds of 0.50, 0.70 and 4.0 pp. The check fails and the failure is published in red. That is the right behaviour, and it is not what this review is about.

This review is about the diagnosis attached to the failure. The case argues that the published literature's high R² values are "not the right expectation here" because they were obtained on an asset index, while this target is a headcount against a nominal, region-specific poverty line — and, in its words, "no satellite can see a price level." A failing model with a correct diagnosis is a scientific result. A failing model with a comfortable diagnosis is a rationalisation. The difference is testable, and the case's own data are enough to test it.

2 Prior art — and where these thresholds actually came from

Begin by granting the case its general point, because the literature grants it. Three studies have measured an asset target and a monetary target on the same data with the same features, which is the only clean way to isolate what the choice of target costs.

asset / wealth index monetary poverty or consumption — same features, same units, same paper Steele et al. 2017 Bangladesh, 600 clusters 0.76 DHS wealth index 0.25 consumption poverty (PPI) −67% of the explained varianceJean et al. 2016 Africa, pooled cross-validated 0.56 DHS asset index 0.45 consumption expenditure −20% of the explained varianceEngstrom et al. 2022 Sri Lanka, GN divisions 0.68 asset index 0.61 poverty headcount −10% of the explained variance r² reported by the paper. In Steele et al.'s urban subset, the consumption-poverty r² is 0.00.
Figure 1. The three studies that report both. Each pair is one paper, one dataset, one feature set, two target definitions. The coral bracket is the drop.

Steele et al. (2017) is the cleanest: across 600 Bangladeshi survey clusters, identical mobile and remote-sensing features explain r² 0.76 of a DHS wealth index and r² 0.25 of a consumption-poverty measure. In their urban subset the consumption-poverty r² is 0.00 — nothing at all. Jean et al. (2016) report pooled cross-validated r² 0.56 for a DHS asset index and 0.45 for consumption expenditure, and note that restricting the sample to households below twice the international poverty line drops r² to about 0.12. Engstrom et al. (2022) find satellite features explain 68% of an asset index and 61% of a poverty headcount over the same Sri Lankan units.

The case is right that a monetary headcount is the harder target. Every study that has measured both agrees, and the penalty runs from a tenth to two thirds of the explained variance.

The rest of the literature is consistent, and worth stating precisely because the page invokes it loosely. Yeh et al. (2020) reach r² 0.67 for asset wealth over 19,669 African villages under held-out-country validation — but fitted separately within urban and rural strata that becomes 0.40 and 0.32, so much of the headline is telling town from country; for change over time it is 0.15; and their own asset index correlates with log consumption at r² 0.50, which caps what even a perfect asset model could reach against a monetary target. Chi et al. (2022) explain 56%–70% of household wealth at 2.4 km, with an index defined as relative to others in the same country.

Yeh et al. 2020 spatial hold-out 0.67 DHS asset wealth indexEngstrom et al. 2022 random CV 0.61 poverty headcount rateJean et al. 2016 random CV 0.56 DHS asset indexChi et al. 2022 spatial hold-out 0.56 relative wealth indexPutri et al. 2022 no hold-out 0.50 BPS poverty rate (P0)Jean et al. 2016 random CV 0.45 consumption expenditure this case, province held out 0.40 this case, random fold 0.65 asset / wealth index monetary poverty or consumption R² reported by the paper →
Figure 2. The same benchmarks by reported R², coloured by what was predicted and labelled by validation design — the variable that decides whether a number is comparable to this case's. The dashed lines are this case's own two results.

Now the thresholds. They did not come from any of that. The case's configuration file records their provenance and names three studies: Putri, Wijayanto & Sakti (2022) — Pearson 0.71, Spearman 0.77, RMSE 3.18 pp — Sartirano et al. (2023) at ρ 0.75 across 513 regencies, and Chi et al. (2022).

Putri et al. did not predict an asset index. They validated against the official poverty data published by BPS at regency level — the same P0 headcount, against the same nominal poverty line, in the same country.

The R² ≥ 0.50 bar and the RMSE ≤ 4.0 pp bar are therefore not asset-wealth numbers imported by mistake. They are Indonesian monetary-poverty numbers, chosen for exactly the right target. Sartirano et al. supplies the ρ ≥ 0.70 bar and is an asset index — but the case's stated defence, that its thresholds belong to a different kind of outcome, is contradicted by the file that sets them.

There is a real mismatch between this case and its three benchmarks, and it is specific: none of them held space out. Putri et al. correlated a constructed index with official rates across the 38 regencies of a single province. Sartirano et al. ranked a published index against a survey index on the units both already cover. Chi et al.'s 0.56 is the low end of a range whose upper end is a within-country census comparison. Not one of them asked what a model knows about a region it has never seen — which is the only question this case asks. §5 shows what that difference is worth.

3 Data, features and the fold design

Target. BPS poverty headcount P0 by kabupaten/kota, 2016–2025, 5,140 regency-years, from the March Susenas via the WebAPI.

Features. 31 intensive features in four families — 11 from Google Open Buildings roof footprints (density, footprint share, roof-size distribution), 6 from NASA Black Marble annual radiance, 3 from WorldPop, 9 from ESA WorldCover class shares, plus 2 binary geography flags. Coordinates and unit identifiers are excluded by design.

Folds. Leave-one-province-out over 38 provinces. A held-out province leaves the training set in every year, so no earlier copy of a test regency can leak back through the panel. We verified this against the fold code, and it holds. A 200 km block scheme and a random 10-fold are computed on the identical rows.

What the design cannot fix. Two of the four feature families are single-vintage — Open Buildings is dated 2023, WorldCover 2021 — and the model carries no year term. Only lights and population vary over time. Any temporal claim this case makes is therefore carried by two of thirty-one features, and the case says so.

4 Finding one — the published numbers are exactly right

Before criticising a result it is worth confirming it. We reloaded the case's stored cross-validated predictions and recomputed all 12 published skill statistics from scratch. The largest disagreement with the published table is 0.0000 in R².

0 0 10 10 20 20 30 30 40 40 perfect agreement fitted slope 0.57 — the model compresses kabupaten kota (cities) official BPS poverty rate 2025 · % of population → held-out prediction · % →
Figure 3. The published headline. Each dot is one regency in 2025, predicted by a model that never saw its province. Cities are the pale markers. The teal diagonal is perfect agreement; the coral line is the fit. Hover any dot.

The picture that emerges is not noise — it is compression. The fitted slope of prediction on truth is 0.57: the model pulls every regency toward the middle, under-calling the poorest and over-calling the least poor. Bias overall is +0.50 pp, so it is not a level error but a range error. A compressed predictor is exactly what a booster produces when the informative signal is weak relative to the target's spread — and it is also what a model produces when it is missing a variable that separates the extremes.

Where the variation in the 514 regency rates lives between provinces · 72% within · 28% Standard deviation 6.9 pp in total, 3.8 pp once each province's own level is removed. The nominal poverty line each province is judged against, 2025 Rp 434,076 Rp 918,778 one tick per province mean · 2.12× end to end (3.64× across all 514 regencies) If the case's defence is right, the 72% on the left is the part a satellite cannot reach, because it is set by a price level the satellite cannot observe. §7 and §8 put that to the test.
Figure 4. Top: how the variance in the 514 regency rates splits between and within provinces. Bottom: the mean poverty line in each of the 38 provinces, one tick each.

72% of the variance in regency poverty rates sits between provinces. That is the structural fact behind everything that follows: a model that cannot set a province's level cannot score well, no matter how well it orders regencies inside one. The case knows this and reports it as 57% of the squared error being a constant province-level offset.

5 Finding two — most of the gap to the literature is fold design

The same model, the same rows and the same target, scored under four fold designs:

threshold 0.50 Random 10-fold 10 folds, space ignored R² 0.653200 km blocks 108 blocks R² 0.577Leave-one-province-out 38 provinces · published R² 0.395Ridge, same folds linear baseline R² 0.185 Identical model, identical rows, identical target — only the fold changes.
Figure 5. Held-out R² under four cross-validation designs. Only the assignment of units to folds differs between the top three bars.

A random 10-fold reports R² 0.653 and ρ 0.70 — it would clear the R² bar outright and miss the rank bar by 0.001. Blocking at 200 km gives 0.577. Holding out whole provinces gives 0.395. 0.257 of R² is pure fold design, on rows that never changed.

This is the documented consequence of spatial autocorrelation, and its size has been measured elsewhere. Roberts et al. (2017) established that non-spatial cross-validation systematically understates predictive error where neighbours resemble each other. Ploton et al. (2020) put a number on it.

random fold space held out — one model, one dataset, both designs this case Ploton et al. 2020 forest biomass, Central Africa 0.53 random fold 0.14 space held out −74% of the explained varianceEngstrom et al. 2022 poverty headcount, Sri Lanka 0.61 random fold 0.45 space held out −26% of the explained varianceThis case poverty headcount, Indonesia 0.65 random fold 0.40 space held out −39% of the explained variance
Figure 6. Studies that scored one model on one dataset under both a random fold and a spatial hold-out. Ploton et al. map forest biomass; Engstrom et al. predict a poverty headcount in Sri Lanka. This case's penalty is the smallest of the three.

Ploton et al. lose 0.53 → 0.14 moving from random 10-fold to spatial 44-fold — 74% of the explained variance, after which their spatially-validated error is barely better than a null model. Engstrom et al. lose 0.61 → 0.45 on a poverty headcount under leave-one-Divisional-Secretariat-out. This case loses 0.653 → 0.395, or 39%. Its spatial penalty is the smallest of the three.

Putri et al. reported no hold-out at all. Sartirano et al. reported none. This case's random-fold 0.653 is the like-for-like comparison with the studies that set its thresholds — and it sits squarely among them.

There is one more comparison, and it is the most important in this article. Engstrom et al. (2022) is, as far as we can find, the only published satellite model of a monetary poverty headcount validated with space held out. Its result is R² 0.45 for a linear model and 0.564 for a random forest, with Spearman ρ 0.70, over 1,291 village-sized units in one small country — and against a ground truth that is a census imputation averaged over a hundred simulations rather than a raw survey estimate. Against the raw survey estimate, the same model scores 0.217.

This case reports R² 0.395 and ρ 0.55 over 514 regencies spanning 5,000 km of archipelago, against an unsmoothed published survey statistic. It is below Engstrom's spatial number, and above the number Engstrom's model achieves when scored against a survey rather than a census imputation. It is, in other words, roughly where the state of the art for this exact problem sits.

So the honest statement is not that the thresholds came from the wrong target. It is that they came from the wrong validation design — and that the case then chose, correctly, to be scored under a harder one than any of its benchmarks used. That decision is the most defensible thing on the page, and it is currently being explained away as somebody else's measurement error.

6 Finding three — what 110 million rooftops actually buy

A sceptic's first objection to any satellite poverty model is that it has rediscovered population density. The case's own answer is its attribution chart, on which land cover accounts for 49% of the model's gain. But gain is an in-sample statistic: it records how often a feature was used to split, not whether those splits transfer to a province the model has never seen. We refitted on single feature families under the identical 38 province folds and asked the out-of-sample question instead.

0 R² — do the levels come out right? ρ — is the order right? National mean (null) 0 features −0.075 Population only 3 features −0.028 0.26Night lights only 6 features 0.394 0.42Buildings only 11 features 0.049 0.49Land cover only 9 features 0.304 0.43Population + lights 9 features 0.412 0.42Population + lights + buildings 20 features 0.277 0.52All 31 features (published) 31 features 0.395 0.55 Leave-one-province-out, 2025 bars left of the line are worse than predicting one number for the whole country
Figure 7. Held-out skill for each feature family alone and for the published stack, in two currencies: R² on the left (are the levels right?) and Spearman ρ on the right (is the order right?). The two disagree, and the disagreement is the finding.

The sceptic is refuted on the narrow point: population alone scores -0.028, worse than predicting one number for the entire country (-0.075). But the wider result is uncomfortable. Night lights alone reach R² 0.394 against the full thirty-one-feature stack's 0.395. Night lights and population together reach 0.412 — higher than the published model, on nine features instead of thirty-one. Land cover alone gives 0.304. Buildings alone give 0.049, and adding buildings to lights and population lowers R² from 0.412 to 0.277.

In level accuracy, the entire Open Buildings layer — 110 million footprints, 315 partitions, 14 GB streamed through the box — is worth nothing at all.

The right-hand column rescues it. In rank terms roofs are the strongest single family: ρ 0.49 alone against night lights' 0.42, and adding them to lights and population lifts ρ from 0.42 to 0.52, with the full stack at 0.55 — the best rank score of any specification tested. Roof geometry tells you which regency in a group is poorer. It does not tell you by how many points.

That distinction happens to be exactly the right one for this product. A benchmarked disaggregator takes the level from BPS and needs only the ordering (§10). So the buildings layer earns its place — but not for the reason the page gives, and not in the currency the failing check is measured in. The attribution chart, read as evidence of predictive contribution, is misleading and should say what it is.

7 Finding four — hand the model the line it says it is missing

The case's defence makes a sharp, falsifiable prediction. If the missing information is the nominal poverty line, then supplying that line should restore the missing skill. BPS publishes it per regency. We added it — as log rupiah, at province level and at regency level — under the identical folds.

This is circular as a shipped feature: the line is computed from the same Susenas the headcount comes from, and the case is right to refuse it in production. As a diagnostic it is not circular at all. It is the experiment that separates "the model is missing the price level" from "the model is missing the poor."

threshold 0.50 Published model ρ 0.55 · RMSE 5.4 pp R² 0.395+ province mean poverty line ρ 0.56 · RMSE 5.3 pp R² 0.403+ regency poverty line ρ 0.56 · RMSE 5.3 pp R² 0.409 total movement +0.014 R², against the 0.105 needed to clear the bar diagnostic only the line comes from the survey being predicted, so it could never ship as a feature
Figure 8. The published model, then the same model with the official poverty line supplied as a feature, under identical leave-one-province-out folds.

Adding the province's own poverty line moves R² from 0.395 to 0.403. Adding the regency's own line gives 0.409. The total movement is +0.014, against the 0.105 the model needs to clear its threshold. Spearman moves 0.55 → 0.56; RMSE 5.4 → 5.3 pp.

Give the model, for free, the one thing it says it is missing, and 14% of the shortfall closes. The nominal poverty line is not the missing information.

-12 -10 -8 -6 -4 -2 0 +2 +4 +6 +8 +10421k684k946k R² 0.014 · 38 provinces official provincial poverty line · rupiah per person per month → province offset · pp →
Figure 9. Each province's constant prediction offset — the quantity the case says the poverty line sets — against that province's official poverty line. Hover any dot for the province, its offset and its internal rank correlation.

The offsets range from -11.9 to +9.7 pp and the poverty line varies 2.12-fold across provinces, from Rp 434,076 to Rp 918,778 per person per month. There is genuine variation on both axes. If the line set the offset, the two would move together. They do not: R² 0.014 in levels, 0.008 in logs — and the weak relationship that exists runs the wrong way, with higher lines attached to more negative offsets.

The offset is far better explained by the province's own poverty level: R² 0.357, slope -0.35. That is the signature of compression, the same fact Figure 3 showed at regency level, not of an unobserved price index. The model over-calls poverty where there is little and under-calls it where there is a lot, and provinces differ in how much there is.

There is also a structural reason the line cannot produce a province-level offset. BPS computes a poverty line for every regency, and separately for urban and rural, so it varies inside a province as well as between them, and on the case's own BPS pull for 2025 it varies inside them a great deal. The median province spans 1.62× internally; 27 of 38 span at least 1.5×. The widest is Jawa Barat, where Garut sits at Rp 407,191 and Kota Depok at Rp 884,633 — a 2.17× range inside a single leave-one-out fold, wider than the 2.12× range across all 38 province means.

A quantity whose median variation within a fold is larger than the variation between folds cannot be what makes each fold's error a constant.

The named cases make the mechanism concrete, and none of them is about prices. Papua has the second-highest poverty line in the country (Rp 835,262) and the model still under-calls it by 11.9 pp — official 24.7%, predicted 12.8%. Aceh, on an ordinary line of Rp 591,016, is over-called by 9.7 pp. Bali is over-called by 9.2 pp — the model reads 13.4% against an official 4.2% — because a tourism economy is invisible in roof geometry and radiance. What the model is missing in these places is not a price level. It is where the income comes from.

8 Finding five — a satellite can see the price level

The strongest form of the case's claim is a physical one: "no satellite can see a price level." It is cheap to check. We kept the 31 features and the 38 province folds and changed only the target, to the official poverty line itself.

The poverty line itself log rupiah · MAPE 13.8% R² 0.378 · ρ 0.62The poverty headcount percentage points · published R² 0.395 · ρ 0.55 the only change Same 31 features, same 38 province folds, same year — only the target moves. A held-out province's cost of living comes out within 13.8% of its true value.
Figure 10. The same features and folds, two targets. The upper bar is the nominal poverty line — the quantity the case says is invisible from orbit.

A model that has never seen a province predicts that province's official cost-of-living line at R² 0.378, ρ 0.62, to within 13.8% of its true value in rupiah. Poverty lines are not mysterious quantities: they track urbanisation, market access and settlement density, which is precisely what roofs, radiance and built-up land measure.

The price level is not a blind spot. This model sees it about as well as it sees the headcount in R² (0.378 against 0.395) and distinctly better in rank (ρ 0.62 against 0.55).

Taken with §7 the argument closes from both ends. The model can already infer the line roughly as well as it infers anything; and when the true line is handed to it outright, the headcount barely improves. Both halves of "the satellite cannot see the line, and that is why the headcount fails" are false on the case's own data.

So why is a headcount harder than an asset index? §2 already gave the answer the case reaches past. An asset index is a smooth function of visible stock — roofs, walls, vehicles, electrification. A headcount is a threshold crossing of a distribution. Two regencies with identical roofs and identical median consumption differ by ten points of headcount if one has more households bunched just below the line. Satellites read the central tendency of a place; a headcount is a property of its lower tail, and no roof reveals a tail.

The literature says exactly this, in numbers. Jean et al. find their r² collapses to about 0.12 once the sample is restricted to households below twice the international poverty line — the satellite loses almost all its skill precisely among the poor. Steele et al. find the same features that give r² 0.76 on a wealth index give 0.00 on consumption poverty in Bangladesh's urban subset — which is the same shape as this case's kota R² of -0.60 (§12). Yeh et al.'s own asset index correlates with log consumption at only r² 0.50, so even a perfect asset model is capped near there against anything monetary.

A tail statistic, not a price index. That limit is real, specific, citable, and it survives every test in this article — which is more than can be said for the one currently published.

9 Finding six — remove the line entirely and the model is still short

There is one setting where the nominal-line problem simply does not exist: inside a single province, where the line is essentially constant. The case computes this as its own defence: strip the province offsets and regencies are ordered at ρ 0.50. It reports that number as "real signal."

It is real. It is also, on the case's own decomposition, worth R² 0.077: once each province's level is removed, the model explains 7.7% of what remains. The page quotes the rank correlation and not the R², and the two tell very different stories about how much resolution is being added.

The comparison that settles it is direct, because Putri et al. (2022) worked in exactly this setting: one Indonesian province, East Java, 38 regencies, the same BPS headcount, one nominal line.

This case · East Java, 38 regencies, held out entirely (2025) Putri et al. 2022 · East Java, 38 regencies, no hold-out (2020) Spearman ρ 0.62 0.77Pearson r 0.61 0.71RMSE, pp — lower is better 3.38 3.18 Correlation bars are scaled 0–1; the RMSE pair is scaled to 8 pp, where a shorter bar is the better result.
Figure 11. This case's held-out performance on East Java against Putri et al. (2022) on the same province and the same target. Their design is more favourable — no hold-out, and they add point-of-interest density and SO₂ that this case does not ingest — so this is an upper bound on them and a lower bound on us.

Held out entirely, this model orders East Java's 38 regencies at ρ 0.62 and r 0.61, against their 0.77 and 0.71. On error the two are close: RMSE 3.38 pp here against their 3.18 pp — and ours is out of sample while theirs is not.

That is the fair summary. This model is a little behind an in-sample Indonesian benchmark on ordering and level with it on error, in the one province where the nominal poverty line cannot possibly be the explanation for anything. Part of the ordering gap is the hold-out. Part of it is that Putri et al. use features this case chose not to ingest: the case's own documentation lists OSM road density, GHSL built-up surface and Sentinel-2 indices as specified but not implemented. Neither part is the price level, because inside one province there is only one price level.

10 Finding seven — the product, tested where it can be tested

None of the above is quite the question a client asks. The case does not ship a regency model; it ships a disaggregator. The official regency rate is taken as given and the model decides only how it is distributed among that regency's kecamatan. There is no ground truth below the regency, which is exactly why the product exists — and exactly why it cannot be validated where it operates.

But it can be validated one level up. Take the province's official rate as given, let the held-out model distribute it among that province's regencies by the same benchmarking rule the case uses, and score the result against the official regency rates. The null is the honest alternative any agency already has: give every regency its province's rate.

mean absolute error Give every regency its province's rate the null this product must beat 2.69 ppModel, benchmarked to the province the shipped procedure, one level up 2.49 ppModel, unbenchmarked for reference 3.94 pp placed closer than the flat rate 54.1% of 514 regencies coin toss
Figure 12. The shipped procedure, run one administrative level up. Mean absolute error against the official 2025 regency rates for 514 regencies, and the share the benchmarked model places closer than the flat provincial rate does.

The benchmarked model cuts mean absolute error from 2.69 pp to 2.49 pp, a reduction of 0.20 pp, and places 54.1% of regencies closer to the truth than the flat rate does. Benchmarking matters: unbenchmarked, the same predictions give MAE 3.94 pp.

This is the case's strongest result and it is not currently on the page. The disaggregator beats the alternative an agency actually has, by a measurable margin, in the only setting where the claim can be checked.

Read the margin honestly. Winning 54.1% of the time means losing 45.9% of the time, and a 2.49 pp typical error is large next to the differences that decide an allocation. The result licenses exploration and prioritisation. It does not license household targeting, and the analogous step down from regency to kecamatan is a longer one than province to regency, over units with thinner feature coverage — 372 of 7,069 kecamatan carry no building footprint at all and 106 are dark in every year.

11 Finding eight — the baseline nobody escapes

At the regency level the model competes with something it can never beat: the number BPS published last year.

0.00 0.25 0.50 0.75 1.00 the satellite model · R² 0.395 0.978 1 yr 0.958 2 yr 0.916 5 yr 0.755 9 yr how old the published rate you fall back on is · R² against 2025 →
Figure 13. R² of an older published regency rate as a predictor of the 2025 rate, by how old it is, against the satellite model's held-out R².

Last year's official rate predicts this year's at R² 0.978 with RMSE 1.04 pp. Even the 9-year-old rate manages 0.755. The satellite model's 0.395 is not in the same contest. Poverty rates are extremely persistent, and the year-to-year movement that remains is close to survey noise: residual scatter about each regency's own smooth trend is 0.41 pp, against the model's 5.4 pp error.

This is why the case's temporal hold-out reads the way it does. Trained on ≤ 2023 and asked for 2025, it reports ρ 0.98 — but the same regencies are in its training set at earlier years and two feature families never change, so it is largely reciting persistence. Hold the province out as well and ρ falls to 0.54. The case publishes both, which is right; the strict number is the only one that means anything.

12 Finding nine — who the model fails, and by how much

R² — against each group's own mean RMSE — in percentage points 0 0.35 = "indicative" Java n = 119 0.36 2.8off-Java n = 395 0.38 5.9kabupaten n = 416 0.37 5.7kota (cities) n = 98 −0.60 3.8 read both Equal R² is not equal error: off-Java error is 2.1× Java's in percentage points.
Figure 14. Held-out R² and RMSE by subgroup. R² is computed against each subgroup's own mean, so equal R² across groups with unequal spread does not mean equal error.

The Java / off-Java disclosure the case promises is made, and it looks reassuring: R² 0.36 on Java against 0.38 off it. But R² is normalised by each group's own variance, and the groups do not have the same variance. In percentage points the off-Java error is 5.9 pp against Java's 2.8 — 2.1 times larger — with a bias of +0.75 pp against Java's -0.30. The model systematically under-calls poverty outside Java, which is where Indonesia's poverty actually concentrates. Reporting the split in R² alone understates that, and the page should carry the percentage points beside it.

Cities are worse and the case says so: kota R² is -0.60, below zero, so the model does worse than the national mean, and every kota estimate is labelled indicative. That disclosure is correct and unusually candid. It is also a structural result rather than a blemish: cities have low, tightly-clustered rates (3.8 pp RMSE on a narrow spread), and roof geometry cannot separate an 8% city from a 5% city.

The sharpest gradient in the whole dataset, though, is not administrative. It is how much of a place the satellite can actually resolve.

0 2 4 6 6.97 1.6 n = 103 3.59 20 n = 103 3.85 56 n = 102 2.99 325 n = 103 2.30 1126 n = 103 mean abs. error · pp → median roofs per km², by quintile of the 514 regencies →
Figure 15. Held-out mean absolute error by quintile of building density. The emptiest fifth of the country has a median of 1.6 roofs per km²; the densest has 1126.

Error runs from 2.30 pp in the densest-built quintile to 6.97 pp in the emptiest — a factor of 3.0. Those 103 regencies are the interior of Kalimantan, Papua and eastern Nusa Tenggara: the places with the highest poverty rates, the thinnest feature coverage, and the least official data. This is a coverage failure with a clear remedy — more and better features for sparsely-built terrain — and it is a far more actionable diagnosis than a nominal poverty line, which admits none.

13 What follows for decisions

Evidence is only worth gathering if it changes an action. Indonesia has specific institutions that would act on a sub-regency poverty surface, and they have different tolerances.

  1. Exploration and prioritisation, inside a benchmarked regency. A provincial planning agency choosing which kecamatan to visit first, or where to commission a survey, is making a reversible decision under a flat-rate alternative that is demonstrably worse (§10). This is the defensible use and it should be the headline.
  2. Survey design. The highest-value output of a model this compressed is not the estimate but the disagreement. The regencies where the satellite and the official rate diverge most — the 46% of 514 where the model loses to the flat provincial rate are a good place to start — are a stratification variable for the next Susenas sample, and a cheaper way to allocate enumeration effort than uniform sampling.
  3. Auditing an existing allocation. Village-fund formulas already distribute money using variables of comparable precision. A satellite surface is a second opinion on those formulas — not a replacement, but a place to look for anomalies.
  4. Not: household or village targeting. The social registry that determines who receives assistance operates on households. Sartirano et al. (2023) ran precisely this experiment on Indonesia, simulating the Kartu Perlindungan Sosial allocation with the best-known satellite wealth index — one that reaches ρ 0.75 against Susenas at exactly this administrative level. It still produced an exclusion error of 32.82%, and 55.66% of the people living in genuinely eligible regencies would not have been targeted. Good rank correlation and unusable allocations are not in tension; they are the normal result. Nothing in this case's numbers suggests it would do better, and its -0.60 urban R² suggests it would do worse in the cities where registries are largest.
  5. Not: anything at regency level. BPS publishes that number and its own prior year predicts it at R² 0.978 (§11).
  6. Not: a substitute for small-area estimation where survey microdata exist. Feriyanto et al. (2024) estimate per-capita expenditure for 626 kecamatan in West Java by EBLUP on Susenas plus PODES with night lights as an auxiliary, reaching a mean relative standard error of 10.4% against BPS's own 25% acceptability threshold. That is the right instrument for this job and it is far more precise than this one. It also requires microdata this case does not have — which is the honest reason this case exists, and should be stated as such.

This also fixes the case's position on the axis FMV works along — see clearly, connect wisely, act with evidence, with the stated priority of closing the gap between accumulating data and acting on it. Presented as an Insight product, this case is a poverty map that fails its own accuracy check. Presented as a Systems & Learning product it is something better: a standing instrument that says where the official statistic is least informative, at a resolution the statistic does not reach, with its own error published beside every number. The deliverable is not the map. It is the shortlist of places where the map and the survey disagree, handed to the institution that can go and look.

14 Corrections applied, and what remains open

Five things the case page said did not survive this review, and have been changed.

  1. The threshold provenance. The page said the thresholds "were set from an asset-wealth literature." The case's own configuration file cites Putri et al. (2022), a study of the BPS poverty rate in East Java, for the R² and RMSE bars. Corrected to name the real mismatch: none of the three source studies held space out, and this one does.
  2. "No satellite can see a price level." False on the case's own data. The same features predict a held-out province's poverty line at R² 0.378 and ρ 0.62 (§8), and supplying the true line raises headcount R² by only 0.014 (§7). Withdrawn, and replaced by the tail-statistic explanation, which the evidence supports.
  3. The province offset as evidence for the line. The offsets are explained by the poverty line at R² 0.014 and by prediction compression at R² 0.357. The page now says which, and names the provinces.
  4. ρ 0.50 reported without its R². The within-province R² of 0.077 is now published beside the rank correlation, and the disaggregation test of §10 replaces both as the headline evidence that the product works.
  5. Attribution read as predictive contribution. The chapter that ranks feature families is computed from in-sample gain. Out of sample the buildings family adds rank skill and costs level skill (§6), and the chart now says so.

Four questions this review opened remain open, and we state them as work not done rather than as caveats.

  1. The missing feature families. Putri et al.'s advantage over this model in East Java is most plausibly point-of-interest density and road access. OSM road density, GHSL built-up surface and Sentinel-2 indices are specified in this case and not implemented. Figure 15 says where they would pay: the sparsely-built quintile, at 6.97 pp of error.
  2. Predict the distribution, not the crossing. If §8 is right that headcounts fail because they are a tail statistic, the fix is to predict a location parameter — the poverty gap P1, or median consumption — and derive the headcount from a fitted distribution. P1 and P2 are already ingested by this pipeline and unused by the model.
  3. Drop the dead weight, or explain it. Nine features (lights and population) score R² 0.412 against thirty-one features' 0.395. Either the extra twenty-two are kept for the rank skill they demonstrably add, and the page says so, or the model is simplified. Both are defensible; the current silence is not.
  4. An independent surface to validate the kecamatan layer. The planned check was dropped on licence grounds — Meta's RWI and the SMERU map are both CC BY-NC. Until something replaces it, §10 is the only evidence the disaggregation works, and it is evidence from one level up.

15 References and reproducibility

  1. Jean, N., Burke, M., Xie, M., Davis, W.M., Lobell, D.B. & Ermon, S. (2016). Combining satellite imagery and machine learning to predict poverty. Science 353(6301), 790–794. doi:10.1126/science.aaf7894
  2. Yeh, C., Perez, A., Driscoll, A., Azzari, G., Tang, Z., Lobell, D., Ermon, S. & Burke, M. (2020). Using publicly available satellite imagery and deep learning to understand economic well-being in Africa. Nature Communications 11, 2583. doi:10.1038/s41467-020-16185-w
  3. Chi, G., Fang, H., Chatterjee, S. & Blumenstock, J.E. (2022). Microestimates of wealth for all low- and middle-income countries. PNAS 119(3), e2113658119. doi:10.1073/pnas.2113658119
  4. Engstrom, R., Hersh, J. & Newhouse, D. (2022). Poverty from space: using high resolution satellite imagery for estimating economic well-being. World Bank Economic Review 36(2), 382–412. doi:10.1093/wber/lhab015
  5. Steele, J.E., Sundsøy, P.R., Pezzulo, C., Alegana, V.A., Bird, T.J., Blumenstock, J., Bjelland, J., Engø-Monsen, K., de Montjoye, Y.-A., Iqbal, A.M. et al. (2017). Mapping poverty using mobile phone and satellite data. Journal of the Royal Society Interface 14(127), 20160690. doi:10.1098/rsif.2016.0690
  6. Feriyanto, N., Wijayanto, A.W., Wulansari, I.Y. & Parwanto, N.B. (2024). Small area estimation approaches using satellite imageries auxiliary data for estimating per capita expenditure in West Java, Indonesia. Jurnal Aplikasi Statistika & Komputasi Statistik 16(2), 205–221. doi:10.34123/jurnalasks.v16i2.799
  7. Putri, S.R., Wijayanto, A.W. & Sakti, A.D. (2022). Developing relative spatial poverty index using integrated remote sensing and geospatial big data approach: a case study of East Java, Indonesia. ISPRS International Journal of Geo-Information 11(5), 275. doi:10.3390/ijgi11050275
  8. Sartirano, D., Kalimeri, K., Cattuto, C., Delamónica, E., Garcia-Herranz, M., Mockler, A., Paolotti, D. & Schifanella, R. (2023). Strengths and limitations of relative wealth indices derived from big data in Indonesia. Frontiers in Big Data 6, 1054156. doi:10.3389/fdata.2023.1054156
  9. Roberts, D.R. et al. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 40(8), 913–929. doi:10.1111/ecog.02881
  10. Ploton, P. et al. (2020). Spatial validation reveals poor predictive performance of large-scale ecological mapping models. Nature Communications 11, 4540. doi:10.1038/s41467-020-18321-y
  11. BPS-Statistics Indonesia. Percentage of poor population (P0), poverty gap (P1), poverty severity (P2) and the poverty line by kabupaten/kota. WebAPI variables 621–624, 2016–2025, March Susenas.

Data vintage 2026-08-30. Literature values in Figures 1 and 2 and in §2 are transcribed from the published papers with their DOIs above, and are the only numbers here not computed from the case's own data. Every threshold was fixed and recorded in advance of the result it judges.