Do night lights measure Indonesia's regional economy?
A replication and stress test at kabupaten level, 2018–2025. We reproduce Gibson, Olivia, Boe-Gibson & Li (2021) on newer data, confirm their result for cities almost exactly, overturn it for rural regencies — and then show why the resulting instrument still cannot be used the way it is most often wanted.
Abstract
Question. Night-lights indices are widely sold as a substitute for slow official statistics. We ask what they can actually support at Indonesia's second sub-national level.
Method. NASA Black Marble monthly composites (2018–2025) over 514 frozen-vintage regencies, matched to BPS real PDRB, with a gas-flare mask and coverage-weighted seasonal adjustment. Thresholds were registered before results.
Findings. Lights explain 67%–73% of the cross-regency variation in output, every year. But the elasticity collapses from 0.718 in the cross-section to 0.025 once place and year are absorbed, and to 0.0012 in growth, where the confidence interval contains zero. Out of sample, the annual model is 41% worse than assuming last year's growth. Running the same specification on annual composites for every year since 2012 shows the rural result is a property of the sensor product, not of the decade: on 2015–16 — the years Gibson et al. used — Black Marble gives rural R² of 0.82 and 0.78 against their 0.01. The fit is not an artefact of scale either: lights add R² +0.15 over population and area. Measured lights growth correlates r = 0.67 with how much of the country the satellite actually saw, so the growth series is partly an observation artefact. The instrument's bias is structured, not random: lights overstate Java by 10 points, and in Teluk Bintuni 92% of all measured light is a single gas flare.
Conclusion. Night lights are a location instrument, not a timing instrument. Used for allocation and targeting they are defensible; used for nowcasting regional growth they are, on this evidence, worse than doing nothing.
1 The claim under test
Indonesia publishes real GDP for all 514 kabupaten and kota, but annually and with a long lag. The appeal of night lights is obvious: a satellite passes nightly, sees every regency equally, and cannot be lobbied. The claim that follows — that radiance can stand in for regional accounts, months ahead of the statistical office — is the one this paper tests.
We separate two questions that are routinely merged. Where is economic activity? and When does it move? They are not the same question, they do not have the same answer, and conflating them is how a defensible instrument becomes an indefensible forecast.
2 Prior art, and the benchmark we replicate
Henderson, Storeygard & Weil (2012) established the method at country level, reporting that a one-percent change in luminosity accompanies roughly a 0.28-percent change in GDP, and proposing lights as a composite input where national accounts are weak. Chen & Nordhaus (2011) made the boundary condition explicit: lights add statistical value only where conventional data are poor.
The decisive prior work for our setting is Gibson, Olivia, Boe-Gibson & Li (2021), who chose Indonesia precisely because it is one of the few developing countries publishing GDP at the second sub-national level. Their verdict was blunt: using VIIRS composites for 2015–16, night lights predicted city GDP well (elasticity 0.936, R² 0.68) and rural kabupaten GDP essentially not at all (elasticity 0.086, R² 0.01). They concluded that lights are "not a suitable source of data to proxy for GDP in non-urban areas in developing countries like Indonesia."
We run their specification — log real GDP on log sum of lights, same administrative level, same country — on a different lights product and a later decade. This is a direct replication, and the two halves of it disagree.
3 Data and method
Radiance. NASA Black Marble monthly composites (VNP46A3 / VJ146A3, Román et al. 2018), 15 arc-second, QA-masked below two cloud-free observations. Black Marble applies atmospheric, terrain, vegetation and lunar BRDF corrections that the annual composites used by earlier work do not — a difference that matters most exactly where the signal is dimmest.
Boundaries. geoBoundaries ADM2, frozen at the 2020 vintage, with a crosswalk for subsequent splits. Without freezing, district pemekaran manufactures growth that never happened.
Flares. 182 persistent gas-flare sites from the VIIRS Nightfire survey, buffered at 3 km, removed before every statistic. The buffer is a measured choice, not an inherited one: 3 km captures about 84% of an isolated flare's excess light, while the 5 km often used swallows whole towns.
Output. BPS real PDRB (ADHK), annual table 2194 and quarterly table 2534, 2010–2025.
Registered in advance. The location check required R² ≥ 0.65; it passed at 0.673. The out-of-sample check required a 60% province win rate; it reached 44.1% and failed. Both are published here, unchanged.
4 Finding one — lights locate activity, and the rural verdict does not replicate
Across every year in the sample, log real PDRB regressed on log sum of lights yields R² between 0.673 and 0.732 over all 514 regencies. That is the result the instrument is built on, and it is stable.
Splitting the sample the way Gibson et al. split it produces a result in two halves. For cities, we replicate them almost exactly: our mean elasticity is 0.966 against their 0.936, with R² 0.772 against their 0.68. An independent pipeline, a different sensor product and a different decade land on the same number — which is the strongest available evidence that the pipeline is sound.
For rural kabupaten, we do not replicate them at all. Where they found an elasticity of 0.086 and R² of 0.01 — statistically indistinguishable from no relationship — we find an elasticity of 0.680 and R² of 0.686.
Two explanations are consistent with this, and they have very different implications. The first is instrumental: Black Marble's BRDF and stray-light corrections rescue a signal that the uncorrected annual composites of 2015–16 buried in noise, and dim rural regencies are precisely where that correction bites. The second is substantive: Indonesia's electrification ratio rose from roughly 88% in 2015 to near-universal by 2024, so rural light may genuinely have crossed the detection floor in the intervening decade. The next section separates them.
5 Finding two — the divergence is the sensor, not the decade
The two explanations make opposite predictions, which makes this cheap to settle. If rural light only became legible as Indonesia electrified, fit should be poor early in the series and improve steadily. If the earlier verdict was a property of the composites, fit should be high in every year — including 2015 and 2016, the exact years Gibson et al. analysed.
We ran the identical specification on NASA Black Marble annual composites for all 14 years from 2012 to 2025, holding boundaries, output data and estimator fixed.
The answer is unambiguous. Rural fit in 2015 is 0.825 and in 2016 is 0.783, against the 0.01 reported for those same years. Across the whole 2012–2025 span rural R² moves only from 0.768 to 0.833 — essentially flat, with no trend resembling an electrification ramp.
The conclusion that night lights cannot proxy rural Indonesian output is a property of the composite product it was measured on, not a property of night lights.
This matters beyond this case. Gibson et al.'s finding is widely cited as the reason not to use lights below the provincial level in developing countries. On this evidence that caution should be re-scoped: it applies to uncorrected annual composites, not to BRDF- and stray-light-corrected products, and the difference between them in rural Indonesia is the difference between R² of 0.01 and R² of 0.83.
What this test does not do. We did not re-run their pipeline — EOG's programmatic access became paid in June 2026 — so this compares our Black Marble estimates to their published table rather than to composites we processed ourselves. Two further differences remain: they used 497 units on the 2010 administrative vintage where we use 514 on the frozen 2020 vintage, and our series crosses a sensor change from Suomi-NPP to NOAA-20 in 2018 (visible in Figure 4, and undramatic). Finer units and a sensor change would both tend to depress fit, so neither explains the gap — but a full head-to-head on identically processed composites remains the cleaner test.
6 Finding three — and it is not an artefact of scale
A sceptic should still object that both sides of that regression are extensive quantities. Big regencies have more output and more light because they are big, so some of R² = 0.8 could be arithmetic rather than evidence. We put the objection to the data directly.
Population alone explains 0.693 of the variation in regency output and adding area moves it only to 0.698 — so the sceptic is partly right, and most of what a raw lights regression captures is scale. But lights are not merely a scale proxy: adding them lifts R² to 0.845, an increment of +0.147, and the conditional lights coefficient remains 0.48. Lights carry information about where output is that population and area do not.
The stronger version of the test removes area from both sides entirely. Regressing log output per km² on log lights per km² gives R² 0.905 with elasticity 0.818 — higher than the levels form, and 0.870 for rural regencies alone. Stripping scale out does not weaken the relationship; it sharpens it.
7 Finding four — the signal dies as the cross-section is removed
A cross-sectional fit answers "where", and it is flattered by scale: big regencies have more light and more output. The question that matters for any timing use is what survives once you stop comparing places to each other and start comparing a place to itself.
The decay is monotone and severe. Absorbing regency fixed effects cuts the elasticity from 0.718 to 0.108. Absorbing year effects as well — that is, asking whether a regency that brightened faster than average also grew faster than average — leaves 0.025, with an interval spanning zero. In growth rates the coefficient is 0.0012 and the within-R² is 4.8e-4.
This is not a failure of the pipeline. It is the well-documented shape of the phenomenon — Gibson et al. observed the same ordering, noting that between-R² "greatly exceed[s] the within-R²". Our contribution is to quantify how far the collapse goes on modern data: a factor of roughly 600 between the cross-sectional and the growth elasticity.
8 Finding five — the published nowcast is, numerically, a constant
If the growth coefficient is 0.0012, then a growth nowcast built on it cannot carry much information. It is worth making that concrete rather than abstract.
National lights moved 19.0% year-on-year. Passed through the calibration, that motion contributes 0.022 percentage points to a headline of 4.66%. The remaining 99.5% is the intercept — a number that would print essentially unchanged if the satellite had been switched off. The 95% band is 65 times wider than the signal inside it.
The same collapse governs the regency board. Across the movers, lights growth ranges from -43.1% to 88.2% — a 131-point spread — which the mapping compresses into 0.15 points of predicted growth. Kolaka brightened 88% and Tanah Laut darkened 43%; the model separates them by 0.15 points.
Out of sample this is not merely uninformative, it is harmful. On the held-out year, mean absolute error using lights was 0.0180 against 0.0128 for assuming last year's growth — the lights model is 41% worse than doing nothing, winning in only 23.5% of provinces. The quarterly variant does better than naive on error (0.0165 against 0.0219) but still wins in under half of provinces (44.1%). This is the out-of-sample check, and it is published failed.
9 Finding six — coverage confounds the growth series
Indonesia is equatorial and cloudy. In this sample 44.2% of regency-months carry no usable composite at all, and 62 of 103 national months are flagged low-coverage. The pipeline divides by coverage to compensate. We tested whether that works.
It does not work well enough. Across years, measured lights growth correlates r = 0.67 with mean coverage. The two years where lights most overshoot official growth are the two years coverage jumped, and the year lights show a contraction that BPS does not is the year coverage fell to its sample minimum. A clear-sky pixel is not a random pixel, so when the satellite sees less of the country it sees a systematically different country, and dividing by the observed fraction does not undo that.
This has a direct consequence for a claim the instrument's own front page makes. The page states that in 2020 "GDP fell while the lights barely blinked." On the case's own numbers that is not what happened: BPS real GDP fell 2.12% and mean measured lights fell 2.07% — they moved together almost exactly. But coverage also fell sharply that year, so the agreement cannot be read as the instrument succeeding either. The honest statement is that 2020 is not identified: signal and artefact moved in the same direction and cannot be separated. We recommend the page be corrected accordingly.
10 Finding seven — the bias is structured, and therefore usable
An instrument with random error is noisy. An instrument with structured error is correctable, and its residuals carry information. This one is the second kind.
Java holds 6.6% of the land and 57.0% of the output, but emits 66.7% of the light. The 10-point gap is the urban-industrial weighting of the sensor made numerical — the same bias Gibson et al. identified qualitatively, here with a magnitude attached to it.
The extractive tilt is sharper still. In 45 regencies a meaningful share of all measured light is gas flaring rather than economic life:
And the residuals themselves are diagnostic. 29 regencies sit beyond two standard deviations of the levels fit in the latest year. They are not scattered at random: the largest positive residuals are dense service economies where output vastly exceeds light (Kota Jakarta Pusat), and the largest negative residuals are sparsely-populated regencies where light vastly exceeds recorded output. A regency that is bright but poor, or dark but rich, is a place where the satellite and the statistical office disagree — and that disagreement is a question worth asking, not a defect to hide.
11 What follows for decisions
Evidence is only worth gathering if it changes an action. Read strictly, this instrument supports three uses and forbids a fourth.
- Allocation and targeting. Where the cross-section is stable at R² ≈ 0.7 and refreshes monthly, lights are a defensible input to spatial targeting between annual PDRB releases — which regencies look under-served relative to their economic footprint.
- Audit and triage. The residual list is the product. Persistent divergence between measured light and reported output flags where regional statistics deserve a second look. This is the highest-value use and it is currently presented as a limitation.
- Structural monitoring. Flare shares, urban expansion and electrification are measured well by lights because they are precisely what lights physically are.
- Not: regional growth nowcasting. On this evidence the growth mapping is 41% worse than a naive rule out of sample and its headline is 100% intercept. It should be withdrawn or relabelled as a research diagnostic, not shipped as a forecast.
The discipline that makes this instrument trustworthy is the same one that forbids its most marketable use. That trade is the point: an organisation that will publish its own failed gate is one whose passed gates mean something.
12 What remains open
Two of the questions this review opened have been answered in §5 and §6. Three remain, and we state them as work not done rather than as caveats.
- A true head-to-head on identically processed composites. §5 compares our estimates to a published table. The cleaner test processes both products ourselves over the same years and units, which needs EOG access that became paid in June 2026.
- Match the pixels before differencing. Compute year-on-year change only on pixels observed in both periods. This removes the composition change behind Figure 8 at its source, rather than dividing by coverage after the fact — and it is the most likely route to a growth signal that survives §7 and §8.
- Interact Ramadan with what it should depend on. The national Ramadan coefficient is 0.33% — indistinguishable from nothing — but that average conceals a spread from -21% to 20% across regencies.
That last one is not a robustness check but an unexploited result. Some Indonesian regencies brighten by a fifth during Ramadan and others darken by a fifth. Whatever explains that — night markets, mudik outflow, industrial shutdown, the local balance of household against industrial lighting — is a measurable feature of Indonesian economic life that no national statistic records, and this pipeline already computes it.
13 References and reproducibility
- Henderson, J.V., Storeygard, A. & Weil, D.N. (2012). Measuring Economic Growth from Outer Space. American Economic Review 102(2), 994–1028. doi:10.1257/aer.102.2.994
- Gibson, J., Olivia, S., Boe-Gibson, G. & Li, C. (2021). Which night lights data should we use in economics, and where? Journal of Development Economics 149, 102602. doi:10.1016/j.jdeveco.2020.102602
- Chen, X. & Nordhaus, W.D. (2011). Using luminosity data as a proxy for economic statistics. PNAS 108(21), 8589–8594. doi:10.1073/pnas.1017031108
- Román, M.O. et al. (2018). NASA's Black Marble nighttime lights product suite. Remote Sensing of Environment 210, 113–143. doi:10.1016/j.rse.2018.03.017
- Elvidge, C.D. & Zhizhin, M. (2021). Global Gas Flare Survey by Infrared Imaging, VIIRS Nightfire 2012–2019. ORNL DAAC. doi:10.3334/ORNLDAAC/1874
- BPS-Statistics Indonesia. Gross Regional Domestic Product, constant prices, kabupaten/kota. WebAPI tables 2194 (annual) and 2534 (quarterly).
Data vintage 2026-07. Every threshold was fixed and recorded in advance of the result it judges. Boundaries are frozen at the 2020 vintage under a published crosswalk.