← The instrument REVIEW ARTICLE · GROUND DATA THROUGH 2026-08-30
METHODS & VALIDATION · JABODETABEK AIR-QUALITY NOWCAST

Can Jakarta forecast air it can no longer measure?

A stress test of a 24–72 hour PM2.5 nowcast for Greater Jakarta, 2023-01-16–2026-08-30. The headline skill number survives its own baseline and fails a better one; the held-out panel shrinks from six sensors to one across the window it is measured on; and the last two instruments in the city disagreed by a factor of six before one of them went dark. The sparsity the case reports as context turns out to be the result.

Abstract

Question. A machine-learning PM2.5 nowcast for Jabodetabek reports +17.7% lower mean absolute error than persistence at 24 hours, and separately reports that only 2 of 24 registered stations still transmit. We ask whether the first number means what it appears to mean, and whether the second is being told as the limitation it is presented as or the finding it actually is.

Method. No model was refitted. Every alternative baseline is scored on exactly the rows the published model was scored on — 23,232 held-out station-hours at 24 hours — fitted on the training period alone. Uncertainty is a day-block bootstrap over 1,031 station-days. Results are read against published tables from Beijing, Trondheim, Accra and the European forecast-validation standard.

Findings. The skill curve is a property of the baseline, not the model. Hourly PM2.5 autocorrelation troughs at 12 hours (0.27) and rebounds at 24 (0.44), so persistence is worst exactly where the case reports its best skill (+31.1% at 12 h). Measured against the strongest trivial baseline at each lead, skill is flat near 11.6% from 6 to 24 hours and decays to +4.5% at 72. At 24 hours the model beats a last 24 h mean by +11.6% [+9.4%, +13.7%] — an interval lying entirely below the case's own 15% threshold, so the gate the case publishes as passed fails on a fairer ruler. 93.8% of the held-out hours fall before 2026 and the panel falls from 6 stations to 1; no station spans the window. The one station never seen in training scores +5.6% against +21.3% for the rest. At the two air-quality categories a city acts on, the forecast issued zero warnings in 267 qualifying hours. And two instruments 43 metres apart agree only to 6.85 µg/m³ — within 0.71 of the model's entire one-hour error.

Conclusion. The forecast is a competent interpolator of a measurement system that has stopped functioning. Its remaining value is not the number it predicts but the fact that nothing in Jabodetabek can now contradict it — and that is the finding worth publishing.

1 The claim under test

Greater Jakarta is a metropolitan region of roughly 32 million people whose air, on this case's own record, sits above the WHO 24-hour guideline of 15 µg/m³ on 94.3% of complete days. A forecast with a horizon is worth more than a dial, because only a horizon lets anyone move a training session, reschedule an outdoor shift, or send a school advisory the night before.

The case ships one. It reports a 24-hour mean absolute error of 16.04 µg/m³ against persistence's 19.49, an improvement of +17.7%, and it registered a threshold of 15% in advance. 2 of its 6 checks pass and the failures are published in red, which is more discipline than most dashboards manage.

This review asks three questions of it. Is +17.7% a measurement of the model or of the baseline? Is the held-out period a held-out period? And is the network finding — 2 live stations of 24 — being reported as the caveat it is written as, or the headline it deserves to be?

2 Prior art, and where this result would sit in it

The nearest published comparator is Wang & Du (2026), who run exactly this design — direct multi-horizon regression on hourly station PM2.5 with a persistence reference — over 12 Beijing stations, 2013–2017. Their Table 3 gives 24-hour persistence MAE 65.17 µg/m³ against a best model of 56.37: an improvement of 13.5%. On RMSE the best model (Elastic Net) reaches 17.6%.

This case's +17.7% on MAE would therefore be better than the best published MAE-basis figure for Beijing — produced from a panel that by the end of its own evaluation window contains one sensor. A result above the state of the art on thinner data is a reason to check the ruler, not to celebrate.

Three further results frame what follows. Murad et al. (2021) obtain 38.4% RMSE improvement at 24 hours in Trondheim against a harder baseline — "the value observed 24 hours earlier" — which is what a 24-hour reference should be. A rolling-origin re-evaluation of PM10 forecasting found gradient boosting scoring +23.1%–29.9% skill under a single chronological split and -19.2% — negative — under 47 rolling folds, of which 34 gave no positive skill at all. And Berlinghieri et al., comparing 6 operational smoke forecasts across the 2023 US fire season, found persistence had the highest precision of all of them (0.908) at identifying high-PM2.5 days, concluding that none offered substantially more than checking the morning's readings.

Europe has settled this into a standard. Vitali et al. (2023) define the forecast quality indicator as RMSE(forecast) / RMSE(persistence), required to be at or below 1.00 at 90% or more of monitoring stations — a per-station test, written that way precisely so a good average cannot hide a bad site. We apply it in §7.

3 Data, method, and what was re-run

The case. OpenAQ v3 hourly PM2.5 for the Jabodetabek box is the target; ERA5 single levels supply boundary-layer height, wind, temperature, humidity, precipitation and radiation; NASA FIRMS VIIRS supplies upwind hotspot counts in three distance rings, lagged a day, with volcanoes and gas flares masked out of both products. 49 features feed one gradient-boosting model per horizon, trained on log1p PM2.5, with 10th/90th-percentile quantile models for the interval. One time-based split at 2025-04-07 placed at the 75th percentile of observed hours: 122,644 rows before, 36,538 after.

This review. Nothing was refitted. The published predictions were joined back to the feature table on station, issue time and horizon, and five additional trivial baselines were attached to those same rows — a trailing 24-hour mean, a diurnal climatology, a seasonal climatology, a seasonal-diurnal climatology, and the meteorological standard of persisting the anomaly from the seasonal-diurnal norm. All are fitted on the training period only. Intervals are a bootstrap over whole station-days, because hourly PM2.5 is autocorrelated for days and resampling hours independently would understate them badly.

What we did not do. We did not re-run the pipeline from source, retrain any model, or refit hyperparameters, so every model number quoted here is the case's own. We did not obtain reference-grade co-located data for these specific instruments; §9 uses the network's own accidental co-location instead. Data vintage is pinned at ground observations through 2026-08-30 and fire detections through 2026-08-30.

4 Finding one — the skill curve is the baseline's diurnal phase

PM2.5 in Jakarta has a large daily cycle: on the case's own archive the mean hour runs from 34.8 µg/m³ in mid-afternoon to 52.6 before dawn, as the nocturnal boundary layer collapses and lifts. A persistence forecast at a 24-hour lead is therefore handed that entire cycle for free — it compares an hour to the same hour — and at a 12-hour lead it is handed the cycle inverted.

The autocorrelation function says this precisely. It decays from 0.85 at one hour to a trough of 0.27 at twelve, then rebounds to 0.44 at twenty-four, troughs again at 0.22, rebounds to 0.38, and so on for three full days.

0.0 0.3 0.6 0.9 24h 48h 72h 12h 36h 60h Autocorrelation of hourly PM2.5, by lag 0 10 20 Mean absolute error at that lead, µg/m³ persistence model 1122436486072 lead time / lag, hours →
Figure 1. Above: the autocorrelation of hourly PM2.5 across all stations, 72 lags, on the case's own complete hourly grid. Amber marks the diurnal echoes, slate the anti-phase troughs. Below: mean absolute error at the same leads. The model's error (cyan) rises monotonically the way a forecast's should. Persistence's (coral) does not — it peaks at 12 hours and falls back at 24, tracing the inverse of the curve above. Hover any point.

Persistence MAE at 12 hours is 22.45 µg/m³ and at 24 hours it is lower, 19.49. No forecast gets easier further out. The baseline simply got better at 24 hours because it was allowed to be in phase.

So the published skill curve — +21.5% at 6 hours, peaking at +31.1% at 12, back to +17.7% at 24 — is not describing where the model is strong. It is describing where persistence is weak. Replacing persistence at each lead with whichever trivial rule is actually best there removes the shape entirely.

0% 10% 20% 30% the case's own threshold · 15% 1h3h6h12h24h48h72h vs persistence, as published vs the best trivial rulebest trivial rule:1h · persistence3h · persistence6h · last 24 h mean12h · last 24 h mean24h · last 24 h mean48h · diurnal clim.72h · diurnal clim. forecast lead time →
Figure 2. The same model, two rulers. Coral is skill against persistence, as the case publishes it. Cyan is skill against the strongest trivial baseline available at each lead — named in the panel at right, and never the model's own machinery. The peak at 12 hours disappears; what remains is a flat band near 12% from 6 to 24 hours, decaying at 48 and 72.

The case's own dashboard already says the important half of this — that persistence "gets this cycle for free, which is exactly why it is a hard baseline to beat." It does not draw the conclusion, which is that a curve measured against a baseline with a phase cannot be read as a statement about the model.

5 Finding two — the registered gate does not survive a fairer baseline

The case's first check requires the 24-hour forecast to beat persistence by 15% on MAE. It reports +17.7% and passes. The threshold was fixed before the run and has not been moved, which is the right discipline. The question is whether the comparison it is applied to is the right comparison.

MAE, µg/m³ model's skill · 95% interval ↑ 15% gateThe model 16.04Last 24 h mean 18.15 Diurnal climatology 18.37 Persistence 19.49 Persist the anomaly 19.57 Seasonal-diurnal climatology 20.46 Seasonal climatology 21.04 0%10%20%30% Held-out 24-hour forecasts, 23,232 station-hours. Intervals: day-block bootstrap, 1,031 blocks.
Figure 3. The 24-hour horizon against every trivial baseline, on the identical 23,232 held-out station-hours. Left: mean absolute error. Right: the model's skill over each, with 95% intervals from a bootstrap over whole station-days. Amber marks the best-performing trivial rule and the case's own 15% threshold.

Persistence is not the hardest trivial baseline at this lead. A last 24 h mean — carrying forward the mean of the sensor's last 24 reported hours, which requires no model, no meteorology and no fire data — reaches MAE 18.15 µg/m³ against persistence's 19.49. Against it the model's skill is +11.6%, with a 95% interval of [+9.4%, +13.7%].

The gate fails on this ruler, and not marginally. The whole interval lies below the 15% threshold; the bootstrap puts the probability of clearing it at 0.05%. On the case's chosen baseline the result is +17.7% [+15.7%, +19.5%], which does clear the bar — but by 0.7 percentage points at the lower bound, on an interval the case does not publish. A pass this narrow, resting this heavily on the choice of comparator, should be reported as a pass with an interval and a named baseline, not as a headline percentage.

Two baselines are worth noting for what they say about the data rather than the model. The seasonal-diurnal climatology is worse (20.46) than the plain diurnal one (18.37), which it should not be if the seasonal cycle were stable and the panel were stationary — a first sign of §6. And persisting the anomaly, the meteorological standard, lands at 19.57: statistically indistinguishable from plain persistence, because at a 24-hour lead the two are nearly the same estimator.

6 Finding three — the held-out future is a panel that dissolves

The split is placed at the 75th percentile of observed hours rather than of the calendar, and the reasoning given for that is sound: splitting the calendar would have handed the test set whichever months the network happened to be dead in. But it has a consequence the case does not follow through. Because observations are concentrated where the network was alive, the "final 25%" of observed hours is not the final 25% of the timeline.

0 1k 2k 3k666445432111111stations25-0425-0625-0825-1025-1226-0226-0626-08 Depok Jakarta South Qoryah Darussalam Bogor Selatan Rusunawa Marunda West Jakarta Mayo… Jakarta Jakarta Central held-out 24-hour forecasts scored, by month
Figure 4. Every scored 24-hour forecast in the held-out window, by month and station; the numerals above each bar are the count of stations scored that month. Band heights within a month are drawn as equal shares — the count of bands and the total height are the exact figures; the split between them is not published per month.

93.8% of the 23,232 scored hours fall before 2026. The final 4 months of the record — the months the dashboard's live panel actually describes — contribute 34 scored hours between them. The panel falls from 6 stations in 2025-04 to 1 at the end, and the intersection of the first month's stations and the last month's is empty: not one station is present throughout, so there is no sub-panel on which a like-for-like comparison over the window can be made at all.

Four stations — Depok, Jakarta South, Qoryah Darussalam, Bogor Selatan — carry 84.6% of the evaluation. Every one of them has since stopped reporting. The headline skill number is therefore a weighted average over a set of instruments that no longer exists, and the weights are set by which of them happened to survive longest, not by any sampling design.

This is the condition that the rolling-origin study was written about. A single chronological split turned -19.2% skill into +23.1% in that paper's PM10 case. We cannot rerun this case's model across rolling folds without refitting it, so we do not claim the same reversal here — but the monthly error in Figure 4 moves between 13.51 and 20.02 µg/m³ as the panel changes shape, which is the variability a single split averages away.

7 Finding four — what the model learned, and what it cannot carry

The case's own driver check reports that the single strongest feature at 24 hours is station identity, ahead of the sensor's current reading and every meteorological term. The case publishes this as a failed check and reads it correctly, as "a real result about a sparse, heterogeneous network." It has a sharper consequence than that, and the panel contains an accidental experiment for it.

Bogor Selatan first reported on 2025-08-29, which is after the training cut. It contributes 4,313 scored hours — 18.6% of the evaluation — and the model has never seen it. It is, unintentionally, the only honest estimate on offer of what this forecast would do at a newly installed sensor.

TRAINING HOURS MQI · MODEL RMSE ÷ PERSISTENCE RMSE 80% BAND COVERS 1.00 — must be at or below Depok AirGradient · 5,925 h scored 7,318 0.80 67%Jakarta South AirNow · 4,906 h scored 7,693 0.78 63%Qoryah Darussalam AirGradient · 4,517 h scored 11,664 0.78 64%Bogor Selatan AirGradient · 4,313 h scored never seen 0.87 54%Rusunawa Marunda Clean Air Catalyst · 1,861 h scored 11,874 0.75 66%West Jakarta Mayor Offi… Clean Air Catalyst · 1,646 h scored 7,161 0.75 65%Jakarta Clarity · 35 h scored 7,986 2.13 23%Jakarta Central AirNow · 29 h scored 3,273 0.76 97% Europe requires MQI ≤ 1.00 at 90% of stations. This panel reaches 87.5%.
Figure 5. Every station in the held-out window. Left: hours available in training. Centre: Europe's forecast quality indicator — the model's RMSE divided by persistence's, which must sit at or below 1.00. Right: how much of the observations the model's nominal 80% prediction interval actually contains. Amber is the station the model never trained on; coral is the station that fails the indicator.

At the station it never saw, skill against persistence is +5.6% against a mean of +21.3% across the stations it trained on, and its 80% interval covers 54% against 65%. Because station identity is the model's strongest lever, most of what looks like forecasting skill is per-sensor calibration that cannot leave the sensor it was learned on. That matters directly: the operational proposal implied by any city nowcast is that new monitors get forecasts, and this is the number that would apply to them.

Applying the European criterion to the panel as a whole: 7 of 8 stations reach MQI ≤ 1.00, which is 87.5% against the 90% required. The case does not meet Europe's acceptance criterion for an operational air-quality forecast, and the one station that fails it — MQI 2.13, more than twice persistence's error — is Jakarta, the only station in Jabodetabek still reporting. §10 explains why.

8 Finding five — the forecast cannot make the calls it exists to make

A nowcast earns its keep on bad days. The case tests this at one threshold, 55.5 µg/m³, reports recall 0.489 at precision 0.534, and publishes the check as failed against its 0.50 target. That is honest, and it understates the problem, because the US air quality index has two rungs above that one and Indonesia's own daily limit sits just below it — 55 µg/m³ under PP 22/2021, Lampiran VII, within 0.5 µg/m³ of the threshold already tested.

35.5 55.5 125.5 225.5 What the air did max 338.0 What the forecast said max 99.1 solid to the 99th percentile, faint to the maximum · concentration µg/m³ → Unhealthy for sensitive groups ≥ 35.5 µg/m³ 14,303 h happened 16,081 h forecast · caught 85%Unhealthy ≥ 55.5 µg/m³ 6,084 h happened 5,582 h forecast · caught 49%Very unhealthy ≥ 125.5 µg/m³ 247 h happened 0 h forecast · caught noneHazardous ≥ 225.5 µg/m³ 20 h happened 0 h forecast · caught none Held-out 24-hour forecasts. The 55.5 µg/m³ rung is, to within 0.5 µg/m³, Indonesia's own daily limit.
Figure 6. Above: how far the observations reach against how far the forecast reaches, on the same concentration axis, with the US AQI categories behind them. Solid to the 99th percentile, faint to the maximum. Below: at each category, the hours that happened and the hours the forecast called.

Across 23,232 held-out hours the model's highest 24-hour prediction is 99.1 µg/m³. The observed maximum is 338.0. There were 247 hours at or above the "very unhealthy" line and 20 at or above "hazardous"; the forecast issued 0 and 0 respectively. Recall at both is exactly zero, not by bad luck but by construction: a model trained on log1p and optimised for absolute error shrinks towards the middle of its conditional distribution, and this one has drawn a ceiling at roughly 99.1 µg/m³ that no input combination crosses.

The same shrinkage shows up as sign-flipped bias. On the 6,084 episode hours the model runs -23.0 µg/m³ low; on the 17,148 clean hours it runs +7.7 high. It is systematically optimistic exactly when being optimistic costs something.

This is not peculiar to this implementation. Wang & Du report that on their extreme-concentration subset, persistence beat every machine-learning model they tested at 6 and 24 hours. Table 4: on rows above 242 µg/m³ persistence beats every machine-learning model at h=6 and h=24. A pooled error metric averages the episode failure into a large majority of ordinary hours and reports the average as skill.

The 80% prediction interval is the other half of this. It covers 62.8% of observations against a 72–88% requirement, and the case publishes the failure. Murad et al. document precisely this failure mode for precisely this estimator: quantile-regression gradient boosting delivering 61% coverage on a nominal 90% interval for PM2.5, where Bayesian neural networks on the same data reached 90%–99%. The band is not irreparable; it is the wrong estimator for the job.

9 Finding six — the error floor belongs to the instrument, not the model

Every error figure above is measured against readings from sensors, and OpenAQ pools regulatory-grade monitors with consumer hardware under one field name. In this bbox that distinction is large and it is measurable.

52.8µg/m³ mean across 5 consumer-grade stations (40,690 h)
38.0µg/m³ mean across the 2 regulatory monitors (16,021 h)
+39%what the consumer units read above the regulatory ones

The case's headline air-quality figure — a mean of 44.7 µg/m³ — is a station-hour average over that mixed panel. Restricted to the regulatory-grade subset it is 38.0, which lands inside the published range for the city: IQAir reports Jakarta at 37.3 for 2023 and 41.7 for 2024, and a peer-reviewed Jabodetabek study reports 42.5. The published number is roughly 18% above its own reference-grade subset, and the difference is instrument class, not geography. (We checked the obvious alternative explanation and it is not that: weighting stations equally rather than by hour moves the mean only from 44.7 to 44.3.)

Two of these instruments happen to sit at the same address.

0 0 50 50 100 100 150 150 Jakarta · Clarity, µg/m³ → Rusunawa Marunda · Clean Air Catalyst → 7,297 hours both reported separation 43 m MAE 6.85 · RMSE 10.91 bias +1.83 · r 0.931 2024-01-25 → 2025-06-28 Mean absolute error, µg/m³ these two, same hour 6.85the model, 1 h ahead 7.57the model, 24 h ahead 16.04raw vs reference, Accra 13.68
Figure 7. Jakarta (Clarity) against Rusunawa Marunda (Clean Air Catalyst), 43 metres apart, every hour both reported. Dashed line is exact agreement. Inset compares that disagreement with the model's own forecast error and with the published uncorrected error of this sensor model against a reference instrument in a comparable tropical climate. Hover any point.

Over 7,297 shared hours the two disagree by 6.85 µg/m³ on average (RMSE 10.91, correlation 0.931). That is the observing system's own error at zero distance and zero lead time, and no forecast can be scored below it. The model's entire one-hour error is 7.57 µg/m³ — only 0.71 above the floor. At 24 hours, 10.91²⁄22.43² of the mean squared error — about 24% — is target noise rather than anything a model could predict.

The literature makes the floor worse, not better. Raheja et al. (2023) co-located this exact sensor family against a Teledyne T640 reference in Accra, a tropical high-humidity city, and found uncorrected MAE of 13.7 µg/m³ (R² 0.69, slope 1.8, over 48–89.9% relative humidity), falling to 2.3 after calibration. The same hardware in semi-arid Lubbock gives 3.4: the 4.0× degradation is the humidity regime, and Jakarta's regime is Accra's. Uncorrected, that error is 85% of the model's entire 24-hour forecast error.

None of this makes the model bad. It makes the reported error only weakly attributable to the model. Before the next tenth of a µg/m³ of forecast skill is worth chasing, the sensors need a correction function — and a correction function needs a reference monitor, which brings us to the real finding.

10 Finding seven — the last witness, and why sparsity is the result

Start with the denominator. The case reports 24 PM2.5 stations in the box, 2 of them live. The registry is noisier than that. Those 24 registrations resolve to 16 distinct addresses — 6 of them are the same coordinate carrying "Cilandek" and "Cilandek" and four more names, none of which ever returned a reading. 13 of the 24 have never delivered an hour of data, including two legacy diplomatic-post entries duplicating feeds that are separately registered. The honest count is 11 instruments that ever reported, of which two are flagged live and one of those has delivered no hours at all.

Told as "2 of 24", this reads as an access problem. Told as "11 instruments ever worked and 1 is reporting now", it is a capability problem — and told properly, it is not about measurement at all.

The two co-located instruments of §9 were tracking each other closely — monthly means within a few µg/m³, both catching the December wet-season drop together — for fourteen months. Then they stopped.

0 20 40 60 they stop agreeing24-0124-0424-0724-1025-0125-043h7h8h21h hours both reported Jakarta still reporting Rusunawa Marunda silent since 2025-06-28 last shared month: 9.9 vs 63.0 µg/m³ monthly mean over the hours both instruments reported
Figure 8. Monthly mean PM2.5 at the two co-located instruments, computed only on the hours both reported, so the comparison is never confounded by different sampling. The shaded region is where they cease to agree; the small figures under it are how many shared hours each of those months rests on. Rusunawa Marunda went silent on 2025-06-28; Jakarta is the station the dashboard's live panel now describes.

From 2025-03 the two diverge and never reconverge. In the last month both were alive, on the hours both reported, one read 9.9 and the other 63.0 µg/m³ — a factor of 6.4, at 43 metres. Then the second instrument went dark. It is the surviving one that reads low, and its held-out quality indicator is 2.13 — the worst in the panel, and the reason the network as a whole fails the European criterion.

How much this rests on. The 14 months of agreement are built on 7,258 shared hours; the 4 months of divergence on only 39, because the surviving instrument's own reporting collapsed at the same time. That is thin, and we say so. But the divergence is not a sampling artefact: the two readings are drawn from identical hours by construction, the gap is 53.1 µg/m³ against a pooled agreement of 6.85, and it holds in the same direction in every one of those months. It is enough to say the surviving instrument needs checking; it is not enough to say which of the two was right — and there is no longer anything in Jabodetabek that could settle it.

We cannot say which of the two is right. That is the point. A network of one cannot be wrong, because nothing can contradict it. The value of a monitoring network is not that it measures more places; it is that instruments check each other, and a single number that no other number can dispute is not evidence. What Jabodetabek has lost is not coverage. It is falsifiability.

For scale: the European framework directive requires 1 PM2.5 urban-background sampling point per million inhabitants summed over agglomerations (2008/50/EC, Annex V §B), which for 32 million is roughly 32. The US federal minimum is flatter and does not scale with size — 3 sites for any metropolitan area above a million (40 CFR 58, Appendix D, Table D-5) — and Jabodetabek does not currently meet that either: the case's own coverage check finds 0 stations at 80% completeness over the trailing 90 days, against a peak of 6 reached in 2024.

0 3 6 9 US minimum · 3 training cut 23·0123·0724·0124·0725·0125·0726·0126·07 reporting at all ≥50% complete ≥80% complete stations returning PM2.5 in each calendar month
Figure 9. The observing system across the whole record. Cyan is stations returning any PM2.5 that month; teal and amber are those clearing 50% and 80% hourly completeness. The coral rule is the 3-site US federal minimum for any metropolitan area above a million people (40 CFR 58, Appendix D, Table D-5); the European per-million rule would put the requirement near 32. The network peaks at 9 in 2024 and ends at 1. It does not decline for a reason anyone published; it simply stops.

11 What follows for decisions

Evidence is only worth gathering if it changes an action. Read strictly, this work supports three uses and forbids a fourth.

  1. An instrumentation case, addressed to the people who fund instruments. The arithmetic in §10 is the product: 11 instruments ever worked, one reports now, and the last independent check available said the survivor and its neighbour differed by a factor of 6.4. For the Ministry of Environment and the DKI Jakarta environment agency, that converts an abstract "we need more monitors" into a costed, falsifiable claim with a European benchmark (32 points) and a US floor (3) attached to it. A reference-grade monitor is the cheapest thing on this list and it unlocks everything else.
  2. Sensor audit and correction. The co-location result is a working method, not a finding to file. Two instruments at one address, 7,297 shared hours, and the divergence is dated to the month. Run deliberately against one reference monitor, the same procedure yields the correction functions the consumer network needs — and the published Accra result says correction takes this hardware from 13.7 to 2.3 µg/m³, which is a larger improvement than any forecasting gain in this entire review.
  3. Climatological planning, not next-day operations. The seasonal shape is robust and is the part of this that survives every test here: 28.2 µg/m³ in the wet-season trough against 55.0 at the peak, and a daily cycle worth 1.5×. For BMKG, for school-calendar planning, and for scheduling outdoor work, a defensible seasonal and diurnal climatology is available now and needs no model.
  4. Not: next-day episode warnings. The forecast called 0 of 247 "very unhealthy" hours and 0 of 20 "hazardous" ones. Its interval covers 63% where it claims 80%. Its skill over a trivial rule is +11.6% at 24 hours, on a panel that no longer exists. No school advisory or outdoor-work suspension should be triggered by it in its current form, and the page should say so in those words.

The most valuable output of this pipeline is not the forecast. It is the evidence that Greater Jakarta can no longer check its own air — produced, unavoidably, by trying to forecast it.

12 What remains open

Three things this review could not settle, stated as work not done rather than as caveats.

  1. A rolling-origin re-evaluation. Everything in §5 and §6 points the same way, but settling it requires refitting the model across expanding folds rather than one cut, which is the single highest-value next experiment and the one the PM10 study shows can reverse a sign.
  2. The fire tail is unexplained and unverifiable. 2026-08 carries 87,541 hotspots across the airshed — 42× the median month and 2.1× the previous record (2023-09, the strong El Niño season). It is geographically coherent rather than scattered, which argues against a simple processing fault, but it sits entirely in the near-real-time product, it feeds the model as a live input, and it arrives in the one month the ground network delivered 44 observed hours. We decline to adjudicate it and recommend the case does not present it as an established fire season either.
  3. Recalibrate the interval before anything else. The point forecast is defensible for climatological use; the band is not, at 62.8% coverage. The literature names both the failure mode and estimators that do not have it.
0 30k 60k02.5k5k2023202420252026 2026-08 · 87,541 42× the median month ▮ VIIRS hotspots in the airshed, per month - - ground PM2.5 hours observed the smoke signal and the capacity to check it, on the same calendar
Figure 10. Monthly VIIRS hotspot counts across the airshed (bars, left axis) against ground PM2.5 hours actually observed (cyan, right axis). The largest smoke signal in the record arrives at the moment the city has almost no capacity to measure whether it reached the ground.

13 References and reproducibility

  1. Wang, Y. & Du, R. (2026). Same-Hour Estimation and Multi-Horizon Forecasting of PM2.5 in Beijing: Model Comparison and Feature-Group Contributions. arXiv:2607.07279. Skill figures read from Table 3; the extreme-concentration result from Table 4.
  2. Murad, A., Kraemer, F.A., Bach, K. & Taylor, G. (2021). Probabilistic Deep Learning to Quantify Uncertainty in Air Quality Forecasting. Sensors 21(23), 8009. doi:10.3390/s21238009
  3. Vitali, L., Cuvelier, K., Piersanti, A., et al. (2023). A standardized methodology for the validation of air quality forecast applications (F-MQO). Geoscientific Model Development 16, 6029–6047. doi:10.5194/gmd-16-6029-2023
  4. Raheja, G., Nimo, J., Appoh, E.K.-E., et al. (2023). Low-Cost Sensor Performance Intercomparison, Correction Factor Development, and 2+ Years of Ambient PM2.5 Monitoring in Accra, Ghana. Environmental Science & Technology 57(29), 10708–10720. doi:10.1021/acs.est.2c09264
  5. Evaluation and Calibration of Clarity Node S Low-Cost Sensors in Lubbock, Texas. EGUsphere preprint egusphere-2025-4300. doi:10.5194/egusphere-2025-4300
  6. Berlinghieri, R., Burt, D.R., Giani, P., Fiore, A.M. & Broderick, T. (2026). Are Hourly PM2.5 Forecasts Sufficiently Accurate to Plan Your Day? Bulletin of the American Meteorological Society 107(3). doi:10.1175/BAMS-D-24-0291.1
  7. Rolling-Origin Validation Reverses Model Rankings in Multi-Step PM10 Forecasting: XGBoost, SARIMA, and Persistence. arXiv:2603.20315. Skill scores read from §3.1.2–3.1.4.
  8. World Health Organization (2021). WHO global air quality guidelines: particulate matter (PM2.5 and PM10), ozone, nitrogen dioxide, sulfur dioxide and carbon monoxide. Annual 5 µg/m³, 24-hour 15 µg/m³.
  9. Republic of Indonesia (2021). Peraturan Pemerintah No. 22 Tahun 2021, Lampiran VII — Baku Mutu Udara Ambien Nasional. PM2.5: 55 µg/m³ over 24 hours, 15 µg/m³ annual.
  10. Directive 2008/50/EC of the European Parliament and of the Council on ambient air quality and cleaner air for Europe, Annex V §B (recast as Directive (EU) 2024/2881, Annex III).
  11. US Environmental Protection Agency. 40 CFR Part 58, Appendix D, Table D-5 — PM2.5 minimum monitoring requirements.
  12. IQAir. World Air Quality Report, 2023 and 2024 editions. Jakarta annual mean PM2.5. Note that IQAir aggregates reference monitors and low-cost sensors together.
  13. Jabodetabek ambient PM2.5 and childhood respiratory outcomes. Annals of Global Health. doi:10.5334/aogh.4623. Exposure from a low-cost sensor network, not reference monitors.
  14. OpenAQ. Platform documentation: sensorType distinguishes reference-grade from low-cost instruments, and both are served under the same measurement field.

Data vintage: ground observations 2023-01-16 → 2026-08-30, fire detections through 2026-08-30, training cut 2025-04-07. The additional baselines are scored against the case's existing predictions, so no model was refitted to produce the comparison. Values taken from the published literature are labelled with the table they were read from; every other number here is the case's own. Every threshold was fixed and recorded in advance of the result it judges, and none has been moved since.