The smoke that closed Singapore's schools — where was that air three days earlier, and what was burning under it?
Indonesia's haze is not weather. It is a land-use decision plus a wind, and the two halves are separable: where a fire is likely to start is a function of drought, peat and fuel; where the smoke goes is a function of the flow aloft. NASA already does detection extremely well. The gap is the two or three days before — and the attribution afterwards.
type filter removes exactly zeroBlame the wind
Pick a receptor and a bad-air day. Seventy-two hours of back-trajectories fan out across the archipelago and land on the fires that were burning where the air came from, with the province attribution assembling beside them. Then flip the dial: the same integrator runs forward, from that day's hotspots to the cities downwind.
Attribution
This is a kinematic trajectory model, not a chemistry-transport model. A single trajectory is a line, not a plume; the spread of the ensemble is a proxy for dispersion, not a simulation of it. Release heights come from CAMS GFAS where GFAS exists and from a parameterised plume rise where it does not — and in this build the fallback carries 96.3% of parcels, because GFAS has no injection height before 2019. The share is printed in the methodology and the review measures what it costs.
How far back this can be read. Across all 9,594 receptor-days in the archive the ensemble's half-width is 24 km at six hours, 82 km at a day, and 177 km at 72 hours, by which point only 55% of parcels are still inside the domain. A province-level answer is inside the instrument's resolution at a day and at the edge of it at three; the top-ranked province changes on 27% of episodes if the window is 24 hours instead of 72. The trajectory says where the air was, not who lit the fire, and the review article puts numbers on how far that can be pushed.
What is actually burning
Roughly one detection in ten inside the Indonesian box is not a landscape fire. Indonesia has ~130 active volcanoes and a working oil-and-gas industry, and both radiate at the wavelengths VIIRS watches. Taken at face value, the hotspot table reports Merapi, Dukono and the Duri flares as fires every single day of the year — a permanent, perfectly seasonal-looking background a risk model will happily learn and then present as skill.
of the near-real-time tail is removed by the static mask — the part of the record where
the FIRMS type field does not exist at all, so the obvious
type == 0 filter removes exactly zero rows and the recent series silently
keeps the volcanoes the historical series drops. That step change at the SP/NRT seam
reads as a trend. It is the whole reason this stage exists.
What was thrown away
The file size tells the story before a model runs
The per-country bulk CSV for Indonesia was 66.1 MB in 2015 and 8.8 MB in 2016. That is El Niño and La Niña, in bytes, with no analysis at all.
Why here, why now
Ignition probability per 0.25° cell per day, from drought, soil moisture, fuel, peat
depth and the ocean state — every predictor lagged to t−1 and the lag
asserted, seasons blocked in cross-validation, and the two anchor years removed from
training entirely.
The obvious objection, tested. Blocking by season does not stop the model memorising where fire lives: every cell is in both train and test, and seven of its features are that cell's own fire history. The review refits the whole model with the split blocked in space as well as season — cells grouped into 73 blocks of 2°, assigned at random to four folds — so that no scored cell was ever seen in training. It costs 0.0054 AUC at one day and 0.0083 at seven. The skill is in the weather and the fuel, not in the memory of the map.
Why this cell is red
Mean |SHAP| aggregated to feature families, so the answer is a sentence rather than forty column names.
A probability a ministry might act on has to mean what it says
0.2 has to burn about one time in five. Isotonic calibration is fitted on a season the model never saw, and the reliability diagram is published rather than asserted.
Against the index everyone uses
Fire is rare and violently seasonal, so raw AUC flatters everyone. The baselines here already contain the easy knowledge: per-cell day-of-year climatology, persistence, and the operational CEMS Canadian Fire Weather Index — an index we did not design and cannot tune, given the fairest possible treatment by being isotonically calibrated to probability on the same folds.
Read the AUC column with the next sentence attached. Fire happens on 3.76% of cell-days, and on a rare event AUC is dominated by easy true negatives. The same predictions score an average precision of 0.292 at one day — 7.8× the base rate, real skill, honestly sized — and at an alert budget of 1% of cell-days the model is right 53 times in 100 at one day and 37 in 100 at seven. AUC and average precision also rank the two internal baselines in opposite orders, which is why the review article argues the gate metric should be the second one.
What the FWI column was scored on, and what changed. Until this review the CEMS record on disk covered only two of the three held-out seasons — 2014, 2015 and 2017 had been recorded as rejected rather than queued — so the "BSS vs CEMS FWI" column was computed on one held-out season and 133,386 cell-days while the AUC column beside it was computed on 2,142,680 across three. The cause was a defect in this pipeline's own Copernicus queue driver: a rejected job's cooling-off timestamp was refreshed on every poll, so it could never be resubmitted. That is fixed, all fifteen CEMS years are now on disk, and the comparison above is re-scored on 266,658 cell-days across two held-out seasons, with an even-handed re-scoring across all three at a 100% join. The claim survives, and it was flattered: the index's best season had been the one it was tested on, so the model's skill against it at one day falls from +0.131 to +0.108. A second, sharper caveat: the composite FWI is not the best member of its own family here — the Build-Up Index beats it at every lead, which is exactly what de Groot et al. (2007) assumed when they built Indonesia's fire-danger system around the Drought Code, and what Mortelmans et al. (2025) measure over this same box.
The honest cost of forecasting rather than hindcasting
The reanalysis path is allowed to see the weather that actually happened over the lead window. The forecast path is not. The gap between them is what it costs not to know the weather, and it is published rather than assumed away.
The plume
Forward trajectories weighted by fire radiative power, released at CAMS GFAS injection heights where GFAS has them and at a parameterised plume rise where it does not — plume height is what decides whether smoke settles over Palangkaraya or joins the 850 hPa flow across the Strait.
In this build the fallback carries almost all of it. GFAS publishes a usable injection height only over the later part of the archive, so 3.7% of parcels were released at a measured height and 96.3% at the parameterised one. That is the single most consequential number in this chapter and it belongs in the first paragraph rather than the methodology — Draxler and Hess showed that releasing at 10, 200 and 750 m changes a trajectory substantially where a half-degree horizontal offset barely does. The CAMS chemistry comparison has not run either: the transport metadata records it as unavailable, and although the forecast years have since landed on disk they arrived after this stage last ran. The physics check the specification calls the real validation is therefore outstanding, not passed, and the only direction check that did run is the integrator's own self-consistency test below, which fails.
Who breathes it
The fire belt is unmonitored. OpenAQ has zero PM2.5 locations in Riau and zero in all of Kalimantan — bbox and 25 km radius searches around Pekanbaru, Palangkaraya and Banjarbaru all return nothing. So every receptor carries a tier, and tier 3 substitutes a reanalysis and calls it a reanalysis on every row.
2015 and 2019
Both anchor years were excluded from training entirely — not held out at scoring time, which is the kind of arrangement that survives one refactor and then quietly stops being true. They are scored blind. If the model cannot rank the two crises everyone remembers, nothing else here is worth reading.
The replay check ships red, and the result underneath it is the strongest thing on this page. The gate fails on arithmetic — a 90th-percentile rule over seven modelled seasons admits one, and two anchors cannot occupy one slot — and the threshold is not moved. But the model reproduces the observed severity ordering of all seven seasons exactly (Spearman ρ = 1.00), having never been trained on two of them, and scores the anchors blind at AUC 0.909 in 2015 and 0.904 in 2019 — higher than its own cross-validated folds. One qualification belongs beside it: the ordering is exact but the top of the distribution is compressed. Observed, 2019 exceeds 2012 by 29%; modelled, by 0.9%. Right in rank, fragile in magnitude.
Explore
Every episode with a precomputed trajectory set, with its attribution and its uncertainty. Exposure is an index, not a concentration — a fire-radiative-power weighted, travel-time-decayed count of parcels reaching within 75 km of the receptor. It is comparable between days and receptors and it is not µg/m³, which is why chapter 05 scores it by rank correlation rather than by error.
Methodology, and where it breaks
Validation gates
Thresholds were fixed before any result was seen and are never moved to fit an outcome. A gate that fails ships red with a diagnosis; a gate whose input has not arrived ships pending with the reason.
Zero retained detections inside the static volcano/flare mask, on every product including NRT — where the FIRMS type field does not exist, so a field-based filter passes while doing nothing. A removal share below 0.5 % also fails: that means the filter is broken, not that Indonesia has no volcanoes.
Held-out AUC ≥ 0.80 and a Brier skill score > 0 against both day-of-year climatology and the CEMS Canadian Fire Weather Index, at 1, 3 and 7 days' lead. Reanalysis and forecast paths scored separately. Two things the threshold does not capture, both measured in the review: on a 3.8% event average precision is the honest companion to AUC (0.292 at one day), and the FWI half of this check is scored on the two held-out seasons the CEMS record covers, not all three.
On episode days, the back-trajectory bearing agrees with the forward run's within ±30° on ≥ 70 % of days. This gates the integrator, not the physics — the physics check is the CAMS comparison, published as divergence rather than as a score we claim to win.
Spearman ρ ≥ 0.5 against Singapore NEA instruments — the only long, clean, commercially-licensed ground record in the region. Every receptor is reported including the ones that fail, each carrying its tier.
2015 and 2019 held out of training entirely; both must land in the top decile of seasonal severity, scored blind. 2015 has no Singapore ground truth — the NEA record begins 2016-03 — so its reference is FIRMS plus CAMS EAC4, and the gate says so.
Six premises that did not survive contact
- The FIRMS
areawindow caps at five days, not ten. - VIIRS does cover 2015, so no MODIS splice is needed for the anchors — on 2015-10-20 VIIRS returned 6,110 detections to MODIS's 1,354.
- The
typefield is absent from every NRT product, so atype == 0filter silently no-ops on the live tail. The fix is to usetypeonly to build a static mask from the archive, then filter everything with the mask. cems-fire-historical-v1is on EWDS, not CDS, which 404s.- ADS and EWDS need no new registration — one ECMWF token authenticates all three Copernicus stores; only a one-time policy acceptance differs.
- The ground-truth problem is not thin Jakarta coverage: OpenAQ has zero PM2.5 locations in Riau and zero in all of Kalimantan.
What this model is not
There is no chemistry, no aerosol microphysics, no secondary organic aerosol formation, no wet or dry deposition beyond a crude exponential decay, and no aerosol–radiation feedback — which in 2015 was strong enough to suppress the boundary layer and make the haze worse than the emissions alone imply. The honest claim is which fires were upwind and which receptors are downwind, at daily resolution, with a stated direction error.
The editorial decision, stated
Attribution is published at province level, with the trajectory ensemble's spread shown beside every share and an explicit “no attributable source” outcome when the air passed over no fire at all. Island level was the conservative alternative and was rejected: “Sumatra” is not an answer anyone can act on, and the commercial premise of this case is that the answer is actionable. A province share is a statement about where the air came from, not about who lit the fire.
Known holes, stated rather than papered over
- Malaysia has no ground truth, but it is not out of the model.
apims.doe.gov.my404s on every path including root, and the only aggregator carrying it forbids commercial use verbatim, so there is no Malaysian receptor to validate against. The trajectory engine still integrates over the whole box, and across 1,284 Singapore episode days it names a Malaysian province first on 29% of them — Johor alone on 25.6%, more often than any Indonesian province. So the honest scope is Indonesia and Malaysia → Singapore, validated only at Singapore. An earlier version of this page said "Indonesia → Singapore only", which the model's own attribution contradicts. - The attribution weighting has no distance term. Province shares are residence time × fire radiative power, the standard receptor-model form. But a back-trajectory spends its first hours near the receptor whatever its origin — the ensemble is still within 24 km of the city at six hours — so any burning land close to the receptor accumulates residence time on almost every parcel and the ranking is biased toward proximity. Johor and Bangka-Belitung are the two nearest land masses to Singapore. Read the ranking as "where the air was", never as "the dominant source".
- Three CEMS FWI years were stranded by a bug, and it is now fixed. 2014, 2015 and 2017 were recorded as rejected across roughly forty submissions because the queue driver refreshed each rejected job's cooling-off timestamp on every poll, so it could never be resubmitted. Dropping the dead job id on rejection recovered all three inside fifteen minutes and the CEMS record is now complete at fifteen years. The external baseline is re-scored above.
- The CAMS chemistry comparison has still not run. The transport metadata records it as unavailable. The forecast years (2015, 2019, 2025) have since landed on disk, but they arrived after the transport stage last ran, so nothing has been compared against them yet. The physics check is outstanding rather than passed, and it is one re-run away.
- Pooling three seasons hides a threshold crossing. At seven days' lead the model scores 0.794 on the 2016 season alone — below the 0.80 the check asks for — against 0.844 on 2018. The pooled 0.822 is a mean over a range that straddles its own threshold, and the spread belongs beside the mean.
- The two halves of the case cover different years. The risk model is fitted on 2012 and 2015–2019 and 2026; the trajectory model is integrated on 2014, 2015, 2019, 2020 and 2024–2026. They share three seasons. That is the Copernicus queue delivering single-level and pressure-level years in different orders, not a design choice, but "where fire starts" and "where the smoke goes" are currently statements about largely different periods.
- 2015 has no Singapore ground truth. The NEA record begins ~2016-03, so that anchor is scored on FIRMS detections plus CAMS EAC4, and the replay check says so.
- There is no operational IOD index. The HadISST DMI runs about three months behind and is stamped “Preliminary”, so it is a historical feature only — and 2019, the strongest positive IOD on record, is unreadable without it.
- No open medium-range FWI forecast exists, so our forecast path is compared against the FWI at analysis time. That flatters us, and the size of the flattery is the reanalysis-vs-forecast gap printed in chapter 03.
- The DLR Sentinel-5P aerosol product is held back. Its collection metadata
says CC-BY-4.0 in one field while the URL in that same field and the
rel:licenselink both point at CC BY-NC 4.0. Not shipped until DLR confirms in writing. - KLHK land cover has no open licence, so ESA WorldCover v200 (CC BY 4.0) is the primary and KLHK is not stored.
Vintages
Attribution
We had this instrument reviewed, adversarially.
An independent read of the same data against the published literature. It refits the whole model with the split blocked in space as well as season — and finds the skill survives, which is the strongest result on this page and one this build never tested for. It also shows that AUC is the wrong gate metric for a 3.8% event; it traces "beats the FWI" back to a single held-out season, finds the queue bug that caused it, fixes it, drains the three missing CEMS years and re-scores; it shows the Build-Up Index beating the composite index the model is measured against; and it replays the attribution over all 1,284 Singapore episode days, where it points at Johor before South Sumatra — which the published dispersion-modelling literature independently supports. Every correction on this page came from it.
Read the review article →