Indonesia → Singapore · live

The smoke that closed Singapore's schools — where was that air three days earlier, and what was burning under it?

Indonesia's haze is not weather. It is a land-use decision plus a wind, and the two halves are separable: where a fire is likely to start is a function of drought, peat and fuel; where the smoke goes is a function of the flow aloft. NASA already does detection extremely well. The gap is the two or three days before — and the attribution afterwards.

2,678,059
Retained detections
VIIRS SNPP, 2012 →, after the static-source mask
10.7%
Removed from the live tail
where a type filter removes exactly zero
0.875
AUC · 1-day lead, forecast path
threshold 0.80, fixed before any result
0.53
Spearman ρ · Singapore
modelled exposure vs NEA instruments

Blame the wind

Pick a receptor and a bad-air day. Seventy-two hours of back-trajectories fan out across the archipelago and land on the fires that were burning where the air came from, with the province attribution assembling beside them. Then flip the dial: the same integrator runs forward, from that day's hotspots to the cities downwind.

Direction
Receptor Episode
2-D canvas · no WebGL
parcel path hotspot receptor

Attribution

Pick an episode.

This is a kinematic trajectory model, not a chemistry-transport model. A single trajectory is a line, not a plume; the spread of the ensemble is a proxy for dispersion, not a simulation of it. Release heights come from CAMS GFAS where GFAS exists and from a parameterised plume rise where it does not — and in this build the fallback carries 96.3% of parcels, because GFAS has no injection height before 2019. The share is printed in the methodology and the review measures what it costs.

How far back this can be read. Across all 9,594 receptor-days in the archive the ensemble's half-width is 24 km at six hours, 82 km at a day, and 177 km at 72 hours, by which point only 55% of parcels are still inside the domain. A province-level answer is inside the instrument's resolution at a day and at the edge of it at three; the top-ranked province changes on 27% of episodes if the window is 24 hours instead of 72. The trajectory says where the air was, not who lit the fire, and the review article puts numbers on how far that can be pushed.

Chapter 01

What is actually burning

Roughly one detection in ten inside the Indonesian box is not a landscape fire. Indonesia has ~130 active volcanoes and a working oil-and-gas industry, and both radiate at the wavelengths VIIRS watches. Taken at face value, the hotspot table reports Merapi, Dukono and the Duri flares as fires every single day of the year — a permanent, perfectly seasonal-looking background a risk model will happily learn and then present as skill.

of the near-real-time tail is removed by the static mask — the part of the record where the FIRMS type field does not exist at all, so the obvious type == 0 filter removes exactly zero rows and the recent series silently keeps the volcanoes the historical series drops. That step change at the SP/NRT seam reads as a trend. It is the whole reason this stage exists.

Static exclusion mask · 0.01° cells

What was thrown away

The file size tells the story before a model runs

The per-country bulk CSV for Indonesia was 66.1 MB in 2015 and 8.8 MB in 2016. That is El Niño and La Niña, in bytes, with no analysis at all.

Chapter 02

Why here, why now

Ignition probability per 0.25° cell per day, from drought, soil moisture, fuel, peat depth and the ocean state — every predictor lagged to t−1 and the lag asserted, seasons blocked in cross-validation, and the two anchor years removed from training entirely.

The obvious objection, tested. Blocking by season does not stop the model memorising where fire lives: every cell is in both train and test, and seven of its features are that cell's own fire history. The review refits the whole model with the split blocked in space as well as season — cells grouped into 73 blocks of 2°, assigned at random to four folds — so that no scored cell was ever seen in training. It costs 0.0054 AUC at one day and 0.0083 at seven. The skill is in the weather and the fuel, not in the memory of the map.

Lead
Day
Ignition probability · 0.25°

Why this cell is red

Mean |SHAP| aggregated to feature families, so the answer is a sentence rather than forty column names.

A probability a ministry might act on has to mean what it says

0.2 has to burn about one time in five. Isotonic calibration is fitted on a season the model never saw, and the reliability diagram is published rather than asserted.

Chapter 03

Against the index everyone uses

Fire is rare and violently seasonal, so raw AUC flatters everyone. The baselines here already contain the easy knowledge: per-cell day-of-year climatology, persistence, and the operational CEMS Canadian Fire Weather Index — an index we did not design and cannot tune, given the fairest possible treatment by being isotonically calibrated to probability on the same folds.

Read the AUC column with the next sentence attached. Fire happens on 3.76% of cell-days, and on a rare event AUC is dominated by easy true negatives. The same predictions score an average precision of 0.292 at one day — 7.8× the base rate, real skill, honestly sized — and at an alert budget of 1% of cell-days the model is right 53 times in 100 at one day and 37 in 100 at seven. AUC and average precision also rank the two internal baselines in opposite orders, which is why the review article argues the gate metric should be the second one.

What the FWI column was scored on, and what changed. Until this review the CEMS record on disk covered only two of the three held-out seasons — 2014, 2015 and 2017 had been recorded as rejected rather than queued — so the "BSS vs CEMS FWI" column was computed on one held-out season and 133,386 cell-days while the AUC column beside it was computed on 2,142,680 across three. The cause was a defect in this pipeline's own Copernicus queue driver: a rejected job's cooling-off timestamp was refreshed on every poll, so it could never be resubmitted. That is fixed, all fifteen CEMS years are now on disk, and the comparison above is re-scored on 266,658 cell-days across two held-out seasons, with an even-handed re-scoring across all three at a 100% join. The claim survives, and it was flattered: the index's best season had been the one it was tested on, so the model's skill against it at one day falls from +0.131 to +0.108. A second, sharper caveat: the composite FWI is not the best member of its own family here — the Build-Up Index beats it at every lead, which is exactly what de Groot et al. (2007) assumed when they built Indonesia's fire-danger system around the Drought Code, and what Mortelmans et al. (2025) measure over this same box.

The honest cost of forecasting rather than hindcasting

The reanalysis path is allowed to see the weather that actually happened over the lead window. The forecast path is not. The gap between them is what it costs not to know the weather, and it is published rather than assumed away.

Chapter 04

The plume

Forward trajectories weighted by fire radiative power, released at CAMS GFAS injection heights where GFAS has them and at a parameterised plume rise where it does not — plume height is what decides whether smoke settles over Palangkaraya or joins the 850 hPa flow across the Strait.

In this build the fallback carries almost all of it. GFAS publishes a usable injection height only over the later part of the archive, so 3.7% of parcels were released at a measured height and 96.3% at the parameterised one. That is the single most consequential number in this chapter and it belongs in the first paragraph rather than the methodology — Draxler and Hess showed that releasing at 10, 200 and 750 m changes a trajectory substantially where a half-degree horizontal offset barely does. The CAMS chemistry comparison has not run either: the transport metadata records it as unavailable, and although the forecast years have since landed on disk they arrived after this stage last ran. The physics check the specification calls the real validation is therefore outstanding, not passed, and the only direction check that did run is the integrator's own self-consistency test below, which fails.

Chapter 05

Who breathes it

The fire belt is unmonitored. OpenAQ has zero PM2.5 locations in Riau and zero in all of Kalimantan — bbox and 25 km radius searches around Pekanbaru, Palangkaraya and Banjarbaru all return nothing. So every receptor carries a tier, and tier 3 substitutes a reanalysis and calls it a reanalysis on every row.

Chapter 06

2015 and 2019

Both anchor years were excluded from training entirely — not held out at scoring time, which is the kind of arrangement that survives one refactor and then quietly stops being true. They are scored blind. If the model cannot rank the two crises everyone remembers, nothing else here is worth reading.

The replay check ships red, and the result underneath it is the strongest thing on this page. The gate fails on arithmetic — a 90th-percentile rule over seven modelled seasons admits one, and two anchors cannot occupy one slot — and the threshold is not moved. But the model reproduces the observed severity ordering of all seven seasons exactly (Spearman ρ = 1.00), having never been trained on two of them, and scores the anchors blind at AUC 0.909 in 2015 and 0.904 in 2019 — higher than its own cross-validated folds. One qualification belongs beside it: the ordering is exact but the top of the distribution is compressed. Observed, 2019 exceeds 2012 by 29%; modelled, by 0.9%. Right in rank, fragile in magnitude.

Chapter 07

Explore

Every episode with a precomputed trajectory set, with its attribution and its uncertainty. Exposure is an index, not a concentration — a fire-radiative-power weighted, travel-time-decayed count of parcels reaching within 75 km of the receptor. It is comparable between days and receptors and it is not µg/m³, which is why chapter 05 scores it by rank correlation rather than by error.

Sort
Chapter 08

Methodology, and where it breaks

Validation gates

Thresholds were fixed before any result was seen and are never moved to fit an outcome. A gate that fails ships red with a diagnosis; a gate whose input has not arrived ships pending with the reason.

CHECK 1
Hotspot hygiene · hard

Zero retained detections inside the static volcano/flare mask, on every product including NRT — where the FIRMS type field does not exist, so a field-based filter passes while doing nothing. A removal share below 0.5 % also fails: that means the filter is broken, not that Indonesia has no volcanoes.

PENDING
CHECK 2
Ignition-risk skill · hard

Held-out AUC ≥ 0.80 and a Brier skill score > 0 against both day-of-year climatology and the CEMS Canadian Fire Weather Index, at 1, 3 and 7 days' lead. Reanalysis and forecast paths scored separately. Two things the threshold does not capture, both measured in the review: on a 3.8% event average precision is the honest companion to AUC (0.292 at one day), and the FWI half of this check is scored on the two held-out seasons the CEMS record covers, not all three.

PENDING
CHECK 3
Transport direction

On episode days, the back-trajectory bearing agrees with the forward run's within ±30° on ≥ 70 % of days. This gates the integrator, not the physics — the physics check is the CAMS comparison, published as divergence rather than as a score we claim to win.

PENDING
CHECK 4
Receptor correlation · hard

Spearman ρ ≥ 0.5 against Singapore NEA instruments — the only long, clean, commercially-licensed ground record in the region. Every receptor is reported including the ones that fail, each carrying its tier.

PENDING
CHECK 5
Anchor-event replay

2015 and 2019 held out of training entirely; both must land in the top decile of seasonal severity, scored blind. 2015 has no Singapore ground truth — the NEA record begins 2016-03 — so its reference is FIRMS plus CAMS EAC4, and the gate says so.

PENDING

Six premises that did not survive contact

  1. The FIRMS area window caps at five days, not ten.
  2. VIIRS does cover 2015, so no MODIS splice is needed for the anchors — on 2015-10-20 VIIRS returned 6,110 detections to MODIS's 1,354.
  3. The type field is absent from every NRT product, so a type == 0 filter silently no-ops on the live tail. The fix is to use type only to build a static mask from the archive, then filter everything with the mask.
  4. cems-fire-historical-v1 is on EWDS, not CDS, which 404s.
  5. ADS and EWDS need no new registration — one ECMWF token authenticates all three Copernicus stores; only a one-time policy acceptance differs.
  6. The ground-truth problem is not thin Jakarta coverage: OpenAQ has zero PM2.5 locations in Riau and zero in all of Kalimantan.

What this model is not

There is no chemistry, no aerosol microphysics, no secondary organic aerosol formation, no wet or dry deposition beyond a crude exponential decay, and no aerosol–radiation feedback — which in 2015 was strong enough to suppress the boundary layer and make the haze worse than the emissions alone imply. The honest claim is which fires were upwind and which receptors are downwind, at daily resolution, with a stated direction error.

The editorial decision, stated

Attribution is published at province level, with the trajectory ensemble's spread shown beside every share and an explicit “no attributable source” outcome when the air passed over no fire at all. Island level was the conservative alternative and was rejected: “Sumatra” is not an answer anyone can act on, and the commercial premise of this case is that the answer is actionable. A province share is a statement about where the air came from, not about who lit the fire.

Known holes, stated rather than papered over

Vintages

Attribution

09 · Review

We had this instrument reviewed, adversarially.

An independent read of the same data against the published literature. It refits the whole model with the split blocked in space as well as season — and finds the skill survives, which is the strongest result on this page and one this build never tested for. It also shows that AUC is the wrong gate metric for a 3.8% event; it traces "beats the FWI" back to a single held-out season, finds the queue bug that caused it, fixes it, drains the three missing CEMS years and re-scores; it shows the Build-Up Index beating the composite index the model is measured against; and it replays the attribution over all 1,284 Singapore episode days, where it points at Johor before South Sumatra — which the published dispersion-modelling literature independently supports. Every correction on this page came from it.

Read the review article →