Can an ignition model beat the Fire Weather Index over Indonesian peat?
A stress test of a 0.25° daily ignition-risk surface and a back-trajectory attribution engine. The model's skill survives the harder test its own design never ran — but the headline metric is the wrong one for a rare event, the comparison that gives the case its title rested on a single season until we found the bug that stranded the rest, and the attribution, replayed over every episode instead of one, points somewhere the page did not.
Abstract
Question. The case claims an ignition-risk surface with AUC 0.875 at one day's lead that "beats the operational Canadian Fire Weather Index", and a 72-hour back-trajectory that attributes a bad-air day in Singapore to fires in South Sumatra. We ask whether either claim survives an independent re-scoring on the case's own data.
Method. We re-derive every headline statistic from the case's published out-of-fold predictions (1,200,000 rows across three leads), audit the CEMS FWI join row by row, refit the model under a split blocked in space as well as season, and replay the trajectory ensemble across all 9,594 receptor-days rather than one.
Findings. The published AUC reproduces exactly (0.874 against 0.875), and the skill is not a spatial artefact: refitting so that no test cell was ever seen in training costs only -0.0054 AUC at one day and -0.0083 at seven. But AUC is the wrong instrument for a 3.76% event. Average precision is 0.292, and at an operational alert budget the model is right 53% of the time at one day and 37% at seven, catching 9.8% of the season's fires. AUC and average precision disagree about which baseline is stronger. The FWI comparison was thinner than it looked: the reviewed build scored it on one held-out season, 133,386 cell-days, because three CEMS seasons — including the 2015 anchor — had been stranded by a defect in the case's own Copernicus queue driver. We diagnosed and fixed it, the record drained to 15 of 15 years, and the comparison now runs on all 3 held-out seasons at a 100% join. The claim survives but had been flattered: the model's recomputed Brier skill against the index at one day falls from +0.131 to +0.108. The index's own value, 0.777, sits inside the 0.75–0.79 band Mortelmans et al. (2025) publish for the same box — but the composite FWI is not the best member of its own family here: the Build-Up Index beats it at every lead, reproducing both the 2007 design decision behind the operational Indonesian fire-danger system and the 2025 literature. On transport, the ensemble half-width reaches 177 km at 72 hours with only 55% of parcels still in the domain; the top attributed province changes on 27% of episodes if the window is 24 hours instead of 72; and across 1,284 Singapore episode days the most frequently attributed upwind province is Johor (25.6%), not South Sumatra (11.1%) — a result that is partly a proximity artefact of the weighting and partly a reproduction of Hansen et al. (2019), whose tagged dispersion modelling concluded that peninsular Malaysia is a large source for Singapore.
Conclusion. This is a ranking instrument for where to put a patrol, not a probability instrument for whether a given cell will burn, and not an attribution instrument at province level beyond about a day of travel. Read that way it is defensible and unusually well built. Read the way the page currently sells it, it promises more than the evidence carries.
1 The claims under test
The case makes two commercial promises, and they are different in kind. The first is anticipation: an ignition probability per 0.25° cell per day, one to seven days ahead, that beats the index fire services already use. The second is attribution: that when the air over a city is bad, a back-trajectory can name the province the air was standing over when the fires were burning.
Anticipation is a forecasting claim and is judged by skill scores. Attribution is a causal claim in a live diplomatic dispute and is judged by whether the answer would survive being contested. This review takes both at face value and tries to break them.
We separate three questions the page merges. Is the model good? Is the metric that says so the right metric? And is the baseline it beats the right baseline? They have different answers.
2 Prior art, and what a fair benchmark looks like
The physics. de Groot et al. (2007) built the Indonesian and Malaysian Fire Danger Rating Systems, and chose the Drought Code as their governing indicator specifically because peat burns: their operational bands — extreme above DC 350 — are still the ones METMalaysia and BMKG publish today, unchanged. They also report that 85% of hotspots in their 1995–2000 record occurred above FFMC 78. That paper is the reason §9's result is a confirmation rather than a surprise.
The benchmark, and it is higher than the case's. Two recent papers score the FWI as a fire-occurrence classifier in ways directly comparable to this case, and they disagree with each other. Mortelmans et al. (2025) score it over almost exactly this bounding box (7°S 95°E to 7.4°N 120°E), 2002–2018, and report FWI AUC 0.75 on drained peat and 0.79 on undrained, with the Drought Code alone doing better at 0.79–0.83; their conclusion is that "DC outperformed both FWIref and FWIpeat, making it the best predictor for fire occurrence." Shmuel et al. (2025), on a 0.25° daily binary ignition grid over Indonesia — this case's exact resolution and country — report the plain Canadian FWI at 0.89, well above the case's baseline. The difference is design: they balance the classes by random undersampling and hold out a random 20%, with no spatial or temporal blocking. §8 places this case's FWI value against both.
The operational verdict. Di Giuseppe et al. (2020), verifying ECMWF's operational fire forecast, single out equatorial Asia as the failure case: "the system seems to have a predictability below 0.2 (only 20% of fires corresponded to the FWI above the 90th percentile) … a fire early-warning system should mostly rely on the drought code." Anyone selling skill over the FWI in this region is pushing on a door the operational centre has already said is open.
Validation design. Roberts et al. (2017) and Ploton et al. (2020) showed that a random or merely temporal split on spatially autocorrelated data reports skill a spatially blocked split removes: Ploton's own case falls from R² 0.53 to 0.14, about three-quarters of the apparent explained variance. Fire is about as spatially clustered as an environmental process gets; Mortelmans et al. measure the FWI's own temporal autocorrelation length at 42 days; and seven of this model's features are the cell's own fire history. The prior expectation was collapse. §7 reports what actually happened, and it is not what we expected.
Metrics. Saito & Rehmsmeier (2015) is the standard reference: because both ROC axes normalise within class, the ROC curve is literally unchanged between a balanced and an imbalanced version of the same problem, while the precision–recall curve moves. Their Table 5 has one tool with the best ROC-AUC in its field (0.886) and a PR-AUC of 0.054, beaten on PR-AUC by a tool with a worse ROC (0.772 / 0.106). Davis & Goadrich (2006) give the same inversion analytically. §5 shows it happening on this case's own numbers, between its own two baselines.
Transport. Stohl (1998) remains the standard review: "errors of 20% of the distance travelled seem to be typical", rising with travel time — 16% at 24 hours, 26% at 48 and 36% at 72 in the one-year set he summarises — and his recommendation is blunt: "single trajectories are hardly sufficient to describe transport processes in the boundary layer … studies that are currently based on the interpretation of forward or back trajectories … should in the future be based on the simulation results of [Lagrangian particle dispersion models]." Engström & Magnusson (2009) add the tropical qualifier that matters most here: a back-trajectory released at 0°N, 120°E — inside this case's domain — split under ensemble analysis into "two different source regions", and their ensemble is under-dispersive in the tropics, so a tropical ensemble spread is a floor on the error, not an estimate of it.
Attribution. Koplitz et al. (2016) is the reference study for exactly this problem, using a GEOS-Chem adjoint over a whole season — and even so it disclaims the granularity this case publishes at: "since our analysis was conducted on a 0.5°×0.67° grid, there is some uncertainty in our attributions of emissions and exposure contributions to specific provinces." Hansen et al. (2019) is the decisive comparator, and §11 shows this case reproducing its most uncomfortable finding.
What this review is not. We did not re-run any external study, and we do not re-derive the case's inputs. Every number below is computed from what the case itself published.
3 Data and method
The instrument under review. 2,678,059 retained VIIRS detections after a static-source mask that removes 11.92% of the archive and 10.69% of the near-real-time tail; a 4,748,695-row panel of 1,955 land cells across 7 ERA5 seasons; gradient-boosted ignition probability at 1, 3 and 7 days' lead, isotonically calibrated; and a Petterssen second-order kinematic trajectory model producing 9,594 receptor-days of 72-hour back-trajectories.
How it validates itself. Cross-validation is blocked by season: each fold holds out one whole calendar year (2016, 2017, 2018), with a separate calibration season inside each fold, and the two anchor years (2015 and 2019) excluded from training entirely. That is already better discipline than most published fire models, and it is the reason this review could go looking for the next problem rather than the obvious one.
What we added. Four tests the case did not run. (a) Average precision and precision at an operational alert budget, alongside AUC. (b) A decomposition of the ROC into its within-day and within-cell parts, plus a deliberately stupid map-only baseline. (c) A complete refit under a split blocked in space and season: cells grouped into 73 blocks of 2°, assigned at random to 4 folds, so that every scored cell-day comes from a model that never saw that cell. (d) The trajectory ensemble replayed across every receptor-day, with the attribution recomputed under 24-, 48- and 72-hour windows.
Sample. The re-scoring in §5 and §6 uses the 1,200,000-row out-of-fold sample the pipeline retains (400,000 cell-days per lead, 15,052 positives at one day). The spatially blocked refit in §7 scores the complete out-of-fold set, 2,142,680 cell-days, so it is directly comparable with the published number rather than with the sample.
4 Finding one — the headline reproduces exactly
Before attacking a number it is worth confirming it. Re-scoring the case's own stored out-of-fold predictions independently gives AUC 0.874 at one day against the published 0.875, 0.846 against 0.848 at three, and 0.823 against 0.822 at seven — the residual difference being the sampling of the retained rows. The arithmetic is sound and the leakage that usually explains a suspiciously good fire model is not present in the obvious places: the folds really are blocked by season, the calibration really is fitted on a season the model never saw, and the anchors really are excluded from training.
That matters, because everything that follows is a criticism of what the number means rather than of whether it is real.
5 Finding two — AUC is the wrong instrument for a 3.8% event
Fire happens in 3.76% of cell-days. On a rare event the ROC curve spends almost all its length in a region no operator will ever visit, because the false positive rate is computed against an enormous pool of easy true negatives. Average precision — the area under the precision–recall curve — is computed against the positives instead, and is the metric that moves when a fire service's day changes.
The gap is not a gotcha; it is the difference between a ranking statement and a probability statement. But it becomes a gotcha when the two metrics disagree about which baseline is stronger, and here they do.
A metric that ranks your two baselines in the opposite order from the metric an operator experiences is not a neutral choice of summary statistic. It is a modelling decision, and it should be declared as one.
The practical version of the same point is what an alert budget buys. Suppose a provincial fire service can inspect the top 1% of cells on a given day — about 20 of 1,955 cells, which is roughly what an air-patrol schedule can visit.
Read from the ROC, the model beats persistence by 13.0 points of AUC and the case is closed. Read from the alert budget, the model's advantage over a rule anyone could implement in an afternoon is 8.3 percentage points of precision at one day and 6.7 at seven. Still a real gain — a fifth more true finds per patrol — but it is a different sales conversation.
The counter-argument, stated fairly. ECMWF's own fire-forecast verification (Di Giuseppe et al. 2020) deliberately declines false-alarm-sensitive metrics: "the identification of false alarms is not meaningful, and the verification should mainly rely on hits and misses", on the reasoning that a high fire danger with no ignition is not a forecast error. That is a defensible position for a danger index. It is not a defensible position for a probability product sold on a patrol budget, because a patrol dispatched to a cell that does not burn has a cost that lands on somebody. Whichever position this case takes, it should state it — and it currently states neither.
6 Finding three — where the skill actually lives
An AUC of 0.874 on a spatially clustered, violently seasonal process invites one obvious suspicion: that the model has learned a map and a calendar. We put that to the data before testing it properly, by removing one source of variation at a time from the ROC itself.
The suspicion is not confirmed. Removing all seasonal variation leaves the model at AUC 0.829; removing all geography leaves it at 0.819. Both baselines fall much further — climatology loses 10.5 points when geography is removed, because a day-of-year climatology is mostly geography. And the pure map baseline reaches only 0.680 AUC with average precision 0.070 — barely above the 3.76% base rate — while its rank correlation with the model's own scores is 0.44.
So the model is answering both questions: given a day it finds the right cells, and given a cell it finds the right days. That is the diagnostic. The proper test is next.
7 Finding four — the decisive test, and it passes
Blocking cross-validation by season removes the leakage between adjacent days. It does nothing about the leakage between adjacent places. Every 0.25° cell in the test year is also in the training years, and the model carries seven fire-history features for that exact cell, so it can in principle learn "this cell burns" rather than "these conditions burn". Roberts et al. (2017) and Ploton et al. (2020) both report that this is where environmental models lose most of their apparent skill.
We refit the whole thing. Cells were grouped into 73 blocks of 2° — far larger than the grid, and larger than the correlation length of a fire season — and the blocks assigned at random to 4 spatial folds. Blocks are assigned at random rather than in stripes deliberately: a striped pattern puts every held-out block against a training block and leaks back across the seam. For each held-out season and each held-out spatial fold, the model was retrained from scratch on the other seasons and the other blocks, recalibrated on a separate season's other blocks, and scored on exactly the rows the published number was scored on.
The result is unambiguous and it goes the case's way. Scoring only cells the model has never seen costs -0.0054 AUC at one day (0.8754 → 0.8700) and -0.0083 at seven (0.8221 → 0.8139). Average precision falls by 0.009 and 0.016. The seven-day surface still clears the gate's 0.80 threshold at 0.814 with no cell of its own history to lean on. Set that against Ploton et al.'s R² 0.53 → 0.14, or against the wildfire preprints that report AUC falling by 0.17 to 0.45 under spatial blocking.
The skill is in the weather and the fuel, not in the memory of where fire lives. On this evidence the model would transfer to peatland it has never been fitted on — which is the property that decides whether this is a product or a demo.
Two reasons the result is credible rather than suspicious. First, the model's features are overwhelmingly meteorological and static-fuel quantities that vary smoothly across a block boundary, so a held-out block is not an out-of-distribution region in the way a held-out country would be — this tests transfer between neighbourhoods, not between biomes. Second, §6 predicted it: a within-day AUC of 0.829 already said the model was ranking cells against each other on a given date, which is the same capability a spatial hold-out asks for.
Why we report a test that failed to break anything. We could find no peer-reviewed fire-occurrence study that reports both a random or season-blocked score and a spatially blocked one on the same data; the two wildfire comparisons we located are unrefereed preprints. That absence is itself a finding about the field, and it makes this the strongest positive statement in this review — a statement the case could not make about itself, because it never ran the test. The published page should carry it.
8 Finding five — "beats the FWI" was scored on one season, because of a bug
The comparison that gives the case its chapter title is the Brier skill score against the CEMS Canadian Fire Weather Index. We audited the join that produces it, row by row, and found the reviewed build had computed it on a single held-out season — 133,386 cell-days, against the 2,142,680 the model's own headline AUC is computed on. Two halves of the same table row, two different samples, no note on the page.
The cause was not the archive. Of the 3 held-out seasons the model is scored on, the CEMS record in that build covered 2016 and 2018; 2017 was absent, and so was 2015, the anchor the case is most about. The ingest metadata described them as "queued — EWDS jobs submitted; rerun to drain", which implies a rerun would fetch them. It would not. The job ledger recorded those seasons as rejected across roughly forty submissions and six and a half hours, and the reason was a defect in the case's own Copernicus queue driver: when a job came back rejected, the poller re-stamped its cooling-off timestamp on every pass, so the submit loop's 180-second retry window never elapsed and the request could never be resubmitted for the life of the ledger. The log said "cooling off" indefinitely and nothing moved.
Dropping the dead job id on rejection is a four-line fix. With it in place all three stranded seasons landed inside fifteen minutes, and the CEMS record went from 12 years to 15 — 31,330,515 cell-days, status complete.
So the comparison can now be made properly, and we made it twice. The case's own routine, re-run against the complete record without refitting anything, now scores the index on 266,658 cell-days across two held-out seasons instead of one. Our own even-handed version scores model and index on every out-of-fold row — all 3 seasons, 400,000 cell-days, a 100% join, identical rows on both sides.
The claim survives, and it is now properly supported — but it had been flattered. On the one season the original comparison happened to land on, the index scored 0.806. On the two seasons it had been missing, the index scores 0.750 and 0.755. The published figure was the index's best of three, which inflated the model's apparent margin: the case's own recomputed Brier skill against the FWI at one day falls from +0.131 to +0.108, and at seven days from +0.055 to +0.051. Smaller, and real.
Two further things the complete record exposes. First, the seasons differ enormously and pooling hides it: at seven days' lead the model scores 0.794 on 2016 alone — below the gate's own 0.80 threshold — against 0.844 on 2018. A pooled AUC of 0.822 is a mean over a range that straddles the threshold it is being judged against, and the spread belongs beside the mean. Second, both anchors now carry an FWI baseline (2015 and 2019), so the comparison the case most wants — how the model and the index behaved in the crisis years — has become answerable for the first time. It has not yet been made.
Is the index's value here plausible? This matters, because a flattering baseline would make the whole comparison meaningless. Mortelmans et al. (2025) report FWI AUC between 0.75 and 0.79 for fire occurrence over 7°S 95°E to 7.4°N 120°E — very nearly this case's box. This case's pooled value, 0.777, sits inside that published band. The index is being scored fairly.
The one comparison that does not agree, and what it means. Shmuel et al. (2025) report the plain Canadian FWI at AUC 0.89 for Indonesia on a 0.25° daily binary ignition grid — the same resolution, the same country, and far above the 0.777 measured here. Three differences account for it and none of them favours this case being quiet about the discrepancy. Their target is the ignition of a new fire from 250 m fire polygons rather than "at least one VIIRS detection in a cell"; their classes are balanced by random undersampling, so the negatives in play are a sample rather than the population; and their test set is a random 20% holdout with no spatial or temporal blocking, on data whose dominant predictor has a 42-day autocorrelation length. Their machine-learning models reach 0.96 on the same split. Since AUC is prevalence-invariant in expectation, the balancing is not the explanation; the split design and the target definition are. A reader is entitled to know that a higher published number exists, and to know exactly which design choice separates it from this one.
9 Finding six — the index being beaten is not the best index available
The Canadian FWI is a composite. Its Build-Up Index (BUI) combines the Duff Moisture Code and the Drought Code and describes how much fuel is available to burn; its Initial Spread Index combines the Fine Fuel Moisture Code with wind and describes how fast a fire would run once lit. The composite multiplies the two. Over equatorial peat that is a strange thing to do, because the limiting factor is almost never spread rate, and the fine fuel moisture code equilibrates in hours to a humidity that in Indonesia barely varies.
The case's data lets us test that directly, because the CEMS product ships every component. We scored all six on identical rows.
At every lead the best member of the family is BUI, not the composite: 0.781 against 0.777 at one day, 0.750 against 0.738 at three, and 0.710 against 0.697 at seven. The drought-and-fuel half of the index carries the signal and the spread half dilutes it.
This is not a novel discovery — it is a reproduction, and that is what makes it trustworthy. de Groot et al. (2007) selected the Drought Code for the Indonesian and Malaysian Fire Danger Rating Systems on exactly this reasoning, because peat fires produced the great majority of particulate emissions in 1997. Mortelmans et al. (2025) reach the same conclusion quantitatively over this box: DC at 0.79–0.83 against FWI at 0.75–0.79, concluding that DC is "the best predictor for fire occurrence". Di Giuseppe et al. (2020) say it operationally: over equatorial Asia "a fire early-warning system should mostly rely on the drought code". The case's own data lands in the same place from an entirely independent direction.
What this costs the case, and what it buys it. It costs a little: the honest margin at one day is 0.0933 AUC over BUI rather than 0.0974 over FWI, so "we beat the FWI" should read "we beat the best drought-code member of the FWI family, by a slightly smaller margin". It buys a great deal more. Three independent lines — a 2007 operational design decision, a 2025 hydrological study over the same box, and this case's own held-out data — agree that the composite index published daily by the ASEAN Specialised Meteorological Centre and by BMKG is mis-specified for peat, and agree on which component should carry the warning. That is a change to a public product rather than a procurement, and it is worth more than the margin it costs.
10 Finding seven — the trajectory ensemble is wider than the answer it gives
The attribution engine releases 30 parcels per receptor-day and integrates them 72 hours backwards on ERA5 winds with a petterssen 2nd-order, 2 corrector passes scheme. A single trajectory is a line; the spread of the ensemble is the only uncertainty statement available. So we measured it, across every stored receptor-day rather than on the one the hero animates.
The ensemble half-width is 24 km at six hours, 82 km at a day, and 177 km at 72 hours, with a 90th percentile of 465 km. Simultaneously the ensemble thins: only 55% of released parcels are still inside the domain at 72 hours, against 88% at 36. Both effects run the same way — the further back you look, the wider and the sparser the evidence — and both are the physically expected behaviour of a kinematic trajectory (Stohl 1998).
And this spread is a floor, not an estimate. The ensemble is generated by perturbing the release point and the release height, so it samples initial-condition uncertainty only. It contains nothing for the error in the wind analysis itself, which Stohl (1998) puts at 36% of travel distance at 72 hours — around 470 km on a typical equatorial trajectory, which is close to this ensemble's 90th percentile of 465 km rather than its median. Engström & Magnusson (2009) add that analysis-error ensembles are under-dispersive in the tropics specifically, and demonstrate a back-trajectory released at 0°N 120°E — inside this domain — splitting into "two different source regions" that are "equally plausible". The honest reading of Figure 7 is therefore that 177 km is the least the 72-hour uncertainty can be.
The consequence is a resolution limit, and it is measurable rather than rhetorical. At 177 km the ensemble is wider than most Sumatran provinces are deep, and wider than the ~200 km separating the Jambi and South Sumatra fire complexes. A province-level statement at 72 hours is being made at a scale the instrument cannot resolve; the same statement at 24 hours, where the half-width is 82 km, is inside it.
Recomputing the attribution under different assumed windows shows what that costs. Holding the parcels, the fires and the residence-time × radiative-power weighting fixed and changing only the window, the top-ranked province is unchanged on 88.8% of 730 episodes at 48 hours and 72.6% at 24. So on roughly one episode in four, the province you name depends on how far back you chose to look.
The case publishes its direction check as failed — 61.3% of episode days agree within ±30° against a 70% threshold — and frames the failure by its median difference of 19.7°, which is comfortably inside the tolerance. That framing is too kind. The 90th percentile of the difference is 106°, and it exceeds 90° at every one of the seven receptors. On the worst tenth of days the forward and backward runs of the same integrator point into different quadrants. The honest statement is not "usually close, occasionally off" but "bimodal: close on most days, meaningless on a minority, and you cannot tell which day you are on from the output".
11 Finding eight — replayed over every episode, the attribution points somewhere else
The case's signature interaction picks one bad-air day in Singapore and lands the back-trajectories on fires in South Sumatra. One episode is an illustration. The archive contains 1,284 Singapore episode days with a computed attribution, so the question of whether that episode is representative has an answer.
It is not representative. The province that most often comes first is Johor, on 329 days (25.6%); second is Bangka-Belitung Islands (18.3%); South Sumatra is 11.1%. 26 different provinces take first place at some point. Across all receptors the median top share is 74.2% carried by a median ensemble agreement of 76.7%, and 8.1% of receptor-days return the case's explicit "no attributable source" outcome — a genuinely good design decision that the page under-advertises.
Three things follow. The first is uncomfortable for the case, the second vindicates it, and the third is the one that should change the page.
First, a geometry artefact. The weighting is residence time × fire radiative power with no correction for distance. But a back-trajectory spends its first hours near the receptor whatever its ultimate origin — the median ensemble is still within 24 km of the city at six hours (Figure 7). So any burning land close to the receptor accumulates residence time on almost every trajectory, and the ranking is biased toward proximity. Johor and Bangka-Belitung Islands are the two nearest land masses to Singapore. Some of this ranking is that bias, and a published province ranking needs a distance term or a far-field restriction before it can be read any other way.
Second — and this is the surprise — the finding is corroborated. Hansen et al. (2019) attributed Singapore's biomass-burning PM10 for 2010–2015 with the NAME Lagrangian dispersion model and emissions tagged by source region, which is a far heavier instrument than this one. Their headline conclusion: "these results challenge the current popular assumption that haze in Singapore is dominated by emissions/burning from only Indonesia … Peninsular Malaysia is a large source for the Maritime Continent off-season biomass burning impact on Singapore." A proximity bias and a real proximity effect look identical in a ranking, and the published dispersion modelling says at least part of this is real. Marlier et al. (2015) give the mechanism: Sumatran fires dominate population-weighted smoke concentrations despite Kalimantan emitting more, "because fire sources were located closer to population centres". Distance is not only an artefact in this problem; it is also physics.
Third, a scope contradiction on the page. The case states plainly that "the transboundary claim here is Indonesia → Singapore only", because Malaysian air-quality data proved unusable. That is a statement about ground truth, and it is correct as such. But the model's own attribution names Malaysian provinces first on 29.0% of Singapore episode days, and the published literature says that is not a mistake. The engine makes Indonesia-and-Malaysia → Singapore statements while the page sold itself as an Indonesia → Singapore instrument. That framing should be corrected — not softened, corrected — and the correction makes the case more interesting rather than less.
What the same literature says about the granularity. Hansen et al.'s two Singapore monitoring stations sit about 25 km apart — less than one cell of this model's grid. During the same September–November 2015 episode they attribute 38.2% and 21.8% of PM10 to South Sumatra, and 31.2% and 41.4% to Central Kalimantan — the two stations reverse the top-ranked source province. The within-receptor spread exceeds the between-province differences any attribution of this kind reports. And Koplitz et al. (2016), running a full adjoint chemistry-transport model over a whole season, still wrote that "there is some uncertainty in our attributions of emissions and exposure contributions to specific provinces". If a tagged dispersion model and an adjoint CTM both decline that precision, a kinematic trajectory should decline it too — while continuing to publish the ranking as what it is, a prioritised list of where to look.
12 Finding nine — three claims the case's own data contradicts
These are not interpretive disagreements. They are statements on the page that the published JSON refutes.
- "Released at CAMS GFAS injection heights." Chapter 04 says the forward trajectories are released at GFAS injection heights. The transport metadata says 3.7% of parcels were — 96.3% used the parameterised plume-rise fallback, because GFAS carries a usable injection height only over the later part of the archive. The hero's smaller print does say heights come from GFAS "where GFAS exists"; the chapter that draws the plume does not. Since release height decides whether smoke settles locally or joins the flow across the Strait — and since Draxler & Hess (1998) showed that starting a trajectory at 10, 200 and 750 m produces "quite substantial" differences where a ±0.5° horizontal offset produces almost none — this is the single most consequential unqualified sentence on the page.
- "Drawn beside the CAMS chemistry forecast." Chapter 04 promises the trajectory result next to a real chemistry-transport model, with divergence as the finding. The published transport metadata records that comparison as unavailable. The forecast data itself has since landed — the ingest reports 3 years (2015, 2019, 2025), status complete — but it arrived after the transport stage last ran, so nothing has yet been compared against it. The physics check the specification calls the real validation is one transport re-run away, and until that happens the chapter is describing a figure that does not exist.
- The two halves of the case run on almost disjoint years. The risk model is fitted on 2012, 2015, 2016, 2017, 2018, 2019, 2026. The transport model is integrated on 2014, 2015, 2019, 2020, 2024, 2025, 2026. They share 2015, 2019, 2026 — 3 of 7. Nothing about that is wrong, and it is a straightforward consequence of the Copernicus queue delivering single-level and pressure-level years in different orders. But it means "where fire starts" and "where the smoke goes" are, at present, statements about largely different periods, and the page reads as though they were one system.
13 Finding ten — the case's best result is published as a failure
The anchor-replay check asks whether 2015 and 2019, held out of training entirely, land in the top decile of modelled seasonal severity. It ships red. The case's own diagnosis is arithmetic: a 90th-percentile rule over 7 seasons admits 1, and two anchors cannot occupy one slot. That is correct, and refusing to move the threshold is right.
What is missing beside it is what the model actually did.
The model reproduces the observed severity ordering of all 7 seasons exactly — Spearman ρ = 1.00 — having never seen two of them, and scores the anchors blind at AUC 0.909 in 2015 and 0.904 in 2019, higher than its own cross-validated folds. A rank correlation of 1.00 over seven seasons would arise by chance about once in five thousand draws.
One qualification the case should also print. The ordering is exact but the top of the distribution is compressed. Observed, 2019 exceeds 2012 by 22%; modelled, it exceeds it by 0.9%. The model gets the order right on a margin of under one percent, so the ordering result is robust in rank and fragile in magnitude. Both halves belong on the page.
Publishing a failed check is a discipline. Publishing it without the successful result standing next to it is an error in the opposite direction — and it is the single most valuable thing this instrument does.
14 What follows for decisions
Evidence is only worth gathering if it changes an action. Before listing the uses it is worth saying what is already published daily over this exact domain, because the case does not say it. The ASEAN Specialised Meteorological Centre publishes the ASEAN Fire Danger Rating System — the full Canadian index set, gridded, free, at analysis plus one to seven days — and BMKG publishes the same thing for Indonesia at observation plus seven days under the SPARTAN banner. Both use de Groot et al.'s 2007 breakpoints unchanged. So a daily 1–7 day fire product is not white space. What is white space is the distinction BMKG itself draws: SPARTAN, in its own words, "does not report fire locations … but a picture of the weather conditions that support a fire occurring if there is an ignition source", and "does not detect heat sources". Fire danger is not fire probability, and nothing operational in this region publishes the second. That, and not the lead time, is this instrument's claim to exist. It should be the first sentence of the case, and it currently appears nowhere on it.
Read strictly, then, the instrument supports four uses and forbids a fifth.
- Patrol allocation, at one to three days. At a 1% alert budget the model returns 53% precision at one day against 45% for the rule a district office already uses. For KLHK's Manggala Agni patrols and BNPB's water-bombing rosters, that is a costable improvement in visits per detection, and §7 says it should hold in districts the model was never fitted on. The caution is that peat fires spread at roughly 135–190 m a day and smoulder for weeks, so the binding operational constraint may be detection latency and dispatch rather than forecast lead — and no published Indonesian dispatch SOP states a target lead time either way. A buyer should be asked which of those two problems they actually have.
- Re-specifying the operational index. §9 is directly actionable by BMKG and by the ASEAN Specialised Meteorological Centre, neither of whom needs to buy anything: over this domain the Build-Up Index outperforms the composite FWI at every lead, reproducing both de Groot et al. (2007) and Mortelmans et al. (2025). That is a change to a published product, not a procurement, and it is the cheapest public good in this review.
- Verification of what is already published. ASMC publishes no skill scores for any of its products, and by ASEAN's own 2023 roadmap it does not publish burned area either. An independently verified, openly scored fire product is defensible even where the variable is not new — and this case's discipline about publishing failed checks is exactly the qualification for that work.
- Seasonal outlook, and near-field warning at 24 hours. The ordering result in §13 — an exact reproduction of seven seasons' severity, blind on two — maps onto the decision that matters at national scale: whether this is a year to pre-position, a BNPB and Ministry of Finance conversation months ahead. At the other end, Figure 7 puts the ensemble half-width at 82 km at a day, so 24-hour downwind warnings for Singapore's NEA and provincial health offices sit inside the instrument's resolution. Seventy-two-hour province-level attribution does not.
- Not: contested attribution. Singapore's Transboundary Haze Pollution Act 2014 is precise about what evidence does what, and a province share fits nowhere in it. The Act engages only above a 24-hour PM2.5 index of 101 sustained for a day; its section 8 presumption requires a specific located fire to be proved, and admits satellite information only to establish that the smoke was moving towards Singapore; ownership is presumed only from maps obtained from a compelled disclosure or a foreign government. The unit of attribution throughout is a parcel of land owned or occupied by a named entity — not an administrative region. In the only two Act cases publicly closed, the ownership presumption was rebutted. A province share from a kinematic back-trajectory with a 177 km ensemble half-width, a proximity-biased weighting and a 27% sensitivity to the choice of window satisfies no limb of it. Under the ASEAN Agreement on Transboundary Haze Pollution the position is softer but no more useful: the treaty obliges monitoring and early warning, it names no attribution mechanism, and its assessment article is hedged twice over. This is a triage instrument that says where to look, not an evidentiary one that says who is responsible.
One thing the case should stop under-selling. Its Singapore receptor correlation is Spearman ρ 0.525 against instruments. Crippa et al. (2016) ran WRF-Chem at 10 km with full aerosol chemistry over the 2015 episode and report a daily PM2.5 correlation of 0.55 at Singapore; Hansen et al. (2019) report 0.35 and 0.32 for the 2015 season with a tagged dispersion model. A parcel-counting exposure index on ERA5 winds, with no chemistry at all, is within the same band as both. The page reports this as a gate that passed; it is a result.
15 What remains open
- Score the anchors against the index. The CEMS backfill is now complete and both 2015 and 2019 carry an FWI baseline for the first time. The model was already scored blind on those years; the index has not been. That comparison — how an operational fire-danger index and this model each behaved in the two crises the case is named for — is the most valuable single number still missing, and it needs no new data.
- Re-run transport against the CAMS forecast. The forecast years (2015, 2019, 2025) are on disk and the comparison code exists; it simply has not been run since they landed. That closes the only outstanding physics check in the case.
- Score against the operational baseline that actually exists. The comparison a buyer will make is not against day-of-year climatology but against the ASEAN FDRS and BMKG products already published daily over this domain at analysis plus one to seven days. Those are the same CEMS-family indices, so the machinery is in place — what is missing is the framing.
- Confront ProbFire directly. Nikonovas et al. (2022) already publish a probabilistic fire-occurrence model on a 0.25° Indonesian grid. It is monthly rather than daily and it publishes no numeric skill table, so a like-for-like comparison is not currently possible — which is itself the argument for building one.
- Give the attribution a distance term. The residence-time × radiative-power weighting is the standard receptor-model form, but with no distance normalisation it ranks the nearest burning province first by construction. A far-field restriction, or a weighting by the parcel's own travel time, would separate "the air came from here" from "here is close".
- Publish an attribution only where the ensemble supports it. The spread is already computed per receptor-day. A province share whose ensemble half-width exceeds the province's own width should be reported at island level or not at all — turning §10's limit into an automatic scope rule rather than a caveat.
- Fit on BUI as well as beat it. If the Build-Up Index is the best single predictor in the CEMS family over peat, the interesting experiment is what the model adds on top of BUI, not what it beats. That is the nested-model version of §9 and it is one refit away.
- Run the spatial test at three days, and by ecosystem rather than by block. §7 blocks geometrically. Blocking by peat/non-peat, or holding out Kalimantan entirely and fitting on Sumatra, tests transfer between the regimes an operator actually cares about.
16 References and reproducibility
- Mortelmans, J., Apers, S., De Lannoy, G.J.M., Veraverbeke, S., Field, R.D., Andela, N., Page, S.E. & Bechtold, M. (2025). Using hydrological modelling to improve the Fire Weather Index system over tropical peatlands of peninsular Malaysia, Sumatra and Borneo. International Journal of Wildland Fire 34(2), WF24057. doi:10.1071/WF24057 — the AUC band in Figure 6, and the 42-day autocorrelation length in §7.
- Shmuel, A., Lazebnik, T., Heifetz, E., Glickman, O. & Price, C. (2025). Fire weather indices tailored to regional patterns outperform global models. npj Natural Hazards 2, 74. doi:10.1038/s44304-025-00126-y — Supplementary Table 1, Indonesia, 0.25° daily grid, class-balanced with a random holdout.
- de Groot, W.J., Field, R.D., Brady, M.A., Roswintiarti, O. & Mohamad, M. (2007). Development of the Indonesian and Malaysian Fire Danger Rating Systems. Mitigation and Adaptation Strategies for Global Change 12(1), 165–180. doi:10.1007/s11027-006-9043-8 — the Drought Code bands still in operational use.
- Di Giuseppe, F., Vitolo, C., Krzeminski, B., Barnard, C., Maciel, P. & San-Miguel-Ayanz, J. (2020). Fire Weather Index: the skill provided by the ECMWF ensemble prediction system. Natural Hazards and Earth System Sciences 20, 2365–2378. doi:10.5194/nhess-20-2365-2020
- Field, R.D., Spessa, A.C., Aziz, N.A. et al. (2015). Development of a Global Fire Weather Database. Natural Hazards and Earth System Sciences 15, 1407–1423. doi:10.5194/nhess-15-1407-2015
- Nikonovas, T., Spessa, A., Doerr, S.H., Clay, G.D. & Mezbahuddin, S. (2022). ProbFire: a probabilistic fire early warning system for Indonesia. Natural Hazards and Earth System Sciences 22(1), 303–322. doi:10.5194/nhess-22-303-2022 — the existing 0.25° probabilistic fire-occurrence model for Indonesia; monthly, and it publishes no numeric skill table.
- Roberts, D.R., Bahn, V., Ciuti, S. et al. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 40(8), 913–929. doi:10.1111/ecog.02881
- Ploton, P., Mortier, F., Réjou-Méchain, M. et al. (2020). Spatial validation reveals poor predictive performance of large-scale remote sensing models. Nature Communications 11, 4540. doi:10.1038/s41467-020-18321-y
- Saito, T. & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE 10(3), e0118432. doi:10.1371/journal.pone.0118432
- Davis, J. & Goadrich, M. (2006). The relationship between precision-recall and ROC curves. Proceedings of the 23rd International Conference on Machine Learning, 233–240. doi:10.1145/1143844.1143874
- Stohl, A. (1998). Computation, accuracy and applications of trajectories — a review and bibliography. Atmospheric Environment 32(6), 947–966. doi:10.1016/S1352-2310(98)00184-8
- Draxler, R.R. & Hess, G.D. (1998). An overview of the HYSPLIT_4 modelling system for trajectories, dispersion and deposition. Australian Meteorological Magazine 47(4), 295–308 — the release-height sensitivity quoted in §12.
- Engström, A. & Magnusson, L. (2009). Estimating trajectory uncertainties due to flow dependent errors in the atmospheric analysis. Atmospheric Chemistry and Physics 9(22), 8857–8867. doi:10.5194/acp-9-8857-2009
- Hansen, A.B., Witham, C.S., Chong, W.M., Kendall, E., Chew, B.N., Gan, C., Hort, M.C. & Lee, S.-Y. (2019). Haze in Singapore — source attribution of biomass burning PM10 from Southeast Asia. Atmospheric Chemistry and Physics 19(8), 5363–5385. doi:10.5194/acp-19-5363-2019 — the two-station reversal in §11.
- Koplitz, S.N., Mickley, L.J., Marlier, M.E. et al. (2016). Public health impacts of the severe haze in Equatorial Asia in September–October 2015 and 2006. Environmental Research Letters 11, 094023. doi:10.1088/1748-9326/11/9/094023
- Marlier, M.E., DeFries, R.S., Kim, P.S., Koplitz, S.N., Jacob, D.J., Mickley, L.J. & Myers, S.S. (2015). Fire emissions and regional air quality impacts from fires in oil palm, timber, and logging concessions in Indonesia. Environmental Research Letters 10(8), 085005. doi:10.1088/1748-9326/10/8/085005
- Crippa, P., Castruccio, S., Archer-Nicholls, S. et al. (2016). Population exposure to hazardous air quality due to the 2015 fires in Equatorial Asia. Scientific Reports 6, 37074. doi:10.1038/srep37074
- Reid, J.S., Hyer, E.J., Johnson, R.S. et al. (2013). Observing and understanding the Southeast Asian aerosol system by remote sensing. Atmospheric Research 122, 403–468. doi:10.1016/j.atmosres.2012.06.005 — "no fire product … provides comprehensive detection of open burning in SEA".
- Van Wagner, C.E. (1987). Development and structure of the Canadian Forest Fire Weather Index System. Canadian Forestry Service, Forestry Technical Report 35. The definition of the Build-Up Index and the Initial Spread Index used in §9.
- Republic of Singapore (2014). Transboundary Haze Pollution Act 2014 (Act 24 of 2014), in force 25 September 2014, with the Transboundary Haze Pollution (Air Quality) Regulations 2014 (S 622/2014). Sections 2(2), 8 and 22 are the ones §14 relies on.
- Association of Southeast Asian Nations (2002). ASEAN Agreement on Transboundary Haze Pollution, Kuala Lumpur, 10 June 2002. Articles 4, 7, 8 and 9. Ratified by Indonesia through UU No. 26 Tahun 2014.
- Copernicus Emergency Management Service. Fire danger indices historical data from the Copernicus Emergency Management Service (CEMS), cems-fire-historical-v1, via the CEMS Early Warning Data Store. doi:10.24381/cds.0e89c522
Data vintage 2026-08-31. Every threshold was fixed and recorded in advance of the result it judges, and none was moved for this review.