JABODETABEK AIR QUALITY NOWCAST CASE E · OPERATIONAL PIPELINE GROUND 2026-08-30 06:00 UTC · BUILT 2026-08-30 14:21 UTC · 2/6 GATES
Sensor fusion + ML forecasting

How bad will
the air be
tomorrow?

Greater Jakarta gets its air quality the way it gets its weather: after the fact. This case turns three open feeds — ERA5 meteorology, NASA fire detections and the handful of ground sensors still reporting — into a PM2.5 forecast with a horizon and an error bar, graded against every trivial rule that could stand in for it: carrying today forward, carrying yesterday's daily mean forward, and the long-run average for this hour and season.

FORECAST ·· µg/m³ loading forecast…
Now+24 h+48 h+72 h
01 · The air today

Jakarta's air is not occasionally bad. It is usually bad.

Of the 3,779 complete days on record at the city's ground monitors, 94.3% exceeded the WHO 24-hour guideline of 15 µg/m³, at a mean of 44.7 µg/m³. The 95th percentile hour reaches 94.4.

Read that mean carefully. It pools instrument classes: across the 16,021 hours from the two reference-grade regulatory monitors the mean is 38.0 µg/m³, against 52.8 across the 40,690 hours from consumer-grade units. The gap is instrument, not geography, and the review page quantifies it.

Below, everything the most recently reporting sensor has actually sent in the last three weeks — not a smoothed line, the raw hours, gaps and all. It is a thin record, and that thinness is the operating condition this whole case is built around. Under it, the shape of the problem across the full archive: the daily and seasonal cycles that a forecast has to get right before it earns any attention at all.

Observed PM2.5 · last 21 days · station

Hover for the hour. Gaps are real: the sensor's own uptime, not smoothing.

The daily rhythm · every observed hour, all stations

So what: a dial showing the current number is worthless to anyone who has to schedule something. The value is in the next 72 hours.

02 · The forecast

Beating “tomorrow looks like today”

At the 24-hour horizon the model's mean absolute error is 16.0 µg/m³ against persistence's 19.5 — a 17.7% improvement on the held-out future, and 11.6% against the strongest trivial rule available at that lead.

Which baseline matters. Persistence is not the hardest trivial rule at a 24-hour lead — a last 24 h mean reaches 18.1 µg/m³ where persistence reaches 19.5, because at a 24-hour lag persistence is handed the entire daily cycle for free. Measured against that stronger rule the improvement is 11.6% [9.4%, 13.7%], which is below the 15% this case registered in advance. The skill check below is reported against persistence, as it was written; on the fairer ruler it would not pass, and the review page works through why.

Every number here comes from a single time-based split: the model never sees an hour later than the cut, and is then asked about the months that follow. A random split would leak tomorrow into today's training set and inflate every figure on this page. Two baselines are drawn below — persistence (carry the last reading forward) and diurnal climatology (the average for this hour of day) — and the review page adds four more.

Error by horizon · held-out period

So what: skill is not a single number, and it is not only a property of the model. The peak at 12 hours is where persistence is at its worst, not where the forecast is at its best — at a 12-hour lead the baseline is comparing night to afternoon. Read the curve against the line it is drawn on.

03 · What drives it

What the model actually leans on

Feature attribution pending.

Permutation importance at the 24-hour horizon: how much worse the model gets when each feature is shuffled. Features are grouped by what they represent, because the interesting question is not which column wins but which physics does — the air's own memory, the depth of the layer it mixes into, the wind that ventilates it, the fires upwind, or the clock.

Feature importance · 24 h model · permutation, 5 repeats

Δ MAE ON LOG SCALE

So what: if this chart were pure autoregression — lagged PM2.5 and nothing else — the model would be a clock, not an air-quality forecast, and it would fail the moment conditions changed. One of the published checks exists to test exactly that.

04 · The airshed

Where the bad air comes from

Jakarta does not breathe alone. The rose below splits every measured hour by the direction the wind was blowing from: petal length is how often, petal colour is the mean PM2.5 that arrived on it. Switch to fires and the same eight sectors show NASA's VIIRS hotspot counts across the airshed — Sumatra's peatlands, the length of Java, the southern edge of Kalimantan — split into three distance rings.

Wind rose · mean PM2.5 by arrival direction

Hover a petal for the numbers. Fires are attributed to a three-sector arc around the bearing from Jakarta and lagged one day, so an afternoon satellite overpass never informs that morning's forecast. Volcanoes, gas flares and other permanently hot ground are excluded — see below.

Fire hotspots detected across the airshed, by month

VIIRS S-NPP · 95–119°E · 9°S–6°N

Not every hot pixel is a fire. Indonesia has around 130 active volcanoes and a great deal of gas flaring, and VIIRS sees all of it. The satellite does label those detections — but only in the standard-processing archive; the near-real-time product that covers the last few months omits the label entirely. Filtering on the label alone would therefore clean the history and leave the recent tail full of volcanoes, putting a step change in this chart exactly where the two products meet — and fire counts feed the forecast, so the step would propagate into the model. Instead the archive is swept across its full record for everything it ever flagged as a volcano, flare or offshore source; those locations are gridded to 0.01° cells with a one-cell buffer for pointing jitter, and the resulting mask (3,908 cells, 708 of them flagged directly) is applied to every row of both products. It removes 5.9% of otherwise-qualifying near-real-time detections against 3.6% of archive ones — the archive needs less help because its own labels already caught most of them, and that asymmetry is exactly what used to land at the seam. In April 2026, the one month the seam falls inside, both products now read as roughly a third static (34% archive, 36% near-real-time); before the fix the near-real-time side read as zero.

So what: the same fire is worth a lot or nothing depending on where the wind is. That is why the model sees upwind fire counts, not a regional total.

05 · The finding nobody publishes

A metropolis of 32 million is watched by 1 working public sensor.

The open record lists 24 PM2.5 stations inside Jabodetabek. Only 11 have ever returned an hour of data. 2 reported in the last fortnight — and 1 of those has delivered no hourly measurements at all.

The denominator is softer than it looks. Those 24 registrations resolve to 16 distinct addresses — 6 of them share one coordinate under six different names, none of which ever reported — and 13 of the 24 have never delivered an hour of data, including legacy diplomatic-post entries that duplicate feeds registered separately. Told honestly it is 11 instruments that ever worked, which is a smaller number and a worse one.

Each line below is one station's life. The solid bar is the hours this build actually holds; the faint rule behind it is the lifetime the provider's metadata claims — the two diverge, and the gap is itself worth knowing before anyone budgets for an analysis. Read left to right and the network does not grow: it flickers. This is not an access problem a bigger API key would fix. The instruments are not there, and the ones that are there are mostly low-cost units run by embassies, NGOs and volunteers.

And a network this thin loses something a bigger one takes for granted: the ability to catch itself. Two of these instruments sat 43 metres apart and tracked each other to 6.9 µg/m³ across 7,297 shared hours — until, in the last months both were alive, they diverged to 9.9 against 63.0 µg/m³. Then one of them went dark. Nothing in Jabodetabek can now say which was right.

Station lifetimes · every PM2.5 station in the bbox

Live Stale Never reported

Where they are

2-D CANVAS · NO WEBGL · HOVER OR TAB FOR DETAIL

So what: a dense city-wide nowcast grid cannot be honestly validated here, so this case does not claim one — and the deeper cost is not coverage but corroboration. A single sensor cannot be shown to be wrong. The station-coverage check is published failed rather than quietly relaxed, and on the review's reading it is the most important output this pipeline has.

06 · Validation

The held-out months, unedited

The 24-hour forecast plotted against what actually happened, for every hour after the training cut. The shaded band is the model's own 80% prediction interval — quantile regression, so it is allowed to be lopsided, which for pollution it always is. Toggle persistence to see the baseline the model has to beat.

24-hour forecast vs observation · held-out period

Gates · thresholds fixed before the first model run

2/6 PASS

One failure the checks above do not capture. The episode check is written at a single threshold. Above it the US index has two more categories, and at those the forecast is not merely weak — it is silent. Across the held-out window 247 hours reached the “very unhealthy” level of 125.5 µg/m³ and the model called 0 of them. Its highest 24-hour prediction anywhere in the record is 99.1 µg/m³ against an observed maximum of 338.0. Training on log1p and scoring on absolute error pulls predictions toward the middle, and the result is a hard ceiling below the categories a city would act on. No school advisory or outdoor-work suspension should be triggered by this forecast in its current form.

So what: a failed gate published with its diagnosis is worth more to a client than six green ticks they cannot audit — and a failure the gates were never written to catch is worth publishing too.

07 · Review

We had this forecast reviewed, adversarially.

An independent read of the same data against the published literature. It re-derives the headline against five baselines this page did not try and finds the skill check fails on the strongest of them; it shows the skill curve's peak at 12 hours is the shape of the pollutant's autocorrelation, not of the model; it dates the month the city's last two co-located instruments stopped agreeing; and it argues the thin network is not this case's limitation but its result.

Read the review article →