Can radar count Java's rice harvest?
A stress test of a Sentinel-1 paddy detector over six of Java's largest rice regencies, 2022-07 to 2026-08. The instrument publishes five failed checks and one diagnosis — that it finds the rice but finds one crop where there are two. We factor the shortfall, and find the diagnosis is the wrong way round: three-quarters of it is fields never found at all. Then we identify what actually sets the detection rate, by taking looks away from a regency that has enough of them.
Abstract
Question. Indonesia's rice harvested area is measured by a field survey with a two-month lag. Sentinel-1 sees through monsoon cloud and could in principle measure the same quantity as it happens. We ask what a published radar paddy detector actually recovers, and whether the reason it falls short is the one its own page gives.
Method. 1,226,885 cells of 100 m over 6 kabupaten in Jawa Barat, Jawa Tengah, Jawa Timur, 2022-07-01–2026-08-31, benchmarked against BPS's area-frame survey (KSA) and, independently, against a published 10 m rice map whose own single/double/triple classes carry a cropping intensity. All statistics below are recomputed from the case's own stored output; one controlled experiment removes radar looks from a single regency and re-runs the unchanged detector.
Findings. The benchmark is not the moving target: the survey's 2018 method change (-19.4%) is 4 years before the record opens, the five-regency series varies by only 4.3% over the whole survey era, and the independent satellite map lands within 1.4% of it. The shortfall is real and it factors: detected area is 35% of the benchmark in 2025, the product of finding 47% of the fields and 79% of the crops on the fields it finds — so 76% of the deficit is missing fields, not missing crops. On cells the independent map calls double-cropping the detector returns 1.49 crops, not one; it detects only 3.6% of single-cropping cells and 12.5% of triple-cropping ones. The deficit is not in the second season either: once a systematic 1-month late bias is removed, the two harvest lobes are recovered at 26% and 28%. What does predict detection is revisit: thinning Karawang's own record from 6 to 12 days between looks, with the fields, the detector and the thresholds unchanged, costs 92% of its crop cycles and 72% of its fields — and reproduces the East Java regencies' observed shortfall. As Sentinel-1C and 1D came online inside the sample, Lamongan's gap fell 12→8 days and its recall rose ×4.36. Finally, the calibrated series the case reports at R² 0.82 on the annual panel has a monthly R² of 0.06, and the satellite contributes 5.0% of its hectares — negatively in 3 of 6 regencies.
Conclusion. This is a sampling instrument, not a resolution one. The case attributes every failure to cell size and proposes a smaller cell; the evidence says the binding constraint is how often the satellite looks. The correct product is not a hectare count at all — it is the harvest date, which the same detector already gets to within one systematic month with a shape correlation of 0.77.
1 The claim under test
Indonesia measures rice harvested area with the KSA area sampling frame: enumerators classify the crop stage at fixed sample points every month, and the classification is grossed up to a regency figure published about two months later. It is a good survey. It is also slow, it cannot be interrogated below its own sample density, and it is the number on which rice import decisions, procurement targets and the national self-sufficiency programme are argued.
Radar is the obvious challenger. Java's rice grows through the monsoon under cloud that defeats optical sensors, and the crop calendar is written directly into the backscatter: a flooded field before transplanting is a mirror and collapses the signal, a tillering canopy is a volume of vertical scatterers and lifts it, and the crop is cut and it drops again. The case under review implements exactly that, over 1,226,885 cells of 100 m and 4 seasons, and publishes 0 of 5 checks passed.
Publishing five failures is the right instinct, and it means the diagnosis is the product. The page's diagnosis is a single sentence: "we find the rice; we find one crop where there are two." That sentence makes an arithmetically testable claim, because harvested area is a flow and factors exactly:
harvested area = the fields that grew rice × the crops each field grew
If the diagnosis is right, the first factor is near complete and the second is near half. Sections 5 to 7 measure both.
2 Prior art, and what a working system looks like
Radar rice mapping is not new and it is not speculative. Nelson et al. (2014), reporting the RIICE programme, mapped 1.6 million hectares of rice across 13 footprints in 6 Asian countries from 127 X-band images, with classification accuracies of 85–95% against more than 1,300 field observations. Torbick et al. (2017) is the closer comparator, because it is C-band Sentinel-1 over a tropical monsoon country: from 632 Sentinel-1A images they mapped 6,652,111 ha of harvested rice across Myanmar and reported R² = 0.78 against the government census.
The third benchmark is not a paper but a product we can test against pixel by pixel. Open-SEA-Rice-10 (Ginting et al., 2025) maps Southeast Asian rice at 10 m for 2021 with an overall accuracy of 98.3% and an F1 of 0.879, and — decisively for this review — labels every rice pixel single, double or triple crop. Its harvested-area estimates agree with official statistics at R² 0.92 region-wide and R² 0.85 for Indonesia. It is therefore simultaneously an independent extent map and an independent statement about cropping intensity, made by other people, from a different method.
Set against that literature, this case's uncalibrated result — R² -11.0889 against the official survey, and -74.65% on the 2023 aggregate — is not a marginal miss. It is a different order of outcome from published work on the same sensor, and that is what makes the diagnosis worth testing rather than accepting.
3 Data, method, and what the instrument published
Radar. Sentinel-1 radiometrically terrain-corrected γ⁰ at 10 m, read at 80 m and aggregated to a 100 m analysis cell, 1,235 acquisitions across 6 regencies, orbits normalised separately and never mixed, averaged in linear power.
Detector. A local minimum of VV at least 3 dB below the cell's own 80th-percentile baseline marks transplanting; a VH rise of at least 4 dB within 45 days confirms it is a crop; the following VH maximum is heading; harvest follows. Events whose defining dates sit inside an observation gap longer than 24 days are refused.
Benchmarks. BPS monthly KSA harvested area per regency for West and East Java, annual for Central Java; and Open-SEA-Rice-10 as an independent extent and intensity map.
Where this review agrees with the case. The thresholds were fixed before results and are not tuned here either. The scope is honestly stated. The flow/stock distinction — that a field yielding two crops contributes its area twice — is correctly implemented, and it is the reason the shortfall can be factored at all. Nothing below is a complaint about rigour; it is a disagreement about a diagnosis.
4 Finding one — the benchmark is not the moving target
The first hypothesis any reviewer should test is that the official series moved rather than the satellite failing. BPS replaced a generation of eye-estimates with the KSA area frame from reference year 2018, and that break is real and large: national harvested area steps from 14.12 Mha under the old method to 11.38 Mha under the new one, a change of -19.4%, with 2016–2017 a moratorium carrying no data at all.
But the break cannot explain this shortfall, because it is 4 years before the radar record opens. Inside the window the benchmark is quiet: over the whole survey era the five regencies with a continuous series move within 14.33% peak to trough, a coefficient of variation of 4.3%.
The stronger test is corroboration by something that is not a survey at all. Summing Open-SEA-Rice-10's own crop classes over exactly these six regencies gives 1,010,219 ha of harvested area against the survey's 996,434 ha — a gap of +1.38%. An area-frame field survey and a 10 m satellite product built by unrelated authors agree to within a point and a half. The benchmark is sound, and the 65% that is missing is ours.
One honest caution on the corroboration. Open-SEA-Rice-10 derives its crop count from NDVI peaks in optical imagery, which is the method this case exists to avoid in a cloudy country. It is therefore an independent benchmark for extent without qualification, and an independent benchmark for intensity that shares no method with the survey but does share a known weakness. That it lands within 1.4% of KSA anyway is what makes it usable here.
5 Finding two — the shortfall factors, and it factors the other way
Because harvested area is a flow, the ratio of our figure to the benchmark is exactly the product of two recalls: the share of the rice fields we find at all, and the share of each field's crops we count where we do find it. Both are measurable against the independent map, and neither has been reported.
In 2025 the instrument reports 35% of the benchmark. That is 47% of the fields multiplied by 79% of the crops. Averaged over the three complete years, 76% of the deficit is the extent term and only 24% is the intensity term. The detector is not finding the rice and missing a season. It is missing most of the rice.
The per-regency picture is starker than the aggregate. Lamongan in 2023 recovers 3.7% of its rice fields — 2.5% of the benchmark — while still counting 1.02 crops on the few fields it does find.
The case's own page already carries the number that says this, and reads it the other way round. It quotes 83% agreement "on our own area" as evidence that the rice is found. That figure is a precision, and precision is the one statistic a detector can always buy by detecting less. Recomputed per year rather than pooled across four seasons against a one-year map, precision is 85%–93% and recall is 31%–39%. Both numbers are true; only one of them is the answer to "do we find the rice".
6 Finding three — asked in the map's own units, the diagnosis fails again
The independent map labels each rice cell with the number of crops it grows. That converts the diagnosis into a prediction with no wriggle room: on cells the map calls double-cropping, a detector that "finds one crop where there are two" should return about one.
On double-cropping cells the detector returns 1.49 cycles, not one — 74% of the truth, which is a real shortfall but not the stated one. It also separates the classes far less than it should: the ratio of what it counts on double-cropping cells to what it counts on single-cropping cells is 1.39 where the map says 2, and on triple-cropping cells it counts 1.15 against three.
The left panel is the finding that matters. The detector sees 37.8% of double-cropping cells, 3.6% of single-cropping cells and 12.5% of triple-cropping cells. It is not a rice detector with a counting problem; it is a detector of one particular signature — deeply flooded, synchronously transplanted, irrigated double-crop paddy — that is nearly blind to everything else rice does in Java. Single-cropped rice is largely rainfed and often direct-seeded, so it never presents the standing-water minimum the first rule requires; triple-cropped fields turn over faster than the 85-day minimum cycle allows.
7 Finding four — nor is the deficit in the second season
"One crop where there are two" has a seasonal signature. Java's calendar is bimodal: the wet-season rendeng harvest peaks around March, the dry-season gadu harvest around August. A detector that finds the first crop and misses the second should recover the rendeng lobe well and the gadu lobe badly.
As published, the pattern is the opposite of the prediction: we recover 18% of the rendeng harvest and 29% of the gadu harvest — the second season is recovered 1.57 times better than the first. That comparison is contaminated, because the detector's harvest dates run late and a late bias moves mass out of the first lobe into the second. Removing the measured 1-month lag flattens it: rendeng 26%, gadu 28%, a ratio of 1.06.
Either way the conclusion is the same and it is the one the diagnosis forbids: the deficit is flat across the calendar. Whatever the detector is missing, it is not a season.
The shoulder months are the exception and they are diagnostic rather than encouraging. January, June, November and December carry 24% of the official harvest but 46% of ours, and we "recover" 60% of them — more than either real lobe. That is not detection; it is the late bias of §12 smearing genuine February–May harvests into the months after them, and it is why a monthly comparison flatters this instrument in exactly the months nobody is procuring against.
8 Finding five — the decisive test: same fields, fewer looks
If the deficit is neither a season nor a crop count, what sets it? The observable that moves with it across regencies is how often the satellite looked. The problem with a cross-regency comparison is that everything else moves too — soil, variety, plot size, irrigation command. So we ran the experiment that holds all of it fixed.
We took Karawang, the densest record in the sample (225 looks, 6 days apart), removed looks from its own time series, rebuilt the series through what survived, and re-ran the identical detector with identical thresholds. Nothing about the fields changed. Only the number of looks did.
Halving the looks — 6 days between them to 12 — costs 92% of the crop cycles and 72% of the fields. At 18 days 1.5% of the fields survive; at 24 days, 0.24%. Detection is not a slope in revisit. It is a cliff, and the edge sits between eight and twelve days — exactly where the East Java regencies sit.
Thinned to a 12-day gap, Karawang's own fields yield 20.2% recall. Bojonegoro, at a 12-day gap, yields 15.8%. The experiment reproduces the case's own cross-section from revisit alone.
There is a second, natural experiment in the same data, and it runs the other way. Sentinel-1B failed on 2021-12-23, six months before this record opens, so the whole series sits on the degraded side of that event and the obvious before/after test is unavailable. But Sentinel-1C and 1D came into service during it, and revisit rose.
Lamongan's median gap fell from 12 to 8 days and its recall rose ×4.36, from 3.7% to 16.3% — with no change to the detector, the thresholds, the crop or the fields. Across all 18 regency-years the correlation between revisit gap and recall is -0.4852.
What the thinning test does not do. It rebuilds the series by interpolating through the surviving looks rather than re-deriving it from the raw scenes, so it reproduces 79% of the published cycle count at full density. Every rung is built the same way, so the ladder is internally consistent and the comparison against the observed regencies is like for like — but a full re-ingest at reduced density would be the cleaner version, and the absolute level of each rung should be read as indicative rather than exact.
9 Finding six — the "damping" the case measures is largely revisit
The case has a mechanism for the shortfall and it is a good one: a 100 m cell holds several 0.3–0.5 ha paddy plots that are not transplanted on the same day, so the cell mean never swings as far as a single plot does, and the second cycle is averaged away. It measures that damping as the cell's seasonal VH range, builds it into the calibration as an interaction term, and reads the cross-section off it.
That mechanism makes a prediction the thinning test can check. If the canopy swing measures plot heterogeneity, removing looks should barely move it.
Thinning Karawang from 6 to 24 days moves its canopy swing from 8.50 to 6.91 dB — 1.59 dB — while costing 99.7% of its fields. So the swing does respond to sampling, but weakly; it cannot be the channel through which revisit destroys detection.
Two conclusions follow, and they point in different directions. First, the case's damping term is not the mechanism: what fewer looks destroy is the shape the detector tests for — a local minimum, then a 4 dB rise inside 45 days — not the amplitude it never tests directly. Second, the East Java regencies do sit genuinely below the thinning curve on swing (5.84 dB against 7.75 dB for Karawang thinned to the same gap), so there is a real amplitude difference on top of the sampling one. The case's error is not that the damping is imaginary. It is that damping was made to carry a cross-section that revisit explains.
10 Finding seven — the calibrated series is almost entirely not the satellite
The case reports a calibrated R² of 0.82 against the official survey. The case is careful to publish the uncalibrated result beside it and to fail the check on that basis, which is exactly right. It is still worth saying what the calibrated number is.
The calibration is fitted on 147 regency-month rows. At that resolution — the resolution it was fitted at — its R² against the survey is 0.06. The 0.82 is what appears after twelve monthly predictions are summed into a year, which is mostly a statement that regencies differ in size.
Across the panel the detected area supplies 5.0% of the calibrated hectares. Worse, the interaction that carries the damping story gives the detected area an effective slope of 1.444 − 9.64 ÷ swing, which changes sign below 6.675 dB. Three of the six regencies — Bojonegoro, Grobogan, Lamongan — sit below that line, so the fitted model says that in those regencies the more rice the satellite finds, the less rice the survey will report. That is not a calibration; it is a set of regency constants with a sign error attached.
11 Finding eight — the detector sits on a knife-edge it does not declare
The case publishes a threshold-sensitivity table, which is more than most do. Read as a stability statement rather than a disclosure, it is alarming.
Moving the canopy-rise threshold by one decibel in either direction changes the answer by +40% and -51%. That is not a robustness range; it is the whole result. The reason is visible in the case's own diagnostics: the median detected rise is 4.44–5.03 dB against a threshold of 4.0 dB, so half of all detections sit within a decibel of rejection. A speckle realisation, an incidence-angle residual or a missing look is enough to move an event across the line — which is precisely why removing looks in §8 is so destructive.
The literature's unchanged single-plot rule (VV below -17 dB) yields 15,161 ha where the scale-adapted criterion yields 744,218 ha. The case is right that this is a fact about the analysis cell rather than about Indonesian rice, and right to publish it. But the restated criterion inherits the fragility rather than curing it.
12 Finding nine — the one output that works is the one presented as a failure
The timing check fails at a median error of 5.02 weeks against a 2-week threshold, and the case attributes part of that to the detector locking onto the wrong lobe of a bimodal harvest. That attribution is testable: an argmax of a two-peaked curve is brittle, but the whole curve is not. We re-scored the same regency-years by sliding our monthly harvest curve around the calendar and taking the shift that best matches the survey's.
The shape-based estimator gives the same median as the argmax — 1 month against 1 — so the median failure is not an artefact of the estimator. What the distribution adds is that the failure has two modes, not one. 10 of 15 regency-years sit at a clean +1-month lag; 4 sit four or more months away, which on a bimodal calendar is the lobe alias the case describes. So both explanations are partly right — and the dominant one is a systematic, one-signed, one-month late bias. Once that single constant is removed, the correlation between our monthly harvest curve and the official one rises from 0.11 to 0.77.
The instrument that cannot count hectares can tell you, three weeks before the survey can, that the harvest is arriving early or late. That is a different product, and it is the one worth selling.
A systematic bias is a calibration constant, not an error; the estimated harvest date is heading plus 30 days, and heading is dated from a maximum that a sparse series finds late. The correction is one number and it is estimable from the same panel. The hectare count, by contrast, needs a different satellite tasking regime.
13 What follows for decisions
Indonesian rice numbers are not academic. BPS's KSA figure feeds the Ministry of Agriculture's production accounting, Bulog's procurement and stock targets, and Badan Pangan Nasional's advice on whether to import — a decision taken on a scale of hundreds of thousands to millions of tonnes, months ahead of the harvest it is meant to cover. An instrument that claims to see the harvest before the survey does is claiming to move that decision.
- Not: an independent hectare count. At 35% of the benchmark uncalibrated, and 5% satellite content once calibrated, this cannot referee a national area figure and should not be offered as a second opinion on one. The independent 10 m map already does that job better and agrees with the survey to 1.4%.
- Yes: harvest timing, with the bias removed. A shape correlation of 0.77 at a fixed one-month offset is a usable early-warning signal for when the peak lands, which is the input procurement scheduling actually needs and the one the survey delivers late. This is the product.
- Yes: a revisit-adequacy map. The single most decision-relevant output of this review is that a 12-day gap costs 72% of the fields. A map of where Java is and is not observed densely enough to monitor is worth having before anyone commissions monitoring, and this pipeline can produce it today.
- Yes: disagreement as a triage list. Cells the independent map calls rice and the radar never sees are either a detector blind spot or a real change in what is grown there. That list is a survey design input, and it is currently discarded.
Where this sits on the FMV axis. FMV's stated priority is closing the gap between accumulating data and acting on it, across Insight, Strategy & Influence, Connection & Engagement and Systems & Learning. As built, this case is an Insight artefact that produces a number nobody can act on. Reframed as an observability assessment — where in Indonesia can rice be monitored from orbit at all, at what revisit, with what lead time, and what would it cost to close the gap — it becomes a Systems & Learning product with a direct decision attached, and one that a statistical agency can adopt without conceding that its own numbers are wrong. That is the version of this case worth taking to BPS.
14 What remains open
- Re-ingest at reduced density, properly. §8's ladder interpolates through surviving looks rather than re-deriving the series from scenes. Re-running the ingest with orbits withheld would give the same curve without the caveat, and would also separate "fewer dates" from "fewer orbits", which the current test conflates.
- A detector that does not need a hard threshold. §11 shows half of all detections sit within a decibel of rejection. A likelihood or shape-matching formulation over the whole cycle, rather than three sequential thresholds, would degrade gracefully with sampling instead of falling off a cliff — and is the change most likely to make the East Java regencies work.
- Estimate and publish the timing bias. One constant, estimated on the calibration years, scored on the hold-out. It converts §12 from a finding into a product.
- Ask BPS for the KSA phase labels. The survey's own vegetatif / generatif / persiapan lahan classification is the exact label set this detector wants, and it is published through the API for one province with almost no rice. A data request costs a letter and would turn every one of these checks into a supervised problem.
- Test the smaller cell against the more frequent look. The case proposes 20–30 m cells as the fix. This review's evidence says revisit binds first. Both are testable on the same six regencies, and the answer decides what a pilot should buy.
15 References and reproducibility
- Nelson, A., Setiyono, T., Rala, A.B. et al. (2014). Towards an Operational SAR-Based Rice Monitoring System in Asia: Examples from 13 Demonstration Sites across Asia in the RIICE Project. Remote Sensing 6(11), 10773–10812. doi:10.3390/rs61110773
- Torbick, N., Chowdhury, D., Salas, W. & Qi, J. (2017). Monitoring Rice Agriculture across Myanmar Using Time Series Sentinel-1 Assisted by Landsat-8 and PALSAR-2. Remote Sensing 9(2), 119. doi:10.3390/rs9020119
- Ginting, F.I., Rudiyanto, Fatchurachman et al. (2025). Open-SEA-Rice-10: high-resolution maps of rice harvested area and cropping intensity in Southeast Asia. Scientific Data 12, 1408. doi:10.1038/s41597-025-05722-1. Dataset: doi:10.5281/zenodo.14627003
- BPS-Statistics Indonesia. Luas Panen Padi, Kerangka Sampel Area (KSA) method from reference year 2018. WebAPI monthly kabupaten tables for Jawa Barat and Jawa Timur; annual for Jawa Tengah. The methodology note attached to the national table states the replacement of the eye-estimate Statistik Pertanian method in BPS's own words.
- Sentinel-1 RTC γ⁰ via Microsoft Planetary Computer (collection
sentinel-1-rtc, CC BY 4.0), produced by Catalyst from ESA Sentinel-1 GRD. Contains modified Copernicus Sentinel data 2022–2026.
No threshold was changed to produce any result here. Data vintage 2026-08-30; radar record 2022-07-01 to 2026-08-31; 4 seasons; benchmark years 2023, 2024, 2025.