← The instrument REVIEW ARTICLE · DATA VINTAGE 2026-08-30
METHODS & VALIDATION · INDONESIA IN THE GLOBAL NARRATIVE

Whose attention? Measuring Indonesia in the world's news

A stress test of a GDELT-derived attention, tone and theme index over 3,529 days. The measured signal survives; three of the four things it is most naturally read to mean do not. We then rebuild the headline test from the case's own archive and show that it changes five of nine verdicts.

Abstract

Question. A country-level media index built on GDELT is routinely read as "how much the world is paying attention to Indonesia, and how warmly." We ask whether the published series can carry that reading.

Method. Two independent layers over one window, 2017-01-01 to 2026-08-30: the GDELT DOC 2.0 API's normalized volume and tone curves, and 8,051,627 CAMEO-coded Indonesia events extracted from 658,416 raw 15-minute export files (49.58 GB). Every headline statistic is re-derived from the case's own published output; the publisher layer is re-derived from the raw archive.

Findings. First, the denominator is doing most of the work. GDELT's crawl fell from 627,594 to 142,709 articles a day (−77.3%); articles naming Indonesia fell −70.8%; the published share therefore rose 28.9%. 60% of that rise (52%–67%) is predicted by crawl shrinkage alone. Second, spikes are partly a news hole: 24 of the 30 loudest days fall on a Saturday, Sunday or Monday against 12.9 expected, and in the median attention spike 34% of the move is a smaller world crawl rather than more Indonesia coverage. The Kanjuruhan stadium disaster passes the attention gate at 1.79× on a numerator that moved 1.03×. Third, the "global" narrative is mostly domestic: 23 of the 25 largest publishers are based in Indonesia, Indonesian outlets carry 63% of the archive in 2017 rising to 76% by 2026, and the state wire alone is 10.3% of everything. Fourth, rebuilding the anchor test on foreign publishers only changes five of nine verdicts: the disaster that passes turns into a 0.75× fall, while three events the case records as failing show foreign spikes of 2.0–3.1×. Fifth, the tone rise is two-thirds real: an exact decomposition splits the +0.90-point warming into +0.60 from the same outlets and +0.30 from a change in which outlets GDELT reads.

Conclusion. The instrument measures what is being said about Indonesia and by whom well, and how much the world is looking badly. Its level series should be withdrawn as a measure of global attention and republished as two: a composition series, which is sound, and a foreign-attention series, which the archive can already produce and which disagrees with the published one in both directions.

1 The claim under test

The instrument's front page opens with a sentence that contains three separate claims: GDELT machine-reads the world's news; this is the share of all of it that mentioned Indonesia; and the flares are the days the world could not look away. The first is a claim about the sample, the second about the measure, the third about what the measure means.

Each is testable, and they fail in different ways. This review takes them one at a time and then rebuilds the headline test with the correction applied — because the archive needed to do that is already on disk.

A share has a numerator and a denominator. When a series is presented as attention, the reader assumes the numerator moved. Over this decade, mostly, the denominator did.

A note on tense. This review was written against the case page as it stood on 2026-08-30, and it quotes that page throughout. The corrections it asks for have since been applied, so a reader following the link back will find the page already changed — the hero rewritten, chapter 01 renamed and carrying both halves of its ratio, chapter 02's headline negative result corrected, every theme panel labelled with its window, and the normalization claim in the methodology withdrawn. The quotations are left as they were: a review that edits the thing it is reviewing is not a review.

2 Prior art, and what it says this instrument can bear

Leetaru & Schrodt (2013) introduced GDELT as a machine-coded event archive over global news; the DOC 2.0 API added the normalized volume, tone and facet timelines this case uses. Its scale is unmatched and its cost is zero, which is exactly why its measurement properties deserve care. There is, as far as we can find, no published longitudinal GDELT study of Indonesia's media profile — so there is no direct replication target, and the benchmarks below have to be assembled from four separate literatures.

How accurate is the event layer? Wang et al. (2016) audited GDELT's protest records and found that after de-duplication and classification only 21% of valid protest URLs indicate a true protest event; of the filtered, de-duplicated set, 49.5% referred to actual protests. They report the same problems across the other nineteen CAMEO categories. Hong et al. (2025), hand-auditing 2021 records, put key-field accuracy near 55% and semantic redundancy at 20.6% — and, directly relevant here, event-type accuracy at 0.532 for non-English sources against 0.562 for English. Since 66% of this case's kept rows come from the translated feed, it sits on the weaker side of that gap.

How well does GDELT agree with anything else? Badly, and the pattern is instructive.

inside this case — same crawl GDELT vs an independent dataset GDELT tone vs a better model the case's own 0.60 bar Event layer vs API volume, all this case · Spearman, daily 0.64…translated feed only this case 0.72…English feed only this case 0.21GDELT vs ACLED, national series Hammond & Weidmann 2014 0.64GDELT vs ICEWS, cross-national Wang et al. 2016 0.32GDELT vs human-coded protests Wang et al. 2016 0.22GDELT vs ACLED, 55 km cell-month Hammond & Weidmann 2014 0.26GDELT tone vs neural, daily mean GDELT, 2019 · best of seven 0.82GDELT tone vs neural, per article GDELT, 2019 · best of seven 0.50
Figure 1. Agreement coefficients. The top three are this case's own two-source coherence check; the rest are published benchmarks. The comparison is indicative, not exact — the case's are Spearman rank correlations on daily series, the published ones Pearson correlations on counts — but the ordering is the point.

Two things follow. First, scale rescues GDELT: Hammond & Weidmann (2014) find GDELT correlates with ACLED at 0.64 as a national time series but only 0.26 at 55 km cell-months, and Wang et al. note agreement improves as time scales get rougher. A national, daily-or-coarser Indonesia series — which is what this case builds — is the regime where GDELT is defensible. Second, the case's own coherence check is not an external validation. Its ρ = 0.635 compares the event export against the DOC API, both drawn from the same crawl of the same articles; that is an internal-consistency test, and it clears its own 0.60 bar by 0.035. Two genuinely independent pipelines aimed at the same events agree at 0.32.

Is volume a measure of what happened? No. Weidmann (2016) found that only 28.5% of Afghan insurgent events with casualties reached the international press, with reporting probability rising in mobile coverage (β = 0.419) and falling in distance from a town (β = −0.407). Kwak & An (2014), on 195,513 GDELT-coded disasters, found the median disaster's coverage reaches 1 country, and that wire-service pickup alone adds 18% of explained variance — more than every national and disaster attribute combined.

Is the level informative at all? Mostly it is structural. Wu (2000) models coverage of 38 countries and reaches adjusted R² = 0.373 for Indonesia from trade volume and news-agency presence alone; Blondheim et al. (2015) find a country's share of world economic news correlates r = 0.895 with its GDP. The informative quantity is the residual from that baseline, not the level. Grasland (2020), from an entirely independent sample of 320,000 stories in 31 dailies in 2015, puts Indonesia at 1.12% of international news, rank 26. This case measures a decade band of 0.86–1.40% of all monitored articles. The denominators differ, so this is an order-of-magnitude check rather than a replication — but it is the one available, and the case passes it.

And the interpretation? The benchmark that matters most is Eisensee & Strömberg (2007). Studying US disaster relief, they showed that coverage depends on what else is in the news that day — a disaster occurring during the Olympics needs 3× the casualties for the same chance of relief — and that relief then follows coverage. Their Table IX prices the geography of attention directly: an Asian disaster needs 43× the casualties of a European one for equal coverage (13% of Asian disasters make the news against 18% of European ones). Their crowd-out mechanism is precisely the one that appears in §5 of this review, running the other way: an empty agenda inflates a coverage share exactly as a crowded one deflates it.

The literature's verdict on this method is not that it is unusable. It is that national, coarse-grained, relative claims survive and everything finer does not — and that a coverage share is a measure of the news agenda at least as much as of its subject.

3 Data and method

Curves. GDELT DOC 2.0 API, 161 cached timelines: TimelineVol (pre-normalized share of monitored articles), TimelineVolRaw (raw counts plus the monitored total), TimelineTone, TimelineSourceCountry, TimelineLang. Daily resolution, 2017-01-01 → 2026-08-30. The window's start is not an analytical choice: the DOC 2.0 API searches back to 1 January 2017 and no further, so the sample begins where the instrument does.

Events. The raw 15-minute export feed, English and translated, streamed in full: 658,416 files, 49.58 GB, 707,272,763 rows scanned, 8,051,627 kept where Indonesia is the action location (FIPS ID) or an actor (ISO3 IDN). Every kept row carries a distinct event id, so nothing is double-counted at the identifier level. That is a weaker guarantee than it sounds: Hong et al. (2025) measure 20.6% semantic redundancy in GDELT — the same event re-emitted under a new id — which this check cannot detect and which no analysis here corrects for.

Publishers. For this review each row's source domain is collapsed to its site (nasional.kompas.comkompas.com) and classified as Indonesia-based or not, using the .id country domain plus a published list of 98 Indonesian outlets curated from the busiest sites in the archive. The list covers 68% of all rows on its own; everything not on it and not .id is counted as foreign, which biases the foreign share up and therefore makes §6's finding conservative.

Published state of the pipeline. The case's own gates stand at PASS on normalization, PENDING on the anchor events (8 pass, 0 fail, 1 pending of 9), PASS on two-source coherence, and PENDING on language honesty — the foreign/domestic curves the API keeps refusing have not landed. 11 of 20 theme curves are in hand. This review is written against that state, not a hoped-for one.

4 Finding one — the attention series is mostly its denominator

GDELT's crawl is not a fixed instrument. Across this window the number of articles it monitors per day fell from 627,594 to 142,709 — a factor of 4.4 — and it fell in almost every year. The case knows this: its normalization gate publishes the drift and states the rule that every attention chart is a share. The question no chart on the page asks is what that rule does to the trend.

Each series indexed to its own 2017 level = 100 0 25 50 75 100 125 1501719212325 23all articles monitored 29articles naming Indonesia 129the share the page charts
Figure 2. The three series, each indexed to its own 2017 level. The crawl and the Indonesia count fall together; the ratio between them rises. Nothing here is a claim about the world's interest in Indonesia.

The numbers are stark. Articles naming Indonesia fell −70.8%, from 5,379 a day to 1,572. Everything GDELT monitored fell −77.3%. Because the second fell faster than the first, the published share rose 28.9%. The case's own normalization check already records that the raw and normalized series are effectively unrelated day to day (Spearman ρ = 0.0836); what it does not say is which of them the page then plots as a decade of attention.

no change Articles naming Indonesia per day, GDELT's own count −70.8%All articles GDELT monitored per day — the denominator −77.3%The share the dashboard charts the first divided by the second +28.9% Share of that rise predicted by the shrinking crawl alone 60% (52–67%)
Figure 3. The same three quantities as percentage changes over the window, and the share of the published rise that crawl shrinkage alone predicts.

Normalizing is the right instinct, and at daily frequency it works: regressed on crawl size, the share has an elasticity of only -0.102 (R² 0.066), against 0.898 (R² 0.845) for the raw count — dividing by the denominator removes most of the mechanical co-movement. But a small elasticity applied to a very large change in the denominator is not small. The crawl's log change over the window is -1.48; multiplied by that elasticity it accounts for 60% of the observed log rise in the share (52%–67% at 95%).

This propagates into the page's first chapter. The five "regimes of attention" read as a rise to a plateau; underneath them, the number of articles about Indonesia falls in every regime:

EraPublished shareIndonesia articles/dayMonitored articles/dayMean tone
Pre-pandemic1.00%5,304541,049−0.29
Pandemic1.28%4,460353,567−0.58
G20 & ASEAN chair1.31%3,013234,798+0.22
Election year1.11%1,808163,328+0.67
Prabowo era & 2025 protests1.08%1,667153,557+0.45

What this does and does not overturn. The share is the right thing to plot for comparing one day with the next. It is the wrong thing to read as a decade-long trend in the world's interest, because over a decade the denominator changed by more than the numerator did. Neither series is a clean attention measure on its own; §6 and §7 build one that is closer.

5 Finding two — the loudest days are partly a news hole

If a share can drift because its denominator drifts, it can also spike because its denominator dips. The world's news volume is strongly periodic: in this archive GDELT monitors 223,381 articles on an average Sunday against 389,375 on a Wednesday, a fall of 43%. Coverage of Indonesia falls on weekends too, which is why the weekly cycle mostly cancels in the share. Individual days are another matter.

0.7×0.8×1.0×1.5×2.2×3.3×1.0×1.3×1.8×2.5×3.3× 2020-11-15 · share 2.20% Indonesia coverage 0.40× · crawl 3.4× smaller ↑ world crawl smaller than usual · → Indonesia covered more than usual · dotted: equal published ratio Indonesia coverage vs its own trailing median → Sat / Sun / Mon (24 of 30) Tue–Fri (6)
Figure 4. The 30 highest-share days in the window. Horizontal axis: coverage of Indonesia against its own trailing 28-day median. Vertical axis: how much smaller the world crawl was than its own trailing median. Dashed diagonals mark equal published attention ratios. Hover any day.

24 of the 30 loudest days in the decade fall on a Saturday, Sunday or Monday, against 12.9 expected if loud days were spread evenly. Across every day whose share reaches 1.5× its trailing median (67 days, 1.9% of the window), the median spike is 34% attributable to the crawl being smaller rather than to Indonesia being covered more; on 7 of them coverage of Indonesia fell while the published share rose.

The same split can be applied to the case's own pre-registered anchor events, which is where it stops being a curiosity.

more coverage of Indonesia a smaller world news crawl published Palu earthquake & tsunami 2018-09-30 · Sun 3.07×Anak Krakatau tsunami 2018-12-23 · Sun 4.07×post-election riots, Jakarta 2019-05-22 · Wed 1.24×first COVID-19 cases announced 2020-03-02 · Mon 1.51×Kanjuruhan stadium disaster 2022-10-02 · Sun 1.79×G20 Bali summit 2022-11-15 · Tue 3.15×presidential election 2024-02-14 · Wed 1.48×PDNS ransomware attack 2024-06-23 · Sun 1.26×Aug–Sep 2025 protests 2025-08-31 · Sun 1.30× baseline log contribution to the published attention ratio →
Figure 5. Each anchor's published attention ratio decomposed into the part contributed by more coverage of Indonesia (violet) and the part contributed by a smaller world crawl (amber). Bars extending left are contributions that worked against the ratio.

Two anchors are clean: the Anak Krakatau tsunami and the G20 Bali summit are carried by real increases in coverage (2.20× and 3.36×). One is not. The Kanjuruhan stadium disaster passes the attention check at 1.79× on coverage that moved 1.03× — the peak fell on a Sun, the world crawl was 1.82× smaller than its own baseline, and 95% of the published ratio is that. The two most recent anchors invert it: for the Aug–Sep 2025 protests the published ratio is 1.30× while coverage of Indonesia fell to 0.77× of baseline.

This is not a threshold problem. Moving the 1.5× bar would change which anchors pass without changing what the number means. The defect is in the measure: a share whose denominator is the volume of unrelated world news cannot separate "Indonesia was covered more" from "everything else was covered less" — which is exactly the crowd-out mechanism Eisensee & Strömberg (2007) identified, running the other way.

6 Finding three — the global narrative is mostly Indonesia's own

The API's curves carry no publisher breakdown, so this question cannot be asked of them. The event archive can answer it directly: every one of the 8,051,596 kept rows carries the domain that published it. Collapsed to sites, the archive holds 22,657 distinct publishers — and is extremely concentrated. The ten largest carry 48.2% of everything.

published in Indonesia published abroad Indonesian state media tribunnews.com 11.98%antaranews.com 9.64%kompas.com 4.72%republika.co.id 4.24%tempo.co 3.87%liputan6.com 3.67%pikiran-rakyat.com 2.92%bisnis.com 2.89%okezone.com 2.36%cnnindonesia.com 1.96%thejakartapost.com 1.90%merdeka.com 1.71%jpnn.com 1.36%sindonews.com 1.36%viva.co.id 1.34%beritasatu.com 0.98%metrotvnews.com 0.94%mediaindonesia.com 0.78%thestar.com.my 0.76%jawapos.com 0.75%beritajatim.com 0.73%rri.co.id 0.70%kontan.co.id 0.61%inilah.com 0.61%msn.com 0.57% share of all 8,051,596 Indonesia events kept →
Figure 6. The 25 largest publishers of Indonesia events, 2017-01-01–2026-08-30. Hover any bar for its mean tone.

23 of the 25 largest publishers are based in Indonesia. The two that are not sit at ranks 19 and 25. The single largest voice in "the world's news about Indonesia" is tribunnews.com at 12.0%; the Indonesian state news agency and state radio together carry 10.3% of the entire archive.

Across the window, Indonesian publishers carry 63.3% of kept rows in 2017, rising to 76.0% in 2026. The API's own facets agree from a different direction: of the languages GDELT reads, Indonesian coverage of Indonesia runs at 33.9 intensity against 0.66 for English, and 65.9% of kept event rows come from the translated feed. The case's two-source coherence check records the same thing without naming it: the API's volume tracks the translated event share at ρ = 0.719 but the English one at only ρ = 0.207.

That composition changes what a rising share can mean. Foreign coverage of Indonesia, measured against the feed's own denominator, moved in the opposite direction to the published series:

Both indexed to 2017 = 100 50 75 100 125 1501719212325 129 the published share 72 coverage published outside Indonesia
Figure 7. The published attention share against foreign-published Indonesia coverage as a share of the whole 15-minute feed, both indexed to 2017. Two measures of the same thing, 29% apart in one direction and −28% in the other.

Over the decade the world's press published 28% less about Indonesia relative to everything else it published. The page's headline series shows the opposite because three-quarters of what it counts is Indonesia talking to itself.

Why this is understated, not overstated. Any site not on the curated Indonesian list and not under .id is counted as foreign, so the foreign share here is an upper bound and the domestic share a lower bound. Note also that a domestic-heavy sample is not a defect of the pipeline — GDELT indexes Indonesian media heavily and that is a real feature of the source. It is a defect of the label.

7 Finding four — the decisive test: re-running the anchors on foreign publishers only

Findings one to three converge on a single experiment. If the published attention measure is contaminated by a drifting denominator and dominated by domestic publishers, then rebuilding it from foreign publishers only, against the feed's own denominator, should change what it says about specific events. The archive needed for that is already on disk, so we ran it: for each pre-registered anchor, the peak of foreign-published Indonesia coverage in the anchor window against its own trailing 28-day median — the identical test statistic the case uses, computed on a cleaner sample.

published, all publishers foreign publishers only 1.5× threshold no change Palu earthquake & tsunami 2018-09-28 4.05×Anak Krakatau tsunami 2018-12-22 4.28×post-election riots, Jakarta 2019-05-22 2.01×first COVID-19 cases announced 2020-03-02 1.36×Kanjuruhan stadium disaster 2022-10-01 0.75×G20 Bali summit 2022-11-15 8.75×presidential election 2024-02-14 2.22×PDNS ransomware attack 2024-06-20 1.09×Aug–Sep 2025 protests 2025-08-28 3.06×
Figure 8. Each anchor's published attention ratio (small grey dot) against the same ratio computed from foreign publishers only (large dot). Teal means the foreign measure is larger, coral smaller.

The result changes the verdict on five of the nine anchors — four of them by a wide margin, one where both measures sit on the threshold — and it changes them in both directions.

AnchorPublishedAll publishersForeign onlyReading
Palu earthquake & tsunami · 2018-09-28 3.07× 1.93× 4.05× agrees
Anak Krakatau tsunami · 2018-12-22 4.07× 1.91× 4.28× agrees
post-election riots, Jakarta · 2019-05-22 1.24× 1.47× 2.01× missed by the published measure
first COVID-19 cases announced · 2020-03-02 1.51× 1.23× 1.36× on the threshold either way
Kanjuruhan stadium disaster · 2022-10-01 1.79× 0.91× 0.75× artefact of the published measure
G20 Bali summit · 2022-11-15 3.15× 3.56× 8.75× agrees
presidential election · 2024-02-14 1.48× 1.17× 2.22× missed by the published measure
PDNS ransomware attack · 2024-06-20 1.26× 1.09× 1.09× agrees
Aug–Sep 2025 protests · 2025-08-28 1.30× 2.13× 3.06× missed by the published measure

The disaster that passes the published attention check is the one event where foreign coverage of Indonesia actually fell — 0.75× its own baseline. Meanwhile three events the case records as failing to lift attention produce large foreign spikes: post-election riots, Jakarta 2.0×, presidential election 2.2×, Aug–Sep 2025 protests 3.1×. The G20 summit, which the published measure scores at 3.15×, reaches 8.75× among foreign publishers — the world did look, far harder than the page reports.

The published instrument does not merely add noise to the world's attention. It systematically mis-ranks events: it credits Indonesia's quiet Sundays and discounts the days the foreign press actually arrived.

What this test is and is not. It counts coded events, not articles, so its ratios are not directly comparable in level with the API's article-share ratios; the comparison that carries the finding is foreign against all publishers within the same measure, which is exactly like-for-like. It also inherits the event layer's known coding error (Wang et al. 2016) — but that error has no reason to differ between domestic and foreign publishers on the same day, which is what the test compares. Monthly, the all-publisher event share tracks the API's volume at r = 0.73 while the foreign-only series tracks it at r = 0.41: the published curve is, quantitatively, the domestic one.

8 Finding five — the warming is two-thirds real

Tone in this dataset rises sharply and monotonically enough to be the page's most eye-catching feature: from −0.44 in 2017 to +0.31 in 2026 on the API curve, and −1.27 to −0.40 on the independent event layer, which tracks it at r = 0.9304. The page attributes it to "a real shift in coverage" while conceding that the shrinking source mix "deserves part of the credit". Both layers also correlate with crawl size at r = -0.6667 and -0.69 respectively, so the concession is not idle — but neither is it quantified, and it is quantifiable.

With a publisher for every row, the change in mean tone decomposes exactly (Griliches & Regev 1995) into four parts: the same outlets writing differently, a reweighting among outlets present in both years, outlets arriving, and outlets dropping out. The four sum to the observed change with no residual.

0 -0.5 -1.0 -1.5 -1.31 2017 tone +0.602 same outlets, warmer copy +0.162 survivors reweighted −0.017 new outlets arriving +0.158 outlets that dropped out -0.40 2026 tone source mix +0.302 · 33% of the move start end
Figure 9. Decomposition of the 2017–2026 change in mean event tone across 1,119,133 and 450,151 rows. The residual is zero by construction.

Mean tone moved +0.905. +0.602 of that (67%) is the same publishers writing more warmly — a real change in what is said about Indonesia. The remaining +0.302 (33%) is a change in who is speaking: +0.162 from reweighting among survivors and +0.158 from outlets leaving the archive altogether. Of the 11,444 sites present in 2017, only 2,998 (26%) are still there in 2026, and the ones that left were harsher than average.

The gap between publisher types is structural, not incidental: Indonesian outlets ran at −0.94 in 2017 against −1.95 for foreign ones, a 1.01-point difference. Any drift toward domestic publishers therefore raises measured tone by itself, and §6 showed that drift is large.

−2 −1 +0 +11719212325antaranews.comstate wire · 9.6% of all rowsbisnis.comrepublika.co.idliputan6.comall publisherstribunnews.comkompas.com mean tone of every Indonesia event a publisher carried that year
Figure 10. The tone of the largest publishers present in every year of the window, against the archive as a whole. The state wire is highlighted.

The within-publisher component is not evenly spread either. The largest single voice in the archive is the state news agency, and its own tone rises from −0.01 in 2017 to +1.02 in 2026 — a publisher-level movement of +1.03 points on 12.0% of all rows, while other large outlets in the same panel are flat. We report this without interpreting it: what a state wire's tone measures is a question about that wire, not about the world.

Scope, and a serious threat to the within component. This decomposition runs on the event layer, which is where publisher identity exists; it cannot be run on the API's tone curve directly. The two move together (r = 0.9304) and by nearly the same amount (+0.76 against +0.86), which is the basis for reading the split across. The harder problem is what tone is. GDELT's score is a dictionary count, and van Atteveldt et al. (2021) found English sentiment dictionaries applied to machine-translated news reach Krippendorff's α of at most 0.34 against human coders — no better than dictionaries in the original language, and far below the ~0.8 a hand-coded benchmark reaches. Worse for a decade-long comparison, GDELT's own documentation states that its translation engine "dynamically modulate[s] the quality of translations" and that accuracy "linearly degraded as needed to cope with increases in volume". 66% of these rows are machine-translated, so if that pipeline improved as ingest volume fell — and volume fell 77% — part of the +0.60 within-publisher warming could be the translator rather than the copy. Nothing on disk separates the two. This is the single largest unresolved threat in this review, and we flag it rather than net it out.

9 Finding six — attention is not sentiment, and the page tests it at the wrong resolution

The page asks a good question — "when it's loud, is it dark?" — and answers it with a monthly scatter, reporting r = +0.14 over 116 months and concluding that the two "barely correlate". That number is right and the conclusion does not follow from it, because the relationship changes sign with the resolution you look at.

Mean tone by attention decile, daily quietest tenth of days → loudest tenth -0.21 quietest loudest Daily correlation of attention with tone, by year negative = louder days are darker monthly r +0.14 — what the page prints 17 18 19 20 21 22 23 24 25 26 0 −0.6
Figure 11. Left: mean tone by decile of daily attention. Right: the daily attention–tone correlation computed separately within each year, against the pooled monthly figure the page reports.

At daily resolution the correlation is −0.136, and once year-to-year drift is removed −0.244, with a slope of -0.61 tone points per point of share. The loudest tenth of days average −0.222 against +0.000 for the rest. In 6 of the 10 years the within-year daily correlation is negative, reaching −0.60.

So loud days are dark and loud months are not, and both are true: days spike on catastrophe, months rise on summits and campaigns. The page half-sees this — it notes that the truly dark moments are single days — but it prints the aggregate correlation as the finding, and the aggregate is the one resolution at which the effect vanishes. The honest headline is that the sign of the attention–tone relationship depends on the timescale, which is a more useful result than either half.

The literature would predict exactly this and warns against collapsing the two into one index. Sheafer (2007) put volume and tone in the same model of what the public calls the most important problem and found them opposite in sign and unequal in size — salience +0.09 against valence −0.23, with tone roughly two and a half times the larger — while Kiousis (2004) found visibility and valence load on separate factors. Volume and tone are two measurements, not two views of one.

Attention is not importance either, and the ranking says so plainly. The single loudest day of the decade is the 2018-12-23 Anak Krakatau tsunami at 4.00% of world coverage; the next three are the G20 Bali summit. A stadium crush that killed 135 people reaches 1.79× and, as §7 showed, none of that was foreign attention. What the series ranks is the coincidence of an event with an empty news agenda.

That coincidence is not a nuisance term. It is the quantity Eisensee & Strömberg (2007) showed decides whether a disaster gets relief — and on their numbers an Indonesian disaster already needs 43 times the casualties of a European one to be covered at all.

Read next to that benchmark, this index is measuring a real and consequential thing: the news agenda's willingness to make room. It is simply not measuring what its label says, and the difference matters most for exactly the events an Indonesian user would care about.

10 What survives, and why it is the more useful half

Three things in this case come through the review intact.

  1. The theme mix. A theme's share is one API share divided by another with the same denominator, so crawl size cancels algebraically. It is the only layer immune to §4 by construction — and it is where the case's substantive results live: COVID at 9.2% of all Indonesia coverage across the window, elections at 6.8%, and an era ranking that reorders itself sensibly around the pandemic, the G20 chairmanship and the election year.
  2. The event typology and the map. Cooperation-versus-conflict mix, protest counts and the self-drawn 0.25° grid are all relative structures within the Indonesia sample. They do not depend on the world denominator at all.
  3. The publisher layer — which the case computes and then does not use. It is the answer to §6 and §7, and on the evidence here it is the most valuable thing in the archive.

Two caveats attach even to the theme layer. Crawl size cancels; crawl composition does not, and §8 showed composition changed a great deal. And the panel is less complete than it looks:

months of data out of 116 in the window COVID-19 116 mo · 9.19%Elections 116 mo · 6.81%Quakes & volcanoes 116 mo · 3.59%Coal & mining 13 mo · 3.53%Protests 104 mo · 3.13%Papua 105 mo · 2.76%Floods 105 mo · 2.59%Terrorism 116 mo · 2.47%Palm oil 116 mo · 0.95%Nickel 116 mo · 0.77%New capital (IKN) 105 mo · 0.08% 9 never landed 0 mo coral = window too short to compare against the others
Figure 12. Months of data behind each theme curve. Coal & mining holds 13 months and is nonetheless ranked and averaged against curves holding 116.

11 of 20 theme queries have landed, and one of those — Coal & mining, 13 months from 2017-01 to 2018-01 — is presented in the theme grid with a decade-mean stamp and ranked against series holding 116 months. Its apparent 3.53% mean is a 13-month figure. Acted on: every panel on the case page now prints its month count, short windows are flagged in the alert colour, and a theme no longer enters an era's ranking unless it covers at least 60% of that era — which removed a one-year curve from a ten-year league table.

One more presentation defect. The "who's talking" heat map ranks publishing countries by the share of their own output that mentions Indonesia. That statistic is not size-controlled, so countries with a tiny GDELT footprint float to the top: the published ranking places North Korea seventh and Sierra Leone tenth, above Thailand. The ordering among the large neighbours (Singapore, Malaysia, Brunei) is informative; the tail is small-sample noise and should be cut or given a minimum-volume floor. Acted on in part: the case page now carries the caveat; the floor itself is still to be implemented.

11 What follows for decisions

An Indonesian institution acting on this dashboard today would be acting on the wrong series. Concretely:

  1. Kementerian Luar Negeri / BKPM — international perception. The published series says global attention to Indonesia rose 29% across the decade. Measured from foreign publishers it fell 28%. A ministry benchmarking its diplomatic communications against the first number would conclude its reach was growing while it was shrinking.
  2. BNPB — disaster communications. §5 and §7 show that whether a disaster registers in this index depends heavily on what day of the week it happened. An agency using the index to judge whether an emergency "broke through" internationally would systematically over-rate weekend events and under-rate weekday ones.
  3. Kominfo / Dewan Pers — media monitoring. The one thing the archive measures unambiguously well is who is publishing. That 48% of Indonesia coverage comes from ten sites and 10.3% from the state wire is a structural fact about the Indonesian information environment, computed from a decade of primary data. Following this review the case page now says so; the table itself still lives only here, and it deserves a chapter of its own.
  4. Not: a global attention index. On this evidence the level series should be withdrawn as a measure of the world's interest. It should be republished as a foreign-publisher series with the feed's own denominator — which §7 shows the pipeline can already produce and which disagrees with the current one on five of nine anchors.

The case's most marketable claim — that this measures the world's attention — is the one its own archive refuses. The claim the archive does support, that it measures who speaks about Indonesia and in what proportions, is the more decision-relevant of the two.

That reframing is also where this case sits on FMV's own axis. See clearly is not a volume chart; it is knowing that 76% of what you are reading is your own reflection. Connect wisely needs the publisher graph, not the pulse. Act with evidence requires a series whose movements survive a change of denominator — and the one built in §7 does, while the published one does not. A media-intelligence product aimed at closing the gap between accumulated data and action should ship §6's publisher concentration and §7's foreign-attention series as its primary screens, and keep the total-volume pulse as a diagnostic behind them.

12 What remains open

  1. The foreign/domestic API curves. The case's language-honesty gate is PENDING because -sourcecountry:ID is the query the API refuses most often. §7's test is a substitute built from the event layer; the direct curve would let the same split be applied to the article-share measure the page actually plots.
  2. A publisher-weighted attention index. §7 counts foreign publishers equally. A reach-weighted version — one that knows the difference between Reuters and a syndication mirror — would be a better instrument still, and the archive already carries the domain, the article count and the mention count needed to build it.
  3. Tone validity on translated text. 66% of kept rows come from GDELT's machine-translated feed, and §8's headline quantity is a dictionary tone score computed downstream of that translation. Whether the +0.60 within-publisher warming partly reflects an improving translation pipeline rather than warmer copy is not separable with what is on disk, and it is the largest unresolved threat to §8.
  4. The nine missing themes. 9 theme curves have never landed, including the one gating the PDNS ransomware attack anchor, which is why the anchor gate stands PENDING rather than resolved.

13 References and reproducibility

  1. Leetaru, K. & Schrodt, P.A. (2013). GDELT: Global Data on Events, Location and Tone, 1979–2012. ISA Annual Convention. Dataset: gdeltproject.org — free for any use with attribution. DOC 2.0 API documentation (20 June 2017) for the normalization rule quoted in §4; GDELT Translingual documentation (19 February 2015) for the translation-quality statement quoted in §8.
  2. Wang, W., Kennedy, R., Lazer, D. & Ramakrishnan, N. (2016). Growing pains for global monitoring of societal events. Science 353(6307), 1502–1503. doi:10.1126/science.aaf6758 — 21% of valid protest URLs are true protests (p. 1503); GDELT vs ICEWS r = 0.317, GDELT vs human-coded ground truth r = 0.222 (p. 1502).
  3. Hong, D., Fu, Z., Zhang, X. & Pan, Y. (2025). Research on the Development and Application of the GDELT Event Database. Data 10(10), 158. doi:10.3390/data10100158 — event-type accuracy 0.532 non-English vs 0.562 English (Table 4); redundancy 0.206 (Table 6).
  4. Hammond, J. & Weidmann, N.B. (2014). Using machine-coded event data for the micro-level study of political violence. Research & Politics 1(2). doi:10.1177/2053168014539924 — GDELT vs ACLED r = 0.64 as a national time series, 0.26 at 55 km cell-months (p. 3).
  5. Weidmann, N.B. (2016). A Closer Look at Reporting Bias in Conflict Event Data. American Journal of Political Science 60(1), 206–218. doi:10.1111/ajps.12196 — 28.5% of events with casualties reach the international press (p. 211); Table 1 for the accessibility coefficients.
  6. Eisensee, T. & Strömberg, D. (2007). News Droughts, News Floods, and U.S. Disaster Relief. Quarterly Journal of Economics 122(2), 693–728. doi:10.1162/qjec.122.2.693 — Olympics 3× casualties for equal relief probability (conclusions); Asia 43× Europe for equal coverage, news shares 0.13 and 0.18 (Table IX).
  7. Kwak, H. & An, J. (2014). A First Look at Global News Coverage of Disasters by Using the GDELT Dataset. Social Informatics, LNCS 8851, 300–308. doi:10.1007/978-3-319-13734-6_22 — median disaster reach 1 country (pp. 303–304); wire-service pickup adds R² 0.18 (Table 1).
  8. Wu, H.D. (2000). Systemic determinants of international news coverage: a comparison of 38 countries. Journal of Communication 50(2), 110–130. doi:10.1111/j.1460-2466.2000.tb02844.x — Indonesia adjusted R² = 0.373, trade β = 0.464 (Table 3, p. 122).
  9. Blondheim, M., Segev, E. & Cabrera, M.-A. (2015). The Prominence of Weak Economies. International Journal of Communication 9, 46–65 — country share of world economic news vs GDP, r = 0.895.
  10. Grasland, C. (2020). International news flows theory revisited through a space-time interaction model. International Communication Gazette. arXiv:1810.04912 — Indonesia 1.12% of international news, rank 26 of all countries, 320,000 stories in 31 dailies, 2015 (Table 3).
  11. Sheafer, T. (2007). How to Evaluate It: The Role of Story-Evaluative Tone in Agenda Setting and Priming. Journal of Communication 57(1), 21–39 — salience 0.09 vs valence -0.23 (Table 1, N = 4,346).
  12. Kiousis, S. (2004). Explicating Media Salience. Journal of Communication 54(1), 71–87. doi:10.1111/j.1460-2466.2004.tb02614.x — visibility and valence emerge as separate dimensions.
  13. van Atteveldt, W., van der Velden, M.A.C.G. & Boukes, M. (2021). The Validity of Sentiment Analysis. Communication Methods and Measures 15(2), 121–140. doi:10.1080/19312458.2020.1869198 — English dictionaries on machine-translated news reach α ≤ 0.34 against a 0.80 human benchmark (Table 2).
  14. Griliches, Z. & Regev, H. (1995). Firm productivity in Israeli industry 1979–1988. Journal of Econometrics 65(1), 175–203. doi:10.1016/0304-4076(94)01601-U — the decomposition used in §8.

Data vintage 2026-08-30. Window 2017-01-01 → 2026-08-30, 3,529 days, 8,051,627 coded events from 658,416 feed files (49.58 GB). The Indonesian-publisher list used in §6 is a fixed, published list of 100 sites plus the .id domain; anything unclassified counts as foreign, so the domestic share is a lower bound. Every threshold was fixed in advance of the result it judges and is unchanged by this review.