Abstract / Summary
Governments increasingly monitor social media to detect emerging crises around large public programmes, yet discourse-based early warning systems (EWS) are usually evaluated on classification accuracy rather than on the validity of their alarms. Following design science research, we built and retrospectively evaluated an EWS for discourse on Indonesia's Makan Bergizi Gratis (MBG) programme on X. The system combines IndoBERT emotion labels, sarcasm labels, a tiered crisis lexicon with emoji markers, and daily volume into a transparent 0–100 index, detects anomalies with a causal seven-day baseline, and dispatches alerts through the Telegram Bot API with de-duplication. The evaluation used 3,389 deduplicated posts collected between 2 March and 24 May 2026. Only 1,134 posts (33.5%) passed a relevance gate. Without the gate, eight of ten alarm days—including the largest spike, a wave of foreign-language e-commerce spam—were dominated by off-topic content. With the gate, alarms concentrated on three episodes, including a food-safety episode involving school kitchens. We derive four design principles and a signal-validity framework spanning collection, temporal, affective, and actionable validity.