Logo Lanfrica

Digital Vitality Index for Indonesian Regional Languages on Twitter: A Lexicon-Based Confirmed-Tweet Measurement Across 32 Cities

Domain:

natural language processing

Record type:

dataset
Creator:
Ami
Publisher:
Zenodo
Host:avatar
Per-language confirmed-tweet counts and rates for ten Indonesian regional languages, measured over a general-topic (non-hate-filtered) Indonesian Twitter corpus of 1,419,641 cleaned, deduplicated tweets collected from 32 Indonesian provincial capital cities (2017–2020). Method. A tweet is counted as "confirmed" for a language if it contains two or more lexically distinctive words or particles from a curated per-language lexicon. Words that are also common in standard or informal Indonesian are excluded from every lexicon, so that shared slang and widely borrowed vocabulary do not generate false positives. The two-marker threshold is deliberately conservative. The confirmed rate is 100 × confirmed_tweets / 1,419,641. Purpose. These figures are cited as an independent scarcity anchor in the paper "Diagnosing a Register-Pragmatic Blind Spot in Javanese Hate Speech Detection via LLM-Generated Register-Stratified Stimuli" (Amien, Kanthi, Sijabat & Yusuf), where they appear as Table 1. This record is deposited so that the cited measurement is publicly verifiable rather than an unpublished personal communication. It is a companion measurement: the paper's primary evidence comes from its own labeling run, and this measurement corroborates that evidence from an independent direction. Contents. dvi_table1_snapshot.csv — ten rows, four columns: language, estimated speaker population (millions), confirmed tweet count, and confirmed rate as a percentage. README.md — full method statement, column definitions, interpretation notes, and limitations. Interpretation. These are lower-bound presence estimates, not speaker or usage statistics. The measure captures written, public language use on one platform in one period, and should not be read as a measure of a language's overall vitality, which depends on spoken and domestic domains this corpus cannot observe. The cross-language comparison is meaningful because the same corpus, threshold, and procedure were applied to all ten languages. Scope note. This deposit contains aggregate per-language counts only. It contains no tweet text, no usernames, no user identifiers, no geographic coordinates, and no personally identifiable information.

Similar