Logo Lanfrica

Language density hotspots

Domain:

natural language processinggeospatial

Record type:

datasetsoftwarepaper
Creator:
Ran
Publisher:
Zenodo
Host:avatar

This code takes the location of each language in the world and counts how many other languages are located within a 39.9 km radius of it (that's a circular area of 5000 sq km). The method was inspired by Bromham et al.'s (2022) concept of a language's 'neighbourhood'. Bromham et al. use an area of 10,000 sq km and calculate values for a number of variables in this neighbourhood. As far as I am aware, that article does not report summarised data on language density hotspots, as defined here. 

The input data is a list of languages in Glottolog 5.2 (Hammarström et al. 2025) with their associated coordinates (Glottolog dialects and other linguoids are excluded).

The code also generates heat maps showing the top language density hotspots based on the count of languages, using 20 languages as a cutoff point. This means that only languages which have 20 or more neighbouring languages within their individual 5000 sq km areas are plotted.

The generated world map shows that there are four main areas of highest language density - the north coast of the Papua New Guinea mainland, Vanuatu, Nigeria-Cameroon and southern Mexico - with some subareas in some of these primary areas. 

These results are not suprising. However, to the best of my knowledge, this has not been presented in this way, using this method before. 

The north coast of PNG stands out as the absolute top language density hotpost. This is well known. The density here is also complemented by genealogical diversity. Two main areas emerge - the larger continous area by far is in Madang and Morobe provinces and neighbouring provinces, itself organised around two centres of density. In PNG, there is also a smaller hotspot in Sandaun and East Sepik provinces. 

In Africa, the Grassfields region of Cameroon emerges as the densest area, with a few smaller hotspots in Nigeria and northernmost Cameroon.

In Vanuatu, two subareas emerge - the two largest islands of Espiritu Santo and Malekula. 

In southern Mexico, there is a single smaller area. 

Changing the area and cutoff point parameters can be used to discover more detail. For example, to help identify more precise subareas (especially within the Madang/Morobe hotspot), the area can be decreased and/or the cutoff point increased. To discover more hotspots, the area can be increased and/or the cutoff point decreased.

The language with the most neighbours is Bau [bauu1244], spoken near Madang in PNG, with 50 neighbours.

The language with most neighbours in the Sandaun/East Sepik area is Agi [agii1245] with 38 neighbours. 

In the secondary subarea in the Madang/Morobe large area, located across the border between the two PNG provinces, the language with most neighbours is Bulgebi [bulg1265] with 32 neighbours.

For the Grassfields (Cameroon) area, this is Kom [komc1235] with 32 neighbours.

For Malekula (Vanuatu), this is Aulua [aulu1238] with 30 neighbours.

For Santo (Vanuatu), it is Ande [moro1286] and Kiai [fort1240] with 24 neighbours each.

In the southern Mexico hotspot, this is El Alto Zapotec [elal1235] with 23 neighbours.

Caveat: These languages do not necessarily constitute some sort of "centre of density" for the identified areas. Languages that are spoken close to a sea coast (or large inland lakes, deserts, mountains or other topographic barriers) naturally have fewer neighbours. The heatmaps may be interepreted as showing the centres of language density hotspots (in brighter colours), but this is less accurate for hotposts that neighbour coastal areas, especially the Vanuatu and PNG areas identified here. 

The files in this upload are:

glottolog-lgs-with-coordinates.tsv -  the input file

Language-Density.R - the R code

glottolog-lgs-with-neighbour-counts.tsv - the output file with list of all languages, the number of their neoghbours and other details inherited from Glottlog

language_neighbour_heatmap_world.png - a map of the world with the identified hotspots, little detail

language_neighbour_heatmap_PNG.png - a zoom in map of the hotspots in Papua New Guinea

language_neighbour_heatmap_Africa.png - a zoom in map of the hotspots in Nigeria and Cameroon

language_neighbour_heatmap_Vanuatu.png - a zoom in map of the hotspots in Vanuatu

language_neighbour_heatmap_Mexico.png - a zoom in map of the hotspot in southern Mexico

Informing the method, composing the code and debugging has been partially facilited by using the GPT-4.1 LLM. All remaining errors are my own.

References:

Bromham, L., Dinnage, R., Skirgård, H. et al. Global predictors of language endangerment and the future of linguistic diversity. Nat Ecol Evol 6, 163–173 (2022). doi.org

Hammarström, Harald & Forkel, Robert & Haspelmath, Martin & Bank, Sebastian. 2025. Glottolog 5.2. Leipzig: Max Planck Institute for Evolutionary Anthropology. doi.org