Logo Lanfrica

maslinych/corbama-ngrams

Domaine:

natural language processing

Type de record:

dataset
Créateur:
mas
Hôte:
Sources for data and paper on positional skipgrams for Bambara Reference Corpus ## Positional skipgrams for Bambara Reference Corpus (Disambiguated subcorpus) This dataset contains linguistically rich n-gram frequency data for Bambara based on the disambiguated part of the Bambara Reference Corpus (corbama-net-tonal). The n-grams in the dataset are *positional skipgrams* that capture information about co-occurrence of lexical items with grammatical categories at various relative positions. These n-grams were constructed with the aim to leverage those types of information that are available in the morphologically annotated corpus of Bambara given the limited amount of textual data. The idea of positional skipgrams is discussed in the accompanying paper, along with the description of methodology and data used for constructing n-grams for Bambara and a brief illustration of how the positional skipgrams data may be employed in corpus-based linguistic research. The paper: * *Kirill Maslinsky*, « Positional skipgrams for Bambara: a resource for corpus-based studies », Mandenkan 62 | 2019. Please cite the accompanying paper if you use this dataset in research. This dataset is made available under the Open Data Commons Attribution License: The current version of the dataset is available at: The source code for the dataset generation, and for the paper is available at: ### Files included in the dataset This version of the data was produced from the corpus (corbama-net-tonal) on March, 4, 2020. For the convenience of dataset users, the skipgram frequency data is presented in several variants. First, the data is split according to the basic lexical item used for building skipgrams that is either an orthographically normalized wordform, or a canonical lemma. Second, frequency data on both wordfrom-based and lemma-based skipgrams are presented in two forms: an aggregated variant showing total counts for a whole corpus, and a disaggregated variant showing document-level frequencies. * `corbama-net-tonal.lemma.tsv` — lemma-based skipgrams, count …