Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

kimgerdes/basaa-treebank

Domain:

natural language processing

Record type:

dataset
Creator:
kim
Host:
First glossed, translated and audio-aligned SUD dependency treebank of Basaa (Bàsàá, A43, ISO bas). Annotations CC-BY-SA-4.0; copyrighted audio/PDFs referenced, not bundled. # Basaa (A43) treebank Materials toward **the first glossed, translated and audio-aligned dependency treebank of Basaa** (Bàsàá, ISO 639‑3 `bas`, Glottolog `basa1284`, Guthrie **A43** / group A.40, Cameroon). - 🌳 **Browse & annotate (with playable audio):** the *Bàsàá* project on Arborator-Grew — - 🔊 **Demo / data hub:** (sentence list with in-browser audio players) - The treebank is annotated in **SUD** (Surface-syntactic Universal Dependencies), with a UD copy of the literature sample for comparison. ## What's here | Path | What | |---|---| | `treebank/` | The treebank: **`basaa_audio.arborator.sud.conllu`** (21 audio-aligned sentences), SUD/UD samples, the forced-alignment scripts, and the syntactic analysis + upload notes | | `glossary/` | `basaa_glosses.tsv` (~90 common words, glosses, noun classes); the *child*-paradigm morphology; the transcription & gloss review | | `rules/` | Annotation rules extracted from the literature, by pipeline stage — transcription, glossing, parsing, and flagged contradictions | | `recordings/` | **Metadata only** (manifests, forced-alignment offsets, Praat TextGrids, provenance READMEs). Audio is **not** in this repo — see §Copyright | | `corpora/` | Provenance/README for the text & parallel data (OPUS); raw dumps not bundled | | `docs/` | `SOURCES.md` + `resources.md` (annotated bibliography). The PDFs themselves are **not** redistributed | | `TREEBANK_PLAN.md` | Gap analysis & roadmap to scale beyond this seed | | `treebank/index.html` | The demo page deployed at elizia.net | ## Status - **21 audio-aligned sentences**: 10 isolated "child" utterances (clean — transcription, Hamlaoui-style glosses, SUD parse, **word-level** forced alignment) + 11 *North Wind & the Sun* clause units (sentence-level timing reliable; word forms & parse **draft**, to re-transcribe by ear). - Audio time-aligned with **torchaudio's MMS** multilingual forced aligner; each token is clickable in Arborator-Grew. - A separate 5-sentence **lite …

Visit

github.com

Tasks

parsing

Languages

Basaa

Licenses

CC-BY-SA-4.0

Similar

Test treebank for the LFG/XLE treebankBasaa-ALCAM-MultimodalDatasetLexique en basaaInteractions en basaaLeMisterIA/basaa-modelsCommon Voice Basaa

Test treebank for the LFG/XLE treebank

A selection of 828 Tswana sentences with their LFG/XLE parse trees

Basaa-ALCAM-MultimodalDataset

This dataset comprises a datasheet of lexical entries in Basaa, accompanied by illustrative sentence

Lexique en basaa

Enregistrement de lexique en basaa

Interactions en basaa

Enregistrement des interactions de base en basaa Conforms to: doi:10.34847/cocoon.49aefa90-8c1f-3ba8

LeMisterIA/basaa-models

Common Voice Basaa

Voice data collection and distribution Interface for Basaa language using Mozilla's Common Voice infrastructure.