Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kabyle Standardized Named Entities Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
bof
Host:
This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments. kabyle_standardized: Target entity string conforming to standardized orthographic regulations. english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns.

Visit

huggingface.co

Tasks

named entity recognitioninformation extraction

Languages

AmazighBerber

Tags

kabyleberbertamazightnamed-entity-recognitionlow-resource-nlp

Licenses

cc-by-4.0

Similar

NERDz: A Preliminary Dataset of Named Entities for AlgerianGhana Named EntitiesGhanaNLP/Ghana-Named-EntitiesGhana Named Entities TTS (Twi)ALIGNMENT OF BILINGUAL NAMED ENTITIES IN FRENCH -ARABIC PARALLEL CORPORANeurals Networks for Projecting Named Entities from English to Ewondo

NERDz: A Preliminary Dataset of Named Entities for Algerian

This paper introduces a first step towards creating the NERDz dataset. A manually annotated dataset

Ghana Named Entities

A curated dataset of named entities extracted from Ghanaian news sources, compiled by the Ghana NLP

GhanaNLP/Ghana-Named-Entities

List of Named Entities in Ghana # 🇬🇭 Ghana Named Entities A curated dataset of named entities extra

Ghana Named Entities TTS (Twi)

A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organ

ALIGNMENT OF BILINGUAL NAMED ENTITIES IN FRENCH -ARABIC PARALLEL CORPORA

International audience Researches in the field of Named Entity recognition and alignm

Neurals Networks for Projecting Named Entities from English to Ewondo

Named entity recognition is an important task in natural language processing. It is very well studied for rich language, but still under explored for low-resource languages. The main reason is that the existing techniques required a lot of annotated data to reach g