Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BVLAC corpus - Extracted Data Corpus BVLAC - Données extraites

Type de record:

dataset
Créateur:
Fiz
Éditeur:
co coFiz
Éditeur:
CIR
Hôte:avatar
[FR] Dans le cadre du projet SONGES sur la mise en correspondance de données textuelles massives et hétérogènes, nous élaborons des modèles de représentation de données ainsi que des mesures de similarité à partir d’indicateurs trouvés dans les textes (thématiques, spatiaux et temporels). L’objectif est d’organiser et valoriser des ensembles de données dans leurs dimensions hétérogènes et massives. Parmi les données exploitées, nous travaillons sur un ensemble de données produites dans le cadre du projet BVLAC, un projet mené par le CIRAD qui promeut des techniques agricoles issues de l’agroécologie à Madagascar.
Ce dépôt rassemble les données brutes extraites à partir du corpus BVLAC. Les données contenus dans l'archive sont :
  • Les données pour chaque document (*.txt)
  • dans "association‧json" : le nom des fichiers originaux pour chaque identifiant
  • dans "association_lang‧json" : langue utilisée dans chaque document

[EN] As part of the SONGES project on the matching of massive and heterogeneous textual data, we are developing data representation models and similarity measures based on indicators found in the texts (thematic, spatial and temporal). The objective is to organize and valorize data sets in their heterogeneous and massive dimensions. Among the data used, we are working on a dataset produced as part of the BVLAC project, a project led by CIRAD that promotes agricultural techniques derived from agroecology in Madagascar.
This repository contains the raw data extracted from the BVLAC corpus. The data contained in the archive are:
  • Data for each document (*.txt)
  • in "association‧json" : the name of the original files for each identifier
  • in "association_lang‧json" : language used in each document

Visit

doi.orgdataverse.cirad.fr

Tags

Agricultural SciencesComputer and Information ScienceEarth and Environmental Sciencesdonnéesdataagroécologieagroecologyaggregate datadonnée agrégéefouille de données+4

Licenses

info:eu-repo/semantics/restrictedAccessCustom terms specific to this datasethttps://dataverse.cirad.fr/api/datasets/:persistentId/versions/1.1/customlicense?persistentId=doi:10.18167/DVN1/8LIG1D

Similaires

DziriOFN Corpus (Dziri Offensive corpus) v1.0The MultiplEYE Text Corpus Data and MaterialsTravail de corpus et modélisation des données : le cas des constructions « distributives » en par AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-CorpusThe SAWA Corpus: a Parallel Corpus English - SwahiliAfroMAFT Corpus: Language Adaptation Corpus for African languages

DziriOFN Corpus (Dziri Offensive corpus) v1.0

Dziri refers to the name of the Algerian dialectal Arabic. DziriOFN is a new corpus dedicated to offensive language detection on this under-resourced language. Dziri dialect if known as a complex socio-linguistic situation, where the latter is known by the code-

The MultiplEYE Text Corpus Data and Materials

Data and materials for the 39 language versions of the MultiplEYE Text Corpus pertaining to Kaspere,

Travail de corpus et modélisation des données : le cas des constructions « distributives » en par

Communication affichée, présentée le 17 juin 2005 au 2ème Colloque Jeunes Chercheurs du Laboratoire

AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-Corpus

Dataset and code for AAAC: Algerian Arabic Adversarial Corpus for dialect-aware hate speech and prom

The SAWA Corpus: a Parallel Corpus English - Swahili

Research in data-driven methods for Machine Translation has greatly benefited from the increasing availability of parallel corpora. Processing the same text in two different languages yields useful information on how words and phrases are translated from a source l

AfroMAFT Corpus: Language Adaptation Corpus for African languages

Language Adaptation Corpus for 17 African languages, English, French, and Arabic.