Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Copyright in the context of tooling up Corsican and other less-resourced languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
KevRet
Éditeur:
Lab
Éditeur:
CCSD
Hôte:avatar
International audience Anyone trying to gather linguistic resources for Natural Language Processing (NLP) will sooner or later be facing the legal aspects, mainly related to copyright, that arise from this activity. These difficulties often occur when collecting corpora, which is generally among the top priorities for processing less-resourced languages. While the current legislative framework is not adequate, it seems that positive developments are emerging. Various actions can also be considered to support this evolution. Toute personne qui essaye de rassembler des ressources linguistiques pour le Traitement Automatique d'une Langue (TAL) sera tôt ou tard confrontée aux aspects légaux, principalement liés au droit d'auteur, que soulève cette activité. Ces difficultés se matérialisent souvent lors de la collecte de corpus, qui se situe généralement parmi les premières priorités pour le traitement des langues peu dotées. Si le cadre législatif actuel n'est effectivement pas adapté, il semble que desévolutions positives se profilent. Différentes actions peuvent aussiêtre envisagées pour accompagner ce changement.

Visit

hal.science

Tags

Corsican languagecopyrightlinguistic resourcescorporaless-resourced languages[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][SHS.LANGUE]Humanities and Social Sciences/Linguistics[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[SHS.DROIT]Humanities and Social Sciences/Law

Licenses

info:eu-repo/semantics/OpenAccess

Similaires

LR-Sum: Summarization for Less-Resourced LanguagesSAGrid-2.0 - Tooling up for African CollaborationTowards a Blended Programme for Arabic and Other Less Commonly Taught Languages (LCTLs) in the South African Higher Education ContextEnd-To-End Multilingual Automatic Speech Recognition For Less-Resourced Languages: The Case Of Four Ethiopian LanguagesToward Computational Processing of Less Resourced Languages: Primarily Experiments for Moroccan Amazigh LanguageOperation LiLi: Using Crowd-Sourced Data and Automatic Alignment to Investigate the Phonetics and Phonology of Less-Resourced Languages

LR-Sum: Summarization for Less-Resourced Languages

This preprint describes work in progress on LR-Sum, a new permissively-licensed dataset created with

SAGrid-2.0 - Tooling up for African Collaboration

The South African National Grid (SAGrid) is a federation of universities, national laborato

Towards a Blended Programme for Arabic and Other Less Commonly Taught Languages (LCTLs) in the South African Higher Education Context

Disruptive technologies are widely used in education today. They aim to develop the knowledge, skill

End-To-End Multilingual Automatic Speech Recognition For Less-Resourced Languages: The Case Of Four Ethiopian Languages

Presenter: Solomon Teferra Abate, Martha Yifiru Tachbelie, Tanja Schultz , ICASSP 20

Toward Computational Processing of Less Resourced Languages: Primarily Experiments for Moroccan Amazigh Language

Operation LiLi: Using Crowd-Sourced Data and Automatic Alignment to Investigate the Phonetics and Phonology of Less-Resourced Languages

International audience Less-resourced languages are usually left out of phonetic stud