Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ALIGNMENT OF BILINGUAL NAMED ENTITIES IN FRENCH -ARABIC PARALLEL CORPORA

Domain:

natural language processing

Record type:

paper
Creator:
AbdKra
Editor:
LIn
Publisher:
CCSD
Host:avatar
International audience Researches in the field of Named Entity recognition and alignment are of strong interest for various applications of natural language processing, such as Cross Lingual Information Retrieval, document management, question-answering systems, data mining etc. But in the processing of Arabic language, the task is particularly difficult and few resources are available to cope with these difficulties. In this paper, we present a simple method of character transcoding - a kind of transliteration that we call character reduction - which could improve an aligning system for Named Entities such as anthroponyms and toponyms. This system has been applied and evaluated on a French-Arabic parallel corpus that has been used during the Arcade 2 evaluation campaign. The purpose of this method is to bring the graphic forms of both languages close together as much as possible, in order to increase aligning precision. An outcome of such aligning is the ability to project on the target language (Arabic) annotations that has been done on the source language, for which more tools and resources are available (French, English, etc.).

Visit

hal.science

Tasks

named entity recognitioninformation extraction

Tags

Natural Language ProcessingNamed Entity RecognitionArabic[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

Ghana Named EntitiesImpact of Parallel vs. Non Parallel Corpora on the Identification of Arabic DialectsGhanaNLP/Ghana-Named-EntitiesCreating a reusable English – Afrikaans parallel corpora for bilingual dictionary constructionLearning to Spot Signs from Named Entities. A study on French Sign LanguageExploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment

Ghana Named Entities

A curated dataset of named entities extracted from Ghanaian news sources, compiled by the Ghana NLP

Impact of Parallel vs. Non Parallel Corpora on the Identification of Arabic Dialects

In this paper, we conduct a study to evaluate the performance of statistical and neural methods to c

GhanaNLP/Ghana-Named-Entities

List of Named Entities in Ghana # 🇬🇭 Ghana Named Entities A curated dataset of named entities extra

Creating a reusable English – Afrikaans parallel corpora for bilingual dictionary construction

This paper investigates the possibilities in creating a bilingual English – Afrikaans dictionary by building a parallel corpus and using the Uplug tool to process it. The resulting parallel corpus with approximately 400,000 words per language was created partly fro

Learning to Spot Signs from Named Entities. A study on French Sign Language

International audience

French Sign Language (LSF) is a low-resourced language

Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment

Multilingual sentence representations pose a great advantage for low-resource languages that do not