Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Large-scale Investigation of French Interrogative Structures Variation on Twitter: a Pilot Study

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
Thi
Éditeur:
IntUniLesANR
Éditeur:
CCSD
Hôte:avatar
International audience This paper presents the foundation of a work which aims at studying direct interrogative structures in French varieties on Twitter in the European and African French-speaking areas within the framework of variationist sociolinguistics. French speakers have a wide range of interrogative structures for producing questions: we can count up to four different structures for Yes-no questions and more than fifteen structures for Wh- questions depending on the presence, the position and the combination of the different elements which compose them. Numerous studies within the framework of variationist sociolinguistics have shown that French interrogative structures are distributed and correlate differently with several variation dimensions such as diatopy (for instance concerning the frequency distribution of Yes-no question vs Wh-question between spoken metropolitan, Belgian and Canadian French (De Cat, 2007)), diastraty (for instance between colloquial bourgeois and worker French (Behnstedt, 1973)), diaphasy (for instance concerning the distribution of formal and less formal questions in SMS (Cougnon, 2015) or more broadly across registers (Coveney, 2011)), and diachrony (Druetta, 2011). The lack of agreement of structures distribution calls into question the scope and the validity of the results which often rely on small datasets which cannot permit to isolate one dimension from another (for instance diatopy from diastraty) or lacks from methodological rigor (for instance extralinguistic factors not distinguished from linguistic ones) was pointed out by Boutin & Rossi-Gensane (2015), Guryev & Delafontaine (2015) or Dagnac (2017). Taking Twitter as a corpus offers a promising new ground for a robust sociolinguistic interrogative structures study based on a large-scale research survey. Under this scope this paper investigates how we can automatically extract interrogative structures from tweets users contents and how we can rely on sociodemographic users information and topologic properties of networks in order to explore how interrogative structures distribution produced by European and African French speakers on Twitter correlate with sociolinguistic and/or linguistic dimensions.

Visit

hal.science

Tags

TwitterNetwork scienceInterrogative structuresVariationist linguisticsComputational sociolinguistics[SHS.LANGUE]Humanities and Social Sciences/Linguistics[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Similaires

Variation across Regions and Demographics in African American Language Morphosyntax: Evidence from Large-Scale Twitter DataRobustness and processing difficulty models. A pilot study for eye-tracking data on the French Treebanknaija-twitter-sentiment-afriberta-largeThe translation of the Vertigo Symptom Scale into Afrikaans: A pilot studyData from: Large-scale variation in biodiversity–ecosystem functioning (BEF) relationships in aquatic metacommunities on terrestrial islandsLow-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Variation across Regions and Demographics in African American Language Morphosyntax: Evidence from Large-Scale Twitter Data

Abstract Social media data and computational tools have become increasingly powe

Robustness and processing difficulty models. A pilot study for eye-tracking data on the French Treebank

International audience We present in this paper a robust method for predicting readin

naija-twitter-sentiment-afriberta-large

naija-twitter-sentiment-afriberta-large is the first multilingual twitter sentiment classification model for four (4) Nigerian languages (Hausa, Igbo, Nigerian Pidgin, and Yorùbá) based on a fine-tuned castorini/afriberta_large large model. It achieves the state-of

The translation of the Vertigo Symptom Scale into Afrikaans: A pilot study

Vertigo is a common clinical problem that is challenging to diagnose and treat. While it has a broad

Data from: Large-scale variation in biodiversity–ecosystem functioning (BEF) relationships in aquatic metacommunities on terrestrial islands

Recent work has shown that the biodiversity of potential colonists in a landscape (the local specie

Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain