Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Robustness and processing difficulty models. A pilot study for eye-tracking data on the French Treebank

Domain:

natural language processing

Record type:

paperdataset
Creator:
RauBla
Editor:
LabANR
Publisher:
CCSD
Host:avatar
International audience We present in this paper a robust method for predicting reading times. Robustness first comes from the conception of the difficulty model, which is based on a morpho-syntactic surprisal index. This metric is not only a good predictor, as shown in the paper, but also intrinsically robust (because relying on POS-tagging instead of parsing). Second, robustness also concerns data analysis: we propose to enlarge the scope of reading processing units by using syntactic chunks instead of words. As a result, words with null reading time do not need any special treatment or filtering. It appears that working at chunks scale smooths out the variability inherent to the different reader's strategy. The pilot study presented in this paper applies this technique to a new resource we have built, enriching a French treebank with eye-tracking data and difficulty prediction measures.

Visit

hal.science

Tags

Linguistic complexitydifficulty modelsmorpho-syntactic surprisalreading time predictionchunks[SHS.STAT]Humanities and Social Sciences/Methods and statistics

Licenses

info:eu-repo/semantics/OpenAccess