Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TxPI-u: A Resource for Personality Identification of Undergraduates

Domain:

natural language processing

Record type:

paperdataset
Creator:
RamVilJim
Host:avatar
Resources such as labeled corpora are necessary to train automatic models within the natural language processing (NLP) field. Historically, a large number of resources regarding a broad number of problems are available mostly in English. One of such problems is known as Personality Identification where based on a psychological model (e.g. The Big Five Model), the goal is to find the traits of a subject's personality given, for instance, a text written by the same subject. In this paper we introduce a new corpus in Spanish called Texts for Personality Identification (TxPI). This corpus will help to develop models to automatically assign a personality trait to an author of a text document. Our corpus, TxPI-u, contains information of 416 Mexican undergraduate students with some demographics information such as, age, gender, and the academic program they are enrolled. Finally, as an additional contribution, we present a set of baselines to provide a comparison scheme for further research.

Visit

arxiv.org

Tasks

text classification

Tags

Computation and Language

Similar

A word‐level language identification strategy for resource‐scarce languagesGlotLID: Language Identification for Low-Resource LanguagesToluwase/Word-Level-Language-Identification-for-Resource-Scarce-ConLID: Supervised Contrastive Learning for Low-Resource Language Identificationdev52003/Biased-News-Identification-For-Low-Resource-Languages-MalayalamThe relationships between Big Five Personality dimensions, harmful psychoactive substance use and academic motivation among undergraduates in Nigeria

A word‐level language identification strategy for resource‐scarce languages

ABSTRACT This study is based on the premise that it is possible to train compute

GlotLID: Language Identification for Low-Resource Languages

International audience Several recent papers have published good solutions for langua

Toluwase/Word-Level-Language-Identification-for-Resource-Scarce-

English, Hausa, Igbo and Yoruba corpora and results (presented in excel files) of word-level languag

ConLID: Supervised Contrastive Learning for Low-Resource Language Identification

Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora fr

dev52003/Biased-News-Identification-For-Low-Resource-Languages-Malayalam

# 📰 Biased News Identification in Malayalam Media This repository contains the research work and to

The relationships between Big Five Personality dimensions, harmful psychoactive substance use and academic motivation among undergraduates in Nigeria

Aims The aim of this study was to determine the relationships between personality traits, stress pe