Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

"A Structured Curriculum Vitae Dataset for CV Generation and Parsing "

Domain:

natural language processing

Record type:

dataset
Creator:
GeoMinMohMin
Publisher:
IEE
Host:avatar
"This dataset consists of structured representations of r\u00e9sum\u00e9s\/CVs collected from real-world users. Each CV is stored in a standardized JSON format, segmented into predefined sections such as personal information, education, work experience, skills, and projects. The dataset was designed to support research in automated CV parsing, structured document understanding, and curriculum vitae generation using Sequence-to-Sequence (Seq2Seq) models. With 133 anonymized entries, this dataset offers a clean and consistent foundation for fine-tuning machine learning models in the domain of natural language processing and information extraction. The CVs were collected as part of the graduation (honours) project \u201cTailored CV Generation\u201d at the Faculty of Computer and Information Sciences, Ain Shams University, Egypt."

Visit

doi.orgieee-dataport.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Paraphrase Generation from Latent-Variable PCFGs for Semantic ParsingPRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented DialogsA bandit approach to curriculum generation for automatic speech recognitionQaamuuska Af-Soomaaliga: A Structured NLP Dataset for Somali languagePre$^3$: Enabling Deterministic Pushdown Automata for Faster Structured LLM GenerationA Hierarchical Transformer for Unsupervised Parsing

Paraphrase Generation from Latent-Variable PCFGs for Semantic Parsing

One of the limitations of semantic parsing approaches to open-domain question answering is the lexic

PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs

Research interest in task-oriented dialogs has increased as systems such as Google Assistant, Alexa

A bandit approach to curriculum generation for automatic speech recognition

The Automated Speech Recognition (ASR) task has been a challenging domain especially for low data sc

Qaamuuska Af-Soomaaliga: A Structured NLP Dataset for Somali language

First machine-readable structured lexicon for Somali NLP research,extracted from the Qaam

Pre$^3$: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation

Extensive LLM applications demand efficient structured generations, particularly for LR(1) grammars,

A Hierarchical Transformer for Unsupervised Parsing

The underlying structure of natural language is hierarchical; words combine into phrases, which in turn form clauses. An awareness of this hierarchical structure can aid machine learning models in performing many linguistic tasks. However, most such models just pro