This dataset provides a benchmark for evaluating the model's ability to leverage richer genetic information from longer sequences to achieve more accurate inference.
Using data from the Human Pangenome Reference Consortium (BioProject ID: PRJNA730823), we designed a population classification task focusing on African, East Asian, and European population groups.
From samples' VCF file and the reference genome sequence, we generated sample pseudo-sequences.