This dataset contains 609 de-identified maternal and neonatal delivery records retrospectively collected from the labor ward register of a government hospital in Bangladesh. It includes maternal demographic characteristics, obstetric history, antenatal care information, admission reason, gestational age, delivery details, neonatal outcomes, and a binary high-risk pregnancy label generated using predefined clinical criteria based on the available maternal clinical information.
The dataset was compiled to support research investigating whether routinely available maternal clinical characteristics—including maternal age, parity, gravida, antenatal care attendance, gestational age, and admission reason—can be used to identify high-risk pregnancies in low-resource healthcare settings where advanced diagnostic resources may be limited. Of the 609 records, 379 (62.2%) were classified as high-risk according to the predefined labeling criteria.
All data were manually transcribed from routine hospital records without modification of the original recorded values. Personally identifiable information was removed prior to dataset preparation to ensure patient anonymity. The high-risk pregnancy label was generated by the research team using predefined clinical criteria rather than directly recorded in the source register. Users should review the accompanying README file for details of the labeling methodology, data quality considerations, and guidance on avoiding potential data leakage when developing predictive models.
A companion Data_Dictionary file provides detailed information on each variable, including data type, description, units, allowed values, and missing values. This dataset is intended to support research in maternal health, obstetric epidemiology, reproductive medicine, public health, and machine learning applications.