Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Data Sheet 1_Methodological guidance for predictor variable selection for adolescent smoking outcomes in Global Youth Tobacco Survey using R and Python.zip

Domaine:

socioeconomic

Type de record:

paper
Créateur:
WinCosLaw
Hôte:avatar
Background

The Global Youth Tobacco Survey is one of the most important sources of data on adolescent tobacco use worldwide. However, studies using these data often apply inconsistent statistical methods, particularly in how they handle complex survey designs and select predictors for analysis. Many analyses use simplified approaches or proprietary software, making results difficult to reproduce or compare across countries and survey years. We developed a clear, open-source workflow to guide researchers in selecting predictors and modeling adolescent smoking outcomes in a way that is transparent, theory-informed, and reproducible in both R and Python.

Results

The framework is designed to incorporate the full two-stage survey design, including weights, stratification, and clustering, with the R implementation serving as the reference platform for design-based inference. Key demographic factors (age, sex, grade, and region or country) are retained in all models, while other modifiable predictors are selected using a constrained stepwise procedure guided by model fit. We demonstrate the approach using Zambia 2021 and pooled data from Ghana, Mauritius, Seychelles, and Togo, 2015–2019. The pooled dataset included 15,914 adolescents, of whom 13,360 had complete information for the complete-case modeling analysis. When identical model specifications were applied, the R and Python implementations selected the same final model and produced nearly identical adjusted odds ratios, with differences below 0.01. However, standard errors and confidence intervals differed slightly because Python's implementation relied on survey weights only and did not fully account for clustering and stratification. In the current cigarette-smoking model, the final model showed adequate event support, with 1,479 current smokers and 33 estimated parameters, giving an events-per-variable value of 44.82. Factors independently associated with current cigarette smoking included intention to use tobacco in the next 5 years, ever trying cigarette smoking, ownership of an item with a tobacco logo, seeing teachers smoking at school, and being taught about the dangers of tobacco use.

Conclusions

This study provides a practical and generalizable framework for analysing adolescent smoking using complex survey data. By combining theoretical grounding, transparent variable selection, and open-source tools, the workflow improves reproducibility, cross-country comparability, and policy relevance. It can be readily adapted to other large-scale health surveys to strengthen evidence for tobacco control and prevention efforts.

Visit

figshare.com

Tags

Applied Mathematics not elsewhere classifiedadolescent smoking behaviorcomplex survey designGYTSmulti-country harmonizationpredictor variable justificationPythonRsocial cognitive theory

Licenses

CC BY 4.0

Similaires

Data Sheet 2_Methodological guidance for predictor variable selection for adolescent smoking outcomes in Global Youth Tobacco Survey using R and Python.zip

Data Sheet 2_Methodological guidance for predictor variable selection for adolescent smoking outcomes in Global Youth Tobacco Survey using R and Python.zip

Background

The Global Youth Tobacco Survey is one of the most important sources of data on adolesc