Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Algorithmic Fairness in Predictive Models of Student Outcomes: A Systematic Review

Domain:

education

Record type:

paper
Creator:
Ant
Editor:
Cen
Publisher:
OSF
Host:avatar
1. Background and Rationale Machine learning is increasingly used in education to predict student outcomes, including academic performance, course pass or fail, dropout, retention, and completion. These predictions inform real institutional decisions: which students get flagged for early intervention, who receives academic support, and how scarce resources are allocated. As deployment expands across schools, colleges, and universities worldwide, an uncomfortable question has gained urgency. Do these models treat all students fairly, or do they reproduce, and sometimes amplify, existing inequalities along lines of gender, race, socioeconomic status, disability, and geography? The technical literature on algorithmic fairness has grown rapidly since 2016, producing formal definitions of fairness, mathematical impossibility results showing that competing fairness criteria cannot be simultaneously satisfied, and a toolkit of intervention techniques applied at different stages of the machine learning pipeline. In parallel, educational researchers have begun asking how these technical advances translate into educational practice. Yet despite this growing body of work, the field lacks a focused synthesis of how fairness-aware machine learning has been applied specifically to the prediction of student outcomes, the task most consequential for individual learners. Existing reviews either survey fairness in machine learning broadly across all application domains, or treat education as a single category that bundles outcome prediction together with very different tasks such as automated essay scoring, plagiarism detection, and recommender systems. A reader interested in the specific question of how fairness has been handled in models that predict whether a student will pass, fail, drop out, or complete cannot easily find a coherent synthesis of the evidence. This systematic review addresses that gap. 2. Purpose of the Study The purpose of this systematic review is to synthesise the published research evidence on algorithmic fairness in machine learning models that predict student outcomes, defined as performance, dropout, retention, completion, success, and failure, and to map both the methodological landscape and the coverage gaps of this body of work. The review pursues four specific objectives. First, to characterise how algorithmic fairness has been conceptualised and operationalised in this literature, including which formal definitions of fairness are used and which demographic attributes are studied as the basis for fairness analysis. Second, to catalogue the fairness interventions that have been applied in educational prediction models and to identify at which stage of the machine learning pipeline, namely pre-processing, in-processing, or post-processing, these interventions are implemented. Third, to synthesise the evidence regarding the trade-off between predictive accuracy and fairness in this domain, and to identify the strategies researchers have used to navigate that trade-off. Fourth, to map the distribution of the literature across educational levels (primary, secondary, higher education), geographic contexts, and prediction tasks, with the explicit aim of surfacing under-represented areas of research. 3. Research Questions The review is guided by four research questions: RQ1. How is algorithmic fairness conceptualised and operationalised in machine learning research on educational prediction, and which demographic attributes are considered? RQ2. What fairness interventions have been applied in educational prediction models, and at which stage of the machine learning pipeline (pre-processing, in-processing, or post-processing) are they implemented? RQ3. What evidence exists regarding the accuracy-fairness trade-off in educational prediction models, and what strategies have been proposed to navigate it? RQ4. How is this literature distributed across educational levels, geographic contexts, and prediction tasks, and where are the most significant coverage gaps? 4. Scope and Boundaries The review is bounded deliberately to keep its synthesis coherent. It includes peer-reviewed studies that apply machine learning to predict student outcomes and that engage explicitly with fairness, bias, or subgroup performance. The time window is 2016 to the present, beginning with the year that produced the foundational technical definitions of fairness in supervised learning. Five databases will be searched: Scopus, Web of Science, IEEE Xplore, ACM Digital Library, and ERIC, providing coverage across both the computer science and education research literatures. Google Scholar will be used as a supplementary check on the first 200 ranked results to identify any relevant records not surfaced by the structured database searches. The review excludes studies that address related but methodologically distinct tasks, including automated essay scoring, plagiarism detection, recommender systems, and admissions screening, because these involve different fairness considerations and different machine learning architectures. It also excludes purely conceptual or opinion pieces, focusing instead on empirical and methodological contributions. 5. Methodology The review will follow the PRISMA 2020 reporting guideline for systematic reviews, with the protocol pre-registered on the Open Science Framework (OSF) before screening begins. The methodology proceeds through nine stages: protocol registration, search-string construction and piloting, full database searches, deduplication and management of records, title and abstract screening, full-text screening, data extraction using a structured template, quality assessment of included studies, and narrative synthesis organised around the four research questions. A PRISMA flow diagram will document the screening process. Where comparable quantitative evidence is available across studies, accuracy and fairness metrics will be tabulated for cross-study comparison; meta-analytic pooling is not anticipated because the methodological heterogeneity of the corpus is expected to preclude it. A structured extraction template will record, for each included study, its bibliographic details, geographic context, educational level, prediction task, dataset, machine learning algorithms, fairness definitions, demographic attributes considered, intervention type and pipeline stage, performance metrics, fairness metrics, baseline and post-intervention values, characterisation of the accuracy-fairness trade-off, and key findings. Quality assessment will follow a five-criterion rubric covering dataset description, justification of sensitive attributes, explicitness of fairness metrics, train-test split documentation, and acknowledgement of analytical limitations. 6. Expected Outcomes and Contributions The review is expected to produce four principal outcomes. A taxonomy of fairness conceptualisation in educational prediction. The synthesis under RQ1 will yield a structured account of how fairness has been defined and measured in this literature, mapping the demographic attributes studied and revealing which sensitive attributes have received attention and which have been neglected. This taxonomy will be useful to researchers selecting fairness criteria for new studies and to practitioners interpreting fairness claims in deployed systems. A catalogue of intervention strategies organised by pipeline stage. The synthesis under RQ2 will produce a structured catalogue of fairness interventions applied in educational prediction, classified by their position in the machine learning pipeline. This catalogue will allow researchers and practitioners to identify which interventions have been tested in educational settings, and at which stage of the pipeline efforts have concentrated. A focused synthesis of evidence on the accuracy-fairness trade-off in educational prediction. The analysis under RQ3 will provide a focused synthesis of how the accuracy-fairness trade-off plays out specifically in educational prediction tasks, and what strategies researchers have proposed to navigate it. This is expected to inform both methodological choices in future studies and policy discussions about acceptable trade-offs in deployed educational systems. A coverage map and research agenda. The analysis under RQ4 will produce a coverage map showing the distribution of the existing literature across educational levels, geographic contexts, and prediction tasks. The map is expected to surface under-representation of research from sub-Saharan Africa and other low-resource educational contexts, and to identify prediction tasks that have received disproportionately little fairness attention. From this map, the review will articulate a forward-looking research agenda for fairness-aware educational machine learning. 7. Significance The significance of this work operates at three levels. For the research community, the review provides a focused synthesis of a fragmented literature, enabling cumulative scholarship rather than continued duplication. By surfacing methodological inconsistencies and coverage gaps, it sets an agenda for the field's next phase of development. For educational practitioners and institutions considering or already deploying predictive models, the review offers an evidence base for evaluating fairness claims and selecting fairness interventions appropriate to their context. The synthesis of trade-off evidence is particularly useful for institutional decision-makers facing the practical question of how much predictive accuracy can responsibly be sacrificed to reduce demographic disparities. For policymakers developing regulatory frameworks for algorithmic decision-making in education, the review provides empirical grounding for understanding what fairness interventions are realistic, what their costs are, and where the research evidence is sufficient or insufficient to support specific regulatory positions. 8. Methodological Rigour and Limitations The review's rigour rests on adherence to PRISMA 2020 reporting standards, pre-registration of the protocol on OSF, and transparent documentation of inclusion and exclusion decisions. Where a second screener is available, dual screening will be conducted on a sample of records and inter-rater agreement reported; where this is not feasible, transparent single-screener procedures with a documented audit trail in Rayyan will be used. A PRISMA flow diagram will provide complete traceability from initial database hits to final included studies, with reasons recorded for all exclusions at the full-text stage. The review nonetheless carries limitations that will be acknowledged transparently. First, the search is bounded by a cutoff date, after which subsequent work will not be reflected. Second, the review is restricted to English-language publications, which may underrepresent research published in other languages. Third, only peer-reviewed publications are included, which excludes potentially valuable preprint work in a fast-moving field. Fourth, the methodological heterogeneity of the corpus precludes formal meta-analysis, restricting synthesis to narrative and descriptive forms. 9. Timeline The project is planned across approximately twenty-six weeks. The first four weeks cover protocol registration, search-string piloting, and full database searches. The following ten weeks are dedicated to title-abstract and full-text screening. Weeks fifteen through twenty cover data extraction. Weeks twenty-one through twenty-four are reserved for synthesis and drafting. The final two weeks cover external review by a senior reader and submission to the target journal. 10. Dissemination The principal output of the project is a peer-reviewed journal article. Target journals include Computers and Education, British Journal of Educational Technology, Education and Information Technologies, and the International Journal of Artificial Intelligence in Education. Beyond the primary publication, the review's coverage map and research agenda will be disseminated through academic conference presentations and made openly available alongside the OSF protocol registration.

Visit

doi.orgosf.io

Tags

Computer EngineeringEducationEngineeringPRISMAalgorithmic fairnessbias mitigationeducational data miningeducational predictionfairness-aware machine learninglearning analytics+2

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode