Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Synthetic Predictions for Real Interventions: Augmenting Behavioral RCTs with Fine-Tuned LLMs

Domaine:

natural language processing

Type de record:

paper
Créateur:
RayTomMarRay
Éditeur:
Cen
Éditeur:
OSF
Hôte:avatar
Randomized controlled trials are the gold standard for causal inference but are costly and slow. Large language models raise the possibility of augmenting human samples with synthetic predictions to recover treatment-effect estimates at lower cost. Existing approaches build foundation models on broad behavioral corpora; we instead test an alternative — specializing an LLM by fine-tuning it on a narrow, context-specific corpus matched to the target trial's behavioral domain, population, and interventions. We benchmark synthetic predictions against two completed field experiments on verified COVID-19 vaccination uptake: Duch et al. (2023), a cluster-randomized trial in rural Ghana, and Campos-Mercade et al. (2021), an individually randomized trial in Sweden. Across six open-weight base models (OLMo 2, Llama 3.1, Qwen 3, at two scale tiers) plus a GPT-4o reference, we evaluate four fine-tuning configurations crossing an external corpus of related studies with a pilot sample drawn from the target trial. We assess four pre-specified dimensions — treatment-effect recovery, within-arm distributional divergence, individual-level predictive accuracy, and precision gain under prediction-powered inference. We adopt an estimation-based pre-registration, specifying estimands, metrics, and comparisons rather than directional hypotheses. All training and inference choices are fully specified in the attached plan.

Visit

doi.orgosf.io

Tasks

language modeling

Tags

Social and Behavioral SciencesCOVID-19 vaccinationGhanaLoRAPPI++computational social sciencefield expeirmentsfine-tuninglarge language modelspredict-then-debias+6

Licenses

Creative Commons Zero v1.0 Universalhttps://creativecommons.org/publicdomain/zero/1.0/legalcode

Similaires

General LLMs Versus Fine-Tuned Models on Algerian Dialect Classification TasksFew-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMsSynthetic Data Generation with Python and LLMspmmlv2-fine-tuned-hausapmmlv2-fine-tuned-yoruba0xnu/pmmlv2-fine-tuned-hausa

General LLMs Versus Fine-Tuned Models on Algerian Dialect Classification Tasks

Few-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMs

This paper presents two effective approaches for Extractive Question Answering (QA) on the Quran. It

Synthetic Data Generation with Python and LLMs

A talk presented on September 4, 2025, at PyCon Somalia 2025.

pmmlv2-fine-tuned-hausa

pmmlv2-fine-tuned-yoruba

0xnu/pmmlv2-fine-tuned-hausa