Randomized controlled trials are the gold standard for causal inference but are costly and slow. Large language models raise the possibility of augmenting human samples with synthetic predictions to recover treatment-effect estimates at lower cost. Existing approaches build foundation models on broad behavioral corpora; we instead test an alternative — specializing an LLM by fine-tuning it on a narrow, context-specific corpus matched to the target trial's behavioral domain, population, and interventions.
We benchmark synthetic predictions against two completed field experiments on verified COVID-19 vaccination uptake: Duch et al. (2023), a cluster-randomized trial in rural Ghana, and Campos-Mercade et al. (2021), an individually randomized trial in Sweden. Across six open-weight base models (OLMo 2, Llama 3.1, Qwen 3, at two scale tiers) plus a GPT-4o reference, we evaluate four fine-tuning configurations crossing an external corpus of related studies with a pilot sample drawn from the target trial.
We assess four pre-specified dimensions — treatment-effect recovery, within-arm distributional divergence, individual-level predictive accuracy, and precision gain under prediction-powered inference. We adopt an estimation-based pre-registration, specifying estimands, metrics, and comparisons rather than directional hypotheses. All training and inference choices are fully specified in the attached plan.