Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

STEERING GENERATIVE AI ON THE FLY: INFERENCE-TIME APPROACHES FOR SAFE, RELIABLE, AND INCLUSIVE LANGUAGE MODELS

Domaine:

natural language processing

Type de record:

paper
Créateur:
Gho
Éditeur:
Dig
Hôte:avatar
Large language models (LLMs) are typically aligned with human values through training-time methods such as reinforcement learning from human feedback. However, these methods produce static policies that cannot adapt to adversarial inputs unseen during training, reasoning challenges that emerge at test time, or the linguistic diversity of low-resource languages. In this thesis, we develop inference-time alignment methods that steer the model's generation during deployment using reward models, without modifying model parameters. We address four complementary challenges: principled decoding for alignment, safety against adversarial attacks, effective test-time reasoning, and inclusive adaptation for low-resource languages. We first establish theoretical foundations for inference-time alignment. Transfer-Q* resolves the central estimation bottleneck of controlled decoding by leveraging existing aligned models, with formal sub-optimality guarantees and consistent improvements across six evaluation setups. SITAlign extends this to multi-faceted alignment using satisficing theory, maximizing a primary objective while enforcing threshold constraints on secondary criteria, outperforming multi-objective decoding baselines by up to 22.3%. We then address safety for multimodal models: Immune reformulates jailbreak defense as an inference-time optimization, reducing attack success rates by 30-60% across five models and four benchmarks, and SafeThink recovers safety in reasoning-augmented models by steering only the first 1-3 reasoning steps, reducing attack success rates by up to 63% while preserving reasoning performance. We next investigate test-time scaling for reasoning models, revealing that extended thinking suffers from diminishing and eventually negative returns due to increasing output variance. We propose "parallel thinking'', which distributes the token budget across independent reasoning paths, achieving up to 22% higher accuracy than sequential scaling, and ThinkRetrieve, which injects retrieved solved exemplars into the reasoning trace at each step, maintaining monotonically increasing accuracy as the thinking budget grows. Lastly, we address the performance gap for low-resource languages. PromptRefine improves generation quality by up to 2.1x through cross-lingual retriever training with diversity-aware fine-tuning, and RELIC improves reward model reliability by up to 24% through pairwise ranking-aligned retrieval. Overall, this thesis develops and evaluates a suite of inference-time methods, spanning controlled decoding, safety steering, retrieval-augmented reasoning, and cross-lingual example selection, that collectively demonstrate that alignment need not be a static property instilled during training but can be a dynamic capability exercised at deployment time. We hope the methods and insights provided by this work will contribute toward building AI systems that are safer, more reliable, and more equitably accessible.

Visit

doi.orgdrum.lib.umd.edu

Tags

Artificial intelligenceInference-time AlignmentLarge Language ModelsLarge Reasoning ModelsTest-time Scaling

Similaires

On the use of generative models for evolutionary inference of malaria vectors from genomic dataOn the use of generative models for demographic inference in malaria vectors from genomic dataChildDiffusion: Unlocking the Potential of Generative AI and Controllable Augmentations for Child Facial Data using Stable Diffusion and Large Language ModelsReliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.IMultimodal Generative AI for African Language Preservation: A Framework for Language Documentation and RevitalizationTowards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications

On the use of generative models for evolutionary inference of malaria vectors from genomic data

Abstract Malaria in sub-Saharan Africa is transmitted by mosqui

On the use of generative models for demographic inference in malaria vectors from genomic data

Abstract Malaria in sub-Saharan Africa is transmitted by mosquitoes from the Ano

ChildDiffusion: Unlocking the Potential of Generative AI and Controllable Augmentations for Child Facial Data using Stable Diffusion and Large Language Models

In this research work we have proposed high-level ChildDiffusion framework capable of generating pho

Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I

The traditional evaluation of information retrieval (IR) systems is generally very costly as it requ

Multimodal Generative AI for African Language Preservation: A Framework for Language Documentation and Revitalization

Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications

Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in spec