Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness

Domain:

natural language processing

Record type:

paper
Creator:
ParSubSenAda
Host:avatar
Restless Multi-Armed Bandits (RMABs) have been successfully applied to resource allocation problems in a variety of settings, including public health. With the rapid development of powerful large language models (LLMs), they are increasingly used to design reward functions to better match human preferences. Recent work has shown that LLMs can be used to tailor automated allocation decisions to community needs using language prompts. However, this has been studied primarily for English prompts and with a focus on task performance only. This can be an issue since grassroots workers, especially in developing countries like India, prefer to work in local languages, some of which are low-resource. Further, given the nature of the problem, biases along population groups unintended by the user are also undesirable. In this work, we study the effects on both task performance and fairness when the DLM algorithm, a recent work on using LLMs to design reward functions for RMABs, is prompted with non-English language commands. Specifically, we run the model on a synthetic environment for various prompts translated into multiple languages. The prompts themselves vary in complexity. Our results show that the LLM-proposed reward functions are significantly better when prompted in English compared to other languages. We also find that the exact phrasing of the prompt impacts task performance. Further, as prompt complexity increases, performance worsens for all languages; however, it is more robust with English prompts than with lower-resource languages. On the fairness side, we find that low-resource languages and more complex prompts are both highly likely to create unfairness along unintended dimensions. Accepted at the AAAI-2025 Deployable AI Workshop

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceMachine LearningMultiagent Systems

Similar

Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization ApproachesEffects of Task-Based Instruction on Students’ Speaking Performance in the EFL Classroomswa8w1r3/RestlessIntermediate-Task Data Scale Effects on Zero-Shot Cross-Lingual Performance in XTREME for Low-Resource LanguagesIntermediate-Task Training Effects on Zero-Shot XTREME Performance Across Low-Resource LanguagesMultimodal Intermediate-Task Training Effects on Zero-Shot Cross-Lingual Sentiment Analysis Performance

Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches

Large Language Models (LLMs) are increasingly employed in high-stakes decision-making tasks, such as

Effects of Task-Based Instruction on Students’ Speaking Performance in the EFL Classrooms

The objective of this study was to investigate the effects of task-based language instruction (herea

wa8w1r3/Restless

Restless Development Tanzania is part of a global youth-led agency. We support young people to lead

Intermediate-Task Data Scale Effects on Zero-Shot Cross-Lingual Performance in XTREME for Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Intermediate-Task Training Effects on Zero-Shot XTREME Performance Across Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Multimodal Intermediate-Task Training Effects on Zero-Shot Cross-Lingual Sentiment Analysis Performance

This paper describes our system developed for the SemEval-2023 Task 12 "Sentiment Analysis for Low-r