Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Limitations of Scalarisation in MORL: A Comparative Study in Discrete Environments

Record type:

paper
Creator:
ShaJee
Host:avatar
Scalarisation functions are widely employed in MORL algorithms to enable intelligent decision-making. However, these functions often struggle to approximate the Pareto front accurately, rendering them unideal in complex, uncertain environments. This study examines selected Multi-Objective Reinforcement Learning (MORL) algorithms across MORL environments with discrete action and observation spaces. We aim to investigate further the limitations associated with scalarisation approaches for decision-making in multi-objective settings. Specifically, we use an outer-loop multi-policy methodology to assess the performance of a seminal single-policy MORL algorithm, MO Q-Learning implemented with linear scalarisation and Chebyshev scalarisation functions. In addition, we explore a pioneering inner-loop multi-policy algorithm, Pareto Q-Learning, which offers a more robust alternative. Our findings reveal that the performance of the scalarisation functions is highly dependent on the environment and the shape of the Pareto front. These functions often fail to retain the solutions uncovered during learning and favour finding solutions in certain regions of the solution space. Moreover, finding the appropriate weight configurations to sample the entire Pareto front is complex, limiting their applicability in uncertain settings. In contrast, inner-loop multi-policy algorithms may provide a more sustainable and generalizable approach and potentially facilitate intelligent decision-making in dynamic and uncertain environments. 15 pages, 4 figures, published in the Proceedings of the 46th Annual Conference of the South African Institute of Computer Scientists and Information Technologists (SAICSIT 2025)

Visit

arxiv.org

Tags

Machine Learning68T05 (Primary) 90C29 (Secondary)I.2.6; G.1.6

Similar

A multi-country comparative study of the perceived police disciplinary environmentsPreserving Digital Cultural Heritage in Resource-Limited Environments: A Comparative Study from South AfricaElectronic government adoption in voluntary environments – a case study of ZimbabweStaging Limitations in Ghanaian Theatre Spaces: An Exploratory StudyA Comparative Study in Shona PhoneticsDiscrete Choice Study Data on Financing Electricity Access in Africa

A multi-country comparative study of the perceived police disciplinary environments

Purpose – The purpose of this paper is to test an aspect of the theory of police i

Preserving Digital Cultural Heritage in Resource-Limited Environments: A Comparative Study from South Africa

Digital cultural heritage preservation is crucial for safeguarding intangible assets such a

Electronic government adoption in voluntary environments – a case study of Zimbabwe

Many governmental organisations across the world are progressively implementing electronic governmen

Staging Limitations in Ghanaian Theatre Spaces: An Exploratory Study

Abstract: This study investigates the staging limitations faced in multiple-set design practiceswith

A Comparative Study in Shona Phonetics

Discrete Choice Study Data on Financing Electricity Access in Africa

CSV data from a survey administered in Accra, Ghana. Altogether about 250 paper-based questionnaires