Thesis of Timon Deschamps


Subject:
Aligning agents with users through multi-objective reinforcement learning

Start date: 07/11/2023
End date (estimated): 07/11/2026

Advisor: Laetitia Matignon
Coadvisor: Mathieu Guillermin, Rémy Chaput

Summary:

Looking back on any outcome, one can retrace the trajectory that led there: a series of choices, each shaped by trade-offs between often competing goals. Multi-objective reinforcement learning (MORL) provides a natural framework to model this kind of sequential decision-making, enabling artificial agents that learn to follow their user's preferences and navigate these competing goals. However, the behavior of autonomous agents is highly sensitive to the objective they try to maximize. When this objective is misspecified, or when user and agent do not have access to the same information, deploying these agents can have unintended consequences, risking undesirable trade-offs or unsafe behaviors. This raises the questions of what it means for an agent to be aligned with its user, and of how to construct such agents, questions that motivate the three contributions of this thesis.

The first contribution reviews what makes an agent ethically aligned. Through a review of the machine ethics and MORL literatures, we discuss how multi-objective reinforcement learning contributes to this goal, and what further properties are needed beyond it. Furthermore, we examine the difficulties of evaluating ethical alignment, illustrating them via a case study on a smart-grid simulator. The two contributions that follow develop points raised by this review. Both focus on the case of non-linear utility functions, which better reflect real user preferences but are considerably harder to handle. The second contribution explores continual adaptation to such preferences as they change over time, formalizing the setting and extending existing metrics to evaluate it. To address this problem, we propose a method that exploits similarities between utility functions to transfer knowledge across them, and evaluate it against baselines on four standard MORL environments. Non-linear utility functions are especially difficult to optimize for when the user cares about every single outcome and outcomes cannot compensate for one another. Existing methods either suffer from high variance, or require representations that scale exponentially in the number of objectives, preventing their use when many objectives must be balanced simultaneously. The third contribution tackles this setting, showing that more compact representations are often sufficient to estimate and optimize the user's utility, empirically matching or improving on existing methods while scaling quadratically.

Together, these contributions suggest that MORL is a suitable foundation to create agents that act according to the values and preferences of the user, but that it needs to be complemented by additional properties to build ethically aligned agents. Autonomous agents trace trajectories on people's behalf, and will increasingly do so as their adoption widens. The goal of this thesis is to help ensure that the choices made along the way, and where they lead, are ones their users would endorse.


Jury:
M. Vamplew PeterProfesseur(e)Federation University AustraliaRapporteur(e)
Mme Beynier AurélieProfesseur(e)Sorbonne UniversitéRapporteur(e)
Mme Rădulescu RoxanaMaître de conférenceUtrecht UniversityExaminateur​(trice)
M. Vercouter LaurentProfesseur(e)INSA Rouen NormandieExaminateur​(trice)
M. Meyer AlexandreProfesseur(e)Université Claude Bernard Lyon 1Examinateur​(trice)
Mme Matignon LaetitiaMaître de conférenceUniversité Claude Bernard Lyon 1Directeur(trice) de thèse
M. Rémy ChaputMaître de conférenceCPE LyonCo-encadrant(e)