Thesis of Alexandre Chapin


Subject:
Disentangled Latent Manipulation Learning for Dexterous Robotic Manipulation

Start date: 01/10/2022
End date (estimated): 01/10/2025

Advisor: Liming Chen
Coadvisor: Emmanuel Dellandréa

Summary:

This thesis investigates the role of visual representations in the learning and generalization capabilities of robotic manipulation policies. While state-of-the-art methods typically rely on global or dense features extracted from large-scale visual foundation models, these representations often entangle task-relevant information with irrelevant background clutter. We address this limitation by proposing a shift toward Object-Centric Representations (OCR), a structured alternative that decomposes visual scenes into discrete, interactive "slots".    Our research establishes a three-stage progression toward generalizable robotic perception. First, we demonstrate the foundational benefits of object-centric inductive biases over holistic representations in multi-object scenes, highlighting their superior robustness to distractor colors and textures. Second, we introduce object-centric representations to real-world robotics, conducting a large-scale systematic comparison across both simulated (MetaWorld, LIBERO) and physical benchmarks. We demonstrate that pre-training these models on large-scale robotic datasets establishes a robust perceptual foundation. However, we reveal new insights regarding the behavior of these models when introducing novel distractor objects, identifying a critical capacity-generalization trade-off: a delicate balance must be struck between the number of allocated slots and the final performance in both in-domain and generalized settings.  To overcome this trade-off, we propose three specific frameworks. We present STORM, a lightweight module that semantically grounds slots using natural language instructions via a novel multi-phase adaptation strategy. We also introduce Slot-RAE, a generative object-centric diffusion model utilizing direct representation auto-encoders. This framework not only streamlines the learning process by bypassing heavy generative biases, but it also unlocks true compositional capacities natively within the continuous feature space of visual foundation models (VFMs). Finally, as an ongoing exploratory direction, we outline ARMOR, a compute-efficient latent-space data augmentation pipeline that extracts permutation-invariant tokens from diverse, task-irrelevant images and dynamically injects them into the state representations to inoculate control policies against visual noise.  This work lays the groundwork for structured, symbol-aware perception in robot learning, bridging the gap between low-level visual input and high-level neuro-symbolic reasoning.


Jury:
M. Marchand EricProfesseur(e)Université de RennesPrésident(e)
M. Vu Ngoc SonProfesseur(e)Université de technologie de TroyesRapporteur(e)
M. Demonceaux CédricProfesseur(e)IUT Le Creusot, Université Bourgogne EuropeRapporteur(e)
Mme. Teulière Céline Professeur(e)Université Clermont AuvergneExaminateur​(trice)
M. Chen LimingProfesseur(e)Ecole Centrale de LyonDirecteur(trice) de thèse
M. Dellandréa EmmanuelMaître de conférenceEcole Centrale de LyonCo-encadrant(e)