Thesis of Alexandre Chapin
Subject:
Start date: 01/10/2022
End date (estimated): 01/10/2025
Advisor: Liming Chen
Coadvisor: Emmanuel Dellandréa
Summary:
This thesis investigates the role of visual representations in the learning and generalization capabilities of robotic manipulation policies. While state-of-the-art methods typically rely on global or dense features extracted from large-scale visual foundation models, these representations often entangle task-relevant information with irrelevant background clutter. We address this limitation by proposing a shift toward Object-Centric Representations (OCR), a structured alternative that decomposes visual scenes into discrete, interactive "slots". Our research establishes a three-stage progression toward generalizable robotic perception. First, we demonstrate the foundational benefits of object-centric inductive biases over holistic representations in multi-object scenes, highlighting their superior robustness to distractor colors and textures. Second, we introduce object-centric representations to real-world robotics, conducting a large-scale systematic comparison across both simulated (MetaWorld, LIBERO) and physical benchmarks. We demonstrate that pre-training these models on large-scale robotic datasets establishes a robust perceptual foundation. However, we reveal new insights regarding the behavior of these models when introducing novel distractor objects, identifying a critical capacity-generalization trade-off: a delicate balance must be struck between the number of allocated slots and the final performance in both in-domain and generalized settings. To overcome this trade-off, we propose three specific frameworks. We present STORM, a lightweight module that semantically grounds slots using natural language instructions via a novel multi-phase adaptation strategy. We also introduce Slot-RAE, a generative object-centric diffusion model utilizing direct representation auto-encoders. This framework not only streamlines the learning process by bypassing heavy generative biases, but it also unlocks true compositional capacities natively within the continuous feature space of visual foundation models (VFMs). Finally, as an ongoing exploratory direction, we outline ARMOR, a compute-efficient latent-space data augmentation pipeline that extracts permutation-invariant tokens from diverse, task-irrelevant images and dynamically injects them into the state representations to inoculate control policies against visual noise. This work lays the groundwork for structured, symbol-aware perception in robot learning, bridging the gap between low-level visual input and high-level neuro-symbolic reasoning.
Jury:
| M. Marchand Eric | Professeur(e) | Université de Rennes | Président(e) |
| M. Vu Ngoc Son | Professeur(e) | Université de technologie de Troyes | Rapporteur(e) |
| M. Demonceaux Cédric | Professeur(e) | IUT Le Creusot, Université Bourgogne Europe | Rapporteur(e) |
| Mme. Teulière Céline | Professeur(e) | Université Clermont Auvergne | Examinateur(trice) |
| M. Chen Liming | Professeur(e) | Ecole Centrale de Lyon | Directeur(trice) de thèse |
| M. Dellandréa Emmanuel | Maître de conférence | Ecole Centrale de Lyon | Co-encadrant(e) |