Thesis of Aloïs Babé
Subject:
Start date: 09/05/2023
Defense date: 09/07/2026
Advisor: Serge Miguet
Coadvisor: Mihaela Scuturici
Summary:
Deploying computer vision systems in industrial environments raises many challenges: variability in acquisition conditions across sites, heterogeneous taxonomies, class imbalance, and scarce or imperfect expert annotations. Under these conditions, transfer robustness and low-cost adaptability become at least as critical as performance in controlled settings.
This thesis, conducted within a CIFRE industrial partnership with Veolia, investigates the extent to which foundation models — large-scale visual models pre-trained through self-supervised or contrastive learning — provide an effective response to these constraints. Two industrial use cases structure the study: visual characterization of solid waste streams on conveyor belts in sorting facilities, and closed-circuit television inspection of sewer networks. Representative of complementary data regimes, they allow the study of, respectively, generalization to new sites from limited annotations, and domain specialization when images are abundant but expert annotations are scarce.
A first contribution evaluates the out-of-distribution robustness of several families of pre-trained models for waste classification, using a cross-dataset transfer protocol in which no retraining on target domains is performed. Results show that performance in the learning domain is not correlated with performance in the target domains and that foundation models maintain a generalization advantage in this industrial context. They further show that this advantage is preserved after adaptation, and that lightweight strategies — such as tuning only the last blocks or using parameter-efficient fine-tuning methods — often provide highly competitive trade-offs for improving transfer.
A second contribution investigates domain specialization in a low-annotation regime, in the context of sewer network inspection evaluated on the Sewer-ML benchmark (1.3 million images, multi-label classification, severe class imbalance). A two-stage strategy is proposed and evaluated: continued self-supervised pretraining on target domain images, followed by supervised fine-tuning. This approach yields systematic gains on rare classes, increasing as annotations become scarcer, and proves competitive with semi-supervised reference methods. It reaches state-of-the-art performance on the benchmark with a compact model, without task-specific architectural modifications.
This work confirms that foundation models offer, in the studied industrial contexts, more transferable representations than standard pretrained approaches, even when images differ substantially from the web corpora used during pretraining. It also shows that their benefit depends strongly on the adaptation strategy employed, and thereby contributes to a better characterization of the conditions under which these models constitute a relevant response to the challenges of industrial computer vision.
Jury:
| Mme Catherine Achard | Professeur(e) | Sorbonne Université | Rapporteur(e) |
| M. Fabrice Meriaudeau | Professeur(e) | Université de Bourgogne | Rapporteur(e) |
| Mme Cécile Favre | Professeur(e) | Université Lumière Lyon 2 | Examinateur(trice) |
| M. Thierry Château | Professeur(e) | Université Clermont-Auvergne | Examinateur(trice) |
| M. Serge Miguet | Professeur(e) | LIRIS Université Lumière Lyon 2 | Directeur(trice) de thèse |
| Mme Mihaela Scuturici | Maître de conférence | LIRIS Université Lumière Lyon 2 | Co-directeur (trice) |
| M. Rémi Cuingnet | Docteur | VEOLIA | Encadrant(e) |