Thesis of Silvia Grosso


Subject:
Towards Interpretable Prototype-Based Federated Learning

Start date: 06/11/2023
End date (estimated): 06/11/2026

Advisor: Sara Bouchenak
Cotutelle: Roberto Esposito, Mirko Polato

Summary:

Machine learning (ML) is applied in many areas to extract knowledge from data and support decision-making processes, such as search engines [7], recommendation systems [2], and disease diagnosis [6]. With the rapid growth of data, ML algorithms have evolved from centralized to distributed solutions. To address data privacy concerns, Federated Learning (FL) has emerged as a paradigm that allows multiple participants to collaboratively train machine learning models without directly sharing their local data [1]. Most current FL methods are limited to datasets from a single modality, such as images or text [9,10]. However, the proliferation of sensor types and data collection methods creates an increasing need to integrate information from multiple modalities while preserving the advantages of FL. At the same time, federated systems face broader challenges related to heterogeneous data distributions and model architectures, communication efficiency, local learning conditions, model reliability, and interpretability.

This project investigates multimodal and heterogeneous federated learning settings, together with approaches that can address these broader challenges. In this context, prototype-based learning is explored as a promising direction for developing lightweight and flexible federated learning methods [3,4]. Prototype-based approaches represent relevant concepts through representative elements, or prototypes, in an embedding space and perform predictions based on the similarity or distance between input representations and these prototypes. In federated settings, prototypes can be computed or learned locally and exchanged or aggregated across clients, providing compact representations that can reduce communication requirements and support heterogeneous data and model conditions [4,5]. A particular focus is the development of interpretable and robust prototype-based federated learning methods. Interpretable prototype-based architectures can associate model predictions with representative and semantically meaningful input patterns, providing explicit evidence for their decisions [8]. The research studies how such representations can be learned and shared across heterogeneous clients while balancing predictive performance and model interpretability. More broadly, it aims to explore prototype-based mechanisms for trustworthy federated learning, with possible extensions to settings involving heterogeneous and multimodal data.

A particular focus is the development of interpretable and robust prototype-based federated learning methods. Beyond their role as compact representations, prototypes can also provide the basis for inherently interpretable architectures, in which model predictions are associated with representative and semantically meaningful input patterns [8]. The project investigates how such representations can be learned and collaboratively exploited across heterogeneous clients while maintaining predictive performance and improving model interpretability and reliability. More broadly, it aims to explore prototype-based mechanisms for trustworthy federated learning, including their potential extension to heterogeneous and multimodal settings.

References:

1. Zhang et al., A Survey on Federated Learning, Knowledge-Based Systems, 2021.

2. Aher (S. B.) et Lobo (L.). – Combination of machine learning algorithms for recommendation of courses in e-learning system based on historical data.Knowledge-Based Systems,vol. 51, 2013.

3. Snell et al., Prototypical Networks for Few-shot Learning, NeurIPS, 2017.

4. Tan et al., FedProto: Federated Prototype Learning across Heterogeneous Clients, AAAI, 2022.

5. Zhang et al., FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning, AAAI, 2024.

6. Kourou (K.), Exarchos (T. P.), Exarchos (K. P.), Karamouzis (M. V.) et Fotiadis (D. I.). – Machine learning applications in cancer prognosis and prediction.Computational and structural biotechnology journal, vol. 13, 2015, pp. 8–17.

7. McCallumzy (A.), Nigamy (K.), Renniey (J.) et Seymorey (K.). – Building domain-specificsearch engines with machine learning techniques. – In Proceedings of the AAAI Spring Symposium on Intelligent Agents in Cyberspace. Citeseer, pp. 28–39. Citeseer, 1999.

8. Chen et al., This Looks Like That: Deep Learning for Interpretable Image Recognition, NeurIPS, 2019.

9. McMahan (H. B.), Moore (E.), Ramage (D.), Hampson (S.) et y Arcas (B. A.). Communication-Efficient Learning of Deep Networks from Decentralized Data, 2017.

10. Chen (Y.), Qin (X.), Wang (J.), Yu (C.) et Gao (W.). FedHealth: A Federated Transfer Learning Framework for Wearable Healthcare. vol. 35, n4, 2020-07, pp. 83–93.

Selected publications of the advisor related to the topic : 
• L. Ferraguig, Y. Djebrouni, S. Bouchenak, and V. Marangozova. Survey of Bias Mitigation in Federated Learning. Conférence sur le Parallélisme/ Architecture/ Système/ Temps Réel (ComPAS’2021), Lyon, France, 5-9 juillet 2021. 
• B. Khalfoun, S. Ben Mokhtar, S. Bouchenak, V. Nitu. EDEN: Enforcing Location Privacy through Re-identification Risk Assessment: A Federated Learning Approach. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, Volume 5, Issue 2, June 2021.
• M. Maouche, S. Ben Mokhtar, S. Bouchenak. HMC: Robust Privacy Protection of Mobility Data Against Multiple Re-Identification Attacks. ACM Journal on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2(3), September 2018.
• R. Talbi, S. Bouchenak, L. Y. Chen. Towards Dynamic End-to-End Privacy Preserving Data Classification. IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2018), Fast Abstract, Luxembourg, June 25-28, 2018.
• S. Cerf, V. Primault, A. Boutet, S. Ben Mokhtar, R. Birke, S. Bouchenak, L.Y. Chen, N. Marchand, B. Robu. PULP: Achieving Privacy and Utility Trade-Off in User Mobility Data. SRDS 2017. The 36th IEEE Symposium on Reliable Distributed Systems (SRDS 2017), Hong Kong, September 26-29, 2017.