Thesis of Mathilde Marcy
Subject:
Start date: 01/11/2022
End date (estimated): 01/11/2025
Advisor: Jean-Marc Petit
Coadvisor: Jocelyn Bonjour, Vasile-Marian Scuturici
Summary:
Integration of data science within temperature-controlled transportation operational processes offers a considerable opportunity to improve the sector's performance and reduce its environmental footprint. However, its adoption within organizations has been hindered by many obstacles such as data-quality issues, which prevents the straightforward leveraging of operational data and demands additional investment for its adequate analysis. Artificially unicity, a complex and pernicious form of redundancy concealed by surrogate keys and their associated surrogate-foreign keys, particularly hinders data exploitation.
Since, to the best of our knowledge, this data-quality issue has not yet been addressed by the database research community, it first had to be formally defined. Furthermore, because conventional methods for detection and correction of duplicates and other data-quality issues are not well-suited to handle artificial unicity, we propose an approach to detect and suppress it from relational databases. Its novelty lies in leveraging elements typically excluded from conventional approaches: surrogate keys, surrogate-foreign keys and database schemas. Our approach explores all relations along the surrogate key--surrogate-foreign key join path, following a specific relation order and a defined sequence of steps for each relation. Beyond outperforming conventional approaches, thanks to blocking improvements derived from previous cleaning steps, it suppresses artificial unicity across an entire database, or a relevant subset therein, in a single pass, producing results that can be reused across multiple data analytics projects, rather than requiring a dedicated duplicate detection task for each one.
Like most other data-quality assessment techniques, our approach requires specific knowledge about the data and its structure that is not always available to the user, such as a list of natural and surrogate keys. However, because we could not infer them with automatic-discovery techniques, as their results are biased by artificial unicity, we propose a key-elicitation method that involves domain experts and relies on simple, easily visualized abstractions based on the so-called redundancy profile that is associated with relations. Quite interestingly, those profiles can be computed very efficiently, in quasi-linear time.
Since this research is grounded in industrial settings and aims to facilitate the integration of data science within temperature-controlled transportation organizations, we also developed RED2hunt, a human-in-the-loop framework supporting the entire cleaning process of operational databases. It implements both proposed approaches for key-elicitation and artificial unicity detection and suppression, alongside various mainstream data fusion and cleaning techniques to address remaining data-quality issues beyond redundancy. RED2Hunt was implemented as a Python library on top of PostgreSQL.
Since publicly available datasets usually adhere to strict quality-control processes before release they are not affected by artificial unicity, and industrial databases affected by it could not be disclosed for confidentiality reasons, we developed SODGAUP, a software for generating and polluting synthetic databases that mimic data and pollution patterns commonly observed within operational databases, including artificial unicity. The FrigoTruck database, used for all experiments presented in this thesis, was generated with SODGAUP.