Thèse Visualization Of The Plausibility And Bias For Data Resources Used In a Geographic Digital Twin H/F
Doctorat.Gouv.Fr
- Paris - 75
- CDD
- Télétravail partiel
- Bac +5
- Service public d'état
À noter sur ce job
Vous avez ces compétences ? Cliquez dessus pour les ajouter rapidement à votre profil !
Détail du poste
Établissement : Université Paris-Saclay GS Informatique et sciences du numérique École doctorale : Sciences et Technologies de l'Information et de la Communication Laboratoire de recherche : Laboratoire Interdisciplinaire des Sciences du Numérique Direction de la thèse : Tobias ISENBERG ORCID 0000000179538644 Début de la thèse : 2027-10-01 Date limite de candidature : 2027-03-31T23:59:59 [Notice: This advertised topic targets students focused on visualization and data analysis, with a background in artificial intelligence. It is not focused purely on AI. Also, please note that we only consider applications sent in English.]
In the field of visualization, questions on the visualization of data uncertainty have been a core part of the research to date [e.g., 1]. Yet past work largely often silently considered data uncertainty to arise (primarily) from measurement uncertainty or as connected to some statistical analysis of captured data values. Only relatively recently have questions of implicit notions of error [4] or data hunches [3] entered the discussions of this general problem-these cover aspects of data imprecision that can arise from various sources: e.g., different forms of measurement or data recording, biases in how people access data, and even intentionally introduced errors. For example, spatial geographic data may not only come from systematically controlled sources but also from, for instance, from multiple sets of data that come from different institutions with different data collection policies, from contributions from the general public, or even by sourcing data from social media. Naturally, this introduces a variety of different levels of data quality and plausibility for spatial data. Sometimes experts are aware of these data caveats (that exist even for professionally collected datasets) and we can try to visualize them [3-5], yet in other cases the data collection happens in a way that this knowledge is not available-e.g., when data is retro-actively collected from sources that were not (primarily) created for recording this data in the first place-and we have to find ways of recovering such information retroactively and without access to the ground truth [2]. For instance, image collection sites such as Flickr partially show geo-located images, from which locations of certain points of interest can be extracted. In this specific example we ourselves recovered species distribution data from images posted on photo sharing platforms, and found that such geographic data is subject to many biases and errors. Yet this data analysis so far [2] has been a manual process and also only recovers anecdotal information about the existence of data errors and biases. In this PhD research project we will investigate if those results generalize to other kinds of spatial data (e.g., automatically detected features in remote sensing imagery, other volunteered geographic information, other forms of 2D and 3D spatial data). Then we will investigate ways to automate the error and bias identification process as well as to quantify the existence of such biases and errors in some form of plausibility measure, to be visualized in the digital twin [6]. Digital twins [6] are virtual representations of real-world products, systems, or processes, enabling simulation, integration, testing, monitoring, and maintenance. They play a pivotal role in optimizing complex systems across a wide range of domains, from industrial manufacturing and energy to environmental monitoring and healthcare.
For the remaining context please see the abstract. * analysis and visualization of the plausibility of crowd-sourced and/or heterogenous geographic data
* use of artificial intelligence to automate the analysis of the data errors and biases Typical requirements analysis-design-implementation-evaluation approach that is used in most of HCI and visualization.
Le profil recherché
* good written and spoken communication in English to be able to interact within the research team, to be able collaborate with domain scientists, and to be able to disseminate the results in scientific publications
* excellent skills in actual manual (non-AI-based) programming
* experience in the use of machine learning and artificial intelligence
* experience (but not necessarily academic) in data visualization
* experience in HCI and empirical evaluation preferred but not required
* past publications preferred but not required
Publiée le 02/09/2026 - Réf : f09fa8d0972e4046d56dc5a070361944