Thèse Apprentissage par Renforcement Social pour Soutenir la Transition Agroécologique H/F
Doctorat.Gouv.Fr
- Toulouse - 31
- CDD
- Bac +5
- Service public d'état
Les compétences pour ce job
Les missions du poste
Les jumeaux numériques fondés sur l'intelligence artificielle offrent un outil puissant pour relever ce défi, en permettant de simuler les processus de décision des agriculteurs et d'évaluer l'impact d'interventions publiques avant leur mise en oeuvre. Des travaux récents ont montré que des jumeaux numériques intégrant des agents d'apprentissage par renforcement peuvent efficacement modéliser les décisions séquentielles des agriculteurs et contribuer à la conception de politiques agricoles [Gautron et al., 2022 ; Vinyals et al., 2023]. L'apprentissage par renforcement [Sutton et Barto, 2018] permet à des agents artificiels d'apprendre en interagissant avec leur environnement. Cette approche a connu des succès remarquables dans la reproduction - voire le dépassement - des capacités humaines d'apprentissage dans des tâches complexes de prise de décision [Perolat et al., 2022 ; Schrittwieser et al., 2020]. Cependant, malgré ces avancées, les agents actuels d'apprentissage par renforcement apprennent principalement à partir de leur propre expérience directe et prennent encore peu en compte les mécanismes d'apprentissage social.
Afin de répondre à cette limite, cette thèse vise à développer des agents avancés d'apprentissage par renforcement social, capables d'apprendre non seulement à partir de leur propre expérience (expérience directe), mais également à partir de l'observation et de l'interaction avec d'autres agents (expérience sociale). Ces agents constitueront les briques fondamentales de jumeaux numériques d'agriculteurs, permettant de simuler et d'analyser les dynamiques sociales complexes qui sous-tendent l'adoption des pratiques agroécologiques. Les agents d'apprentissage par renforcement social permettront ainsi de mieux comprendre les facteurs sociaux, économiques et environnementaux qui influencent les décisions des agriculteurs, et de produire des connaissances opérationnelles pour soutenir la conception de politiques publiques plus efficaces en faveur de la transition agroécologique. Reinforcement learning (RL) [Sutton and Barto, 2018] provides a natural computational framework for modelling how humans learn from experience. By studying how artificial agents acquire behaviours through interactions with their environment, RL is one of the machine learning fields most closely connected to theories of human learning [Fan et al., 2023]. Consequently, advanced RL-based AI agents represent promising candidates for building digital twins capable of reproducing human adaptive decision-making.
However, classical RL architectures are designed to learn exclusively from direct experience, i.e., from the consequences of the learner's own actions. This contrasts with decades of research in the social sciences showing that human learning is not only individual but also fundamentally social. Humans acquire knowledge and adapt their behaviour through social experiences, including observational learning (learning by observing others interacting with their environment) and argumentative learning (learning through exchanges of reasons, explanations, and justifications) [Boyd et al., 2011].
Despite their importance in human learning, social learning mechanisms remain largely unexplored in reinforcement learning. Existing approaches to social reinforcement learning, including observational learning and argumentative learning, are still at an early stage.
Related work on observational reinforcement learning. A few recent studies have extended standard RL architectures to incorporate observational learning, where agents learn from observing the actions and outcomes of other agents [Borsa et al., 2017; Ndousse et al., 2021]. However, these approaches are mainly based on deep neural architectures, and their performance is often highly dependent on specific experimental settings, limiting their generalization. A notable exception is [Lupu et al., 2019], who proposed a formal framework for observational learning, but only in the simpler multi-armed bandit setting (i.e. a reinforcement learning framework where a single state of the environment is considered).
Related work on argumentative reinforcement learning. Humans do not only learn by observing others; they also improve their decisions through social exchanges involving arguments, explanations, and justifications. Such argumentative experiences have been shown to enhance both individual and collective decision-making, with groups often outperforming the average individual performance and sometimes even exceeding the best individual decision-maker [Mercier and Caidière, 2022]. The Interactionist Theory of Reasoning [Mercier and Sperber, 2011] provides a cognitive framework explaining the role of argumentative exchanges in human reasoning. With RL agents, the generation of such arguments could be supported by explainable RL approaches, which aim to extract and communicate the rationale behind agents' decisions [Saulieres, 2025]. Although recent computational models have explored argument exchange inspired by the Interactionist Theory of Reasoning [Baccini et al., 2023], most of these approaches study argumentation independently from the broader learning process. They therefore do not address the problem of how agents can integrate argumentative experiences into their own decision-making and adaptation processes.
Beyond fundamental AI research, reinforcement learning has increasingly been applied to support decision-making in agricultural systems, alongside other machine learning approaches [Bal and Kayaalp, 2021; Benos et al., 2021]. Recent surveys highlight the growing potential of RL for modelling complex agricultural decision processes [Gautron et al., 2022]. For example, Farm-Gym [Maillard et al., 2023] provides a gamified agronomic simulator for RL that models the sequential decisions farmers face on their individual farms, including those related to agroecological practices. Similarly, [Gautron, 2024] demonstrated the practical benefits of RL for agronomic decision-making using the DSSAT crop simulation platform. However, without exception, existing RL applications in agrosystems remain limited to the classical paradigm of a single-farmer agent learning exclusively from its own direct experience.
Against this background, this thesis aims to enable a paradigm shift from the individual reinforcement learner to the social reinforcement learner by developing RL architectures capable of exploiting not only direct experience, but also observational and argumentative experiences. These social RL agents will provide the foundations for a new generation of AI systems capable of learning through social interactions.
Le profil recherché
Bienvenue chez Doctorat.Gouv.Fr
Publiée le 28/07/2026 - Réf : ecd165c3b5d2f47eb0d010c0824da0ea