Phd Position F - M Phd Thesis Extracting From And Verbalizing Rdf To Text To Support Collaborative Co-Editing Of Wikis And Their Corresponding Linked Data H/F INRIA
- Nice - 06
- CDD
- Bac +5
- Service public des collectivités territoriales
Détail du poste
PhD Position F/M [PhD Thesis] Extracting from and verbalizing RDF to text to support collaborative co-editing of wikis and their corresponding linked data
Le descriptif de l'offre ci-dessous est en Anglais
Type de contrat : CDD
Niveau de diplôme exigé : Bac +5 ou équivalent
Fonction : Doctorant
Niveau d'expérience souhaité : Jeune diplômé
A propos du centre ou de la direction fonctionnelle
Inria is the French National Institute for Research in Digital Science, of which the Inria Côte d'Azur University Center is a part. With strong expertise in computer science and applied mathematics, the research projects of the Inria Côte d'Azur University Center cover all aspects of digital science and technology and generate innovation. Based mainly in Sophia Antipolis, but also in Nice and Montpellier, it brings together 47 research teams and nine support services. It is active in the fields of artificial intelligence, data science, IT system security, robotics, network engineering, natural risk prevention, ecological transition, digital biology, computational neuroscience, health data, and more. The Inria Center at Université Côte d'Azur is a major player in terms of scientific excellence, thanks to the results it has achieved and its collaborations at both European and international level.
Contexte et atouts du poste
INRIA is the French national research institute dedicated to computer science and applied mathematics and is a founding member of the World-Wide Web Consortium (W3C). The Inria centre at Université Côte d'Azur includes 42 research teams and 9 support services. The centre's staff (about 500 people) is made up of scientists of dierent nationalities, engineers, technicians, and administrative staff. The teams are mainly located on the university campuses of Sophia Antipolis and Nice as well as Montpellier, in close collaboration with research and higher education laboratories and establishments (Université Côte d'Azur, CNRS, INRAE, INSERM ...), but also with the regional economic players.
The Wimmics team works on the topic of AI on the Web, in particular knowledge graphs and (linked) data representation and processing in the Semantic Web. Wimmics contributes to knowledge formalization and semantic-based methods to extract, control, query, validate, infer, explain and interact with knowledge in epistemic communities on the Web. Wimmics has been involved in a large number of European research projects and national projects. This PhD position takes place within the context of the national ANR project KGTWIN, with partners in Nantes and Nancy.
Mission confiée
Large-scale collaborative platforms such as Wikipedia have demonstrated that distributed communities can maintain shared knowledge at scale. Yet the emergence of Large Language Models (LLMs) redefines the conditions of collaboration: trained on human knowledge, these models can now produce and revise it, raising new questions about coherence, authorship, and trust. The project KGTWIN addresses these questions through a new paradigm of neuro-symbolic collaborative editing, in which free text and structured Knowledge Graphs (KGs) co-evolve as complementary representations of the same knowledge. The shared KG acts as the semantic backbone of human-AI collaboration. It maintains consistency across documents, exposes contradictions between texts that refer to the same concepts, and links contributors working on related entities. In this context, the KG provides a verifiable, traceable foundation for shared knowledge, transforming human-AI interaction from sequential exchange into continuous co-evolution.
KGTWIN models this co-evolution as an iterative cycle combining controlled extraction, semantic synchronization, and grounded generation. Controlled extraction uses LLMs to instantiate ontologiesand thesauri from collaboratively written texts while preserving provenance. Semantic synchronization detects inconsistencies and propagates updates between textual and graph representations. Grounded generation produces text fragments from coherent subgraphs. Together, these processes maintain alignment between natural-language narratives and structured knowledge, establishing a framework where reasoning, learning, and collaboration converge.
In the KGTWIN project, one of the work packages will design a language-model-based pipeline able to extract structured knowledge from unstructured text while keeping users aware of the extraction process and its uncertainties.
This work package aims to develop principled methods for extracting RDF knowledge from free text under the guidance of a predefined ontology. The goal is to ensure semantic fidelity, ontological compliance, and traceability of the extracted triples. It builds on the emergence of generative models guided by explicit semantic constraints. Language Models (LMs) can produce rich relational information from unstructured text, but they remain prone to hallucinations, schema violations, and reasoning errors. Conversely, traditional rule-based and supervised extraction systems offer strong guarantees of correctness but are costly to maintain and brittle when applied to new domains. This work package investigates how to combine symbolic knowledge with the language models, to obtain the best of both worlds: scalable extraction, but still traceable. The goal is to systematically evaluate extraction accuracy and recall according to the target ontologies and thesauri, the chosen model, and the pipeline used. The work package will produce metrics and visualization interfaces to expose what is extracted, what is missing, extraction confidence and highlight potential semantic conflicts with the current KG.
Principales activités
The PhD subject is about extracting from and verbalizing RDF to text to support collaborative co-editing of wikis and their corresponding linked data with the core questions:
- Methods to establish SHACL shapes describing targeted graph patterns, including domain and range restrictions, property cardinalities, and integrity constraints. The method should help define the formal targets that will guide knowledge extraction, starting from ontologies and targeted texts.
- Methods exploiting targeted graph patterns in providing training and validation means for the extraction task. We envision a hybrid pipeline that combines traditional knowledge representation and reasoning techniques and language model-based generation for extracting knowledge graphs from texts. Extracted triples will be verified using the targeted graph patterns and other techniques to ensure that they are compliant with the knowledge graph being built and are entailed by the source text, thereby mitigating hallucinations while maintaining high recall.
- Verbalization techniques capable of transforming RDF subgraphs into fluent natural-language consistent with the ontology semantics. These verbalizations make the entire process explainable to users and can be directly reinserted as editable text within collaborative documents, closing the co-evolution loop between text and structured knowledge.
Compétences
Technical skills and level required : expertise in knowledge graphs, linked data, semantic Web, machine learning, natural language processing, neuro-symbolic appraoches, etc.
Languages : English and French
Relational skills : team player, open source community player
Othervalued qualities : humour, kindness and open-mindedness
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking (after 6 months of employment) and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage
Rémunération
Gross Salary per month: 2300 €
Bienvenue chez INRIA
A propos d'Inria
Inria, l'institut national de recherche dans les sciences et technologies du numérique, est en appui de l'État pour les stratégies nationales de recherche et d'innovation du numérique en tant qu'Agence de programmes. Inria mène plus de 300 projets de recherche et d'innovation avec ses 3500 scientifiques, ingénieurs et personnels d'appui, en partenariat avec les universités et l'écosystème numérique (entreprises, entrepreneurs, acteurs publics). Ensemble, nous explorons des domaines clés comme l'intelligence artificielle, la cybersécurité, l'informatique quantique, le Cloud, la transformation numérique de la santé, les jumeaux numériques ou encore les technologies numériques pour la défense. Nous construisons des solutions concrètes telles que des logiciels, des startups technologiques, des partenariats avec les entreprises du tissu national et des formations de pointe. Notre objectif : l'impact scientifique, technologique et industriel au service de la souveraineté numérique de la France.
Publiée le 10/09/2026 - Réf : c09ac82094b094b6c0b04f00d2501911