Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Blog Post 1
Mis à jour :
<!– — title: ‘Blog Post number 1’ date: 2012-08-14 permalink: /posts/2012/08/blog-post-1/ tags:
- cool posts
- category1
category2
news
Our paper The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations has been accepted at ACL 2026!
I gave a talk at ISIR Young Scientists Day on our work on the GDN-CC dataset
Presented our work on the GDN-CC dataset representing the MLIA team during the ISIR Young Scientists Day. Check out the talk details.
I presented our work on The GDN-CC Dataset at ACL in San Diego, Calfifornia!
Our paper Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders has been accepted at EMNLP 2026! see you there :)
We just published a new Preprint: Toward Collective-Centric Evaluation of Preference Inference.
Our new preprint Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy is now available on arXiv! We propose a collective-centric evaluation framework to analyze how preference inference models affect the collective opinion landscape in participatory democracy platforms. Check out the paper details.
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
publications
Comment mesurer les biais politiques des grands modèles de langue multilingues?
Published in TALN (EALM workshop), 2025
Paper - Position paper on political bias in multilingual models.
Recommended citation: Lequeu, P. A., Labat, L., Cave, L., Lejeune, G., Yvon, F., & Piwowarski, B. (2026). The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations. arXiv preprint arXiv:2601.14944.
Fairbeat: Assessing and Mitigating Bias with the Composite Balance Score
Published in ECML-PKKD (Demo Track), 2025
Proceedings - A user interface for social bias evaluation in tabular datasets.
Recommended citation: Lequeu, P. A., Lagraa, S., Robin, G., & Ouedraogo, M. (2025, September). Fairbeat: Assessing and Mitigating Bias with the Composite Balance Score. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 475-480). Cham: Springer Nature Switzerland.
Download Paper | Download Bibtex
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
Published in ACL, 2026
Paper - We introduce Corpus Clarification, a preprocessing framework of citizens consultation data which allow for ethical downstream analysis. We share a manually-annotated dataset based on the 2019 French consultation ‘Grand Débat National’, and a large automatically annotated dataset using SLMs finetuned for the task.
Recommended citation: Pierre-Antoine Lequeu, Léo Labat, Laurène Cave, Gaël Lejeune, François Yvon, and Benjamin Piwowarski. 2026. The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32976–33006, San Diego, California, United States. Association for Computational Linguistics.
Download Paper | Download Bibtex
Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy
Published in preprint, 2026
Preprint - Designed a new evaluation paradigm for preference inference and showed that recommender systems strongly distort the opinion landscape despite showing good results on standard metrics such as accuracy.
Recommended citation: Lequeu, P. A., Hafid, S., Lerner, P., Shafiabadi, N., Cave, L., Mas, D., ... & Yvon, F. (2026). Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy. arXiv preprint arXiv:2609.02990.
Download Paper
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
Published in EMNLP, 2026
Preprint - explored how encoder-based models use absolute positional (AP) and relative positional (RP) information by explicitly disentangling positional and semantic representations. We find that the learned AP representations are low dimensional and used to encode document structure, while RP information is used as complementary to semantic matching.
Recommended citation: Lequeu, P. A., Barboule, C., & Piwowarski, B. (2026). Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders. arXiv preprint arXiv:2605.30022.
Download Paper
talks
Grounding Synthesis in Human Votes: Improving Sparse Matrix Completion for Opinion Selection in Large-Scale Consultations
Mis à jour :
Invited by the ParliView research team @ UC Dublin.
The GDN-CC dataset
Mis à jour :
Presented our work “The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations” representing the MLIA team for a day-long seminar on ISIR’s research.
teaching
Teaching 2024-2025
Grad & Undergrad, Sorbonne University, 2024
Introduction to Programming (1st-year bachelor), Introduction to Relational Databases (2nd-year bachelor), Industrial Project (1st-year master)
Teaching 2025-2026
Grad & Undergrad, Sorbonne University, 2025
Data Science (1st-year bachelor), Deep Learning (2nd-year Master)
Teaching 2026-2027
Grad & Undergrad, Sorbonne University, 2026
Data Science (1st-year bachelor), Deep Learning (2nd-year Master)
