Posts by Collection

news

portfolio

publications

Fairbeat: Assessing and Mitigating Bias with the Composite Balance Score

Published in ECML-PKKD (Demo Track), 2025

Proceedings - A user interface for social bias evaluation in tabular datasets.

Recommended citation: Lequeu, P. A., Lagraa, S., Robin, G., & Ouedraogo, M. (2025, September). Fairbeat: Assessing and Mitigating Bias with the Composite Balance Score. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 475-480). Cham: Springer Nature Switzerland.
Download Paper | Download Bibtex

The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations

Published in ACL, 2026

Paper - We introduce Corpus Clarification, a preprocessing framework of citizens consultation data which allow for ethical downstream analysis. We share a manually-annotated dataset based on the 2019 French consultation ‘Grand Débat National’, and a large automatically annotated dataset using SLMs finetuned for the task.

Recommended citation: Pierre-Antoine Lequeu, Léo Labat, Laurène Cave, Gaël Lejeune, François Yvon, and Benjamin Piwowarski. 2026. The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32976–33006, San Diego, California, United States. Association for Computational Linguistics.
Download Paper | Download Bibtex

Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy

Published in preprint, 2026

Preprint - Designed a new evaluation paradigm for preference inference and showed that recommender systems strongly distort the opinion landscape despite showing good results on standard metrics such as accuracy.

Recommended citation: Lequeu, P. A., Hafid, S., Lerner, P., Shafiabadi, N., Cave, L., Mas, D., ... & Yvon, F. (2026). Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy. arXiv preprint arXiv:2609.02990.
Download Paper

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

Published in EMNLP, 2026

Preprint - explored how encoder-based models use absolute positional (AP) and relative positional (RP) information by explicitly disentangling positional and semantic representations. We find that the learned AP representations are low dimensional and used to encode document structure, while RP information is used as complementary to semantic matching.

Recommended citation: Lequeu, P. A., Barboule, C., & Piwowarski, B. (2026). Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders. arXiv preprint arXiv:2605.30022.
Download Paper

talks

The GDN-CC dataset

Mis à jour :

Presented our work “The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations” representing the MLIA team for a day-long seminar on ISIR’s research.

teaching

Teaching 2024-2025

Grad & Undergrad, Sorbonne University, 2024

Introduction to Programming (1st-year bachelor), Introduction to Relational Databases (2nd-year bachelor), Industrial Project (1st-year master)

Teaching 2025-2026

Grad & Undergrad, Sorbonne University, 2025

Data Science (1st-year bachelor), Deep Learning (2nd-year Master)

Teaching 2026-2027

Grad & Undergrad, Sorbonne University, 2026

Data Science (1st-year bachelor), Deep Learning (2nd-year Master)