The AI Epistemic Deference Index: A Continuous Measure of Sycophancy Permalink
Alejandro Botas, Paul de Font-Reaulx, and Luke Hewitt. "The AI Epistemic Deference Index: A Continuous Measure of Sycophancy." arXiv preprint arXiv:2606.07897, 2026.
Alejandro Botas, Paul de Font-Reaulx, and Luke Hewitt. "The AI Epistemic Deference Index: A Continuous Measure of Sycophancy." arXiv preprint arXiv:2606.07897, 2026.
Alexander K. Saeri, Jess Graham, Michael Noetel, Peter Slattery, ... Paul de Font-Reaulx et al. "Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts." MIT AI Risk Initiative, 2026.
Paul de Font-Reaulx. "Reward is Evidence of Value." Philosophy of Science, Forthcoming.
Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx et al. "MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes." Proceedings of ICLR 2026, 2026.
Luke Hewitt, Max Kroner Dale, and Paul de Font-Reaulx. "DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users' Views." Proceedings of IASEAI 2026, 2026.
Paul de Font-Reaulx and Chandra Sripada. "Motivation, Pleasure, and Valence." Philosophia (Symposium), Forthcoming.
Paul de Font-Reaulx. "Do Expected Utility Maximizers Have Commitment Issues?" Philosophy and Phenomenological Research, 112(1): 1-23, 2026.
Paul de Font-Reaulx. "Machine Theory of Mind and the Structure of Human Values." NeurIPS MP2 Workshop, 2023.
Paul de Font-Reaulx. "Generative Theory of Mind and the Value Misgeneralization Problem." 2023. AI Alignment Awards, Final Prize Winner.
Paul de Font-Reaulx. "Alignment as a Dynamic Process." NeurIPS ML Safety Workshop, 2022. AI Risk Analysis Award Winner.
Paul de Font-Reaulx. "What Makes Discrimination Wrong?" Journal of Practical Ethics, 5(2):105-113, 2017.