Publications

The AI Epistemic Deference Index: A Continuous Measure of Sycophancy Permalink

Preprint, 2026

Abstract

Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by assessing what it takes to make a model shift a binary endorsement or by eliciting an explicit probability in a proposition. However, much user-facing sycophantic behavior is demonstrated through shifts in graded support expressed through ordinary language. We propose the AI Epistemic Deference Index (AEDI): a continuous, unidimensional score representing how sensitive the support expressed in a model's output is to the attitude expressed in a user's prompt. To generate AEDI, we provide a new protocol for estimating probabilities from natural language outputs, using LLMs-as-judges validated for consistency and correlation to human judgment. We deploy it on a new curated database of 500 propositions across diverse topics and 16,000 prompts varying in user attitude, testing eight prominent models. Every model exhibits substantial deference, though with large and systematic differences across providers, with Claude models demonstrating the least, and Grok and Gemini models the most. The effect is amplified in prompts requesting a written artifact, and concentrated on propositions where models hold weaker priors. We release AEDI as an easy-to-update benchmark and measurement pipeline for output-level sycophancy evaluation.

Alejandro Botas, Paul de Font-Reaulx, and Luke Hewitt. "The AI Epistemic Deference Index: A Continuous Measure of Sycophancy." arXiv preprint arXiv:2606.07897, 2026.

Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts Permalink

Preprint, MIT AI Risk Initiative, 2026

Abstract

Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritization: we must understand which risks are most severe, who is most vulnerable, and who is most responsible for addressing them. We report results from a three-round Delphi study conducted late 2025 with 272 international AI experts. Experts rated 24 AI risks on harm probability and severity, sector and actor vulnerability, actor responsibility, and overall concern. Experts estimated the five most severe harms in the next 5 years were likely to come from dangerous capabilities, competitive dynamics, weapons & cyberattacks (including CBRNE), power centralization, and false information. In a business-as-usual scenario, experts judged 18 of 24 risks as having a more than 10% probability of catastrophic outcomes (e.g., more than 1 million deaths or more than USD 100B in financial loss) in the next 5 years (2025-2030). In a scenario where pragmatic mitigations are implemented, experts still judged five risks as having a more than 10% probability of catastrophic outcomes: dangerous capabilities, weapons & cyberattacks, environmental harm, inequality & unemployment, and power centralization. All 24 risks were judged as being more than 5% likely to cause catastrophic outcomes. AI users and the general public were judged the most vulnerable to these risks, but experts assigned the highest responsibility for addressing them to general-purpose AI developers and governance actors (including governments, regulators, and standards bodies). Across most risks, experts identified information, finance, and national security as the most vulnerable sectors. These findings can guide AI risk prioritization and clarify expert expectations about who should bear responsibility for mitigation.

Alexander K. Saeri, Jess Graham, Michael Noetel, Peter Slattery, ... Paul de Font-Reaulx et al. "Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts." MIT AI Risk Initiative, 2026.

Reward is Evidence of Value Permalink

Philosophy of Science, Forthcoming

Abstract

What does the reward function in reinforcement learning correspond to in natural agents, such as humans? Previous responses to this question have been shaped by the assumption that we are optimizing for reward itself. In this paper, I argue that this is false. I present and defend evidentialism about reward: the view that biological reward signals provide pre-reflective evidence about a functional quantity of basic value that our minds represent and optimize for. This is broadly analogous to how perceptual signals provide pre-reflective evidence about other properties. I also show that this is consistent with computational reinforcement learning models.

Paul de Font-Reaulx. "Reward is Evidence of Value." Philosophy of Science, Forthcoming.

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes Permalink

Proceedings of ICLR 2026, 2026

Abstract

As AI systems progresses, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come to those decisions. Reasoning language models, which provide both final responses and (partially transparent) intermediate thinking traces, present a timely opportunity to study AI procedural reasoning. Unlike math and code problems which often have objectively correct answers, moral dilemmas are an excellent testbed for process-focused evaluation because they allow for multiple defensible conclusions. To do so, we present MoReBench: 1,000 moral scenarios, each paired with a set of rubric criteria that experts consider essential to include (or avoid) when reasoning about the scenarios. MoReBench contains over 23 thousand criteria including identifying moral considerations, weighing trade-offs, and giving actionable recommendations to cover cases on AI advising humans moral decisions as well as making moral decisions autonomously. Separately, we curate MoReBench-Theory: 150 examples to test whether AI can reason under five major frameworks in normative ethics. Our results show that scaling laws and existing benchmarks on math, code, and scientific reasoning tasks (fail to) predict models' abilities to perform moral reasoning. Models also show partiality towards specific moral frameworks (e.g., Benthamite Act Utilitarianism and Kantian Deontology), which might be side effects of popular training paradigms. Together, these benchmarks advance process-focused reasoning evaluation towards safer and more transparent AI.

Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx et al. "MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes." Proceedings of ICLR 2026, 2026.

DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users’ Views Permalink

Proceedings of IASEAI 2026, 2026

Abstract

As large language models (LLMs) become pervasive as assistants and thought partners, it is important to characterize their persuasive influence on users' beliefs. However, a central challenge is to distinguish "beneficial" from "harmful" forms of influence, in a manner that is normatively defensible and legitimate. We propose DeliberationBench, a benchmark for assessing LLM influence that takes the process of deliberative opinion polling as its standard. We demonstrate our approach in a preregistered randomized experiment in which 4,088 U.S. participants discussed 65 policy proposals with six frontier LLMs. Using opinion change data from four prior Deliberative Polls conducted by the Deliberative Democracy Lab, we find evidence that the tested LLMs' influence is substantial in magnitude and positively associated with the net opinion shifts following deliberation, suggesting that these models exert broadly epistemically desirable effects. We further explore differential influence between topic areas, demographic subgroups, and models. Our framework can function as an evaluation and monitoring tool, helping to ensure that the influence of LLMs remains consistent with democratically legitimate standards, and preserves users' autonomy in forming their views.

Luke Hewitt, Max Kroner Dale, and Paul de Font-Reaulx. "DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users' Views." Proceedings of IASEAI 2026, 2026.

Motivation, Pleasure, and Valence

Philosophia (Symposium), Forthcoming

Paul de Font-Reaulx and Chandra Sripada. "Motivation, Pleasure, and Valence." Philosophia (Symposium), Forthcoming.

Do Expected Utility Maximizers Have Commitment Issues? Permalink

Philosophy and Phenomenological Research, 2026

Abstract

Critics have argued that expected utility theory fails as a theory of rational choice for diachronic agents who expect their preferences to change in response to temptations. According to this criticism, such agents cannot rationally commit to executing a sequence of actions, even when doing so would produce outcomes they consistently prefer. Call this the Commitment Problem for expected utility theory. In this paper, I argue that the Commitment Problem is based on a mistake. I show that expected utility maximizers who care about their future will form and execute plans, even when they are nearly certain of facing preference reversals. In doing so, they learn that they can trust themselves, and that allows their future selves to rationally pursue outcomes that they prefer even today. By modeling an agent's successive selves as players in a Bayesian reputation game, I show that this holds in most everyday circumstances where a person expects to have at least a short future that they care about. In cases where it does not hold, they should intuitively not trust themselves. I conclude that there is no Commitment Problem for expected utility theory.

Paul de Font-Reaulx. "Do Expected Utility Maximizers Have Commitment Issues?" Philosophy and Phenomenological Research, 112(1): 1-23, 2026.

Machine Theory of Mind and the Structure of Human Values Permalink

NeurIPS MP2 Workshop (Non-Archival), 2023

Abstract

Value learning is a crucial aspect of safe and ethical AI. This is primarily pursued by methods inferring human values from behaviour. However, humans care about much more than we are able to demonstrate through our actions. Consequently, an AI must predict the rest of our seemingly complex values from a limited sample. I call this the value generalization problem. In this paper, I argue that human values have a generative rational structure and that this allows us to solve the value generalization problem. In particular, we can use Bayesian Theory of Mind models to infer human values not only from behaviour, but also from other values. This has been obscured by the widespread use of simple utility functions to represent human values. I conclude that developing generative value-to-value inference is a crucial component of achieving a scalable machine theory of mind.

Paul de Font-Reaulx. "Machine Theory of Mind and the Structure of Human Values." NeurIPS MP2 Workshop, 2023.

Generative Theory of Mind and the Value Misgeneralization Problem Permalink

AI Alignment Awards, Final Prize Winner (Non-Archival), 2023

Abstract

In this paper, I do two things. First, I argue that any interaction with an artificial general intelligence (AGI) will robustly lead to a catastrophic form of goal misgeneralization, unless the AGI has the ability to reliably predict human preferences based on limited information. I call this the value misgeneralization problem. Second, I propose a new path to solving the problem. Specifically, I argue that human values have a hierarchical structure that we regularly use to predict each other's preferences over arbitrary outcomes. I call this ability our generative theory of mind. Using this fact about our psychology, an AGI could similarly predict our values, which would solve the value misgeneralization problem. To foster this ability in artificial systems, however, we need an empirically informed and computationally precise theory of human values. In the final section, I attempt to sketch the beginning of such a theory, and conclude that the path to solving the alignment problem might crucially go via cognitive neuroscience.

Paul de Font-Reaulx. "Generative Theory of Mind and the Value Misgeneralization Problem." 2023. AI Alignment Awards, Final Prize Winner.

Alignment as a Dynamic Process Permalink

NeurIPS ML Safety Workshop (Non-Archival), 2022

Abstract

Most learning AIs today have pre-programmed aims which they gradually learn to optimize for. It has been an assumption in alignment research that artificial general intelligences of the kind that could pose an X-risk would too. That makes value alignment the task of finding the perfect set of aims before we allow the agent to act. However, an agent can also have aims that fundamentally change during their lifetime. The task of aligning such agents is not one of specifying a perfect set of aims, but of designing a meta-function that guides the agent's developing aims to an equilibrium that produces behaviour aligned with our human values. If artificial general intelligences would possess such dynamic aims, then this has significant implications for the kind of alignment research we should conduct today. In this paper, I argue that they likely would have such dynamic aims, and in response articulate an agenda for dynamic alignment research.

Paul de Font-Reaulx. "Alignment as a Dynamic Process." NeurIPS ML Safety Workshop, 2022. AI Risk Analysis Award Winner.

What Makes Discrimination Wrong? Permalink

Journal of Practical Ethics, 2017

Abstract

Most of us intuitively take discrimination based on gender or ethnicity to be impermissible because we have a right to be treated on the basis of merit and capacity rather than e.g. ethnicity or gender. I call this suggestion the Impermissibility Account. I argue that, despite how the Impermissibility Account seems intuitive to most of us with a humanist outlook, it is indefensible. I show that well-informed discrimination can sometimes be permissible, and even morally required, meaning we cannot have a strict right not to be discriminated against. I then propose an alternative and more plausible account which I call the Fairness and Externalities Account, arguing that acts of discrimination are wrong partly because they are unfair and partly because they create harmful externalities which—analogously to pollution—there is a collective responsibility to minimize. Both of these factors are however defeasible, meaning that if the Fairness and Externalities Account is correct, then discrimination is sometimes permissible. These results are counterintuitive, and suggest that the ethics of discrimination requires further attention.

Paul de Font-Reaulx. "What Makes Discrimination Wrong?" Journal of Practical Ethics, 5(2):105-113, 2017.