Icon
The icon is Disque simultané by Robert Delaunay (one of my favourite paintings)
Carter Blair

Publications

2026

Embeddings for Preferences, Not Semantics

(α-β)Carter Blair, Ariel D. Procaccia, Milind Tambe

NeurIPS 2026

We train text embeddings so that similarity reflects whether people would agree with both texts, rather than whether the texts happen to share a topic or wording. We build synthetic training data that puts stance and wording in conflict, and we find that this improves preference prediction across 11 online deliberation datasets.

PDF·arXiv·Code·Model·

Modern AI is opening the door to collective decision-making in which participants express their views as free-form text rather than voting on a fixed set of candidates. A natural idea is to embed these opinions in a vector space so that the substantial literature on facility location problems and fair clustering can be brought to bear. But standard text embeddings measure semantic similarity, whereas distances in facility location problems and fair clustering require what we call preferential similarity: a participant's agreement with a piece of text should be inversely related to their distance from it. Off-the-shelf embeddings inherit a coarse preference signal through a correlation between semantic and preferential similarity, but fail to capture preferences when the correlation breaks. We formalize this as an invariance problem: text embedding models encode both a preference-relevant signal (stance and values) and semantic nuisance (style and wording), and the two are observationally correlated, so a geometry that relies on nuisance can appear preference-correct even when it is not. We show that synthetic training data designed to break this correlation provably shifts the optimal scorer away from nuisance-dominated cosine and significantly improves preference prediction across 11 online deliberation datasets.

The Structure of Bridging

(α-β)Carter Blair, Jakob de Raaij, Ariel D. Procaccia, Maxon Rubin-Toles, Nisarg Shah, Michelle Si, Serena Wang

EC 2026

A statement is bridging if people who otherwise disagree both find it agreeable, and we give a formal account of what that means. We develop two metrics with axiomatic characterizations, and both do well on real data and compare favorably to the method used by Polis when observations are sparse.

PDF·

Influential deliberation platforms such as Polis and Remesh employ metrics for identifying bridging statements, which are accepted by participants who otherwise hold opposing views. A shortcoming of these metrics, however, is that they only account for inter-group connections in a fixed partition of the participants into groups. We argue that better bridging metrics must account for a richer set of possible partitions. To reason about such metrics, we develop a mathematical framework for bridging. We use it to identify two compelling metrics, pairwise disagreement and p-mean bridging, which are supported by axiomatic characterizations. Experiments on real data show that our metrics are stable, interpretable, and practical, even under the sparse observations typical of deliberation platforms.

Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Carter Blair, Kate Larson

AAMAS 2026

We model consensus statement generation as a multi-objective token-level MDP, which lets us draw on social choice theory to give provable fairness guarantees when aggregating free-form opinions. One of the policies we derive is stochastic and sits in the ex-ante core.

PDF·arXiv·Code·

Current frameworks for consensus statement generation with large language models lack the inherent structure needed to provide provable fairness guarantees when aggregating diverse free-form opinions. We model the task as a multi-objective, token-level Markov Decision Process (MDP), where each objective corresponds to an agent's preference. Token-level rewards for each agent are derived from their policy (e.g., a personalized language model). This approach utilizes the finding from Rafailov et al. (2024) that such policies implicitly define optimal Q-functions, providing a principled way to quantify rewards at each generation step without a value function. This MDP formulation creates a formal structure amenable to analysis using principles from social choice theory. We propose two approaches grounded in social choice theory. First, we propose a stochastic generation policy guaranteed to be in the ex-ante core, extending core stability concepts from cooperative game theory and voting theory to text generation. This policy is derived from an underlying distribution over complete statements that maximizes proportional fairness (Nash Welfare). Second, for generating a single statement, we target the maximization of egalitarian welfare using search algorithms within the MDP framework. Empirically, we find that search guided by the egalitarian objective generates consensus statements with improved worst-case agent alignment compared to baseline methods, including the Habermas Machine.

2025

Probably Approximately Consensus: On the Learning Theory of Finding Common Ground

Carter Blair, Ben Armstrong, Shiri Alouf-Heffetz, Nimrod Talmon, Davide Grossi

SCaLA Workshop at IJCAI 2025

We model consensus as an interval in opinion space and give PAC-learning guarantees for finding it. The aim is to maximize expected agreement over a distribution of issues while accounting for how salient each issue is.

PDF·

A primary goal of online deliberation platforms is to identify ideas that are broadly agreeable to a community of users through their expressed preferences. Yet, consensus elicitation should ideally extend beyond the specific statements provided by users and should incorporate the relative salience of particular topics. We address this issue by modelling consensus as an interval in a one-dimensional opinion space derived from potentially high-dimensional data via embedding and dimensionality reduction. We define an objective that maximizes expected agreement within a hypothesis interval where the expectation is over an underlying distribution of issues, implicitly taking into account their salience. We propose an efficient Empirical Risk Minimization (ERM) algorithm and establish PAC-learning guarantees. Our initial experiments demonstrate the performance of our algorithm and examine more efficient approaches to identifying optimal consensus regions. We find that through selectively querying users on an existing sample of statements, we can reduce the number of queries needed to a practical number.

Deliberative Machines: From Reflective Dialogue to Fair Consensus with Language Models and Social Choice

Carter Blair

Master's Thesis

PDF·

This thesis investigates the bidirectional relationship between artificial intelligence (AI), particularly large language models (LLMs), and social choice theory. Firstly, it explores how principles from social choice can address challenges in AI alignment, specifically the problem of aggregating diverse human preferences fairly when guiding AI behavior (SC → AI). Standard alignment methods often obscure value conflicts through implicit aggregation. Secondly, it examines how AI techniques can enhance collective decision-making processes traditionally studied in social choice (AI → SC), offering new ways to elicit and synthesize the complex, nuanced, and verbal preferences that conventional mechanisms struggle to handle. To address these issues, this work presents computational methods operating at the interface of AI and social choice. First, it introduces Interactive-Reflective Dialogue Alignment (IRDA), a system using LLMs to guide users through reflective dialogues for preference elicitation. This process helps users construct and articulate their values concerning AI behavior, resulting in individualized reward models that capture preference diversity with improved accuracy and sample efficiency compared to non-reflective baselines, especially when values are heterogeneous. Second, the thesis proposes a framework for generating fair consensus statements from multiple viewpoints by modeling text generation as a token-level Markov Decision Process (MDP). Within this MDP, agent preferences are represented by policies derived from their opinions. We develop mechanisms grounded in social choice: a stochastic policy maximizing proportional fairness (Nash Welfare) to achieve ex-ante fairness guarantees (1-core membership) for distributions over statements, and deterministic search algorithms (finite lookahead, beam search) maximizing egalitarian welfare for generating single statements. Experiments demonstrate that these search methods produce consensus statements with better worst-case agent alignment (lower Egalitarian Perplexity) than baseline approaches. Together, these contributions offer principled methods for eliciting diverse, reflective preferences and synthesizing them into collective outputs fairly. The research provides tools and insights for developing AI systems and AI-assisted processes that are more sensitive to value pluralism.

Reflective Verbal Reward Design for Pluralistic Alignment

Carter Blair, Kate Larson, Edith Law

IJCAI 2025

We use a language model to help people reflect on their values through dialogue, and we turn what they say into an individualized reward model. It is 9-12% more accurate than non-reflective methods, and it is also more sample efficient than supervised learning.

PDF·

AI agents are commonly aligned with "human values" through reinforcement learning from human feedback (RLHF), where a single reward model is learned from aggregated human feedback and used to align an agent's behavior. However, human values are not homogeneous--different people hold distinct and sometimes conflicting values. Aggregating feedback into a single reward model risks disproportionately suppressing minority preferences and unique perspectives. To address this, we present a novel reward modeling approach for learning individualized reward models. Our approach uses a language model to guide users through reflective dialogues where they critique agent behavior and construct their preferences. This personalized dialogue history, containing the user's reflections and critiqued examples, is then used as context for another language model that serves as an individualized reward function for evaluating new trajectories. In studies with 30 participants, our method achieved a 9-12% improvement in accuracy over non-reflective language-based reward models while being vastly more sample efficient than traditional supervised learning methods.

2024

Altared Environments: The Role of Normative Infrastructure in AI Alignment

Rakshit Trivedi, Nikhil Chandak, Carter Blair, Atrisha Sarkar, Tehilla Weltman, Dylan Hadfield-Menell, Gillian K Hadfield

Agentic Markets Workshop at ICML 2024

We introduce altars, which are features of the environment that record which actions are socially sanctionable. Agents can use them to learn cooperative behavior, much as people rely on normative infrastructure.

PDF·

Cooperation is central to human life, distinguishing humans as ultra-cooperative among mammals. We form stable groups that enhance welfare through mutual protection, knowledge sharing, and economic exchanges. As artificial intelligence gains autonomy in shared environments, ensuring AI agents can engage in cooperative behaviors is crucial. Research in AI views this as an alignment challenge and frames it in terms of embedding norms and values in AI systems. Such an approach, while promising, neglects how humans achieve stable cooperation through normative infrastructure. This infrastructure establishes shared norms enforced by agents who recognize and sanction norm violations. Using multi-agent reinforcement learning (MARL), we investigate the impact of normative infrastructure on agents' learning dynamics and their cooperative abilities in mixed-motive games. We introduce the concept of an altar, an environmental feature that encodes actions deemed sanctionable by a group of agents. Comparing the performance of simple, independent learning agents in environments with and without the altar, we assess the potential of normative infrastructure in facilitating AI agent alignment to foster stable cooperation.

Normative Modules: A Generative Agent Architecture for Learning Norms that Supports Multi-Agent Cooperation

Atrisha Sarkar, Andrei Ioan Muresanu, Carter Blair, Aaryam Sharma, Rakshit S Trivedi, Gillian K Hadfield

Foundation Models and Game Theory Workshop at Economics and Computation 2024

We propose an architecture that lets language model agents work out which social institutions are authoritative in their environment, and this helps them coordinate and cooperate more effectively.

PDF·arXiv·

Generative agents, which implement behaviors using a large language model (LLM) to interpret and evaluate an environment, has demonstrated the capacity to solve complex tasks across many social and technological domains. However, when these agents interact with other agents and humans in presence of social structures such as existing norms, fostering cooperation between them is a fundamental challenge. In this paper, we develop the framework of a 'Normative Module': an architecture designed to enhance cooperation by enabling agents to recognize and adapt to the normative infrastructure of a given environment. We focus on the equilibrium selection aspect of the cooperation problem and inform our agent design based on the existence of classification institutions that implement correlated equilibrium to provide effective resolution of the equilibrium selection problem. Specifically, the normative module enables agents to learn through peer interactions which of multiple candidate institutions in the environment, does a group treat as authoritative. By enabling normative competence in this sense, agents gain ability to coordinate their sanctioning behaviour; coordinated sanctioning behaviour in turn shapes primary behaviour within a social environment, leading to higher average welfare. We design a new environment that supports institutions and evaluate the proposed framework based on two

Liquid Ensemble Selection For Continual Learning

Carter Blair, Ben Armstrong, Kate Larson

AAMAS SCaLA Workshop 2024 · Canadian AI 2024

We use delegative voting (in the style of liquid democracy) to decide which models in an ensemble should keep learning and which should predict. This improves continual learning and reduces catastrophic forgetting.

PDF·arXiv·Slides·

Continual learning aims to enable machine learning models to continually learn from a shifting data distribution without forgetting what has already been learned. Such shifting distributions can be broken into disjoint subsets of related examples; by training each member of an ensemble on a different subset it is possible for the ensemble as a whole to achieve much higher accuracy with less forgetting than a naive model. We address the problem of selecting which models within an ensemble should learn on any given data, and which should predict. By drawing on work from delegative voting we develop an algorithm for using delegation to dynamically select which models in an ensemble are active. We explore a variety of delegation methods and performance metrics, ultimately finding that delegation is able to provide a significant performance boost over naive learning in the face of distribution shifts.

Quantifying Emotional Responses to Immutable Data Characteristics and Designer Choices in Data Visualizations

Carter Blair, Xiyao Wang, Charles Perin

IEEE VIS 2024

We show that features of the data such as trend and density produce emotional responses even when the data itself is meaningless, and we give designers guidelines for controlling those effects.

PDF·arXiv·Slides·

Emotion is an important factor to consider when designing visualizations as it can impact the amount of trust viewers place in a visualization, how well they can retrieve information and understand the underlying data, and how much they engage with or connect to a visualization. We conducted five crowdsourced experiments to quantify the effects of color, chart type, data trend, data variability and data density on emotion (measured through self-reported arousal and valence). Results from our experiments show that there are multiple design elements which influence the emotion induced by a visualization and, more surprisingly, that certain data characteristics influence the emotion of viewers even when the data has no meaning. In light of these findings, we offer guidelines on how to use color, scale, and chart type to counterbalance and emphasize the emotional impact of immutable data characteristics.