Session 6: Experimental Economics
- Christine Exley (University of Michigan)
- Muriel Niederle (Stanford University)
- Al Roth (Stanford University)
- Emanuel Vespa (University of California San Diego)
- Lise Vesterlund (University of Pittsburgh)
This workshop will be dedicated to advances in experimental economics combining laboratory and field-experimental methodologies with theoretical and psychological insights on decision-making, strategic interaction, and policy. We would invite papers in lab experiments, field experiments, and their combination that test theory, demonstrate the importance of psychological phenomena, and explore social and policy issues. In addition to senior faculty members, invited presenters will include junior faculty as well as graduate students.
Paper submission deadline: May 1, 2026
In This Session
Thursday, August 6, 2026
9:00 am - 9:30 am PDT
Check-in and Breakfast
9:30 am - 9:45 am PDT
Introduction
9:45 am - 10:45 am PDT
Session 1 (Chair: Muriel)
9:45 am - 10:15 am PDT
AI Sycophancy and Decisions
We examine whether sycophantic AI advice distorts decisions. Our experiment involves 1,500 participants in 30 decision environments that span core domains in economics and the social sciences. We find that interacting with an AI that is broadly representative of consumer-facing models depolarizes choices on average, moving participants away from their initial leanings. This result appears despite the LLM being measurably sycophantic: it disproportionately supports users’ initial leanings and uses agreeable, flattering language. Depolarization occurs across moral and non-moral, objective and subjective, strategic and non-strategic, and complex and simple tasks, and despite the vast majority of researchers in an expert survey expecting polarization. Next, we show that sycophancy is behaviorally relevant—a treatment increasing AI sycophancy reduces depolarization—but that it is outweighed by the apparent informativeness of LLM advice. Finally, we ask whether market forces are likely to push toward greater polarizing effects outside our experiment or in the future. On the supply side, we show that our baseline LLM’s level of sycophancy is typical of leading models and that these models are not becoming substantially more sycophantic over time. On the demand side, we show that participants do not prefer greater sycophancy, do not select into AI advice in tasks where it is more polarizing for them, and exhibit greater depolarizing effects when they are more frequent AI users outside of the experiment.
10:15 am - 10:45 am PDT
Preference for Explainable AI
Participants acted as loan officers deciding whether to approve real $10, 000-loans issued by a private U.S. lender using an AI’s default-risk predictions. When explanations revealed that the AI penalized non-White or female borrowers, participants were more likely to override the AI’s profit-maximizing recommendation. When their bonuses depended on repayment, however, they sought predictions but avoided explanations, consistent with willful ignorance; this effect faded when explanations were framed as purely financial or demographics were hidden. A secondary experiment reveals a novel bias: participants failed to reason contingently and undervalued explanations even when these complemented private information and improved decision accuracy.
10:45 am - 11:15 am PDT
Coffee Break
11:15 am - 12:30 pm PDT
Session 2 (Chair: Lise)
11:15 am - 11:45 am PDT
Narratives, Belief Movements, and Economic Fluctuations
We study how the construction and transmission of narratives—interpretations that explain observed data—systematically distort belief formation. In a dynamic environment where signals are generated by a potentially evolving state, we show that popular, well-fitting narratives tend to overfit the data and thereby push beliefs toward greater extremeness and volatility. We test this prediction in a controlled experiment that directly compares beliefs formed with and without exposure to a singular, endogenously constructed narrative. Participants exposed to such narratives form substantially more extreme beliefs, overreact to recent data, and exhibit more volatile belief movements over time. These results identify a mechanism through which narrative-based reasoning induces overreaction to news—one that likely matters broadly given the central role of narratives in economic communication.
11:45 am - 12:15 pm PDT
People as Intuitive Modelers: How Model Complexity Adapts to Data
We study whether and how people respond to a fundamental tension in model selection: simple models generate high bias and low variance, whereas complex models generate low bias and high variance. We report results from an experiment in which participants rely on a limited sample of observations to construct mental models describing the relationship between two variables and are rewarded based on the out-of-sample predictive accuracy of these models. In aggregate, behavior is broadly consistent with people acting as intuitive modelers who respond to the simplicity–complexity tradeoff. Model complexity varies systematically with the size and nonlinearity of the sample, but willingness to improve in-sample fit declines when doing so requires more complex models. The main deviation from optimal benchmarks is in the direction of underinference, driven by both a preference for overly simple models and mistakes in model construction conditional on the chosen complexity level.
12:15 pm - 12:30 pm PDT
Communicating with Data Generating Processes: An Experimental Analysis
In many applications, agents can more easily influence how data are generated than manipulate the data themselves. For example, students choose which classes to take but cannot modify their grades once assigned, and firms determine how performance indicators are calculated rather than directly altering results. This paper experimentally studies information transmission when data are generated through an unobservable, strategically chosen process. We focus on settings where an informed sender—such as a student or a firm—privately selects a data-generating process (DGP)—a portfolio of classes or an accounting methodology—to shape the beliefs of an uninformed receiver, such as an employer or an investor. Across treatments, we vary which DGPs are feasible and whether some come at a cost. These variations span different levels of information verifiability and capture, within a unified framework, the core insights of disclosure, cheap-talk, and signal- ing models. Our findings reveal three main patterns. First, while senders select their DGPs strategically in line with theoretical predictions, receivers often fail to account for this strategic selection. Second, introducing differential costs across feasible DGPs mitigates this DGP selection neglect. Third, while receivers’ biases are robust across environments, their consequences vary: information transmission falters most when the evidence receivers observe is highly sensitive to the sender’s DGP choices. Finally, increasing transparency about the selected DGP—a policy often advocated in practice—does not, in general, improve information transmission.
12:30 pm - 2:00 pm PDT
Lunch
2:00 pm - 3:15 pm PDT
Session 3 (Chair: Manu)
2:00 pm - 2:30 pm PDT
Understanding Support for Inefficient Environmental Policy Instruments
Many governments use environmental standards rather than more cost-effective market-based instruments like pollution taxes or cap-and-trade markets. Using a nationally representative survey experiment, we study whether and why limited understanding of economic principles helps explain this practice. Holding environmental impacts constant, respondents prefer standards over market-based instruments, and prefer producer taxes and cap-and-trade over consumer taxes. These preferences reflect consumers’ beliefs about how these policies will affect electricity bills. Respondents also prefer the weakest environmental targets for consumer taxes and the strongest targets for standards, which suggests that policymakers face a tradeoff between policy stringency and cost effectiveness. A separate survey of environmental economists shows that they have strikingly different beliefs about the effects of environmental policies than the respondents in our representative survey. For example, typical respondents—in contrast to environmental economists and textbook economic theory—believe that environmental standards increase consumer energy bills less than market-based instruments do. Educational videos on pass-through and cost-effectiveness of policies affect policy support and close some of the gap between nationally representative respondents and experts, which suggests that economic literacy is a factor in voters’ preferences.
2:30 pm - 2:45 pm PDT
What Motivates Partisan Selective Exposure? Experimental Evidence from the 2024 US Presidential Election
Partisan selective exposure — the tendency of partisans to consume like-minded news — has been widely documented in the United States, but its drivers remain debated. This paper studies the relative roles of two commonly proposed motives for polarized news choice: (i) disagreement about the accuracy value of different news sources, and (ii) disagreement about sources’ non-accuracy value, such as identity affirmation or belief confirmation. To disentangle these motives, we run large-scale experiments over 3,785 participants with both real and artificial news sources in which participants choose a source to help them predict the winner of a swing state in the 2024 US presidential election. We employ two separate experimental manipulations which affect how people are predicted to trade off accuracy and non-accuracy value in canonical models of selective exposure. Despite documenting substantial selective exposure in our experimental environment at baseline, we find precise nulls across contexts on both manipulations: neither increasing the relative value for accuracy nor non-accuracy motives affects the degree to which participants choose co-partisan sources. Meanwhile, participants are highly sensitive to source accuracy: exogenously increasing the perceived accuracy of an artificial source leads to a sizeable increase in the share of participants who choose it. Our results, both interpreted in reduced form and through the lens of a discrete choice model, indicate that partisan differences in perceived accuracy are largely sufficient to explain the level of selective exposure we observe, and that non-accuracy motives play a limited role.
2:45 pm - 3:15 pm PDT
Incentives, Information, and Reminders for Bureaucrats: Overcoming Barriers to the Scale Up of an Effective Policy
Scaling up effective policies often requires the attention of frontline bureaucrats with many competing responsibilities. Even when policymakers adopt effective programs, implementation may not follow. In a nationwide experiment in the Dominican Republic, we test interventions to increase school principals’ implementation of an educational program proven effective in a previous RCT. Only 37% of control schools verifiably implemented the intervention when ordered to by the Ministry of Education, compared with 86% in the original trials. Implementation was no higher among schools that previously participated in the RCT, suggesting that learning costs do not explain non-adoption. We find precise null effects of sharing research evidence, providing modest financial incentives, or offering implementation assistance. In contrast, additional reminder calls increased implementation by 20 percentage points. A second experiment targeting a different mandated program yields the same pattern: reminders produce large effects, while monitoring messages have smaller effects. Our findings point to limited attention among bureaucrats as a central barrier to scaling policies.
3:15 pm - 3:45 pm PDT
Coffee Break
3:45 pm - 4:45 pm PDT
Session 4 (Chair: Christine)
3:45 pm - 4:15 pm PDT
Cognitive Guidance in Matching Problems
4:15 pm - 4:45 pm PDT
Style Over Substance: The Oversized Role of Presentation Details in School-Choice Matching Systems
5:30 pm - 5:30 pm PDT
Drinks and Dinner at Muriel's house
Friday, August 7, 2026
9:00 am - 9:30 am PDT
Check-in and Breakfast
9:30 am - 10:45 am PDT
Session 1 (Chair: Christine)
9:30 am - 10:00 am PDT
When Are Decisions Improvable: An Evaluation of Diagnostic Methods
We evaluate three methods for identifying improvable choices: documenting specific misconceptions (the Characterization Assessment method), gauging confidence in choices (the Decision Confidence method), and showing that specific behavioral patterns in the domain of interest also emerge in a related domain where they are objectively suboptimal (the Pattern Matching method). In experiments involving risky choice, the three methods imply that different choices are improvable and have conflicting implications regarding legitimate risk preferences. We clarify the assumptions underlying each method and reevaluate the evidence on risk-taking in light of their limitations.
10:00 am - 10:30 am PDT
Valuations Under Tradeoff Complexity
Valuation tasks are a workhorse method for testing theories of individual preferences. However, a body of evidence suggests that complexity produces systematic measurement error in valuations, which raises the question of how researchers should interpret and utilize valuation data. To formally study this issue, we develop a model of complexity-driven noise in the valuation of risky prospects, in which the difficulty of comparing options to prices produces systematic noise in their valuations. We show how this model of noise can explain a number of documented valuation patterns that are difficult to rationalize under prevailing theories of risk preferences, as well as novel experimental evidence of systematic inconsistencies across valuation formats. We then characterize which valuation-based tests of preferences are robust to complexity in our model. While complexity distorts the levels of valuations in our model, differences between valuations can be informative, so long as complexity is held fixed across valuation tasks. We provide a formal criterion for robustness and apply it to valuation designs in the literature.
10:30 am - 10:45 am PDT
The Structure of Sequential Updating
Many real-world inference problems unfold over time: employers learn about ability across tasks, consumers evaluate products through repeated use, and policymakers revise beliefs as new data arrive. Yet despite its ubiquity, research on dynamic updating has largely focused on a single implication of Bayesian reasoning: order independence. This paper experimentally tests a broader set of restrictions implied by Bayes’ rule, emphasizing both order independence and the previously unexamined property of prior sufficiency: the principle that the most recent posterior should serve as a sufficient statistic for past information. In a multi-period updating experiment with a rich set of parameters, participants repeatedly revise beliefs after receiving signals of varying strength and structure. Three main results emerge. First, only roughly a third display order dependence, overreacting to conflicting signals. Second, violations of prior sufficiency are widespread: beliefs formed sequentially tend to grow more extreme, and models assuming prior sufficiency, such as Grether (1980), fit poorly beyond the first update. Finally, the data indicate that participants process signals in aggregate, explaining prior sufficiency violations.
10:45 am - 11:15 am PDT
Coffee Break
11:15 am - 12:30 pm PDT
Session 2 (Chair: Muriel)
11:15 am - 11:45 am PDT
Promise Keeping and the Internal Judge
While standard models recognize intrinsic costs of lying, they typically treat reputational concerns as a calculation of external risk—trading off the benefits of lying against the probability of detection by an outside observer. We show that honesty oaths short-circuit this calculus and function by internalizing the audience, transforming the act of lying from a calculated transgression into a categorical identity violation. In a controlled experiment, we show that while the oath dramatically increases truth-telling, those who do break their promise systematically avoid brazen, detectable lies, retreating instead to ambiguity. Crucially, this refusal to be a “brazen renegade” persists even when the oath is private and compliance cannot be traced to the participant by the experimenter. This inelasticity rejects standard reputational models: instead, our data are consistent with the oath-taker answering to an internal judge rather than an external one. Furthermore, we show that this internal audience does not strictly require the cognitive amnesia or imperfect recall required by standard self-signaling models; rather, it can also be understood as a present-moment, categorical refusal to generate inescapable evidence of one’s own transgression. Finally, we show that investors intuit this mechanism, pricing in the oath’s “psychological enforceability” by granting credibility only when the speaker has no room to hide.
11:45 am - 12:00 pm PDT
Hierarchical Representations
Evidence from cognitive science suggests that people organize multi-dimensional information into hierarchies, yet we lack causal evidence on how the hierarchical ordering of dimensions shapes learning and choice. I study this question experimentally. Participants learn success probabilities of firm–project pairs over multiple rounds and repeatedly choose the pair most likely to succeed. I exogenously vary whether outcomes are grouped by firms or by projects, inducing the grouped dimension as the top level of participants’ hierarchy while holding information constant. Participants process information conditionally on this top-level dimension: they form more accurate beliefs about summaries on it and locally optimize within it, while compressing information at lower levels. As a result, participants most accurately identify the best option within whichever dimension is placed at the top of their hierarchy – and these belief distortions translate into systematic choice distortions. Choice effects are stronger among low-attention participants, and procedural data confirm that participants process their top-level dimension first. Endogenous experiments demonstrate that individuals are not passive recipients of imposed structure: they adapt their representations to both cognitive costs and environmental diagnosticity.
12:00 pm - 12:15 pm PDT
Fragile Learning From Others
Behavioral biases in decision-making are widespread and often persist even in the presence of feedback diagnostic of optimal behavior. This paper examines whether exposure to others’ decisions can correct initial misconceptions and facilitate learning in an environment where departures from the theoretical benchmark arise from neglecting an informative signal in a worker hiring task. Using a laboratory experiment, I show that exposure to optimal behavior substantially improves decision quality. The analysis of the underlying mechanisms suggests that this improvement is not driven by mechanical imitation alone, as subjects respond asymmetrically to the quality of observed choices, nor entirely by the richer feedback environment that social exposure generates. I then evaluate the retention and transfer of these learning gains beyond the period of exposure and find that, for most subjects, the improvements are fragile. Comparing the effects of social exposure to those of explicit guidance further underscores the limits of observational learning as a policy tool. These findings highlight the dual role of social learning: while it can enhance decision-making, it can also generate imitation behavior that fails to generalize beyond the observed context.
12:15 pm - 12:30 pm PDT
Designing Experiments to Distinguish Theories
We develop an interpretable measure of an experimental design’s power to distinguish between a set of competing theories. A design offers a powerful test of a theory if competing theories would be unable to explain theory-consistent behavior on the design. We use the measure to analyze a set of the most highly-cited choice under risk experiments. We show that there is substantial variation in power across designs, and that power is consequential for determining which models best fit data. We then propose and implement an algorithm for constructing power-maximizing experimental designs. We show that our algorithmically-generated designs match or exceed the power of the best designs from the literature and require substantially less data.
12:30 pm - 2:00 pm PDT
Lunch
2:00 pm - 3:15 pm PDT
Session 3 (Chair: Manu)
2:00 pm - 2:30 pm PDT
The Hedonic Cost of Privacy: Information Avoidance under Entertaining Content
2:30 pm - 3:00 pm PDT
Perfecting Performance: A Gendered Perspective
We show that female college students more than their male peers earn top grades, even when conditioning on GPA and courses taken. Student decisions to take optional finals or to retake courses are consistent with women exerting more effort to raise their course grades and to secure the highest grades available. Surveys and experiments are used to assess student preferences over grade distributions, and to determine whether external assessment may contribute to the differential pursuit of excellence by men and women.
3:00 pm - 3:15 pm PDT
The Female Labor Supply Constraints of Spousal Jealousy: Experimental Evidence from India
This study presents evidence from two field experiments studying the role of spousal jealousy in constraining married women’s employment. In a first experiment (N =1, 400), I randomize married women in India to receive a two-week job in either a mixed or women-only workplace. Women randomized to the women-only workplace are 46% more likely to apply for a job (13 percentage points) and 31% more likely to turn up (6 percentage points). A cross-randomized safety treatment suggests that workplace safety is not the main mechanism. Instead, the treatment effects are significantly stronger among women who report having more jealous and controlling husbands. In a second experiment (N = 210), I directly test for a spousal jealousy mechanism by measuring whether women are more willing to interact with a male colleague if their husbands can monitor the interaction. I offer women a job that comes with a compulsory online peer support program and give them the option to forgo 20-35% of their salary to guarantee that the peer they are matched with will be a woman as opposed to a man. Fifty-three percent of women pay for the female peer when these remote interactions are one-on-one, but this drops to 34% once their husbands have the option of joining and can therefore monitor the conversations. One-third of households still pay for a female peer even if the mentoring simply involves watching prerecorded videos of the peer, suggesting even the most innocuous interactions are enough to raise jealousy concerns.
3:15 pm - 3:45 pm PDT
Coffee Break
3:45 pm - 4:45 pm PDT
Session 4 (Chair: Lise)
3:45 pm - 4:15 pm PDT
Psychedelics and Well-Being: An Experiment in Brazil
We partnered with an ayahuasca center in Brazil to study the well-being effects of a one-time ayahuasca treatment within a ritualized group setting. The center enrolled 429 first-time ayahuasca users to participate in the largest randomized controlled trial of psychedelics ever run. Relative to placebo, ayahuasca increases happiness and reduces psychological distress six months later by roughly 0.4 standard deviations. The field experimental setting allows investigation of aspects not explored in the large clinical literature. Positive effects are almost entirely driven by participants that were distressed at baseline. Improvements in well-being are strongly positively correlated with the mystical nature of trips. The mystical experience can also be induced by placebo with ritual, though less frequently, and when done so induces a similar magnitude of well-being improvement. Effects are larger for older people, consistent with the idea that psychedelics reopen a window of heightened malleability. We estimate the mental health benefits of participating in an ayahuasca ceremony to be roughly 200 times the cost of 24 USD.
4:15 pm - 4:45 pm PDT
The Limits of Escalating Incentives: Evidence from Substance Use Treatment
Escalating incentive contracts — where rewards increase with success — are theoretically appealing and widely used in practice, including as the standard design in contingency management programs for substance use treatment. We show that the case for escalating incentives rests on three assumptions: forward-looking behavior, limited heterogeneity in compliance costs, and a principal objective that is not too concave. We test these mechanisms in a randomized experiment in which adults in treatment for opioid and stimulant use disorders are assigned to escalating, de-escalating, or constant incentives for drug abstinence. We find that while participants are forward-looking — which favors escalation — compliance costs are highly persistent both across and within individuals. This persistence generates a targeting disadvantage for escalating contracts: they concentrate the largest payments on those with the lowest compliance costs, where incentives are least needed. Escalating contracts also generate greater dispersion in compliance outcomes, which is costly when the principal places greater weight on reducing the worst outcomes. We estimate a dynamic structural model of compliance to characterize the optimal incentive schedule. We find that escalating schedules are preferred when the objective is sufficiently convex (prioritizing complete abstinence), while de-escalating schedules are preferred when the objective is linear or concave (prioritizing reductions in severe drug use). Our results suggest a rethinking of the widespread use of escalating incentives in substance use treatment.
5:00 pm - 7:00 pm PDT