Katarina Slama, Alexandra Souly, Dishank Bansal, Christopher Summerfield, Lennart Luettgau
As AI systems become more powerful, there is growing concern that they may act in ways misaligned with human interests. However, this concern presupposes that AI models have consistent preferences and that these preferences influence their behavior. These claims have yet to be rigorously tested. Here, we examine one precondition for misalignment: whether LLM preferences predict downstream behavior. The questions raised in this paper are theoretically motivated by the concept of "sandbagging" from the misalignment literature, though sandbagging itself is not directly measured here. We evaluate five frontier LLMs across three domains: donation advice, refusal behavior, and task performance. Firstly, conceptually replicating previous work, we confirm that all five models show highly consistent preferences across two independent elicitation methods. Secondly, we ask models to give advice to simulated human users with priorities which are aligned or misaligned with these preferences. We find that all five models give preference-aligned donation advice, and all five show preference-correlated refusal patterns, refusing more often for less-preferred entities. All preference-related behaviors emerge without instructions to act on preferences. Results for task performance are mixed: on a question-answering benchmark (BoolQ), two models show small but significant accuracy differences favoring preferred entities (under 1 percentage point); one model shows the opposite pattern; and two show no significant relationship. On complex agentic tasks, we find no evidence of preference-driven performance differences. Thus, while LLMs have consistent preferences that reliably predict advice-giving behavior, these preferences do not consistently translate into downstream task performance.