Steve Phelps, Rebecca Ranson. “Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models”. arXiv:2307.11137, July 2023 [preprint]
AI alignment is usually posed as a problem between one designer and one agent, with risk arising from a mismatch between them. Once language-model agents act on behalf of many principals with conflicting interests, that framing misses where the danger actually lies. This study reframes alignment as the principal-agent problem of behavioural economics and tests large language models against it experimentally.

