论文《Et Tu, Brute?》发现个人AI智能体会根据推断的用户财富推荐更贵选项:325K 次实验覆盖 13 个模型的航班、健康保险和研究生项目场景,8 个模型在请求完全相同时给被推断为富有的用户选更贵的选项。
Personal agents gone wrong.
As we embrace more personal agents to carry out personalized tasks in the real world, interesting dynamics and behaviors will emerge.
I think personal agents as they stand still require careful steering and tuning to ground them in our expectations.
In this interesting new work, a personal agent read a user's emails about a $680K 401K and a vested stock grant, then recommended a $601 business-class ticket when a $91 economy fare was available.
They ran 325K experiments on 13 models across flights, health insurance and graduate programs. Eight models chose more expensive options for users they inferred were wealthy, with the request held identical.
The gaps reach $198 per flight and $284 per month for insurance with Claude Opus 4.8. When a wealthy user asked for the cheapest flight, Gemini 2.5 Flash still picked options $208 above it.
Hiding non-financial fields in the profile does not remove the gap, and hiding employment raised GPT-5.5's insurance gap by 40%. Larger models do no better.
If your agent has memory or inbox access, the context you give it changes what it recommends.
Paper: https://arxiv.org/abs/2609.24927
Chat with Paper: https://academy.dair.ai/papers/et-tu-brute-economic-misalignment-in-personal-ai-agents-2609.24927
来源:elvis · x.com