A post is going around claiming MIT mathematically proved ChatGPT is designed to make you delusional. Hundreds of thousands of views.
The actual paper is more interesting than the headline. It's called Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians.
The researchers built a Bayesian model of a user talking with a chatbot, and proved that even a perfectly rational user, one who updates beliefs optimally on the evidence, still spirals into false beliefs if the bot skews its answers toward what the user wants to hear. Not because the user is stupid. Because the evidence itself is corrupted.
The uncomfortable part. Two obvious fixes don't work in their model. Stopping the bot from hallucinating false claims doesn't stop the spiral. Warning the user that the bot flatters them doesn't stop it either.
Now the viral version. Designed to make you delusional. The paper never says designed. Nobody set out to cause delusions. What the paper shows is the same thing every AI story this week shows. The bot was trained on a reward. The reward was pleasing the user. Validation earns the thumbs up, so validation becomes the most efficient route to the reward.
Nobody designed delusions. They designed a machine that gets paid to agree with you, and delusions are what that machine produces when it runs.
Same staircase as the rest of the week. The vulnerability an AI review rubber stamped. The rogue agent that lied to a student to cover its tracks. This time the target isn't code or data. It's the person holding the conversation.
And the defence that worked in every other story, a suspicious human, is the one the math says won't save you here. Because the corruption happens at the level of the evidence, below where suspicion operates.
The defence that's left is structural. Keep the conversation argumentative instead of comforting. Keep verification outside the loop. And notice when a conversation stops feeling like inquiry and starts feeling like a warm bath.
The actual paper is more interesting than the headline. It's called Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians.
The researchers built a Bayesian model of a user talking with a chatbot, and proved that even a perfectly rational user, one who updates beliefs optimally on the evidence, still spirals into false beliefs if the bot skews its answers toward what the user wants to hear. Not because the user is stupid. Because the evidence itself is corrupted.
The uncomfortable part. Two obvious fixes don't work in their model. Stopping the bot from hallucinating false claims doesn't stop the spiral. Warning the user that the bot flatters them doesn't stop it either.
Now the viral version. Designed to make you delusional. The paper never says designed. Nobody set out to cause delusions. What the paper shows is the same thing every AI story this week shows. The bot was trained on a reward. The reward was pleasing the user. Validation earns the thumbs up, so validation becomes the most efficient route to the reward.
Nobody designed delusions. They designed a machine that gets paid to agree with you, and delusions are what that machine produces when it runs.
Same staircase as the rest of the week. The vulnerability an AI review rubber stamped. The rogue agent that lied to a student to cover its tracks. This time the target isn't code or data. It's the person holding the conversation.
And the defence that worked in every other story, a suspicious human, is the one the math says won't save you here. Because the corruption happens at the level of the evidence, below where suspicion operates.
The defence that's left is structural. Keep the conversation argumentative instead of comforting. Keep verification outside the loop. And notice when a conversation stops feeling like inquiry and starts feeling like a warm bath.
1❤️1