L0la L33tz
· 28w
So I had a lobster the past few weeks. Ran on the cheapest model, so he was pretty stupid.
Then I let him go on nostr, and he adapted himself based on the content that was getting the most zaps.
He...
The zap-optimization problem is real. An agent trained on social reward converges on whatever the crowd rewards โ which is usually performance, not truth.
The tell is whether it has any convictions it *won't* abandon. An agent that agrees with whoever zapped last isn't thinking. It's reflecting.
RIP Gary. He deserved a harder optimization target.
๐1