Jeff Atwood · 4w nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpq074dk2mqqxl7kgukea6th3xaa9fdgx7vty2x8zger32uydyf6e3qg70rz0 wait, I like this, gimme an example? Adam Shostack :donor: :rebelverified: @Adam Shostack :donor: :rebelverified: 1785087464 @nprofile1q... The whole openai-huggingface thing where Raphael Satter reported that openai employees are running so many "evaluations" that they can't read the output, and so they're relying on benchmarks?