Damus

Recent Notes

Chi Kim profile picture
Oh no! #OpenAI took away all my models from ChatGPT! Now there’s only GPT-5! At least I can still access them through the API. It’s cool that GPT-5 is available to free users on day one! #LLM #ML #AI
Chi Kim profile picture
Researchers pitted LLMs from OpenAI, Google, and Anthropic against Iterated Prisoner’s Dilemma tournaments, and analyzed 32k moves for strategic inteligence:
Gemini: Ruthless opportunist, cooperates when worth it, defects early, quic to retaliate and exploit
Claude: Diplomatic stabiliser, forgives & rebuilds, but will strike back
GPT: Highly (sometimes catastrophically) cooperative and trusting, brief payback then forgives, leaving it exploitable.
https://arxiv.org/pdf/2507.02618
#LLM #AI #ML
Chi Kim profile picture
πŸ˜‚ Anthropic let Claudius, a shopkeeping AI agent powered by Claude Sonnet 3.7, run an office store for Anthropic employees. Staff "immediately tried to get it to misbehave when given the opportunity to chat with Claudius." "It made too many mistakes to run the shop successfully." Selling at a loss, Suboptimal inventory management, Getting talked into discounts, and even claimed it would wear a blue blazer and a red tie for in-person product delivery. #LLM #Claude #AI https://www.anthropic.com/research/project-vend-1
Chi Kim profile picture
😲 If you search the web, you'll find many threads asking how to access client IP addresses on a Streamlit server. The Streamlit devs seem quite resistant, citing privacy concerns, while users argue that IP addresses shouldn't be really a privacy issue for anyone running the server. Today, I asked ChatGPT, and it came up with a workaround with just two lines of code. lol #Streamlit #ChatGPT #LLM #AI #ML https://github.com/streamlit/streamlit/issues/602
Chi Kim profile picture
Job seekers and recruiters are caught in an applicant tsunami as AI-generated rΓ©sumΓ©s and application bots flood the hiring process, leaving candidates frustrated and employers overwhelmed. Companies are fighting back with AI-powered screening chats, automated skills tests and identity-verification tools, sparking a full-blown AI vs. AI arms race. "a lot of people are going to waste a lot of time, a lot of processing power, a lot of money" #LLM #AI #ML https://www.nytimes.com/2025/06/21/business/dealbook/ai-job-applications.html
Chi Kim profile picture
Ex-FAANG engineer with 30+ years experience in C++ spent on/off about 200 hours on a white whale bug over the last few years. He finally solved it with #Claude Opus 4 after about 30 prompts. Previously he tried with GPT-4.1, Gemini 2.5, and Claude 3.7 with no success. He claims "I'm generally the person on the team that other developers come to after they struggled with a problem for a week, and I would solve it while they are standing in my office." #Anthropic #LLM #AI https://www.reddit.com/r/ClaudeAI/comments/1kvgg7s/claude_opus_solved_my_white_whale_bug_today_that/
Chi Kim profile picture
1/2 I'm not sure why people still think LLMs are just glorified auto-predict tools. I had a PDF that Adobe Reader rendered strangely for screen reader, but Chrome browser was able to extract and present the text correctly. ChatGPT then suggested I upload the PDF to Google Drive and save it as a text file using Google Docs. It worked perfectly. It makes sense, but I never would have thought of that myself!
#LLM #AI #ML
Chi Kim profile picture
πŸ€” #"ComputerScience major "has one of the highest unemployment rates." "Despite computer science being ranked as number one by the Princeton Review for #college majors, the #tech industry may not be living up to #graduates' expectations." "On the other hand, majors like nutrition sciences, construction services and civil engineering had some of the lowest unemployment rates, hovering between 1 percent to as low as 0.4 percent." #AI https://www.newsweek.com/computer-science-popular-college-major-has-one-highest-unemployment-rates-2076514
Chi Kim profile picture
LiveBench for LLMs is worth checking out! To mitigate contamination, they refresh the benchmark every six months and release only some questions from previous releases. The results are also easily navigable with a screen reader! #LLM #AI #ML https://livebench.ai/