I tested the top local LLMs against ChatGPT and Claude, and one absurd prompt revealed the gap

Open original source
Back