Writing

Trust the benchmarks? Or not.

  • ai
  • benchmarks

The real question is how reliable these benchmarks actually are. Because after reading this (though already expected it), my skepticism toward LLM benchmarks went up again.

“ok slopus solve this”
Validator: “hmmm… ok there is a message.. SOLVED!”

That is essentially what could be happening? The models are optimizing for the validator, not for our actual use case.

Though they are “not claiming that current leaderboard leaders are cheating” but the possibility is there.

But whether it is intentional or not, the gap between a top benchmark score and actual day-to-day utility is still massive.

That’s why I always trust my own experience and everyone’s sentiments with the model over the leaderboard. Always!

Yes, I am looking at you, Gemini 3.1 Pro… [1]