Open weight models closed a big chunk of the gap on frontier models this year, here's the actual movement
Full disclosure, we're the GitKraken team, and one of our devs has been tracking model benchmarks for a while now. Sharing because the trend line surprised us and figured this sub would have opinions. The frontier labs (Anthropic, OpenAI) keep trading the top spot by a point or two every couple months. Nothing new there. What's actually moving is the open weight tier underneath them: GLM went 5.1 to 5.2 and picked up 11 points of intelligence in about two months Kimi K3 picked up 13 points over K2.6 in three months and is now sitting in 4th place overall, a few points off the leaders Deepseek's V4 Flash 0731 update leapfrogged their own V4 Pro model, and it was a fine tune, not a full retrain Kimi K3 is popular enough right now that some providers are rate limiting or waitlisting it. Deepseek's approach (fine tuning an existing model instead of training a new one from scratch) seems like the more interesting engineering story here, honestly, more than the leaderboard position itself. Curious what this sub is actually running day to day. Anyone switched their default agent model in the last quarter because a newer open weight release actually caught up, or is everyone still defaulting to whatever frontier model they started with out of inertia? submitted by /u/GitKraken [link] [留言]