#
DIFF.BLOG
New
Following
Discover
Jobs
More
Top Writers
Suggest a blog
Upvotes plugin
Report bug
Contact
About
Sign up
The home for great developer writing.
We surface the best developer writing from thousands of independent blogs, updated daily.
Join now
→
Learn more
TOPICS
An Engineer's Guide to AI Code Model Evals
280
·
Addy Osmani
·
July 25, 2025, 8:02 p.m.
AI Models
Code evaluation
Machine Learning
software development
Summary
This blog post provides an in-depth analysis of evaluation methods such as evals, goldens, and hill climbing to enhance the performance of AI models capable of coding. It offers insights valuable for developers interested in AI and machine learning.
Read full post on addyosmani.com →
MORE POSTS LIKE THIS
I Tried Kimi K3 Inside Claude Code
Philipp D. Dubach ·
Jul 19, 2026
AI Models
Machine Learning
Z.ai's GLM 5.2 is a great model, but is it good value?
Blog Kronis ·
Jul 7, 2026
Machine Learning
AI Models
Reading the agent traces is how you make the call your eval can't
Sentry ·
Jul 1, 2026
AI Models
Machine Learning
Open models don't need to be OpenAI
Joseph E. Gonzalez ·
Jun 25, 2026
Machine Learning
Open Source
Mythos-class Claude Fable 5 arrives on GitLab Duo Agent Platform
GitLabBlog ·
Jun 9, 2026
AI Models
Gitlab
What is Qwen AI?
Zapier ·
Jun 9, 2025
AI Models
Qwen AI
Discover more posts →
AUTHOR
BLOG POST FEATURED ON
r/Frontend
4 points
Hacker News
2 points
r/webdev
1 points
r/programming
0 points
Add this plugin to your blog
RECENT POSTS FROM THE AUTHOR
Choose how you want to continue.
Continue with GitHub
Continue with Google