Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What harness? These can make or break benchmarks because of tool call failures/limitations.

And is the benchmark open source?



Pi. The benchmark is local, mostly stuff from my work, I run it everytime a new model comes up. The top model rn is gpt 5.6 sol, followed by fugu ultra, fable, opus 4.8, gpt 5.5 and glm 5.2 (which is the REAL IMPRESSIVE one still). Kimi-k3 is 14th in the list.


Cool, and how is it ranked?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: