If you're going to use OpenRouter to test reasoning levels, always make sure you are locking to the official provider instead of third party providers.
Pretrained dictionaries have never been intended to help with book sized or bigger compression. zstd automatically learns the most efficient dictionary it can within a few kilobytes. Pretrained dictionaries are only useful when you're independently compressing very small records.
Please run GLM-5.3 and GLM-5.3-Flash. I would love to see how they do. On the smaller end of things, Qwen3.8-27B and Ling-3.0-Flash would also be interesting.
In the benchmark, have you considered instructing the models to build their own SPICE simulations to test their work? Simply asking them to write and run simulations could improve performance, even without telling them what to simulate.
FunctionGemma never worked well for me (without fine tuning). Liquid has released 230M and 350M models that work far, far better in my testing: https://huggingface.co/LiquidAI/LFM2.5-230M
I really look forward to a hypothetical LFM3-230M, because LFM2.5-230M is so close to being usable, while FunctionGemma is miles away from being usable.
I fully expect Meta will release other, smaller Muse models in the near future too.
The 5090 is also supposed to be a $2000 GPU, not a $5000 one. The entire market is utterly distorted right now, which will impact cloud inference more and more over time too. They are not immune to the absurdly high RAM prices, so their prices will have to go up over time too until the RAM supply chain goes back to normal.
I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for.
At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.
The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.
Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.
Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.
(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
If you're going to use OpenRouter to test reasoning levels, always make sure you are locking to the official provider instead of third party providers.
reply