Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k context size it means it wont do much before filling it context
We don't actually know how much thinking GPT and Opus do, the labs won't show us anymore. And they certainly take their time before starting to answer.
I got downvoted yesterday for complaining about the name/brand confusion from all these Chinese models with similar names all claiming they are the best. While I'm an AI enthusiast, it's being hard to keep track.
Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k context size it means it wont do much before filling it context
It’s definitely a heavy thinker, like most Qwens.
I see strings like “write, now.” In the thinning traces then it goes on to think for a lot longer, so it’s kind of weird.
I haven’t figured out hope to use this effectively yet on my strix halo.
We don't actually know how much thinking GPT and Opus do, the labs won't show us anymore. And they certainly take their time before starting to answer.
Maybe a bit of context for this post? Some people have a life outside of AI.
I got downvoted yesterday for complaining about the name/brand confusion from all these Chinese models with similar names all claiming they are the best. While I'm an AI enthusiast, it's being hard to keep track.
[dead]