Of course, I only tested on my system and my use case. Other people will have different experiences when testing on different criteria and frameworks.
I am developing Agentic workflow in my own AI system, every model is graded the same. But I'm sure tweaking the parameters and fine-tuning the model would yield different results.
I haven't tested Nanbeige yet, that involves installing it's fork of llama.cpp and building it in Jetson which could take over an hour.