NanoFlare Test Results
I have published the MMLU Pro scores for NanoFlare over on the HuggingFace page. It did okay, but clearly it is at a diminished quality compared to its bigger, older brother, microFlare. But at half the file size, it still performed pretty well. Scoring an average of 0.61 across 280 tests (only the first 20 questions of each of the 14 categories were ran for timeliness's sake).
Additionally I have released an alternate version that also includes the MTP draft predictor if you have the room for it and would like some more speed. It's too big to fit on my 8GB card, but perhaps someone could get it running on a headless server without the desktop overhead eating into VRAM. If you have a 10GB GPU, this is the version I would suggest you use.
These tests were ran using the latest revision, NanoFlare v1.c. Tests were ran with a 5 shot and an 8k context. Note, I did not mess with temperature nor top-k settings when running the test suite. If I had it set to top-k: 20 and temperature: 0.95, like I usually run it, it would probably have scored even higher. Additionally setting the model to medium reasoning levels instead of it's default xhigh might have also yielded some better results. But I left these at the default so it could be a cleaner comparison against how microFlare scored.
I think this will be the end of testing and revisions for NanoFlare for the time being. Until I'm ready to use a different source model for V2 at least (Qwen 4 is moving away from open source licensing sadly).
I'm considering making a NanoFlare Plus that targets a 12GB GPU rather than an 8, so I can reclaim some more quaility and get it more on par with microFlare. However, I don't have such a card to test with right now. Let me know if anyone would be interested in that. But for now, if you have 12GB or more available in your VRAM I'd suggest you go with one of Unsloth's UD-3.0 builds of Qwen 3.8 27B.
PS: I've got a new revision of microFlare in the works that will hopefully be published fairly soon.