for those looking for something small AND powerful, there is a new 1B (they claim, it looks more like 1.7B ...) model that claims to beat qwen 3.5 0.8B & 2B and gemma 4 E2B on a range of benchmarks. the model seems to be english and danish only. math and coding seem to be quite ok-ish. apparently, it builds on sapient's hrm-text model, which does some weird layer-recurrence magic. paper: https://huggingface.co/papers/2608.13517 hf: https://huggingface.co/danish-foundation-models/DFM-Mimir submitted by /u/ZookeepergameCool173 [link] [comments]

Read original ↗ Content from Reddit r/LocalLLaMA(Community