DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Building a powerful AI model using only ethically sourced training data
Researchers created Mimir v1, a 1-billion-parameter language model trained entirely on permissible datasets, and showed it performs as well as models twice its size. The model sets a new benchmark for Danish language tasks and matches larger competitors across 20 different tests covering English, math, code, and Danish—all without relying on scraped or questionable data sources.
Most cutting-edge language models train on massive datasets of unclear origin, creating legal and ethical risks. Mimir v1 proves you can build competitive AI using only legally permissible data, potentially opening doors for researchers and companies who want powerful models without copyright or privacy concerns. The model is freely available, lowering barriers for smaller teams and non-English-speaking communities to develop their own language AI.