AMÁLIA is a €5.5M Portuguese government-funded open source LLM for European Portuguese, built as a continuation of EuroLLM's pre-training rather than trained from scratch. The author analyzes the technical report and raises concerns: the model weights, data, and benchmarks are not yet publicly available despite open source claims; only ~5.5% of pre-training tokens are clearly European Portuguese; and the new benchmarks don't measure intrinsic knowledge about Portugal. Despite beating models like Qwen 3-8B on most Portuguese benchmarks, the author questions whether the approach is fully optimized and calls for more Portuguese data, genuine openness of weights and datasets, and better evaluation dimensions.
Table of contents
AMÁLIA in a nutshellHow open source, really?How much Portuguese data for a Portuguese model?What should we be optimizing for?Final thoughts686 Impressions