EleutherAI releases Llemma, open-weight language models (7B and 34B parameters) specialized for mathematics. Initialized from Code Llama weights and trained on Proof-Pile II — a 55B token dataset of mathematical and scientific content including arXiv papers, web math pages, and mathematical code — Llemma outperforms Minerva on an equi-parameter basis and achieves state-of-the-art results among open models on MATH and GSM8k benchmarks. Notably, Llemma is the first open base model to demonstrate in-context formal theorem proving, surpassing GPT-4-based approaches on the miniF2F benchmark. Models, dataset, and training code are all publicly released. A memorization analysis shows Llemma's performance is not explained by training set overlap.

6m read timeFrom blog.eleuther.ai
Post cover image
Table of contents
Dataset : Proof-Pile II #Training #Evaluation #Memorization #Future directions #Citation #
5 Impressions