EleutherAI researcher Leo Gao fine-tuned GPT-Neo 2.7B for 1,100 iterations across all eval harness tasks that have training sets, training a single model on all tasks simultaneously. Results show the tuned model doesn't uniformly dominate the baseline — performance is roughly a tossup across most tasks. Notable gains appear on ANLI (adversarial NLI), MNLI, QNLI, QQP, RTE, and SST. However, tasks excluded from tuning (lambada and pubmedqa) show significant performance degradation, confirming catastrophic forgetting. Detailed zero-shot and one-shot benchmark tables are provided for both the baseline 2.7B and the fine-tuned model.

15m read timeFrom blog.eleuther.ai
Post cover image
Table of contents
Zero shot #One shot #
3 Impressions