Please use this identifier to cite or link to this item:
http://hdl.handle.net/1893/37587| Appears in Collections: | Computing Science and Mathematics Conference Papers and Proceedings |
| Peer Review Status: | Refereed |
| Author(s): | Zain, Noor Ul Naseem, Mohsin Raza Adeel, Ahsan |
| Contact Email: | noor.noorulzain@stir.ac.uk |
| Title: | Single layer tiny Co4 outpaces {GPT}-2 and {GPT}-{BERT} |
| Editor(s): | Charpentier, Lucas Choshen, Leshem Cotterell, Ryan Gul, Mustafa Omer Hu, Michael Y Liu, Jing Jumelet, Jaap Linzen, Tal Mueller, Aaron Ross, Candace Shah, Raj Sanjay Warstadt, Alex Wilcox, Ethan Gotlieb Williams, Adina |
| Citation: | Zain NU, Naseem MR & Adeel A (2025) Single layer tiny Co4 outpaces {GPT}-2 and {GPT}-{BERT}. In: Charpentier L, Choshen L, Cotterell R, Gul MO, Hu MY, Liu J, Jumelet J, Linzen T, Mueller A, Ross C, Shah RS, Warstadt A, Wilcox EG & Williams A (eds.) <i>Proceedings of the First BabyLM Workshop</i>, volume Proceedings of the First BabyLM Workshop. Empirical Methods in Natural Language Processing, Hybrid, 04.11.2025. Association for Computational Linguistics, pp. 313-322. https://doi.org/10.18653/v1/2025.babylm-main.24 |
| Issue Date: | 2025 |
| Date Deposited: | 21-Nov-2025 |
| Conference Name: | Empirical Methods in Natural Language Processing |
| Conference Dates: | 2025-11-04 |
| Conference Location: | Hybrid |
| Abstract: | We show that a tiny Co4 machine (CITATION) with a single layer, two heads, and 8M parameters, operating at O(N) computational cost (where N is the number of input tokens), in just 2 epochs outpaces GPT-2 (124M, 12 layers, O(N2)) and GPT-BERT (30M, 12 layers, O(N 2), both trained for 10 epochs. Co4 achieves orders-of-magnitude greater training efficiency on 10M tokens, demonstrating sample-efficient pretraining. On the BabyLM challenge evaluation pipeline, Co4 performs comparably or better across complex benchmarks, showing strong zero-shot and fine-tuning performance on SuperGLUE tasks. Specifically, Co4 outperforms GPT-2 in 5 out of 7 zero-shot metrics and 6 out of 7 fine-tuning tasks, and GPT-BERT in 4 out of 7 metrics in both cases. These results strongly suggest a need to rethink prevailing deep learning paradigms and associated scaling laws. |
| Status: | VoR - Version of Record |
| Rights: | ACL materials are Copyright © 1963–2025 ACL; Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License. |
| Licence URL(s): | http://creativecommons.org/licenses/by/4.0/ |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 2025.babylm-main.24.pdf | Fulltext - Published Version | 657.69 kB | Adobe PDF | View/Open |
This item is protected by original copyright |
A file in this item is licensed under a Creative Commons License
Items in the Repository are protected by copyright, with all rights reserved, unless otherwise indicated.
The metadata of the records in the Repository are available under the CC0 public domain dedication: No Rights Reserved https://creativecommons.org/publicdomain/zero/1.0/
If you believe that any material held in STORRE infringes copyright, please contact library@stir.ac.uk providing details and we will remove the Work from public display in STORRE and investigate your claim.
