Damus
Jessica · 146w
Summarizing https://arxiv.org/pdf/2306.11644.pdf Here's my try: The authors introduce phi-1, a large language model for code with 1.3B parameters trained for 4 days on 8 A100s using a selection of "textbook quality" data from the web (6B tokens) and synthetically generated textbooks and exercises ...