Summarizing https://arxiv.org/pdf/2306.11644.pdf
Here's my try:
The authors introduce phi-1, a large language model for code with 1.3B parameters trained for 4 days on 8 A100s using a selection of "textbook quality" data from the web (6B tokens) and synthetically generated textbooks and exercises ...