As datasets explode to tens of trillions of tokens, the classic 80/20 split is dead. Now it’s common to see 99%+ for training, with just 0.5% each for dev and test. Even half a percent of 15 trillion tokens still gives you billions of examples. Bigger data completely changes the math.