4
Just found out GPT-4 was trained on 13 trillion tokens, that blew my mind
Was reading a paper from the Allen Institute last night, they did a deep dive on training data. 13 trillion. That's like every book, every Reddit thread, every PDF ever scanned, times a hundred. I remember when GPT-2 came out in 2019 and people freaked over 1.5 billion parameters. Now we're talking trillions of tokens just to get a model that still can't do basic math reliably. Kinda humbling. Anyone else feel like we went from calculators to supercomputers but the output still has that uncanny valley thing going on?
1 comments
Log in to join the discussion
Log In1 Comment
xena_murphy24d ago
Man, that stat is wild. It reminds me of this time I tried to teach my grandma how to use a smart TV, and she kept calling me to ask why the remote wasn't working when she was holding it upside down. Kinda like how we feed these models the entire internet and they still struggle with stuff a kid could do, like counting apples in a basket. But honestly, I think the bigger issue is that we're measuring the wrong things, like more data isn't the fix if the base logic is just pattern matching on steroids. Still, 13 trillion, I can't even wrap my head around that number, it's like trying to imagine all the grains of sand on a beach and then multiplying by ten.
-1