Don’t forget where the I in AI comes from

Yet they have erred, we feel, in working on a foundation principle that words and sentences mean something. In fact, it is only persons who mean something; language is their instrument.

-Wilfred Cantwell Smith, The Meaning and End of Religion (emphasis mine)

To properly train an LLM (like ChatGPT) you need vast quantities of data. I need to clarify what I mean when I say “vast quantities of data.“

GPT-3 (the precursor to the first ChatGPT model to really go viral, which was 3.5) used 300 billion tokens in training (think of a token as a syllable — a portion of a word, or one small word) which was about 570 GB of data. There’s an average of roughly 1.3 tokens per word (for english) which means 300B tokens is about 230B words, and if the average novel is around 100,000 words that gives us roughly 2.3 million books. That’s a lot of books but the library of congress has about 39 million books.

Even though that first version of ChatGPT seemed amazing, it kind of sucks compared to what we have today. But to get to where we are today they needed more training data. WAY more training data. Labs can be pretty tight lipped about how they train their models, but Meta released details about how they trained Llama 3.1 and it required roughly 15 trillion tokens. Estimates of recent models put the number closer to 25 trillion.

If we do our conversion of tokens to average number of books what we get is … hold on a second while I pull out my calculator … 25 trillion / 1.3 tokens per word / 100,000 words per novel puts us at … 192.3 million books. Books is a pretty unwieldy unit of measurement so let’s simplify that a little.

Training the first big model took roughly .06 libraries of congress. Six one hundredths of one library of congress.

Training the latest models took roughly five libraries of congress.

The technology behind LLMs is pretty incredible when you dig into it. I mean, I couldn’t have thought it up. Those are some smart scientists and computer programmers working away in those labs. Yessir.

But, no matter how impressive the technology, it would be functionally useless without the training data. The secret sauce in an LLM is not the math. It’s the english (bet that drives those math majors nuts). And it requires lots and lots and lots and lots of that writing to make the math useful. How much? Five libraries of congress. At least.

And where did all that writing come from?

LLMs are a testament to the power of writing

It was all written down. LLMs map language — they create statistical links between words. So when you get an insightful or even competent response to a prompt posed to an LLM who should get the credit? Well, not the labs that made them. As we’ve already discussed, they would be useless without all the training data they, uh … acquired.

Not the mythology of silicon valley. As much as I like a good book about “Infrastructure as Code” the majority of writing doesn’t come out of silicon valley.

It belongs to the people who put human thought down on paper. I’m currently looking at my copy of “The Teachings of Ptahhotep” — an english translation of an Egyptian book where a Vizier who is 96 years old attempts to teach his son how to be a good Vizier. It was likely written roughly 2000 BCE. You want a taste of 4000 year old wisdom? Here’s the first maxim:

Be not proud because thou art learned; but discourse with the ignorant man, as with the sage. For no limit can be set to skill, neither is there any craftsman that possesseth full advantages. Fair speech is more rare than the emerald that is found by slave-maidens on the pebbles.

Some advice as applicable today as it was 4000 years ago, plus some casual slavery.

Was “The Teachings of Ptahhotep” used to train LLMs? Almost certainly! It’s good english (and Egyptian) and it’s in the public domain, plus quotes from it are all over the internet.

LLMs only have value of any kind due to the last 5000 years of human ingenuity and our stubborn persistence to write down stuff we think is important. If it weren’t for that, an LLM would be a cool academic paper that no one outside a small cadre of computer scientists would’ve heard of.

When you use an LLM who deserves the credit? Anyone who has ever written, going all the way back to Ptahhotep (at least).

But writing isn’t everything

Plato is famous for his works that have survived thousands of years. He wrote books, dialogues, I’m assuming monologues at some point — he and his students were prolific.

But like many philosophers in his time, Plato was actually somewhat wary of writing. It was OK for transmitting some thoughts and ideas, but to REALLY understand some of the more complicated concepts you couldn’t count on writing them down. You had to transmit them directly, teacher to student, through dialogue (man, he loved dialogues).

This is the basis for what are called “Plato’s Unwritten Doctrines.” In his seventh letter Plato explains (pretty complicatedly so I’m just going to quote wikipedia here) that “one attains knowledge only from the combination of verbal description and sense perception.” To understand some things it is not enough to read about them. You must experience them. There are some things Plato didn’t write down (though later students alluded to them in writing) because he felt they couldn’t be transmitted effectively through the written word.

Plato (or whomever may have written the seventh letter) is right here, and if you think about it you recognize it too. How sweet is a strawberry? How sweet is a banana? You can easily say “A banana is sweeter than a strawberry” but you don’t really understand the magnitude or quality of the difference unless you have eaten a strawberry and a banana. Even if you broke it down and said “A banana has roughly 2.5 times the sugar density by weight of a strawberry” it still wouldn’t transmit the truth of how much sweeter a banana is than a strawberry.

To understand some things requires sense perception. This is obviously true of food, but it wasn’t actually food that Plato was talking about. He was more interested in that aha moment, “like a light kindled from a leaping spark,” when we mull things over in our mind, or when we talk through them with someone knowledgeable, and suddenly they just click into place and confusion is replaced with understanding.

This, to me, is the biggest problem with LLMs. They are incredibly impressive and useful. But they are limited solely to language, and when you are limited to language there is an entire aspect of human intelligence that is missed. As I (and others much better credentialed than me) have discussed before, an LLM cannot actually UNDERSTAND anything. It is a statistical language model, but the most important things to be understood live outside of language.

Plato knew that. We know it too, when we stop to think about it.

Or, put another way:

Yet they have erred, we feel, in working on a foundation principle that words and sentences mean something. In fact, it is only persons who mean something; language is their instrument.


Leave a comment