Grinding Old Math, Burning New Cash

The year was 1936. Alan Turing asked a boring question. Can every mathematical problem be solved by an algorithm? He proved the answer was no. But to do it, he had to define what an algorithm even is. So he imagined a simple machine. An infinite tape, a read-write head, and a tiny table of rules. That abstract doodle became the blueprint for every computer you have ever owned.
Twelve years later, Claude Shannon asked an even stranger question. What is information, as a thing you can measure? In his 1948 paper, he stripped meaning from language entirely. "I love you" and "the cat is on fire" carry the same information if they are equally surprising. He called that surprise entropy and proved all information could be reduced to ones and zeros. Then, to measure how predictable English was, he simply asked people to guess the next letter in a sentence. Easy letters had low entropy. Hard ones had high entropy. That experiment, humans guessing the next token, is exactly what modern AI does today. Shannon was not trying to build AI. But he gave us the math for uncertainty, prediction, and compression. One of today's leading AI labs even named their flagship model Claude.
Then came the perceptron. In 1958, a psychologist built a machine that could learn by adjusting weights whenever it was wrong. The hype was immediate. Newspapers predicted consciousness. Then, 11 years later, two researchers published a book proving a single-layer perceptron could not even learn XOR. Funding evaporated. AI winter began. But buried in the fine print, they figured out that stacking layers fixes everything. The problem was nobody knew how to train a stack. It would take 17 years.
Before that, we needed to solve a different problem. In 1978, Leslie Lamport asked how separate computers with no shared clock could agree on the order of events. When you have thousands of machines, there is no universal now. His answer was the happens-before relation. Stop trusting the wall clock. Order events by causality instead. He built logical clocks that let machines stay in agreement without ever looking at a real clock. This paper became the bedrock for databases, blockchains, and every massive AI training run. You need thousands of GPUs to stay in sync without dissolving into chaos. Lamport made that possible.
Seventeen years after neural networks were left for dead, a group of researchers cracked training stacked layers. The answer was backpropagation. Run data forward, measure the error, and push it backward through every layer using the chain rule. Nudge each weight toward less wrong. Do that millions of times and the network teaches itself. The hidden layers started inventing their own features. Edges, shapes, concepts nobody programmed. XOR became trivial. But backprop still sucked. We lacked data and compute.
That changed in 1998. Two graduate students published a paper describing a search engine that treated every link as a vote, weighted by trustworthiness. They built it in a dorm room. It became Google. But the paper's real legacy was assembling the largest structured pile of human text ever created. That pile would become training data for future AI.
In 2012, a small team wired up a deep convolutional neural network, trained it on a couple of gaming GPUs in a bedroom, and walked it into an image classification contest. AlexNet crushed everyone, dropping the error rate by 10 points in a single year. Deep learning was suddenly real. It just needed more data, more compute, and the right architecture.
That architecture arrived in 2017. "Attention Is All You Need" introduced the transformer. Instead of reading tokens one by one, every word looked at every other word at once. Models stopped forgetting context and scaled better too. The lab gave it away for free. Now every AI lab uses it. That is the T in GPT.
Then in 2020, one lab asked the dumbest question possible. What if we just make it enormous? One hundred and seventy-five billion parameters, fed the entire internet. They bet that intelligence is not a secret algorithm. It simply emerges at scale. The model translated, summarized, and wrote code without being taught. Two years later, it became one of the most valuable products in tech history.
Here is what strikes me after tracing this 85-year arc. When you strip everything down, what is a modern large language model actually doing? It is predicting the next token. Just like Claude Shannon was doing in 1948 with volunteers guessing the next letter. The math has not changed. The entropy he borrowed from thermodynamics is still baked into every loss function and attention head.
Turing defined the machine. Shannon gave it currency. The perceptron gave it a neuron. Backpropagation taught it how to learn. PageRank gave it data. The transformer gave it architecture. And the latest labs turned the dial to maximum. That is the honest history of AI in 10 papers.
None of these breakthroughs required new math. They required new budgets. The perceptron failed in 1969 because nobody could train multi-layer networks yet. The math was there. The algorithm was not. Backpropagation fixed that in 1986, but needed a gaming graphics card to prove itself in 2012. The transformer could have been discovered earlier. Nobody had the compute to make it sing.
This should humble every engineer working in AI today. Before you reach for the latest framework, maybe reach for the oldest paper. The next breakthrough is probably sitting in a dusty journal from 1978, waiting for a bigger GPU.
Disclaimer: All content reflects my personal views only and does not represent the positions, strategies, or opinions of any entity I am or have been associated with.


