oosioo
다음을 맞히는 기계들. 왼쪽에는 레이더 점을 따라가다 점선으로 다음 자리를 짐작하는 궤적, 오른쪽에는 ‘누구도 다음에 어떤 단어가 주어질지’ 뒤 빈칸에 올 낱말의 확률 막대
SeriesA Page for Today · Ep. 2

The Machines That Guess What Comes Next

Starting from a line that says no one knows the next word, this piece follows the thread from the Kalman filter that predicts a missile's next position, through weather forecasting and grocery carts, to the LLMs that choose the next word. The second installment of A Page for Today.

· SYSOP · 3 views

The machines that guess what comes next. On the left, a trajectory following radar dots with a dotted line guessing the next position; on the right, probability bars for the word that fills the blank after 'no one knows what word comes next'

Today, too, I opened my idea notebook to a random page. Where my finger landed, there was this sentence:

We are all like someone writing a new sentence of life. No one knows what word comes next.

I laughed the moment I read it. Because what we now call AI — large language models like ChatGPT — were built to do exactly that. To look at the text in front of them and guess the word that comes next. Today's LLMs are trained to do this an enormous number of times, and extremely well. People live without knowing the next word, while machines make a living by guessing it.

But LLMs are not the first machines built to guess what comes next. The next position of an incoming missile, tomorrow afternoon's air pressure, even what fills a shopping cart before a typhoon hits — people have been calculating "next" for a long time. Today, let's meet those machines one by one.

The Next Position of Something in Flight

In 1960, the Hungarian-born mathematician Rudolf Kálmán published a paper titled "A New Approach to Linear Filtering and Prediction Problems." At first, the response was lukewarm. That fall, Kálmán visited NASA's Ames Research Center, where he met a researcher named Stanley Schmidt. Schmidt's team was wrestling with the problem of calculating exactly where a spacecraft bound for the Moon was at any given moment. The onboard computer was small and slow, and ground observations were always slightly off.

Schmidt adapted Kálmán's method to the lunar trajectory problem. This calculation method went into the Apollo spacecraft's navigation computer and was used to send people to the Moon and bring them back.

The Apollo Guidance Computer unit (left) and the DSKY control panel with its numeric keypad (right). The navigation calculations that included the Kalman filter ran on machines like these

What the Kalman filter does can be boiled down to three steps.

First, it predicts. If you know the position and velocity a moment ago, physics lets you calculate where something will be 0.1 seconds later. But wind blows and engines wobble slightly, so this prediction grows fuzzier over time.

Second, it measures. Radar looks at the object. This, too, can't be fully trusted — the position radar reports jitters slightly each time.

Third, it blends. This is the heart of the Kalman filter. It weighs which is more certain — the prediction or the measurement — and leans toward whichever is more trustworthy. If radar is sharp, it leans toward radar; if radar is fuzzy, it leans toward the prediction. The blended value ends up with a smaller error than trusting either one alone. This sounds strange at first, but mixing two pieces of wrong information the right way produces information that is less wrong.

The three steps of the Kalman filter. A blurry ellipse predicted by physical laws, a jittery ellipse measured by radar, and a narrower estimate ellipse formed by blending the two

It's similar to finding your way to the bathroom in a dark room. You estimate where you are by counting steps (prediction), and when your fingertips touch a wall (measurement), you correct your mental map. If the touch is clear, you correct a lot; if it's something ambiguous, you correct only a little.

We use this filter every day, too. When you enter a tunnel and lose GPS signal, the blue dot on your phone's map keeps moving forward for a while. It keeps predicting the next position using the last known speed and direction. When you exit the tunnel and the signal returns, the dot jumps slightly and settles into place. That's the moment prediction and measurement blend again.

Words alone don't quite convey it. So I built something you can try yourself.

The Kalman filter playground. Radar watches an object flying in a parabolic arc, with a jittery reading every 0.2 seconds (orange dot). The solid blue line is the estimate the Kalman filter blends together, and the dotted blue line is the prediction rolled forward 1.5 seconds from the current estimate, with the circle at its end showing how much that prediction has spread. The green square is the window radar will look at next. The controls below let you change radar jitter, how much to trust physics versus radar, the strength of wind the filter doesn't know about, and how far the clock has drifted. The gray band in the middle is a stretch of cloud radar can't see through. In the upper left, the average error of a single radar reading and the average error of the Kalman estimate are shown side by side.

A few things are worth trying. Crank radar jitter all the way up, and the orange dots scatter in every direction, but the blue line wobbles less than you'd expect. Check the numbers in the upper left, and the estimate's error is smaller than a single radar reading's error. Enter the cloud stretch, and measurement cuts out — the filter has to get by on prediction alone. Meanwhile the blue circle swells, which means the filter itself knows how much it doesn't know. Push "what to trust more" all the way toward physics and crank up the wind, and you can watch the filter trust its own math too much and drift further and further from the actual trajectory. Trusting the prediction too much is bad, and so is trusting the measurement too much.

The last control, "clock drift," is there for the next part of the story.

The 687 Meters Made by 0.34 Seconds

February 25, 1991, Dhahran, Saudi Arabia, in the middle of the Gulf War. An Iraqi Scud missile came flying in. Stationed there was a Patriot missile battery designed to intercept it.

A Patriot missile launching at a training range in Texas, USA, in 1997

When the Patriot radar sweeps the sky and finds a suspicious object, it calculates where that object will be next. Then it opens a small window (a range gate) at that spot and looks again. If the object appears inside the window, the system treats it as a real missile and begins tracking. This is the green square from the Kalman filter we just saw.

The problem was the clock. The Patriot's computer, designed in the 1970s, did its math in 24-bit numbers and counted time in units of 0.1 seconds. To us, 0.1 is a clean number, but in binary it's an endlessly repeating fraction that doesn't fit neatly into 24 bits. So every time it counted off 0.1 seconds, a tiny error dropped out. An error that was nothing on its own accumulated for as long as the system stayed on.

That day, the battery at Dhahran had been running continuously for over 100 hours. The accumulated error came to 0.3433 seconds. For a Scud falling at nearly five times the speed of sound, 0.34 seconds is a long time. The window the radar opened looked at a spot 687 meters off from the missile's actual position, and there was nothing there. The system concluded that what it had first seen was a false alarm. The Scud struck a U.S. Army barracks, killing 28 people.

Clock error and the resulting shift in the tracking window as a function of how long the Patriot battery had been running. 1 hour: 7 meters. 8 hours: 55 meters. 20 hours: 137 meters. 100 hours: 687 meters. Past roughly 20 hours, it misses the Scud

According to a 1992 report by the U.S. General Accounting Office (GAO), leaving the system running for as little as about 20 hours could shift the window enough to miss a Scud entirely. The fix was simple: turning the system off and back on reset the error to zero. Software that corrected the clock reached Dhahran the day after the attack.

Push the playground's "clock drift" up to 0.34 seconds, and the green square turns red and slides into empty space behind the object. A prediction depends as much on the clock underlying the calculation as on the calculation itself. A machine that guesses what comes next first has to know exactly what time it is now.

Even With Laws, You Can't See Far: Weather

Guessing what the sky would do next also began with calculation. Britain's Lewis Fry Richardson drove an ambulance during World War I and did weather calculations in his spare time. He wrote the atmosphere's motion as equations and worked out, by hand, the air pressure six hours later at one point in Europe. The calculation took about six weeks. The result was that pressure would change by 145 hectopascals over six hours — an absurd figure, more than even the strongest storm could produce. Still, Richardson published this failure as it was in a 1922 book. He then imagined a "forecast factory": 64,000 people seated around a vast circular hall, each calculating their assigned patch of sky. That, he reasoned, was how many people it would take to compute faster than the weather itself.

It was in 1950 that a machine, rather than people, became that factory. The early electronic computer ENIAC ran the first computerized weather forecast. Forecasting 24 hours ahead took about 24 hours of computation — barely keeping pace with the weather itself, but the direction was right.

People programming ENIAC. They plugged in cables and set switches in front of a calculator that filled an entire room

Then in 1961, MIT meteorologist Edward Lorenz noticed something strange. Trying to rerun a weather simulation partway through, he typed in the numbers printed on a paper printout. Inside the computer, the number had been 0.506127, but on paper, to save space, it had been printed only as far as 0.506 — a difference of roughly one part in ten thousand. He went for coffee, and when he came back, the two weather runs, which had started out overlapping, had gradually diverged and ended up as completely different weather.

The result of running the Lorenz equations twice, changing only the starting value between 0.506127 and 0.506. The two curves overlap for a long stretch, then at some point split apart completely

The figure above redoes the same experiment using the three-line equations Lorenz wrote in his 1963 paper. For a while, the blue and orange lines track together as if they were one line, then at some point they go their separate ways. This is the phenomenon later named the "butterfly effect." Even if you know the laws perfectly, you can't predict the distant future unless you can measure the present perfectly too.

That's why modern forecasts don't run the model just once. They run it dozens of times with slightly different starting values and count how many of those runs produce rain. That's where "60% chance of rain" comes from. Just as the Kalman filter draws a circle to depict its own uncertainty, a forecast uses probability to state how much it doesn't know. Built up this way, weather forecasting has gained roughly one extra day of reliable range for every decade over the past 40 years — so that today's six-day forecast is about as accurate as a five-day forecast was ten years ago.

Without Laws, There's Still Repetition: The Shopping Cart

For some kinds of "next," there's no physical law at all. What people choose to buy is like that. But there is repetition.

In 2004, Hurricane Frances was bearing down on Florida. Walmart dug through purchase records from Hurricane Charley, which had passed through a few weeks earlier, to see what people swept off the shelves before a storm. Beyond the obvious items like flashlights and bottled water, something stood out: beer, and strawberry Pop-Tarts. Pop-Tarts sold seven times their usual volume ahead of the storm. Trucks loaded with Pop-Tarts and beer rolled down the highway to stores in the storm's path.

This kind of prediction doesn't know the reason why. Whether it's because Pop-Tarts don't need cooking and keep well, or because kids like them, the machine doesn't care. It simply reasons that because it happened last time, it will happen again this time. Where the Kalman filter calculates the next step from physical laws, this kind of prediction guesses the next step from the repetition of the past. And LLMs are distant descendants of this latter approach.

The Next Character: From Pushkin to ChatGPT

In January 1913, the Russian mathematician Andrey Markov gave a rather odd presentation at the Academy of Sciences. He had gone through the first 20,000 characters of Pushkin's verse novel Eugene Onegin, classifying each one as a vowel or a consonant and counting them. A consonant was likely to follow a vowel, and a vowel was likely to follow a consonant. The next letter leaned on the one before it. The calculation did nothing to help anyone understand the poem, but out of it was born the probability theory known as the "Markov chain" — a method of guessing what comes next by looking only at what came immediately before.

In 1951, Claude Shannon, the creator of information theory, played a guessing game with people. He showed them the preceding 100 characters of an English text and had them keep guessing the next character until they got it right. People guessed far better than expected. From this experiment, Shannon calculated that a single English letter carries about 1 bit of information on average — meaning that choosing among 27 possibilities (the alphabet plus the space) shrinks, once you know the preceding context, to roughly the uncertainty of a single coin flip. Text is far more predictable than it seems.

Autocomplete on a phone keyboard is an extension of this. For a long time, autocomplete chose the next word by looking only at the two or three words right before it. It kept a tally of what most often followed "For lunch today," and offered up the most common option. So no matter what you'd said earlier in the sentence, roughly the same candidates kept showing up.

For the same blank in a sentence, autocomplete that only looks at the two preceding words ranks "something, deliciously, simply" highly, while an LLM that reads the whole sentence ranks "something hot, kalguksu (noodle soup), gukbap (rice-and-soup)" highly. Probabilities are illustrative

An LLM does the same thing: it guesses the next word. What's different is its field of view. Thanks to the transformer architecture that came into use around 2017, a model can scan an entire block of text at once and focus its attention on whichever part matters most right now. A model that has read "it rained all night and I've got chills" will rank soup-based dishes highly after "For lunch today." Feed a model more text than a person could read in a lifetime, and grammar, facts, and tone accumulate inside it on their own, all in service of guessing the next word well. Even when ChatGPT writes an answer, what it's doing is, in the end, choosing one word, appending it to the end of the sentence, and choosing the next one — over and over.

This is also where mistakes come from. What an LLM picks is the most plausible next word, not the true next word. Usually the two coincide, but when asked about something it doesn't know, the model doesn't say so — it strings together plausible-sounding words instead. That's how nonexistent paper titles and menus for restaurants that were never visited come into being. Just as the Patriot's window looked at empty space and concluded nothing was there, a prediction is only as accurate as the ground it stands on.

Even If It's Not the Highest Probability

Looking back, all these machines that guess what comes next do roughly the same thing. Where there's a law, they narrow things down using the law; where there isn't, they narrow things down using repetition. The Kalman filter reveals how much it doesn't know with a circle, forecasts with a chance of rain, LLMs with probability bars. The better the prediction, the more honestly it states its uncertainty rather than its confidence.

What's interesting is that even LLMs don't always pick the top word. Stringing together only the single most probable word every time makes the text predictable and repetitive. So, deliberately, they sometimes pick the second or third candidate instead. That's why the same question produces a slightly different answer each time. Even a machine, it turns out, has learned that always choosing the most plausible next thing kills the writing.

Back to the sentence in my notebook. No one knows what word comes next. That's true. And maybe that's exactly why it's all right. In the sentence of a life, the next word doesn't have to be the one with the highest probability. Sometimes a 3% word is what saves the sentence.

A Page for Today continues this way. I open the notebook to a random page and lay that day's thoughts on top of the line written there. Even I don't know what will be on the next page.

References

Read this series from the start: A Page for Today.