In the last issue, I wrote about the AI companies calling for the industry to slow down. I’m still not sure whether they will. But if they do, what happens to the tools we’re already using? The workflows we’ve built? The software our companies bought that depends on AI?
I think a slowdown could actually benefit us. But it depends on what we slow down.
To explain that, we first need to understand why these companies say a slowdown is necessary. That means covering a couple of basics about the types of AI and how they’re trained.
For this discussion, think of AI in two broad groups.
Narrow AI focuses on a particular job. Navigating a road. Predicting the structure of a protein. Forecasting the weather. Its training is focused on the problem it’s being built to solve. Some of these systems incorporate general-purpose technology, but their purpose is specific.
General-purpose AI powers tools like ChatGPT. You can ask the same model to explain a mortgage, write a marketing plan, analyze a spreadsheet, or help build an app. It’s trained on a much broader collection of material because it’s being prepared for many different kinds of work.
So how does that training happen?
Imagine giving a model the beginning of a sentence: “Once upon a…”
Before it has been trained, its predictions are largely random. For our simplified example, imagine it produces: “Once upon a starship.”
During pretraining, the model works through enormous collections of text, including web pages, books, and code. This is where you hear that AI is trained on “the entire internet.” That’s an exaggeration, but the amount of material is enormous.
The model repeatedly predicts what comes next, compares its prediction with the actual text, and adjusts numbers inside the network called weights.
Think of those weights as dials that influence which words and patterns the model considers likely. Training tunes many of these dials together.
Picture the output changing from “Once upon a starship” to “Once upon a princess” and eventually “Once upon a time.” That’s an illustration of the adjustment, not a literal sequence every model goes through. With enough examples, training makes the familiar completion more likely.
Now imagine that process happening across an enormous range of language, ideas, and problems.
That takes specialized chips doing calculations around the clock, housed in data centers that need electricity and cooling. Training a leading model can run for months. The International Energy Agency compares the electricity use of a typical AI-focused data center with that of 100,000 households. That’s the facility’s consumption, rather than a bill for one model, but it gives you a sense of the scale.
One way companies have improved these models is by making them bigger and giving them more resources for training.
More adjustable dials, called parameters. More examples. More computing power to work through them.
For perspective, GPT-3 had 175 billion parameters. Meta later previewed Behemoth, an unreleased model with nearly two trillion total parameters. These are different model designs, but the numbers illustrate how large these systems have become.
Scaling has been a major driver of progress. It isn’t as simple as adding parameters and guaranteeing a better model, and companies have other ways to improve training. But putting more resources into the process has helped produce increasingly capable systems.
The concern some AI leaders raise is where that continues to lead. Could these systems eventually become better than humans across most demanding intellectual work? That’s the idea behind artificial superintelligence, or ASI.
Then there’s another possibility: AI becoming capable enough to build the next generation of AI.
This is called recursive self-improvement. Imagine an engineer building a more capable AI engineer. That new system finds a faster, more effective way to build the next one. Each improvement helps produce another.
AI is already helping companies write code and run research experiments. The larger concern is what happens if it can take over more of the process, finding better ways to design and train its successors. Progress could accelerate because the thing being improved is also doing the improving. Anthropic says it sees movement in that direction, while acknowledging that fully autonomous self-improvement is not yet here or inevitable.
That’s the argument behind the calls to slow down: the technology could advance faster than our ability to understand, test, and control it. It is a warning about what could happen, not proof of an inevitable outcome.
Which brings me back to what I would like a slowdown to accomplish.
I would support slowing, or temporarily pausing, the push toward more powerful general-purpose systems while we build the safeguards to manage their growth. That would need to cover more than simply making models bigger. It would also need to address AI accelerating its own development.
At the same time, I’d want focused work to continue. We’ve discussed protein folding and Isomorphic Labs in past issues. There are also opportunities in transportation, weather forecasting, manufacturing, and detecting equipment problems before something breaks. Those applications still need testing and oversight. But there are benefits worth pursuing while we work through the larger questions.
What does that mean for the general-purpose AI we already use?
Stopping the training of the next model doesn’t switch off the current one. The capability developed through its existing training remains available to use, provided the company continues operating the service.
That last part matters. Someone still has to pay to run it. The Ed Zitron discussion last issue raised questions about whether the economics hold up. A sharp price increase or loss of access could disrupt a workflow regardless of whether training slows down.
But assuming access remains available at a workable cost, I think there’s an enormous amount left to do.
From what I’ve seen, we haven’t scratched the surface of how we could reimagine work with the models we already have. We’re still connecting the right information, building processes, figuring out what needs review, and teaching people how to use the tools well.
Meanwhile, another model arrives. We stop to evaluate it, figure out what changed, and decide whether to rebuild around it.
Some of those improvements are worth the effort. Some solve problems the previous model couldn’t. But keeping up with releases can become its own work, even while the opportunity in front of us remains underused.
A slowdown could give us time to realize more of that opportunity. To finish building the workflow. To get people comfortable using it. To measure whether it actually saves time or improves the result.
If no more capable model arrived for the next year, what could your company still do with the ones we already have?
My answer is quite a lot.
I don’t know whether the AI companies will slow down. But if they do, I’d like us to use that time to build the safeguards we need and get more out of what already exists. Slower development of the next model could give us room to make faster progress with the current one.