For less than a thousand dollars, a modern consumer laptop may be capable of trillions of calculations per second. But consumers’ applications of this hardware are often disappointing. Many are products of Wirth’s law, the observation that over time, software gets less efficient faster than hardware gets more efficient.
The occasional video stream may use a significant amount of compute, but modern browsers spend truly vast amounts of resources to display little more than formatted text and images on the screen, often with performance hiccups. Meanwhile, a video game from twenty years ago may require a fraction of the resources of today’s internet browsers such as Google Chrome, even as it draws myriads of polygons across millions of pixels, with elaborate lighting and shading effects computed on top.
The text editor in which this article is being written is doing very little that could not be done just as well on a computer from thirty years ago—a machine with thousands of times less memory and a million times less computing power. For the average person, compute is available in such great abundance that it exceeds our knowledge of how to put it all to productive use.
This abundance is famously the product of Moore’s Law, which asserts that transistor densities, and presumably performance, grow at an exponential rate. It has been true for the past six decades, and for those working in tech, this rapid growth is grounds for optimism about the future of humanity in general. They believe the endpoint of industrial society is post-scarcity: an edenic condition in which we are no longer constrained by resources and all hard problems of material life are solved, allowing human creativity to flourish and actualize anything it may desire.
While Moore’s Law has exceeded expectations, there is nothing at all to suggest any flaw in Gordon Moore’s central warning that his law was a result of economics, not physics. Accordingly, the future of computing will be determined not by our ability to innovate technologically, but by our ability to conceptualize what this compute would be used for.
For the moment, artificial intelligence seems to have supplied the answer. The large language model is the first consumer application in decades whose appetite for hardware has no visible ceiling: hundreds of millions of people now route their questions, drafts, and small work tasks through chatbots that are absorbing the functions of the search engine, customer service, and a growing share of clerical labor. This may sustain the industry at something like its present scale once the infrastructure buildout is complete, despite its current unprofitability. Firms are increasingly using customer-service agents, automated Excel spreadsheets, and low-level software engineering, but the claim that every firm will see autonomous agents that consume tokens by the billions sits uneasily with these prevailing uses.
The current buildout is a wager that there is more to come. If we find ourselves in a state where consumers and firms have more computational resources than we know what to do with, and struggle to name problems that we care about but lack the resources to immediately address, then it will indicate a bottleneck not in the production of goods but in the production of ideas—a much harder problem to solve. A world in which humanity lacks the imagination or ambition to strive for anything even greater than what is implied by the idea of “post-scarcity” is, upon closer inspection, a pessimistic prospect.
The Start of the Curve
Much of early computing was focused on automating various forms of bureaucracy. Bureaucracy is overwhelmingly linear: simple records for each citizen or customer. Even if millions of computations are required to process a single person, a modern laptop may be able to process facts about the entire world’s population in less than an hour. Exponential growth in computation has rapidly overwhelmed the demands of bureaucratic systems, yet pen-and-paper processes like signatures remain everywhere in the digital world. Public-key cryptography solved this problem properly decades ago; DocuSign still has 7,000 employees today.
Exciting advances in computing were for a long time downstream of those working to improve the graphics of video games, not administration. In the 1970s and 1980s, the number of pixels on the screen often greatly exceeded the number of bytes in memory, requiring a variety of methods for specifying images to display with grids of simpler components. The number of pixels was often even comparable to the number of operations per second, and various coprocessors such as “blitters” were used to offload often repetitive graphics work from the main processor, allowing it to do more complex tasks.
In the 1990s, graphics processors began to incorporate dedicated hardware for specific common operations, such as drawing polygons for 3D graphics, though still leaving more complex or special-purpose tasks to the main processor. In the early 2000s, the increasing complexity of “shaders”—mathematical operations applied in bulk to every pixel, or every polygon, or every instance of some other graphical primitive, often for complex lighting or stylistic effects—began to justify building processors for the bulk processing of simple but arbitrary functions.
In 1999, NVIDIA released the GeForce 256, marketing it as the world’s first GPU and a revolution in video game graphics. Within a few years, these gaming accelerators were being applied to a wide range of problems across science and engineering. Protein-folding simulations became at least twenty times faster thanks to general-purpose gaming technology.
In recent years, the holy grail of photorealistic video game graphics, ray tracing, became achievable on consumer hardware. Now the difficulty of photorealism in video game graphics is more a matter of game studios being able to pay small armies of artists to create realistic-looking art for the games, which is getting prohibitively expensive for no obviously greater spectacle. Meanwhile, the past twenty years have seen the growth of the indie game industry, which often uses more stylized art that is vastly cheaper, both for artists to produce and for computers to display. Technical advances in graphics no longer excite the imagination of technologists and consumers, though it was simpler games that attracted their efforts in the first place.
Because of this, there is no obvious remaining opportunity for video games to push computing demand. Higher resolutions and more complex geometries and animations may be possible, and may require more resources, but the demand for this likely does not exist on any scale that can justify it, and the art costs are not obviously sustainable as-is.
The technology that fueled the graphical development of gaming found a second life in machine learning, which is largely enabled by general-purpose GPUs. The initial work here was first explored in the 1980s, but was held back by the insufficient compute resources of the time. Machine learning will not be the last discarded or forgotten technology to be revived by hardware that finally caught up with it—but such technologies only ever reemerge if other breakthroughs make them worthwhile.
Moore’s Law from the Inside
It is not just the history of gaming that serves as a warning for what happens to innovation when technology gets ahead of demand. Ten years ago, a fabrication plant for microchips may have cost a few billion dollars to build. Today, modern fab campuses are getting pushed into the hundreds of billions. With this in mind, it is easy to forget that in the 1970s and 1980s, it was not uncommon for small chip startups to have their own fab facilities. In fact, the first open fabs focused on manufacturing other people’s chips took a while to appear. TSMC, an early pioneer in this business model, was not founded until 1987, a whole nineteen years after the founding of Intel.
Jerry Sanders, founder of AMD, famously used to claim that “real men own fabs” as a critique of the fabless model, for he believed that semiconductor companies could only hone the performance of their product if design and production were integrated into one process. However, as these fabs further pushed the limits of physics and engineering, their operating costs grew exponentially. In the semiconductor space, it is common for companies that bet wrong to be permanently removed from the cutting edge, and those that have survived the decades are to a large extent merely those who have survived the most rounds of industrial Russian roulette as manufacturing processes changed.
While there used to be dozens of companies at the bleeding edge, today, the leading chip fabs include only TSMC and Samsung, with Intel having fallen behind a few years ago and receiving heavy subsidies from the U.S. government in the hopes that it can catch back up. Five years after Sanders stepped down from his role as chairman, AMD spun off its fabs into GlobalFoundries, which went on to announce a strategic decision to no longer compete on the most advanced manufacturing processes.
The masks used in the photolithography processes that manufacture these chips have grown dramatically in price in recent years: from $1 million for 28-nanometer chips, to $10 million for 7 nm, to $40 million for 3 nm. Raising $1 million for a 28 nm chip tapeout is not unrealistic for a startup—raising $40 million is far more difficult. Upcoming transistor nodes can be expected to be even more expensive, and foundries like TSMC may eventually be forced to charge prices that few if any of their fabless customers can afford.
With fabs being as expensive as they are and almost every cost factor growing exponentially, small manufacturers are long dead, small designers may soon be priced out completely, and it is a serious possibility that pushing semiconductor technology further may exceed what even the most deep-pocketed nation-states are capable of paying. AI companies are currently very good at raising large sums of money on the hype of their product, but the actual profitability of these companies is questionable—some industry leaders have accused them of overcharging—and it’s not clear that LLMs alone will be nearly sufficient to justify building a trillion-dollar superfab.
Even though scaling laws provide a clear roadmap to justify such a concept, their exponential nature may simply grow too fast for fab technology to keep up. The compute requirements, and by extension the costs, may grow faster than revenue. There are many problems we can apply computation to, but if our imagination and ambition yield diminishing returns past a certain quantity of compute, these applications may be insufficient to fund the development of greater compute.
Branching Veins
Advances in technology aren’t always accompanied by a boom in scaled applications. The reason is that ideas, like minerals, come in veins. Any idea with a degree of generality that allows it to reflect meaning beyond a single confined event can be applied to a wide range of similar problems. In many cases, there may be a theoretically infinite range of such similarly shaped problems to which a given idea applies. The simplest forms of this often present as arbitrary scaling: throwing more resources at a given formula allows one to solve bigger problems of the same kind. When it comes to LLMs, this is why inference—the “thinking” that goes into the generation of an answer—now exceeds training in terms of compute costs.
This may also support scaling in multiple different dimensions, but there are limitations. The value of larger problems may diminish, or the cost of solving them may rise too fast to justify it. There is no shortage of deposits throughout the world containing valuable minerals that are too low in concentration or too difficult to extract at reasonable prices with existing technology. There are many large oil deposits around the world that go untapped because it is simply cheaper to extract oil from locations that are more convenient.
If we simply rephrase “post-scarcity” as “overproduction,” then this idea turns out to be rather old, and was a problem addressed by economic theorists and industrialists ranging from John Stuart Mill to Taiichi Ohno, the founder of Toyota’s production philosophy. The latter recognized overproduction as a greater problem than underproduction, because it reflects inefficiencies in the demand-production pipeline that lead to compounding waste through overprocessing, transport, waiting, and defects.
Ohno’s preference was to let demand pull production, modeled on the American supermarket shelf restocked only as customers empty it. AI scaling laws invite the opposite. Compute yields more capability, yes, but that says nothing about whether the capability will find a buyer. The trillion-dollar fab is production running far ahead of any order.
What separates useful from overbuilt technology is that the former sees widespread adoption even while primitive. For example, if virtual reality were seriously useful, it would have gained wide adoption during the “metaverse boom.” If the benefits outweigh the costs, people will generally adopt something regardless of its drawbacks, but if the technology must be made much more convenient before significant adoption can occur, as so many insist, the effort to diminish its inconveniences suggests a lack of genuine value to begin with. VR is an example of production exceeding our current ability to find a meaningful use for a technology, and this may change in coming decades as deeper thought is applied to the various design problems around virtual reality.
The cost of building a competitive semiconductor fab has grown exponentially over time, while the number of companies competing at the cutting edge has dropped exponentially. Extrapolating these trends another decade suggests fab costs in the trillions of dollars and approximately zero companies that can keep up. Even if theoretically possible, advances in semiconductors may simply slow down in response to a lack of profitable scalability, which is downstream of industrial society’s aggregate demand. This is a recipe for stagnation.
Moore’s Law can and will grind to a halt if human innovation cannot find applications for the vast amount of compute we have produced. The main fallacy of the post-scarcity thesis is that human needs could have been said to be fully satisfied after the agricultural revolutions of the twentieth century, or the widespread availability of cars and air-conditioning. After a certain point, technology develops only to account for desires beyond human needs. The future will be brought into existence only by people with a concrete desire to outstrip those needs.