You trained the AI. Big Tech got paid
Original ↗

You trained the AI. Big Tech got paid

Tech fastcompany.com · 4h ago· Cameron Armstrong Read on fastcompany.com ↗
Article not available — AI summary
On January 24, 1956, the American Telephone and Telegraph Company was the largest private company in the world. Its revenues amounted to almost 2% of the U.S. gross domestic product. It employed 746,000 people. It owned Bell Labs, the fabled research division that had already produced the transistor, the solar cell, information theory, and radio astronomy, while also laying the first transatlantic telephone cable. In the following decades, it would add UNIX, modern cellular telephony, the CCD image sensor, and the first active communications satellite to its long list of scientific milestones. This singular stretch of intellectual output paved the way for Bell scientists to eventually collect five Turing Awards and 10 Nobel Prizes. By many metrics, life as a regulated monopoly was very good for AT&T. Yet by the end of that day, AT&T had signed away exclusive rights to every single one of its 7,820 unexpired patents, royalty-free, to any American company that asked. AT&T would also license any future patents it filed at “reasonable rates.” A bleeding-edge intellectual property treasure hoard was suddenly and irrevocably opened to the free market. A technician at Bell Telephone, 1922 [Photo: Bell Telephone Magazine / Wikimedia ] Antitrust officials initially sold the agreement as a triumph. The Justice Department called it a major victory, with one DOJ lawyer hailing it as “miraculous.” Despite AT&T already existing for decades as a regulated monopoly, with its returns constrained to a relatively conservative (by today’s standards) ~7% per year, government regulators had pursued and established a landmark set of additional restrictions to curtail AT&T’s monopoly power. Soon, however, public sentiment started to shift. Business Week called the consent decree “hardly more than a slap on the wrist.” A House congressional subcommittee would later deem it “a blot on the enforcement history of antitrust laws” for its perceived lenience on AT&T’s exclusive supply chains and vertical integration. Both the ratepayers, who subsidized AT&T’s vast research budget through its rate contracts, and many in the federal government believed this unprecedented economic concentration to still be far too dangerous for the Republic to continue unabated. The now-infamous 1956 patent decree was just half of a settlement negotiated over seven years between AT&T and the federal government. AT&T wanted to continue manufacturing telephone equipment through its subsidiary Western Electric, but regulators believed the vertical integration was foreclosing competition within the industry. The federal government itself was so conflicted about this issue that the secretary of defense under President Eisenhower, Charles Wilson, pleaded with litigators that severing AT&T from Western Electric was “contrary to the vital interests of our nation.” The second half of the settlement barred Bell from pursuing any business other than telecommunications. A later analysis of the historical record revealed that 69% of Bell’s patents had little to do with telecom. Rather, they ranged from chemistry to computing to semiconductors to metalworking, lighting, optics, and more. The two halves of the settlement combined to ensure that this rich intellectual corpus (roughly 1.3% of all unexpired American patents at the time) became freely available essentially overnight— and included a guarantee from Uncle Sam that the big, bad legal wolf would not come knocking. Within just a few years, these released patents would generate an estimated $5.7 billion in follow-on patent value outside the telecom industry. About $3.5 billion of that value came from patents filed by young, startup companies. One famous branch of that startup explosion ran through Shockley Semiconductor, then Fairchild Semiconductor , and eventually into the storied company known as Intel . Intel’s cofounder, Gordon Moore (of Moore’s Law fame), would later describe this consent-decree-driven innovation cascade as “one of the most important developments for the commercial semiconductor industry”: ‘[It] allowed the merchant semiconductor industry to really get started in the United States. There is a direct connection between the liberal licensing policies of Bell Labs and people such as Gordon Teal leaving Bell Labs to start Texas Instruments and William Shockley doing the same thing to start Shockley Semiconductor in Palo Alto. This started the growth of Silicon Valley.’ A generation of brilliant, publicly subsidized scientists built one of the most impactful clusters of technical genius the world has ever seen. Bell generated patents, invented products, and became the undisputed epicenter of American frontier science for decades. But how? Sediment Imagine a carefully crafted rice paddy, terraced by exacting farmers who spent years precisely engineering a fertile environment. It looks like just a flooded field, but it turns out that rice is one of the few major crops that tolerates submerged roots. Since most weeds can’t tolerate submersion either, the water does the weeding. The deliberate flooding also cuts off the oxygen required for organic decomposition, so the soil retains more of its nutrients rather than burning them off like a dry, aerated field does. And the warm, waterlogged mud triples as an excellent habitat for nitrogen-fixing microbes. A well-tended paddy largely fertilizes itself, season after season, sometimes for centuries. This humble mud pond is one of the most productive growing systems humans ever designed. AT&T’s unique economic position as a monopoly set the conditions for Bell Labs’ culture of deliberate experimentation, patient exploration, and delayed harvesting. Bell drew from an enormous and stable nationwide revenue base that didn’t have to be rejustified every budget cycle. American regulators set this revenue base through AT&T’s prices by using a fixed percentage return calculation on the capital it invested in the network. Here, invested capital means switches, cables, buildings, and the like. Bell Laboratories’ Eero Saarinen-designed headquarters, Holmdel, New Jersey. 2000 [Photo: Library of Congress] At a normal company, research is a cost you minimize, but not at AT&T. Every dollar spent on research at Bell Labs did two things at once. First and foremost, it was a no-risk, recoverable cost subsidized by U.S. telephone ratepayers under contract. Second, it was a wellspring of new, capital-intensive technology for AT&T to build and deploy. This capital expenditure expanded the very rate base on which its guaranteed return was calculated. The more money spent on these new technologies, the larger the absolute profit gained by the same regulated ~7% return. This arrangement worked out very well for all parties for decades, but is not necessarily replicable. Nor is it obvious we should even try to recreate it, because it came with real costs too. Inefficient overinvestment, lack of price discipline, and most importantly an incentive to hoard inventions behind a monopoly wall all hurt ratepayers. But for much of the twentieth century, these guaranteed profits did objectively create an expansive paddy field in which one technological innovation after another could flourish. Frontier science looks different today. It’s rooted in model weights and GPUs. It is flooded with token spend and agentic loops. It blooms in data centers. While AI -assisted research is still young as a field, usage statistics show something big is happening in and around the major AI labs. Serious people are using this new technology to solve real problems , sometimes entire classes of problems, that were previously unsolvable. Protein structures, research mathematics , material design, drug discovery, and complex systems analysis are just a few of the fields where AI models are tangibly improving researchers’ abilities to clear humanity’s scientific roadblocks. But from where does this rich soil come? It’s not really a secret. OpenAI says it “primarily rel[ies] on publicly available information to teach [its] models how to be helpful.” Anthropic attempted to build a “central library of ‘all the books in the world’” to train its models. Sam Altman himself elaborates that their frontier models are trained on “the collective experience, knowledge [and] learnings of humanity.” Strip the euphemisms and you’re left with the stark reality that these unprecedented capabilities were assembled out of the self-expression of every person across the globe who ever wrote anything down. And the product built from this reality—at least according to the frontier labs’ own revenue, projections, and usage numbers—is the most valuable thing built in a generation. Maybe in history. Anthropic’s annualized revenue run rate rocketed from $87 million in January 2024 to $1 billion by year-end, grew roughly tenfold through 2025, and in May 2026, hit $47 billion. This makes it the fastest-compounding enterprise software company in history. OpenAI isn’t that far behind. An estimated 80% of the American workforce now holds a job where some portion of the work is exposed to these models. All of this impact was made possible by multiweek training runs over a data corpus measured in the lifetimes of billions. This is the private capture of public genius. A frontier model is the compression of a massive amount of training data into numerical weights. It’s staggering to even think about the combined collection of books, forums, code repositories, manuals, papers, chat logs, transcripts, court cases, essays, comment sections, articles, tutorials, and every errant thought scrapeable by the frontier labs’ army of spiders crawling across the internet and beyond. In a way, its incomprehensibility is almost like psychic armor. It’s too big to understand directly. Consider a wild river delta. As water runs from highlands to the sea, it erodes the land it travels through and carries the debris downstream as sediment. Silt, sand, clay, and all manner of organic material, scoured from every inch of tributary and riverbank, from plowed fields to rugged hillsides, end up aggregated in the delta. So does the richness of every life the river supports along the way. A continental watershed, swirling, accumulating, and ultimately settling at its terminus. The vast volume of disparate material combines in the delta to form something lush, strange, and alive. The Yukon Delta National Wildlife Refuge , Alaska. Captured by Landsat-7 on September 22, 2002. [Photo: NASA / USGS / Landsat ] And what is the sum of all human knowledge if not this? Every cluster of letters scraped from the pages of history (the literal tokens an AI model ingests) is a single grain of silt deposited by the ever-flowing river of man’s exploration. Pile enough grains and you understand the movement of the stars. Stare long enough at the mud and you see the structures of logic itself. The large language model’s transubstantiation of alluvial soil into answers is the grand harvest of the society that grew it. But subtract the dirt and there is no delta. Subtract the corpus and there is no harvest. There is nothing. The model did not learn to reason in a vacuum. It absorbed rationality by observing rationality over and over and over again. Its powers of generalization are downstream of every example, correction, and argument it subsumed. A human decision somewhere in the echoes of history, culture, and science set the stage for today’s chatbot response. This cultivated intelligence grows from the sediment of human sensemaking, but there is no sediment here that was not deposited by someone . Many of those someones are dead. They wrote the ancient texts, tested the baseline science, and recorded the history of the world from antiquity for the benefit of all of us still here. But many of those someones are alive. They are writing the working code that the model spits out. They’re pushing that baseline science past its frontier. They’re organizing and investigating and acting upon and reacting to the infinite feed of current events. Any response germane to today is borrowed from somebody . In fact, you’re one of those somebodies. Literally. Your 2 a.m. shitpost. That eloquent reply to a stranger’s essay. The scathing restaurant review you left. Your captions, comments, inside jokes, and all of your public conversations. Every contribution you ever made to the infinitely branching stream of digital communication, big and small, has settled somewhere in the delta. The Nile River delta fed Egypt for 5,000 years. The Mekong and the Ganges regions still feed hundreds of millions today. It’s no coincidence that every cradle of civilization owes its formation in whole or part to the floodplains and deltas of great rivers. These areas supported humanity through our most primitive eras with little more than the inherent richness of their raw materials. This dirt is begging to burst forth with life, yet somehow the richest farmland on earth is, almost without exception, accidental. So too goes the internet. We myriad digital denizens of the information superhighway did not set out to create a training corpus. We wrote for ourselves and for each other. We joked, argued, taught, complained, flirted, and debugged our way into this aggregated mass of interrelational raw material now harvested by private capital. The field of economics (which is also in the corpus) has a vocabulary for this. To categorize any resource, economists ask two questions: Is it excludable, and is it rivalrous? More plainly, can you stop people from using it? And does one person using it diminish what’s left for everyone else? There are caveats and subcategories, but this simple test gives us a map. If a good is excludable and rivalrous, it is a private good. Think about a sandwich. If I eat it, it is gone, and the law protects me from sandwich thieves. If a good is excludable but mostly nonrivalrous, it is a club good. A Netflix subscription is a club good. If I watch a movie, you can still watch it too, but only if we both pay to have access. If a good is hard to exclude people from using and rivalrous, it is a common-pool good. A pasture is the classic example. Many farmers can access the pasture, and while one cow grazing does not destroy the field, add enough cows and they’ll eventually gnaw the grass down to dirt. This is the infamous “tragedy of the commons” problem. Finally, if a good is hard to exclude people from using and nonrivalrous, it is a public good. Streetlights are public goods. Once the street is lit, all of us can walk beneath the light, and my doing so does not darken the road for you. Private and club goods are typically governed by profit-seeking actors and the legal system in which they operate. Public goods are primarily governed by governments or nobody, and common-pool goods tend to exist in a liminal space where everybody seeks the benefit and nobody wants to own the costs of upkeep. The frontier labs generally argue that data on the internet is open for training under fair use copyright regimes. In economic terms, this argument implies the internet is a public good. The mass scraping, ingestion, and use of internet data for training does not destroy that original data. Every blog post, tweet, and flame war is indeed still there and for the most part accessible. Nobody clearly owns it. Does the platform you post on own your posts? Do you share ownership with the platform? Can this relationship change over time? You did post it online for free after all. Except granting access is not the same thing as giving license. A library card gets you access to read a book, not to photocopy the entire library. Buying a national park pass does not confer logging rights. Visiting an open store does not entitle you to steal its inventory. Public access to work on the internet does not automatically confer usage rights. And there is a second, deeper problem with “you posted it, you accepted this.” Until very recently, the LLM training data use case did not exist and could not have been reasonably foreseen by a party posting online. A blogger from 2008 could not have consented to their work being used to train a language model today, because that wasn’t conceivable back then. Consent can’t be assigned backwards in time, least of all for a sci-fi subplot turned real. The current legal battleground for LLMs is a story of nonresolution. Notably, despite our moral intuition, access and consent are irrelevant to the frontier labs’ primary legal defense claims of “fair use.” Instead, courts evaluate four criteria as they rule on a fair use defense. They look at the purpose of the use of copyrighted material, the nature of the work, the amount used, and the effect on the market for the original. In practice, these four criteria for fair use generally collapse to two important questions: Is the new work transformative? And does it harm the market for the original? In June of 2025, Judge William Alsup, senior district court judge for the U.S. District of Northern California, ruled in Bartz v. Anthropic that training on legally acquired books was “quintessentially transformative,” but building a library from pirated books was “inherently, irredeemably infringing.” With this mixed victory, Anthropic faced a theoretical exposure of up to $70 billion in copyright damages and quickly settled the case for $1.5 billion a few months later. It’s the largest copyright settlement in U.S. history—so far—and granted no future licenses to Anthropic, nor did it clarify any law going forward. In a related ruling, Kadrey v. Meta , Judge Chhabria, a judge in the same U.S. district court, found LLM training similarly transformative and grudgingly ruled the evidence of market harm insufficient. In his ruling, he criticized the plaintiffs for putting forth almost no evidence of market dilution and suggested that LLMs’ ability to flood a market with AI work similar to the training data “will often cause plaintiffs to decisively win the fourth factor—and thus win the fair use question overall—in cases like this.” Complicating the discussion further, the U.S. Copyright Office issued a nonbinding report in 2025 concluding that public availability does not inherently allow fair use model training. As of this writing, there is no settled legal standard for measuring LLM-driven market dilution, but this is primed to be a major confrontation in future legal decisions. Already, dozens of lawsuits and policy fights are testing the frontier labs’ evolving training-data defenses. The labs’ most seductive defense is also the simplest. “It’s just reading” is a common refrain among technologists defending AI model training, and it is a compelling argument. Every writer alive is built from the books they consumed. Nobody sends Hemingway’s estate a check for being inspired by The Old Man and the Sea . If the model is just another reader, it owes what every reader owes: nothing. A person who reads 10,000 books in their lifetime becomes one more writer, working at human speed, publishing at human volume, and returns their sediment to the delta one grain at a time. A model that reads everything becomes a printing press that prints more printing presses. It spits out work at industrial volume, trains its successors, and competes with the very writers it consumed, at the push of a button. Inspiration never diluted a market, but printing does. A printing press. Engraving by Wilson Lowry after John Farey. 1819 [Photo: Wikimedia ] The invention of the Gutenberg press around the year 1440 ultimately led to the passage of the Statute of Anne in 1710. A cartel of powerful book publishers lobbied the British Parliament to restore their monopoly rights over the book trade, and Parliament instead vested the right in authors as legal owners. The incumbents asked for protection and the public’s representatives handed ownership to the creators. This statute established the basis o
Read on fastcompany.com ↗
0:00 / 0:00