All articles

No. 001 · · 29 min read

You Are the Oil

Everything humanity wrote, sang, filmed and photographed was drilled, refined into artificial intelligence and sold back to us by the token. Do you wonder how it was done, what the machines are really doing, or what to do about it?

Snow

The bright smile of a young boy holding a peace sign, surrounded by his fellow classmates.

What did you want to be when you were twelve?

Maybe a lawyer. A doctor, an architect, a musician, a writer. Somebody who knew things other people needed, and got paid well to know them.

A common misconception about oil is that it’s made of dinosaurs. Crude oil formed from ancient marine life, microscopic plankton and algae, life that drifted through ancient seas, died, sank and was buried before the dinosaurs existed.1 Countless tiny and not so tiny lives, pressed under sediment for millions of years and cooked by heat and pressure into the black liquid that built the twentieth century. Like you, none of them were asked. None of them were paid.

The most valuable new resource on Earth was made the same way. This time, the plankton – its you. You are the oil.

Every forum answer you typed at 2 a.m. Every Google review, every blog post, every wordpress, every reddit, every tweet, every caption (yes, you spent a decade labelling data), every IG photo dump, every song you uploaded, every video essay on YouTube, every line of public code, every book on every library shelf. For thirty years, billions of people poured their thoughts, experiences, their lives and intellect, into the internet. That deposit was drilled, refined and compressed into artificial intelligence, and it is now sold back to the people who made it at about $20 a month.

In African Solarpunk we followed the metal that AI is built from. This time we follow the other raw material: the part that came from you. Our journey begins at the drill, to the refinery, to the inside of the machine, to your bank account, ending with you at twelve-years-old.

The Drill

Whether the drilling was theft is being argued in courtrooms right now. You being asked for permission to be turned into a commodity, is not up for debate. To be clear, you were not. We recommend the film, There Will Be Blood, (feat. Daniel Day Lewis), a story about an ambitious oil prospector.

The open web. Since 2008 a small nonprofit called Common Crawl has sent automated programs, called crawlers, across the internet to copy public web pages. Its archive now holds more than 300 billion pages and grows by 3 to 5 billion a month.2 When OpenAI trained GPT-3 in 2020, the ancestor of ChatGPT, 60% of its training mix came from a filtered copy of that archive.3

AI companies now run crawlers of their own, and they broke the old bargain of the web. A search engine copied your page and sent you readers in return. In 2025 Cloudflare counted how many pages each AI company’s crawlers took for every visitor they sent back.4

A field of 3,700 small grey dots and a single gold dot.
Pages OpenAI’s crawlers copied for every visitor they sent back, 2025. For Anthropic, the maker of Claude, the crop field would be 135 times bigger.

Books. Books are long, edited and dense, which makes them prized training material. Several AI companies downloaded millions of them from pirate libraries. In 2025 a US judge ruled that training on books you have bought can be fair use, but that a library of pirated copies was not: “Anthropic had no entitlement to use pirated copies for a central library.”5 Anthropic settled for $1.5 billion, roughly $3,000 for each of about 482,000 books, and the court gave final approval in July 2026.6 In a similar case against Meta, a different judge sided with the company, while noting that the authors had made the wrong arguments.7

Video. In 2024 Proof News found subtitles from 173,536 YouTube videos, taken from more than 48,000 channels, inside a training dataset used by Apple, Nvidia, Anthropic and others. Crash Course, Khan Academy, MrBeast and PewDiePie were all in there.8 The same year, The New York Times reported that OpenAI had transcribed more than a million hours of YouTube video to train GPT-4.9

Music. Sued by the major labels, the song generator Suno told the court that its training data “includes essentially all music files of reasonable quality that are accessible on the open internet.”10 In July 2026 a court in Munich found that Suno had ripped songs from YouTube to train on them, and ordered it to stop.11 The same court had already ruled, against OpenAI, that song lyrics memorised by a model still count as copies, even when they exist only “in the form of probability values.”12 Warner and Universal have since settled with the music generators in exchange for licensed versions.13

Images. The image generators of 2022 learned from LAION-5B, a list of 5.85 billion pictures and captions gathered from the open web.14 In the UK, Getty Images lost most of its case against Stability AI in November 2025, in part because the training had happened in another country.15

Conversations. Reddit now licenses its users’ posts to AI companies. Google pays it about $60 million a year.16 The company was paid. The people who wrote the posts were not.

You can guess that some of this was legal. Some of it wasn’t. Much of it is still undecided: The New York Times’ case against OpenAI and Microsoft, filed in 2023, is still waiting for a ruling.17 The pattern is plain all the same. The people who made the work were rarely asked and almost never paid. Where money did change hands, it went mostly to whoever held the rights: labels, publishers, platforms, i.e. Not You!

The Refinery

A refinery for oil is a forest of steel columns. A refinery for intelligence is a warehouse.

Speaking from experience, we can walk you through one. You pass a fence, a guard, a badge reader and a small airlock of a room that only opens one door at a time. You then risk hearing loss from prolonged exposure to… thousands of fans all at once, picture a billion screaming mosquitoes directly in your ear. In the cold aisle, chilled air pours toward the fronts of the machines. Step round to the hot aisle behind them and it feels like a hairdryer in your face. Either side of you stand racks, steel cabinets taller than a doorway, packed with flat metal trays and rows of blinking lights. Bundles of fragile fibre run overhead, carrying data as pulses of light. Truly fascinating technology.

Every tray is a computer. Most run Linux, the free operating system that also runs most of the internet. The prized ones hold GPUs, chips first designed for video games. A GPU does one job extremely well, and in parallel: it multiplies long lists of numbers and adds up the results, billions of times a second. At the level of the silicon, that is all artificial intelligence is. Multiplication and addition, at a scale that is extremely difficult to picture.

The machines do two kinds of work. Training is how a model is made: thousands of GPUs working in lockstep for weeks or months on one job. Meta trained its largest Llama 3.1 model on more than 16,000 of them.18 Inference is how a model is used. When you send a question, a server slots it into a batch alongside other people’s questions, and the GPUs work through them together. Your private conversation is processed side by side with strangers’, in a building you will never enter.

Today’s flagship AI racks hold 72 GPUs wired so tightly together that they behave like one enormous chip. Each rack draws around 120 to 140 kilowatts, day and night.19 A household kettle draws about 2.

Bar chart of power drawn: a kettle 2 kilowatts, an average server rack in 2021 7 kilowatts, an AI rack of 72 GPUs 120 to 140.
One AI rack draws as much power as sixty kettles, around the clock.

Almost every watt that enters the hall leaves it as heat. Air alone cannot carry that much away from the newest racks, so coolant is piped straight onto the chips. In most places the heat then goes up through the roof and into the sky. Solarpunk asks the obvious question: why throw it away, when a greenhouse next door would pay for it?

It adds up. In 2024 data centres used about 415 terawatt-hours of electricity, around 1.5% of the world’s supply. The International Energy Agency expects that to more than double by 2030, to about as much as Japan uses today.20

Bar chart: data centres used 415 terawatt-hours in 2024 and are projected to use 945 in 2030.
Electricity used by the world’s data centres. Source: International Energy Agency.

Inside the Machine

This is the part where most people think all this stuff is too hard for them. It is not.

The crux of the idea: a model is a huge collection of adjustable numbers, called weights (or parameters), and “learning” means nudging those numbers until the model’s predictions get “better”.

“Better”, of course, is according to their

Trust Me Bro Benchmarks.

Fireship

One neuron

The clearest demonstration comes from Andrej Karpathy, a founding member of OpenAI and former head of AI at Tesla, who wrote a teaching tool called micrograd in about a hundred lines of code. (We highly recommend you view it on YouTube.)21 Take one artificial “neuron”. It receives two inputs, multiplies each by a weight, adds a bias, and squashes the result with a gentle S-shaped curve called tanh:

n = x₁·w₁ + x₂·w₂ + b
o = tanh(n)

With Karpathy’s numbers the inputs are 2 and 0, the weights are −3 and 1, and the bias is 6.88. Work it through and you get n = 0.88 and an output of 0.71. That is the forward pass: numbers go in, an answer comes out.

Suppose the answer was incorrect. Which numbers should change, in which direction, and by how much? Calculus comes to the rescue with the chain rule, working backwards from the output. This backward pass gives every number a gradient: how much the output moves when that number moves. Nudge the weight w₁ up a tiny amount and the output rises by exactly the same amount, so its gradient is 1.0.

Computation graph of one neuron, showing each value and, in red, its gradient.
One neuron, from Karpathy’s micrograd. Black numbers are the values going forward. Red numbers are the gradients coming back: how much the output would move if that number moved. Blue boxes are the weights, the numbers learning changes.

Learning is walking downhill

Picture the model’s error, called the loss, as a valley. Again, by error, we mean how “intelligent” the model seems to you as the user. High error in some models, is why you say the model hallucinates. Every weight is a direction you can walk in, and the gradient tells you which way is downhill. So every weight takes a small step down the slope:

w ← w − η · ∂L/∂w

L is the loss and η is the size of the step. Forward, backward, step. Repeat. That loop is the whole of deep learning.

A U-shaped loss curve with a gold ball stepping down the slope, each step smaller than the last, until it settles at the lowest point.
Each step goes downhill. The steps shrink as the slope flattens near the bottom.

A real model has more than one weight, so the valley becomes a landscape. With two weights you can draw it, the way a map draws mountains: height is the loss, and the contour lines join points of equal loss. Gradient descent is a walker in fog who can only feel which way the ground slopes underfoot, and keeps stepping downhill.

A 3D loss landscape shaded like a relief map, with contour lines on the surface and on the floor, and a gold path stepping from a hillside down into the lowest valley.
The loss landscape for two weights, drawn like a relief map. A real model walks the same kind of landscape in billions of dimensions.

Words become numbers

A chatbot is the same loop, scaled until the numbers stop meaning anything to a human brain.

First, text is chopped into tokens, pieces of words about three-quarters of a word long on average. Each token is swapped for a list of several thousand numbers, which the model learns during training. Mathematicians call a list like that a vector: think of it as an arrow pointing somewhere in a space with thousands of directions. Words with related meanings end up pointing in related directions.

Five tokens drawn as arrows from the origin in three dimensions, with each token written out beside the plot as its list of numbers, from dimension 1 up to n.
Each token is a vector: an arrow pointing somewhere in space. We can draw three dimensions; a real model uses n of them, 4,096 in Meta’s smallest Llama 3.1. ‘The’ and ‘the’ point almost the same way. Values are illustrative.

Paying attention

Those columns pass through dozens of layers. In each one, a step called attention lets every token look back at the earlier ones and decide which matter most for what comes next. Then a block of neurons, exactly like the one above, transforms the result.

A triangular grid of attention weights, each row adding up to one.
Attention weights. Each row is a word looking back at the words before it, and each row adds up to one. Values are illustrative.

Choosing the next word

The last layer gives a score to every token the model knows, often around 100,000 of them. A formula called softmax turns the scores into probabilities that add up to one:

pᵢ = e^zᵢ / Σⱼ e^zⱼ

Try it by hand. Imagine a model that knows three words, “sun”, “rain” and “snow”, and scores them 2.0, 1.0 and 0.1. Raise e (about 2.718) to each score: 7.39, 2.72 and 1.11. They add up to 11.21. Divide each by the total and you get 66%, 24% and 10%.

Scores of 2.0, 1.0 and 0.1 for sun, rain and snow become probabilities of 66%, 24% and 10%.
Softmax turns raw scores into probabilities that add up to one.

Now suppose that in the sentence the model is learning from, the real next word was “snow”. The model gave it only 10%, so its loss is −ln(0.10) ≈ 2.3: badly wrong. The backward pass nudges every weight so that, next time, “snow” scores higher.

That sentence came from somebody, probably you. Every correction in training is a small lesson taken from a person’s writing. Multiply it by trillions and you have a model.

Using a model runs the same machinery forward only. It picks a token, adds it to the text and runs the whole stack again for the next one. A 500-word answer takes roughly 670 trips around that loop.

The size of it

Researchers size this work with two rules of thumb. If a model has N weights and reads D tokens while training:

training   ≈ 6 × N × D   operations
one token  ≈ 2 × N       operations

Meta’s Llama 3.1 has 405 billion weights and read more than 15 trillion tokens.18 That comes to about 4 × 10²⁵ operations, a 4 followed by 25 zeros. Every token it writes afterwards costs about 800 billion more.

Stored at 16 bits each, those 405 billion weights make a file of about 800 gigabytes.

That file is the refined product. You were its raw material.

The Five Senses

The same two tricks, turn everything into numbers and then learn by nudging weights, power nearly everything else AI does.

Images. Image generators learn to remove noise. In training, they take a real picture, bury it in static, and learn to predict what static was added. To make a new picture, they start from pure static and remove predicted noise, step by step, steered by the words of your prompt. Every one of those learned steps came from studying billions of pictures made by people. Probably your facebook wall from 10 years ago.

Six squares from pure noise to a clean picture of a sun over mountains.
The same picture at six points of a real diffusion noise schedule. A model learns to walk this row from left to right.

Music. Every note you hear is several pure tones sounding at once. Split them apart and each moment of sound becomes a short list of numbers: how strong each pitch is. Line those lists up over time and you have a spectrogram, a picture of the sound. From there, prediction and denoising work just as they do for words and images. An AI song carries the texture of a genre because it learned (stole) that texture from people who spent their lives making it.

A four-note melody as one wiggly line, and the same melody as a 3D landscape of pitch over time with a row of peaks for each note.
The same four notes twice. A microphone records one line. A model sees a landscape: each note is a row of peaks, and the peaks step to the right as the melody climbs.

3D. Newer methods describe a scene as millions of soft, coloured blobs floating in space. Gradient descent, the same downhill walk as above, adjusts every blob’s position, size and colour until the scene looks right from every camera angle in a set of photographs.

Using a computer. An AI agent is shown a screenshot, decides what to do, and replies with an instruction such as click at 412, 188. Software carries it out, takes a new screenshot and hands it back. It learned what buttons mean from years of human screenshots, tutorials and instructions.

A loop from screenshot to model to action and back.
How an AI agent uses a computer.

Crawling. A crawler fetches a page, copies it, collects every link on it and repeats. A website can post a file called robots.txt asking crawlers to stay away, but obeying it is voluntary. In July 2025 Cloudflare began blocking AI crawlers by default for new customers and testing a way to charge them per page.22

Follow the $20

We built Harvest because none of this can be fixed while it stays invisible.

Harvest draws the AI economy as a river. Choose the tour Where your $20 goes and one subscription payment flows downstream in front of you: first to the company that built the model, then mostly to the data centres it rents, then to chipmakers, power companies and builders, then to the factories in Taiwan, and finally, as a thread, to the mines.

Sankey diagram of where a $20 subscription goes, with a dashed branch carrying $0.00 to the people whose work trained the model.
Where $20 goes, simplified from Harvest’s sample numbers. The dashed red branch is the one that does not exist. Unless of course, you won your lawsuit against OpenAI, Anthropic, Alphabet, Suno, Spotify, xAI, how exhausting.

Look for the branch that flows back to the people whose words, songs and pictures trained the model. Well would you look at that. It does not exist. In the real economy it barely exists. The $1.5 billion book settlement was a one-off payment for past piracy. The licensing deals mostly pay platforms and rights holders. The money from every subscription, every month, flows straight past the people who made the raw material.

This is why seeing the flow matters so much to solarpunk. Solarpunk claims that the systems we depend on can be made visible, local and shared: energy from the sun overhead, food from the greenhouse down the road, heat passed along from the data hall next door. None of that can be organised while the flows stay hidden. So Harvest is written for optimistic future-humanist readers on purpose. These are not boomers. A flow of money that only specialists can follow will only ever be steered by specialists. Cyberpunk depends on that kind of darkness. Solarpunk starts by literally turning on the lights.

The Rent

Perhaps by now, you remember what you wanted to be when you were twelve.

Mike Ross was a great character, so say you wanted to be a lawyer. You pictured the sharp suit, the case nobody thought could be won. The path was long but clear: get the degree, start as a junior, spend years reviewing thousands of documents, drafting memos and researching precedent, and slowly become good. Those junior years were your training. They were also how you paid the rent while you learned.

A model can now do a large share of that junior work overnight. It learned how by reading court judgments, textbooks, law-firm blog posts and legal forum threads, written by the very people whose footsteps you meant to follow.

Swap lawyer for designer, translator, illustrator, copywriter, junior programmer, session musician. The shape is the same.

It already shows in the numbers. Researchers at Stanford tracked payroll records for millions of American workers. Since generative AI arrived, employment of 22- to 25-year-olds in the jobs most exposed to AI has fallen 19% relative to their peers in less exposed jobs. The gap comes mainly from fewer young people being hired. Firings have barely changed.23

So here it is in plain brutal english. Nobody stole your brain. You live in a world where you’re being told that there is a 10% change that you’re going to die, and your brain is worthless. The knowledge you would have spent a decade building was already written down by others, copied without asking, compressed into a file, and is now rented back to you, and to your future employer, by the month.

Medieval serfs farmed land they did not own and paid their lord in labour and harvest. The economist Yanis Varoufakis argues that the digital economy has quietly rebuilt that arrangement, with cloud platforms as the new estates.24 You do not have to accept all of his argument to recognise the shape. We produce the raw material for free. A handful of companies own the refinery. We pay rent to use what comes out.

If you are young, you have probably spent the past few years hearing, in adverts and interviews, that the machines will soon know more than you ever will, and that the sensible thing is to become a good customer. Much of that message comes from the people who own the refinery. Notice who benefits when you believe it.

Science fiction has a name for this future: cyberpunk. Dazzling technology, owned by a few, rented to everyone else. William Gibson, who helped invent the genre, put it best: “The future is already here. It’s just not very evenly distributed.”

What It Is Good For

None of this makes the technology bad. It is extraordinary, and we use it every day, including to research this piece.

In a study of more than 5,000 customer-support agents, an AI assistant raised the number of problems solved per hour by 14% on average, and by 34% for the newest, least experienced workers.25 Used well, it narrows gaps. Used carelessly, it misleads: in a 2025 trial, experienced programmers were 19% slower with AI tools, while believing they were 20% faster.26

A single question costs little energy. Google measured its median Gemini text prompt at 0.24 watt-hours, about as much as nine seconds of television.27 It is the billions of prompts a day, and the ever-larger training runs behind them, that add up to a country’s worth of electricity.

So the machine itself is really the easy part. The hard questions are who owns it, who gave permission, who gets paid, where does the power come from, and where the heat goes. Those are choices, that really belong to the future of our species.

The Fork

Every tool that changes how we think has changed us.

Around 370 BC, Plato wrote down an old warning about a new invention. Writing, an Egyptian king declares, “will create forgetfulness in the learners’ souls, because they will not use their memories.”28 He was right about the forgetting. He was wrong that it would make us smaller. Writing let our species remember far more than any one brain could hold.

The changes are physical, and they can be measured. When adults in Portugal and Brazil who had never been to school learned to read, brain scans showed that part of their visual system had been rebuilt: a patch of cortex now answered to letters, and answered a little less to faces.29 Go back further and tools changed our genes. About 7,500 years ago, dairy farmers in central Europe began drinking milk, and a gene that lets adults digest it spread with them. In northern Europe today, most people carry it.30 A technology rewrote part of human DNA in a few hundred generations.

Artificial intelligence is the first tool built to do the thing our species is named for: thinking. And it is spreading faster than anything before it. ChatGPT reached an estimated 100 million users within two months of its launch.31 By the middle of 2026, about one in five working-age people on Earth used generative AI, and in wealthier countries nearly three in ten.32

You can already feel what this kind of change does. Your brain is the same organ your great-grandparents had, but it is wired by use. London taxi drivers, who memorise some 25,000 streets, grow a measurably larger rear hippocampus, the region that handles spatial memory.33 It works in reverse too: people who lean harder on GPS have worse spatial memory when finding their own way, and those who increase their GPS use decline faster.34 If you grew up with fast internet, games and a phone in your pocket:

  • You probably know three phone numbers by heart. Your grandparents knew dozens.
  • You can find any fact in ten seconds, and forget it just as fast, because your brain knows it can always look again.
  • If you played action games, you likely track moving objects and spot changes on a busy screen better than people who didn’t.35
  • If you juggle many streams at once, you may find it harder to shut out what doesn’t matter.36

None of this makes you worse than your grandparents. Brains adapt to the tools around them. The new question is what happens when the tool can do the thinking.

The early evidence points in two directions at once. At the MIT Media Lab, people who wrote essays with ChatGPT showed the weakest brain connectivity of three groups, and many could not accurately quote the essays they had just “written”. But people who wrote on their own first, and only then used the AI, showed more engagement than those who started with it.37 In a survey of 319 knowledge workers by Microsoft and Carnegie Mellon, the more people trusted the AI, the less critical thinking they reported; the more they trusted themselves, the more they did.38 A study of 666 people in the UK found heavier AI use linked to weaker critical thinking, driven by handing mental work to the machine, and the youngest participants leaned on it most.39 These are early studies, and some only show correlation. Together they sketch a fork with three paths.

Hand the thinking over, and capability rises for a while, then quietly decays, because the skill was never built. Never touch it, and life goes on at the old, slow rate. Think first, then build with the machine, as a sparring partner, a lab assistant, an extra pair of hands, and your abilities compound. Compounding is merciless:

C(t) = C₀ · (1 + r)ᵗ

Grow 2% a year for thirty years and you end up about 1.8 times where you started. Grow 10% a year and you end up about 17 times. Two people begin in the same place. A generation later they barely recognise each other.

A spacetime-style grid stretching into the distance, bent into a deep well by a single glowing point labelled AI, with a timeline from 2026 to 3026. From "you, now", a red path spirals into the point, a grey path runs straight on, and a green path swings around the well and heads off in a new direction.
AI bends the path of everyone who comes near it. Some are pulled in. Some pass by unchanged. Some use the pull to go further. A metaphor, not a measurement; the timeline runs on a log scale.

A fork between two people is a difference in careers. A fork across hundreds of millions is a difference between societies, and the gap in who uses these tools is already widening between rich and poor countries.32 Carried across generations, it becomes something larger still. Children who grow up thinking alongside machines, children who grow up letting machines think for them, and children who never meet them will not end up with the same minds. That is the same process that turned a patch of visual cortex into a letter detector and spread a milk gene across a continent, running faster than it ever has.

For the first time since writing, the direction our minds evolve in is a choice. It is being made right now, mostly by default, mostly by the people who own the refinery. What that fork means for the future of our species is the subject of our next piece.

Take Back the Refinery

You can accept the arrangement, or you can change it. Alone, changing it is very close to hopeless. Together, however, it has already started to work.

Look at who has actually won anything so far. The owners of nearly half a million books, acting as one class. GEMA, a collective of German songwriters, which has now beaten both OpenAI and Suno in the first round of court. Every real concession in this story came from people who organised.

Understand the machine. Now you can explain a data hall, a gradient and a token to anyone at dinner, which already puts you ahead of many of the people making decisions about them.

Run it yourself. Open-weight models, which anyone can download, (the best ones originate from China. Ranked, it’s Kimi K3, DeepSeek V4-Pro, GLM-5.2, MiniMax M3, and Qwen3.6-35B-A3B) now run on an ordinary laptop with free tools. No subscription, and your words never leave your house. A tool you own cannot be taken away or repriced.

Decide who crawls you. If you publish anything, you can block AI crawlers or charge them. Many large publishers already do.

Join or build a collective. Creators’ unions, collecting societies, data cooperatives, community-owned compute. The labels and publishers got paid because they negotiated as blocs. So can you.

Demand to see the ledger. Since August 2025, the EU has required the makers of large AI models to publish a summary of what they trained on.40 Push for the same everywhere, and for the money trail to be as public as the training data.

Bring the refinery home. Data halls do not have to sit behind fences, dumping heat into the sky. You definitely do not want them in space either. They can be small, owned by the communities around them, powered by the local sun, wind or earth, and piped into the greenhouse next door. That is the world Snow is working toward.

The plankton never had a say in what became of them. You do. And the twelve-year-old who wanted to know things worth knowing is still in there. What you know still has a price. Together, we can make sure it gets paid.

Sources and notes

Figures were checked against the sources below in September 2026. Harvest’s river uses illustrative sample numbers, labelled as such in the tool.

Notes

  1. U.S. Energy Information Administration, “Oil and petroleum products explained.” https://www.eia.gov/energyexplained/oil-and-petroleum-products/ ↩
  2. “Selecting Language Models for Social Science: Start Small, Start Open, and Validate,” arXiv:2601.10926 (2026), on the size and growth of Common Crawl. https://arxiv.org/abs/2601.10926 See also Mozilla Foundation, Training Data for the Price of a Sandwich (2024). https://www.mozillafoundation.org/en/research/library/generative-ai-training-data/common-crawl/ ↩
  3. Brown et al., “Language Models are Few-Shot Learners,” arXiv:2005.14165 (2020), Table 2.2. https://arxiv.org/abs/2005.14165 ↩
  4. Cloudflare Radar 2025 Year in Review, as reported by InfoQ (December 2025): about 3,700 pages crawled per referral for OpenAI and about 500,000 for Anthropic. Ratios move month to month. https://infoq.com/news/2025/12/cloudflare-2025-ai-bots/ ↩
  5. JURIST, “Judge approves record $1.5 billion AI copyright settlement involving Anthropic” (July 2026), quoting Judge William Alsup’s June 2025 order. https://www.jurist.org/news/2026/07/judge-approves-record-1-5-billion-settlement-involving-anthropic/ ↩
  6. The Authors Guild, “Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement” (July 2026). https://authorsguild.org/news/court-grants-final-approval-anthropic-copyright-settlement/ On the per-title figure: Princeton University Press (September 2025). https://press.princeton.edu/news/bartz-v-anthropic ↩
  7. TechCrunch, “Federal judge sides with Meta in lawsuit over training AI models on copyrighted books” (June 2025). https://techcrunch.com/2025/06/25/federal-judge-sides-with-meta-in-lawsuit-over-training-ai-models-on-copyrighted-books/ ↩
  8. Proof News with Wired, “Apple, Nvidia, Anthropic Used Thousands of Swiped YouTube Videos to Train AI” (July 2024). https://www.proofnews.org/apple-nvidia-anthropic-used-thousands-of-swiped-youtube-videos-to-train-ai/ ↩
  9. Cade Metz et al., “How Tech Giants Cut Corners to Harvest Data for A.I.,” The New York Times (April 2024). https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html ↩
  10. 404 Media, “AI Music Generator Suno Admits It Was Trained on ‘Essentially All Music Files on the Internet’” (August 2024). https://www.404media.co/ai-music-generator-suno-admits-it-was-trained-on-essentially-all-music-files-on-the-internet/ ↩
  11. JUVE Patent, “Munich Regional Court stops Suno using GEMA-protected music” (2026). The judgment is not yet final. https://www.juve-patent.com/cases/munich-regional-court-stops-suno-using-gema-protected-music/ ↩
  12. CMS, “GEMA vs. OpenAI: Munich Regional Court I issues landmark copyright decision,” on case 42 O 14139/24 (11 November 2025). https://cms.law/en/deu/legal-updates/gema-vs.-openai-munich-regional-court-i-issues-landmark-copyright-decision ↩
  13. Music Business Worldwide, “Warner Music Group settles with Suno” (November 2025), which also reports Universal’s earlier settlement with Udio. https://www.musicbusinessworldwide.com/warner-music-group-settles-with-suno-strikes-first-of-its-kind-deal-with-ai-song-generator/ ↩
  14. Schuhmann et al., “LAION-5B: An open large-scale dataset for training next generation image-text models,” arXiv:2210.08402 (2022). https://arxiv.org/abs/2210.08402 ↩
  15. Mayer Brown, “Getty Images v Stability AI: What the High Court’s Decision Means” (November 2025). https://www.mayerbrown.com/en/insights/publications/2025/11/getty-images-v-stability-ai-what-the-high-courts-decision-means-for-rights-holders-and-ai-developers ↩
  16. The Register, on Reddit’s licensing agreement with Google (February 2024), reporting Reuters’ figure of about $60 million a year. https://www.theregister.com/2024/02/22/reddit_google_license_ipo_altman/ ↩
  17. Axios, “Historic NYT v. OpenAI copyright battle heats up” (September 2026). https://www.axios.com/2026/09/08/nyt-openai-microsoft-copyright-lawsuit ↩
  18. Meta, “Introducing Llama 3.1: Our most capable models to date” (July 2024). https://ai.meta.com/blog/meta-llama-3-1/ ↩
  19. Pantheon, “GB200 & GB300 NVL72 Power and Cooling Requirements,” summarising NVIDIA and OEM specifications. https://pantheon.run/learn/nvidia-gb200-nvl72-power-and-cooling The 2021 average rack figure is from AFCOM’s State of the Data Center report, as cited by CoreSite (2025). https://www.coresite.com/blog/how-colocation-data-centers-are-helping-solve-the-power-density-challenge ↩
  20. International Energy Agency, Energy and AI (2025), executive summary. https://www.iea.org/reports/energy-and-ai/executive-summary ↩
  21. Andrej Karpathy, micrograd. https://github.com/karpathy/micrograd His video walkthrough, “The spelled-out intro to neural networks and backpropagation,” uses the same neuron. https://www.youtube.com/watch?v=VMj-3S1tku0 ↩
  22. Cloudflare, “Content Independence Day: no AI crawl without compensation!” (July 2025). https://blog.cloudflare.com/content-independence-day-no-ai-crawl-without-compensation/ ↩
  23. Brynjolfsson, Chandar and Chen, “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” Stanford Digital Economy Lab (updated August 2026). https://digitaleconomy.stanford.edu/app/uploads/2026/08/Canaries_August2026.pdf ↩
  24. Yanis Varoufakis, Technofeudalism: What Killed Capitalism (Bodley Head, 2023). ↩
  25. Brynjolfsson, Li and Raymond, “Generative AI at Work,” The Quarterly Journal of Economics (2025). https://doi.org/10.1093/qje/qjae044 ↩
  26. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (July 2025). https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ ↩
  27. Google Cloud, “Measuring the environmental impact of AI inference” (August 2025). https://cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference ↩
  28. Plato, Phaedrus 275a, in Benjamin Jowett’s translation (1892). Public domain. https://www.gutenberg.org/ebooks/1636 ↩
  29. Dehaene et al., “How Learning to Read Changes the Cortical Networks for Vision and Language,” Science (2010). https://pubmed.ncbi.nlm.nih.gov/21071632/ ↩
  30. Itan et al., “The Origins of Lactase Persistence in Europe,” PLOS Computational Biology (2009), and ScienceDaily’s summary, “Milk Drinking Started Around 7,500 Years Ago In Central Europe.” https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1000491 and https://www.sciencedaily.com/releases/2009/08/090827202513.htm ↩
  31. UBS analyst estimate, as reported by Reuters via The Globe and Mail, “ChatGPT sets record for fastest-growing user base” (February 2023). https://www.theglobeandmail.com/business/article-chatgpt-sets-record-for-fastest-growing-user-base-analyst-note-says/ ↩
  32. Microsoft On the Issues, “The continued state of global AI diffusion in 2026” (September 2026): 18.8% of the world’s working-age population, 28.8% in the Global North and 16.2% in the Global South, with the gap widening. https://blogs.microsoft.com/on-the-issues/2026/09/21/the-continued-state-of-global-ai-diffusion-in-2026/ ↩
  33. Maguire et al., “Navigation-related structural change in the hippocampi of taxi drivers,” PNAS (2000). https://www.pnas.org/doi/10.1073/pnas.070039597 ↩
  34. Dahmani and Bohbot, “Habitual use of GPS negatively impacts spatial memory during self-guided navigation,” Scientific Reports (2020). https://www.nature.com/articles/s41598-020-62877-0 ↩
  35. Green and Bavelier, “Action video game modifies visual selective attention,” Nature (2003) https://www.nature.com/articles/nature01647; and the meta-analysis by Bediou et al., Psychological Bulletin (2018). https://doi.org/10.1037/bul0000130 ↩
  36. Ophir, Nass and Wagner, “Cognitive control in media multitaskers,” PNAS (2009). Later studies have found smaller effects. https://www.pnas.org/doi/10.1073/pnas.0903620106 ↩
  37. Kosmyna et al., “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task,” arXiv:2506.08872 (2025). A preprint with 54 participants. https://arxiv.org/abs/2506.08872 ↩
  38. Lee et al., “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers,” CHI 2025. https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/ ↩
  39. Michael Gerlich, “AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking,” Societies (2025). A correlational study. https://www.mdpi.com/2075-4698/15/1/6 ↩
  40. Mayer Brown, “EU AI Act News: Rules on General-Purpose AI Start Applying, Guidelines and Template for Summary of Training Data Finalized” (August 2025). https://www.mayerbrown.com/en/insights/publications/2025/08/eu-ai-act-news-rules-on-general-purpose-ai-start-applying-guidelines-and-template-for-summary-of-training-data-finalized ↩