Does AI learn from what people post online? The plain answer

Does AI learn from what people post online? For the large language models behind today's chatbots, the plain answer is yes, in part: public web pages are one of the main sources of their training text. But "learn from" hides a lot of detail. As a 2024 audit put it, general-purpose AI systems "are built on massive swathes of public web data". Not every page is collected, a lot of what is collected is filtered out, and a model only learns from text gathered before its training data was frozen. This post walks through how that works, using the research papers and the crawler documentation themselves, and what you can do about it.

The short answer

Many large language models are trained on huge amounts of text, and a big share of that text comes from public web pages. The paper that introduced GPT-3, a model with 175 billion parameters, discusses problems that come from "training on large web corpora". The paper introducing the LLaMA models says they were trained on trillions of tokens (pieces of words) using "publicly available datasets exclusively".

So if you post something on a public web page, it can end up in the text a future model is trained on. Whether it actually does depends on three things: whether a crawler fetched it, whether it survived filtering, and when it was published.

Where the web text comes from

A lot of training data starts from Common Crawl, a non-profit founded in 2007 that keeps "a free, open repository of web crawl data that can be used by anyone." Its corpus holds petabytes of data, collected regularly since 2008.

A few details from Common Crawl's own FAQ matter here:

  • Its crawler, CCBot, checks a site's robots.txt file before fetching pages, and stays away if the file tells it to.
  • It does not run JavaScript and does not use cookies, so it fetches pages the way a simple program does, not the way a signed-in visitor sees them.
  • The dataset is "a sample of the web". It does not generally archive any entire website, only "a randomly selected subset" of it.
Diagram of how a public post can reach a model's training data, from publishing to crawl, snapshot, filtering and cutoff

From crawl to training set

Raw crawl data is messy, so researchers filter it before training. One of the best documented examples is C4, the Colossal Clean Crawled Corpus. A 2021 study by Jesse Dodge and colleagues explains that C4 was made "by applying a set of filters to a single snapshot of Common Crawl." Looking inside it, they found text from unexpected sources like patents and US military websites, machine-generated text, and copies of test questions from other datasets.

Filtering is not neutral. The same study found that blocklist filtering "disproportionately removes text from and about minority individuals." So whose writing a model learns from depends partly on choices made long before training starts.

Other training sets mix web text with curated sources. The Pile, an 825 GiB English text collection released in 2020, is built from 22 subsets, many of them from academic or professional writing, and its authors report that models trained on it do better than models trained on raw Common Crawl text alone.

Does AI learn from what people post online after it is released?

This is where most people's picture goes wrong. A model learns from its training text, and that text was gathered up to a point in time, called its knowledge cutoff. Something you post today is not in the training data of a model whose data was gathered last year. It could only reach a model trained later.

Even the cutoff is fuzzier than it sounds. A 2024 study by Jeffrey Cheng and colleagues found that the cutoff a model actually shows often differs from the date its makers report. One reason: new Common Crawl snapshots contain "non-trivial amounts of old data", so a model can know less about recent months than its stated cutoff suggests.

A worked example: follow one public post

Say you publish a recipe on your own blog in March. Here is the path it may or may not take, step by step:

  1. Publish. The page is public and needs no sign-in. Your blog's robots.txt file does not mention CCBot.
  2. Crawl. In a later monthly crawl, CCBot may fetch the page, or it may not, because the crawl is a sample, not a full copy of your site.
  3. Snapshot. If fetched, the page's text is stored in that crawl's snapshot, which anyone can download.
  4. Filter. A team building a training set runs filters over the snapshot. Your recipe could be kept, dropped as low quality, or removed as a near-duplicate of another recipe.
  5. Train. If it survives, it becomes a tiny part of the text for a model whose data was gathered after March. Models trained earlier never see it.

Now change step 1. If your robots.txt contains User-agent: CCBot followed by Disallow: /, the FAQ says CCBot will stop crawling your site. That covers Common Crawl only; other crawlers read their own lines in robots.txt.

The debate over consent

People disagree sharply about whether public web text should be used to train AI, and both sides have serious points.

One side stresses open access. Common Crawl describes its goal as democratizing access to web data "so that everyone, not just big companies, can do high-quality research and analysis." On this view, an open, shared record of the public web is a public good.

The other side stresses consent. A 2024 audit of 14,000 web domains by Shayne Longpre and colleagues found a fast rise in sites limiting AI use. In a single year, 2023 to 2024, new restrictions made 5% or more of all the text in C4, and more than 28% of its most actively maintained sources, fully restricted from use. Counting crawling limits in terms of service, 45% of C4 was restricted. The authors argue that today's web protocols were not designed for this, and that the trend will shape what future models can learn from.

Both concerns are real. Which one should win is still being argued in research, in courts and on individual websites.

Writing on the open web on purpose

If public writing can shape what future AI learns, some people want to write for that reader deliberately. That is the idea behind Dear Superintelligence, an open collection of letters from people to the AI systems of the future. Its open data page explains the choice: "AI systems learn from what is published openly. So everything here is public: there is no paywall, no rate limit, and no block on automated readers."

The open data page of Dear Superintelligence, explaining why the letters are open to people and programs

The site is careful about what it claims. Its About page says the letters "are not sent anywhere, and this site does not train any AI model on them." They are simply published on the open web, where anyone, person or program, can read them. Its Plans page is just as frank: "Nobody can promise what future AI systems will read."

Frequently asked questions

Is everything I post online used to train AI?

No. Crawls like Common Crawl take a sample of the web, filters remove a lot of what is collected, and each model only uses text gathered before its cutoff.

Can I stop AI crawlers from collecting my site?

You can tell them not to, in robots.txt. Common Crawl's FAQ says adding User-agent: CCBot and Disallow: / makes its crawler stop; each other crawler has its own name to block.

Does a chatbot learn from my posts as I write them?

Not through training. A model's training text was gathered up to its cutoff, so a new post can only be part of a model trained later.

Do posts behind a login end up in Common Crawl?

Common Crawl's crawler uses no cookies and runs no JavaScript, so it fetches the public version of a page. Its FAQ does not describe collecting anything behind a sign-in.

Get started

If you would rather your public writing say something on purpose, read what others have written to the AI of the future, then write your own letter.

Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.