{"id":75802,"date":"2024-02-14T14:00:00","date_gmt":"2024-02-14T14:00:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2024\/02\/14\/the-rise-and-fall-of-robots-txt\/"},"modified":"2024-02-14T14:00:00","modified_gmt":"2024-02-14T14:00:00","slug":"the-rise-and-fall-of-robots-txt","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2024\/02\/14\/the-rise-and-fall-of-robots-txt\/","title":{"rendered":"The rise and fall of robots.txt"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"content\">\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup mb-20 font-fkroman text-22 leading-150 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-franklin [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white first-letter:float-left first-letter:mr-18 first-letter:font-polysans-mono first-letter:text-[117px] first-letter:font-medium first-letter:leading-[.72] dark:first-letter:text-franklin\">For three decades, a tiny text file has kept the internet from chaos. This text file has no particular legal or technical authority, and it\u2019s not even particularly complicated. It represents a handshake deal between some of the earliest pioneers of the internet to respect each other\u2019s wishes and build the internet in a way that benefitted everybody. It\u2019s a mini constitution for the internet, written in code.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">It\u2019s called robots.txt and is usually located at yourwebsite.com\/robots.txt. That file allows anyone who runs a website \u2014 big or small, cooking blog or multinational corporation \u2014\u00a0to tell the web who\u2019s allowed in and who isn\u2019t. Which search engines can index your site? What archival projects can grab a version of your page and save it? Can competitors keep tabs on your pages for their own files? You get to decide and declare that to the web.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">It\u2019s not a perfect system, but it works. Used to, anyway. For decades, the main focus of robots.txt was on search engines; you\u2019d let them scrape your site and in exchange they\u2019d promise to send people back to you. Now AI has changed the equation: companies around the web are using your site and its data to build massive sets of training data, in order to build models and products that may not acknowledge your existence at all.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">The robots.txt file governs a give and take; AI feels to many like all take and no give. But there\u2019s now so much money in AI, and the technological state of the art is changing so fast that many site owners can\u2019t keep up. And the fundamental agreement behind robots.txt, and the web as a whole \u2014 which for so long amounted to \u201ceverybody just be cool\u201d \u2014\u00a0may not be able to keep up either.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white min-h-[80px] first-letter:float-left first-letter:mr-18 first-letter:font-polysans-mono first-letter:text-100 first-letter:font-medium first-letter:leading-[.72]  first-letter:selection:bg-franklin-20 dark:first-letter:text-franklin\">In the early days of the internet, robots went by many names: spiders, crawlers, worms, WebAnts, web crawlers. Most of the time, they were built with good intentions. Usually it was a developer trying to build a directory of cool new websites, make sure their own site was working properly, or build a research database \u2014 this was 1993 or so, long before search engines were everywhere and in the days when you could fit most of the internet on your computer\u2019s hard drive.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">The only real problem then was the traffic: accessing the internet was slow and expensive both for the person seeing a website and the one hosting it. If you hosted your website on your computer, as many people did, or on hastily constructed server software run through your home internet connection, all it took was a few robots overzealously downloading your pages for things to break and the phone bill to spike.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">Over the course of a few months in 1994, a software engineer and developer named Martijn Koster, along with a group of other web administrators and developers, came up with a solution they called the Robots Exclusion Protocol. The proposal was straightforward enough: it asked web developers to add a plain-text file to their domain specifying which robots were not allowed to scour their site, or listing pages that are off limits to all robots. (Again, this was a time when you could maintain a list of every single robot in existence \u2014 Koster and a few others helpfully did just that.) For robot makers, the deal was even simpler: respect the wishes of the text file.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">From the beginning, Koster made clear that he didn\u2019t hate robots, nor did he intend to get rid of them. \u201cRobots are one of the few aspects of the web that cause operational problems and cause people grief,\u201d he said in an initial email to a mailing list called WWW-Talk (which included early-internet pioneers like Tim Berners-Lee and Marc Andreessen) in early 1994. \u201cAt the same time they do provide useful services.\u201d Koster cautioned against arguing about whether robots are good or bad \u2014 because it doesn\u2019t matter, they\u2019re here and not going away. He was simply trying to design a system that might \u201cminimise the problems and may well maximize the benefits.\u201d\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component clear-both block md:float-left md:mr-30 md:w-[320px] lg:-ml-100\">\n<div class=\"duet--article--article-pullquote mb-20\">\n<p class=\"duet--article--dangerously-set-cms-markup relative bg-repeating-lines-dark bg-[length:1px_1.2em] pb-8 font-polysans text-28 font-medium leading-120 tracking-1 selection:bg-franklin-20  dark:bg-repeating-lines-light dark:text-white dark:selection:bg-blurple\">\u201cRobots are one of the few aspects of the web that cause operational problems and cause people grief. At the same time, they do provide useful services.\u201d<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">By the summer of that year, his proposal had become a standard \u2014\u00a0not an official one, but more or less a universally accepted one. Koster pinged the WWW-Talk group again in June with an update. \u201cIn short it is a method of guiding robots away from certain areas in a Web server\u2019s URL space, by providing a simple text file on the server,\u201d he wrote. \u201cThis is especially handy if you have large archives, CGI scripts with massive URL subtrees, temporary information, or you simply don\u2019t want to serve robots.\u201d He\u2019d set up a topic-specific mailing list, where its members had agreed on some basic syntax and structure for those text files, changed the file\u2019s name from RobotsNotWanted.txt to a simple robots.txt, and pretty much all agreed to support it.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">And for most of the next 30 years, that worked pretty well.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">But the internet doesn\u2019t fit on a hard drive anymore, and the robots are vastly more powerful. Google uses them to crawl and index the entire web for its search engine, which has become the interface to the web and brings the company billions of dollars a year. Bing\u2019s crawlers do the same, and Microsoft licenses its database to other search engines and companies. The Internet Archive uses a crawler to store webpages for posterity. Amazon\u2019s crawlers traipse the web looking for product information, and according to a recent antitrust suit, the company uses that information to punish sellers who offer better deals away from Amazon. AI companies like OpenAI are crawling the web in order to train large language models that could once again fundamentally change the way we access and share information.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">The ability to download, store, organize, and query the modern internet\u00a0gives any company or developer something like the world\u2019s accumulated knowledge to work with. In the last year or so, the rise of AI products like ChatGPT, and the large language models underlying them, have made high-quality training data one of the internet\u2019s most valuable commodities. That has caused internet providers of all sorts to reconsider the value of the data on their servers, and rethink who gets access to what. Being too permissive can bleed your website of all its value; being too restrictive can make you invisible. And you have to keep making that choice with new companies, new partners, and new stakes all the time.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white min-h-[80px] first-letter:float-left first-letter:mr-18 first-letter:font-polysans-mono first-letter:text-100 first-letter:font-medium first-letter:leading-[.72]  first-letter:selection:bg-franklin-20 dark:first-letter:text-franklin\">There are a few breeds of internet robot. You might build a totally innocent one to crawl around and make sure all your on-page links still lead to other live pages; you might send a much sketchier one around the web harvesting every email address or phone number you can find. But the most common one, and the most currently controversial, is a simple web crawler. Its job is to find, and download, as much of the internet as it possibly can.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">Web crawlers are generally fairly simple. They start on a well-known website, like cnn.com or wikipedia.org or health.gov. (If you\u2019re running a general search engine, you\u2019ll start with lots of high-quality domains across various subjects; if all you care about is sports or cars, you\u2019ll just start with car sites.) The crawler downloads that first page and stores it somewhere, then automatically clicks on every link on that page, downloads all those, clicks all the links on every one, and spreads around the web that way. With enough time and enough computing resources, a crawler will eventually find and download billions of webpages.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component clear-both block md:float-left md:mr-30 md:w-[320px] lg:-ml-100\">\n<div class=\"duet--article--article-pullquote mb-20\">\n<p class=\"duet--article--dangerously-set-cms-markup relative bg-repeating-lines-dark bg-[length:1px_1.2em] pb-8 font-polysans text-28 font-medium leading-120 tracking-1 selection:bg-franklin-20  dark:bg-repeating-lines-light dark:text-white dark:selection:bg-blurple\">The tradeoff is fairly straightforward: if Google can crawl your page, it can index it and show it in search results.<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">Google estimated in 2019 that more than 500 million websites had a robots.txt page dictating whether and what these crawlers are allowed to access. The structure of those pages is usually roughly the same: it names a \u201cUser-agent,\u201d which refers to the name a crawler uses when it identifies itself to a server. Google\u2019s agent is Googlebot; Amazon\u2019s is Amazonbot; Bing\u2019s is Bingbot; OpenAI\u2019s is GPTBot. Pinterest, LinkedIn, Twitter, and many other sites and services have bots of their own, not all of which get mentioned on every page. (<a href=\"https:\/\/en.wikipedia.org\/robots.txt\" target=\"_blank\" rel=\"noopener\">Wikipedia<\/a> and <a href=\"https:\/\/facebook.com\/robots.txt\" target=\"_blank\" rel=\"noopener\">Facebook<\/a> are two platforms with particularly thorough robot accounting.) Underneath, the robots.txt page lists sections or pages of the site that a given agent is not allowed to access, along with specific exceptions that are allowed. If the line just reads \u201cDisallow: \/\u201d the crawler is not welcome at all.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">It\u2019s been a while since \u201coverloaded servers\u201d were a real concern for most people. \u201cNowadays, it\u2019s usually less about the resources that are used on the website and more about personal preferences,\u201d says John Mueller, a search advocate at Google. \u201cWhat do you want to have crawled and indexed and whatnot?\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">The biggest question most website owners historically had to answer was whether to allow Googlebot to crawl their site. The tradeoff is fairly straightforward: if Google can crawl your page, it can index it and show it in search results. Any page you want to be Googleable, Googlebot needs to see. (How and where Google actually displays that page in search results is of course a completely different story.) The question is whether you\u2019re willing to let Google eat some of your bandwidth and download a copy of your site in exchange for the visibility that comes with search.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">For most websites, this was an easy trade. \u201cGoogle is our most important spider,\u201d says Medium CEO Tony Stubblebine. Google gets to download all of Medium\u2019s pages, \u201cand in exchange we get a significant amount of traffic. It\u2019s win-win. Everyone thinks that.\u201d This is the bargain Google made with the internet as a whole, to funnel traffic to other websites while selling ads against the search results. And Google has, by all accounts, been a good citizen of robots.txt. \u201cPretty much all of the well-known search engines comply with it,\u201d Google\u2019s Mueller says. \u201cThey\u2019re happy to be able to crawl the web, but they don\u2019t want to annoy people with it\u2026 it just makes life easier for everyone.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white min-h-[80px] first-letter:float-left first-letter:mr-18 first-letter:font-polysans-mono first-letter:text-100 first-letter:font-medium first-letter:leading-[.72]  first-letter:selection:bg-franklin-20 dark:first-letter:text-franklin\">In the last year or so, though, the rise of AI has upended that equation. For many publishers and platforms, having their data crawled for training data felt less like trading and more like stealing. \u201cWhat we found pretty quickly with the AI companies,\u201d Stubblebine says, \u201cis not only was it not an exchange of value, we\u2019re getting nothing in return. Literally zero.\u201d When Stubblebine announced last fall that Medium <a href=\"https:\/\/blog.medium.com\/default-no-to-ai-training-on-your-stories-abb5b4589c8\" target=\"_blank\" rel=\"noopener\">would be blocking AI crawlers<\/a>, he wrote that \u201cAI companies have leached value from writers in order to spam Internet readers.\u201d\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">Over the last year, a large chunk of the media industry has echoed Stubblebine\u2019s sentiment. \u201cWe do not believe the current \u2018scraping\u2019 of BBC data without our permission in order to train Gen AI models is in the public interest,\u201d BBC director of nations Rhodri Talfan Davies <a href=\"https:\/\/www.bbc.co.uk\/mediacentre\/articles\/2023\/generative-ai-at-the-bbc\/\" target=\"_blank\" rel=\"noopener\">wrote last fall<\/a>, announcing that the BBC would also be blocking OpenAI\u2019s crawler. <em>The New York Times<\/em> blocked GPTBot as well, months before launching a suit against OpenAI alleging that OpenAI\u2019s models \u201cwere built by copying and using millions of <em>The Times<\/em>\u2019s copyrighted news articles, in-depth investigations, opinion pieces, reviews, how-to guides, and more.\u201d <a href=\"https:\/\/palewi.re\/docs\/news-homepages\/openai-gptbot-robotstxt.html\" target=\"_blank\" rel=\"noopener\">A study by Ben Welsh<\/a>, the news applications editor at <em>Reuters<\/em>, found that 606 of 1,156 surveyed publishers had blocked GPTBot in their robots.txt file.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">It\u2019s not just publishers, either. Amazon, Facebook, Pinterest, WikiHow, WebMD, and many other platforms explicitly block GPTBot from accessing some or all of their websites. On most of these robots.txt pages, OpenAI\u2019s GPTBot is the only crawler explicitly and completely disallowed. But there are plenty of other AI-specific bots beginning to crawl the web, like Anthropic\u2019s anthropic-ai and Google\u2019s new Google-Extended. According to a study from last fall by Originality.AI, 306 of the top 1,000 sites on the web blocked GPTBot, but only 85 blocked Google-Extended and 28 blocked anthropic-ai.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">There are also crawlers used for both web search and AI. CCBot, which is run by the organization Common Crawl, scours the web for search engine purposes, but its data is also used by OpenAI, Google, and others to train their models. Microsoft\u2019s Bingbot is both a search crawler and an AI crawler. And those are just the crawlers that identify themselves \u2014\u00a0many others attempt to operate in relative secrecy, making it hard to stop or even find them in a sea of other web traffic. For any sufficiently popular website, finding a sneaky crawler is needle-in-haystack stuff.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">In large part, GPTBot has become the main villain of robots.txt because OpenAI allowed it to happen. The company published and promoted a page about how to block GPTBot and built its crawler to loudly identify itself every time it approaches a website. Of course, it did all of this <em>after <\/em>training the underlying models that have made it so powerful, and only once it became an important part of the tech ecosystem. But OpenAI\u2019s chief strategy officer Jason Kwon says that\u2019s sort of the point. \u201cWe are a player in an ecosystem,\u201d he says. \u201cIf you want to participate in this ecosystem in a way that is open, then this is the reciprocal trade that everybody\u2019s interested in.\u201d Without this trade, he says, the web begins to retract, to close \u2014\u00a0and that\u2019s bad for OpenAI and everyone. \u201cWe do all this so the web can stay open.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">By default, the Robots Exclusion Protocol has always been permissive. It believes, as Koster did 30 years ago, that most robots are good and are made by good people, and thus allows them by default. That was, by and large, the right call. \u201cI think the internet is fundamentally a social creature,\u201d OpenAI\u2019s Kwon says, \u201cand this handshake that has persisted over many decades seems to have worked.\u201d OpenAI\u2019s role in keeping that agreement, he says, includes keeping ChatGPT free to most users \u2014 thus delivering that value back \u2014\u00a0and respecting the rules of the robots.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component clear-both block md:float-left md:mr-30 md:w-[320px] lg:-ml-100\">\n<div class=\"duet--article--article-pullquote mb-20\">\n<p class=\"duet--article--dangerously-set-cms-markup relative bg-repeating-lines-dark bg-[length:1px_1.2em] pb-8 font-polysans text-28 font-medium leading-120 tracking-1 selection:bg-franklin-20  dark:bg-repeating-lines-light dark:text-white dark:selection:bg-blurple\">But robots.txt is not a legal document \u2014 and 30 years after its creation, it still relies on the good will of all parties involved.<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">But robots.txt is not a legal document \u2014 and 30 years after its creation, it still relies on the good will of all parties involved. Disallowing a bot on your robots.txt page is like putting up a \u201cNo Girls Allowed\u201d sign on your treehouse \u2014 it sends a message, but it\u2019s not going to stand up in court. Any crawler that wants to ignore robots.txt can simply do so, with little fear of repercussions. (There is some legal precedent around web scraping in general, though even that can be complicated and mostly lands on crawling and scraping being allowed.) The Internet Archive, for example, simply announced in 2017 that it was no longer abiding by the rules of robots.txt. \u201cOver time we have observed that the robots.txt files that are geared toward search engine crawlers do not necessarily serve our archival purposes,\u201d Mark Graham, the director of the Internet Archive\u2019s Wayback Machine, <a href=\"https:\/\/blog.archive.org\/2017\/04\/17\/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives\/\" target=\"_blank\" rel=\"noopener\">wrote at the time<\/a>. And that was that.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">As the AI companies continue to multiply, and their crawlers grow more unscrupulous, anyone wanting to sit out or wait out the AI takeover has to take on an endless game of whac-a-mole. They have to stop each robot and crawler individually, if that\u2019s even possible, while also reckoning with the side effects. If AI is in fact the future of search, as Google and others have predicted, blocking AI crawlers could be a short-term win but a long-term disaster.\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">There are people on both sides who believe we need better, stronger, more rigid tools for managing crawlers. They argue that there\u2019s too much money at stake, and too many new and unregulated use cases, to rely on everyone just agreeing to do the right thing. \u201cThough many actors have some rules self-governing their use of crawlers,\u201d two tech-focused attorneys wrote in <a href=\"https:\/\/digitalcommons.law.uw.edu\/wjlta\/vol13\/iss3\/4\/\" target=\"_blank\" rel=\"noopener\">a 2019 paper<\/a> on the legality of web crawlers, \u201cthe rules as a whole are too weak, and holding them accountable is too difficult.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white\">Some publishers would like more detailed controls over both what is crawled and what it\u2019s used for, instead of robots.txt\u2019s blanket yes-or-no permissions. Google, which a few years ago made an effort to make the Robots Exclusion Protocol an official formalized standard, has also pushed to deemphasize robots.txt on the grounds that it\u2019s an old standard and too many sites don\u2019t pay attention to it. \u201cWe recognize that existing web publisher controls were developed before new AI and research use cases,\u201d Google\u2019s VP of trust Danielle Romain <a href=\"https:\/\/blog.google\/technology\/ai\/ai-web-publisher-controls-sign-up\/\" target=\"_blank\" rel=\"noopener\">wrote last year<\/a>. \u201cWe believe it\u2019s time for the web and AI communities to explore additional machine-readable means for web publisher choice and control for emerging AI and research use cases.\u201d\u00a0<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph mb-20 font-fkroman text-18 leading-160 -tracking-1 selection:bg-franklin-20 dark:text-white dark:selection:bg-blurple [&amp;_a:hover]:shadow-highlight-franklin dark:[&amp;_a:hover]:shadow-highlight-blurple [&amp;_a]:shadow-underline-black dark:[&amp;_a]:shadow-underline-white after:absolute after:ml-8 after:mt-2 after:content-[url(\/icons\/endmark.svg)]\">Even as AI companies face regulatory and legal questions over how they build and train their models, those models continue to improve and new companies seem to start every day. Websites large and small are faced with a decision: submit to the AI revolution or stand their ground against it. For those that choose to opt out, their most powerful weapon is an agreement made three decades ago by some of the web\u2019s earliest and most optimistic true believers. They believed that the internet was a good place, filled with good people, who above all wanted the internet to be a good thing. In that world, and on that internet, explaining your wishes in a text file was governance enough. Now, as AI stands to reshape the culture and economy of the internet all over again, a humble plain-text file is starting to look a little old-fashioned.<\/p>\n<\/div>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.theverge.com\/24067997\/robots-txt-ai-text-file-web-crawlers-spiders\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>For three decades, a tiny text file has kept the internet from chaos. This text file has no particular legal or technical authority, and it\u2019s not even particularly complicated. It represents a handshake deal between some of the earliest pioneers of the internet to respect each other\u2019s wishes and build the internet in a way [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":75803,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-75802","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/75802","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=75802"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/75802\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/75803"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=75802"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=75802"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=75802"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}