Right now, there's a good chance an AI crawler is reading your blog. And this isn't its first visit — it's the second. According to data Cloudflare has released, more than half of all AI crawler traffic worldwide is spent re-visiting pages the crawler has already scraped once before. Your newsletter archive, your columns, your blog posts are almost certainly sitting on the servers of multiple AI companies by now — and you've never been paid a cent for any of it.
On July 1, Cloudflare announced a date that could change that math. The deadline is September 15, 2026.
Note: Cloudflare is a US-based global ICT company that speeds up websites and shields servers from cyberattacks. More than 20% of the world's websites reportedly use its service, which runs on a reverse-proxy setup that acts as a "shield" between users and the origin server.
On September 15, Crawlers Get Sorted Into Two Camps
Cloudflare's new policy requires AI companies to clearly separate the crawlers they use for search indexing from the crawlers they use for AI training and agents. Any company that hasn't made that separation by the deadline will have its crawlers blocked by default on ad-supported publisher sites. The policy applies to new Cloudflare sign-ups, new sites from existing customers, and everyone on the free plan. Given how much of the world's web traffic passes through Cloudflare's infrastructure, this policy reaches further than it might first appear.
Until now, most AI training crawlers have used identifiers identical or nearly identical to search-indexing bots. Publishers had no easy way to tell which bot was crawling for search visibility and which was harvesting data to train a large language model. Unless a site explicitly blocked a specific crawler identifier in robots.txt, the door was effectively open to almost everything. Cloudflare is using that opacity as the basis for its policy: let identifiable crawlers through, and block the rest.
Alongside this, Cloudflare has proposed a "Pay Per Use" revenue model. The AI search services Ceramic.ai and You.com are on board as early partners, and publishers get a share of revenue whenever their content is actually used in an AI search result or triggers premium access. In effect, Cloudflare is positioning itself as more than a CDN provider — a broker for content transactions. Cloudflare also disclosed that Google accesses roughly twice as much data as other AI companies, a figure clearly meant to let publishers judge for themselves whether today's playing field is actually a fair one.
Why Google Isn't Buying Into This Framework
Google is the one that has publicly pushed back on this policy. The company says its "Google Extended" bot already gives publishers a separate opt-out from AI training, and that opting out doesn't affect a site's search visibility. Its argument: it already runs the kind of separated system Cloudflare is demanding, so Cloudflare's unilateral deadline is unfairly lumping it in with everyone else.
There's a reason this pushback is hard to wave off as simple self-defense. For Cloudflare's policy to actually function, AI companies need to follow the classification standard Cloudflare has set. But if a major player like Google can claim it has its own separate standard and stay outside that framework, the policy's real effect on the biggest data collector is limited. A publisher-protection policy that fails to cover the largest participant isn't much of a protection at all.
Cloudflare's own business context is worth reading into this as well. If content-distribution deals get routed through Cloudflare's infrastructure, Cloudflare collects a new stream of brokerage fees. This is a policy where the stated cause of protecting publishers and the company's own revenue experiment are operating side by side. I don't think those two motives necessarily conflict. But it's worth keeping in mind that Cloudflare siding with publishers isn't rooted in pure goodwill alone.
What Independent Korean Publishers Should Check Before September 15
There's an observation that's been discussed for a long time in competitive-strategy theory: businesses that accurately recognize how valuable their assets are, and are willing to pay a deliberate cost to defend that value, tend to survive longer at the negotiating table. On the other side are the ones who try to save on the cost of holding a position and end up losing the position itself. The fact that your content keeps getting scraped repeatedly by AI companies is itself a signal that it's valuable. Whether you receive that signal passively or act on it is a separate choice.
Plenty of Korean newsletter publishers, bloggers, and content directors either don't use Cloudflare, or use it without ever having checked its AI-crawler settings.
If your site runs on Cloudflare, look for the "AI Scrapers and Crawlers" section under the Security menu in your dashboard. It's been available on the free plan since the second half of 2024, and a single toggle lets you block the major AI training crawlers all at once. Unseparated crawlers will be blocked by default after September 15 regardless, but checking the setting yourself beforehand is the more proactive move.
Even without Cloudflare, you can use robots.txt. The identifiers of major crawlers — GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), and others — are publicly known, and adding them to a block list is something you can do without any technical background. This isn't a complete defense. There's no technical way to stop a crawler that ignores robots.txt, and any newly launched crawler won't be caught by your block list. But there's a meaningful difference between having done nothing and having explicitly stated your position — one that matters if you ever end up at a negotiating table over this later on.
It's unlikely the Pay Per Use model will become an immediately realistic revenue stream for small Korean publishers. Ceramic.ai and You.com are services built around the English-speaking market, and it will take time before Cloudflare's revenue-sharing network extends to cover Korean-language content. But without at least attempting to put a price on your own content, it's hard to stop the market from fixing that price at zero.
If you write the kind of content AI keeps wanting to scrape again and again, it's time to ask yourself a question at least once: Is it okay if my writing gets used to train AI? And if it is, under what conditions? If you don't settle on an answer in advance, someone else will eventually settle it for you.
September 15 is a deadline for a technical policy. But the question that date raises reaches far beyond any settings menu. What is your content worth right now?



