I'd like to see more of a legal response.<p>First, find out who's on the other end of a few hundred IP addresses. Start with ones in the US. Sue for damages. Use discovery to find out what's on the other end. Sue the maker of that device. If it turns out to be an appliance or smart TV, it may be possible to consolidate cases into one case against the manufacturer. Criminal negligence, tort interference with contract, harassment, Computer Fraud and Abuse act violation...
Maybe a restraining order prohibiting the sale of "smart TV" known to be able to host attacks. Have imports seized by Customs and Border Protection. That would get a manufacturer's attention.<p>The manufacturer's EULA will not help the manufacturer, because the plaintiff, the party being attacked, is not a party to the EULA at all.
There’s an assumption that turning on Cloudflare’s “under attack” mode would mitigate the attack.<p>Given how adaptive the rest of the attack was, I would be very curious to find out how it would approach that obstacle.
That has only partially mitigated much smaller attacks (residential proxy scraping etc) on my employer's site.<p>We're currently on the "Business" plan, but I'm coming to the conclusion that we need to upgrade to the "Enterprise Advantage" plan for the JA3/4 fingerprinting and detection ID features.<p>I get put off by "Contact Sales" pricing.
My assumption would be that it would drastically reduce the attack to be borderline irrelevant. I've never turned on Under Attack so somebody else may have more insight and the docs[1] don't describe precisely what happens besides a JS interstitial.<p>I know that JS challenges, both interactive and non-interactive, can be solved by bots. I've seen it. However, I suspect that the challenges just get harder and harder until the attack levels drop.<p>[1] <a href="https://developers.cloudflare.com/fundamentals/reference/under-attack-mode/" rel="nofollow">https://developers.cloudflare.com/fundamentals/reference/und...</a>
It changes the economics of the attack because it requires the attacker to do compute before they can make requests. Depending on how good the fingerprinting is, it can also get very expensive (eg. requiring you to run a full browser, rather than merely computing a few sha256 hashes)
My naive take on a Cloudflare perspective wants to combine "three times is enemy action" with toddler-speed block dropping and manual clearance. What's the money reason this problem isn't handled at the ISP level?
I don't quite understand your post, but is your question why don't the ISPs of the sources of the abusive traffic sort it out?<p>The distributed nature of DDoS means each participating host isn't sending <i>that</i> much traffic, and there are often tens or hundreds of thousands of participating hosts. An ISP should verify claims of abuse before cutting off customers, and since most of the customers are presumably unaware of what their systems are doing, there will be a lot of unhappy customers and then you've got to spend a lot of customer support time on helping them clean up their systems so they can get back online.<p>I spent a fair amount of time sending out abuse reports for phishing / malware senders about a decade ago, and most abuse reporting addresses are a black hole. Even if you do get to someone who will do something about abuse, they won't do it quickly.<p>There's be a few high profile longer term DDoS attacks lately, but when I was running infra that got a lot of stuff, it was mostly people kicking the tires on DDoS as a service offerings and most attacks were 90 seconds long ... there's no way I'm convincing an ISP to drop a pwned customer over that.<p>Starting from there, this DDoS sounds like layer 7 DDoS which is easy to track to the immediate senders, but a ton of DDoS is volumetric stuff, often volumetric reflection attacks where the senders spoof your address. If you're getting that, best you can do is get the reflectors kicked off (or cleaned up) ... tracing back to the sending hosts means getting a reflector (and their ISPs) engaged to do a lot of labor intensive work.<p>All of that investigation stuff takes qualified people lots of time, that's your money reason it doesn't happen.
A more interesting question is, what exactly do the attackers gain from hitting read the docs? Most of their docs hosting is static/easily CDN cached. Unlike database bound sites, you would need a lot more traffic to overload pure/mostly static hosting. Maybe it's a malicious AI lab looking to deny their competitors training data? As far as infosec profiling goes, this is probably the oddest case I have heard of.<p>I am thinking it's probably an AI lab that misconfigured their data scraper (made it too agentic) and it ended up looking like a DDoS.<p>The new generation of scrapers are all agentic and self healing. (As an example see YC's <a href="https://parse.bot">https://parse.bot</a>)
Author here. This was not a misconfigured data scraper. We see those every week[1]. This attack wasn't scraping useful content. It was almost entirely 404s and 302s and pulled virtually zero real docs. It specifically looked for URLs not served by the CDN and when it found a pattern, did millions of variations of it. Whether built by an AI or not, it was designed to cause outages and financial damage from autoscaling. However, as others have suggested, we may have been a test run for a real target.<p>[1] <a href="https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse/" rel="nofollow">https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse...</a>
> Most of their docs hosting is static/easily CDN cached<p>The article says<p>> and it purposefully attacked areas that bypassed caching<p>So that doesn't work. Also, it seems that they were trying to cause financial harm, not to take down the infrastructure but to make it costly for the org itself. That's smart.
Interesting that the Under Attack Mode wasn’t used at all here. I understand not wanting to break APIs but I feel temporarily challenging non-API usage could have at least helped without impacting users too much?
I'm curious if anybody could speculate who would be attacking a documentation silo, and to what end?
I'm the author of the blog. I don't know. Internally, we were half joking that we were going to get ransom notice, but we never did.<p>The only thing that sort of correlates with this attack is that before it started, we began rolling out some slightly more aggressive rate limits one by one. This was mostly because anytime any new "company" thinks they're going to catchup with Claude/OpenAI, they scrape us very aggressively (and they're not respectful about it). My guess is that the attackers behind this attack were already probing us (they were) and they thought the window of opportunity might be closing.
Good to know. I use your site (with a manual transmission user-agent) often, and it's fantastic. Thanks for your work and the writeup!
Just curious, if you're tolerant of scraping, do you make an archive of all your content available so that scraping is unnecessary, and if so do the scrapers prefer that?
Current evidence is that scrapers mostly aren't nearly considerate or sophisticated enough to take an "archive of all content" option if one exists.<p>See <a href="https://people.kernel.org/monsieuricon/creepy-crawlies" rel="nofollow">https://people.kernel.org/monsieuricon/creepy-crawlies</a> which describes how the <a href="https://git.kernel.org" rel="nofollow">https://git.kernel.org</a> gets hammered by crawlers all the time even though you could run a single `git clone` and get the data that way instead.
This is exactly the problem, unfortunately.<p>For somebody who knows a bit how things are set up, or is willing to spend 10 minutes researching, it's a no-brainer that you can just "git clone" entire linux kernel development history, or download entire wikipedia [0].<p>Alas, large number of scrapers are not willing to spend those 10 minutes, it would appear. So, here we are.<p>[0] <a href="https://dumps.wikimedia.org/" rel="nofollow">https://dumps.wikimedia.org/</a>
Is there a standard for exposing such sitedata dumps? If not, it's not really surprising that they don't.
It's terabytes of content and other than we're the host not really related to each other. However, for most projects, it's possible to download a zip file of all the HTML docs for that project. We have a lower rate limit to pull these, but a scraper can pull thousands of docs at once. We only host a few hundred thousand projects so pulling a zip of the latest docs for all of them could be done in a day or two at a very reasonable rate.<p>It's also possible to request the docs already processed into markdown[1]. Lastly, basically all of the docs come from Git. A smart scraper could just clone a project's repo.<p>[1] <a href="https://docs.readthedocs.com/platform/stable/reference/markdown-for-agents.html" rel="nofollow">https://docs.readthedocs.com/platform/stable/reference/markd...</a>
Could be testing in preparation for attacking something more critical?
I run a similar service, and we get almost daily attacks like this. Sometimes it's a specific high-profile customer, other times it's broader.<p>I can't speak for RTD, but I think it's less "documentation site" and more just that we sit on the domains of high-profile products and the tools are just looking for any hole they can find?<p>Often it's even the company themselves, for whatever reason (security research, etc).
Either testing for something bigger OR demonstrating their power to a 3rd party with minimal real disruption
Could have been a live-fire exercise by a nation state.<p>Edit: why the down vote? That is literally in the realm of possibility!
> One decision we made is to always give real users an escape hatch. Read the Docs very rarely issues outright blocks or bans to specific IPs or user agents. Instead, our "worst" is a JavaScript challenge, and if a user solves a challenge, they are very unlikely to get challenged again for the next day or so.<p>Finally, a competent response that doesn't leave the users hang out to dry.<p>I'm so tired of seeing incompetents with measures like "blackhole 2 continents" deployed even outside active attacks.
Yippee, free load test!
Sure it wasn't a AI crawler?