AI crawler blocked by ModSecurity at the gates of a business website, illustrating an AI Observatory case study

AI crawler blocked by ModSecurity: featured image for an AI Observatory case study showing ClaudeBot stopped at a business website origin

The Problem

AI Crawler Blocked by ModSecurity — While the Website Looked Completely Healthy

A client website was operating normally. Human visitors could browse it, the site was online, Cloudflare was serving traffic correctly, and there was no obvious indication that anything was wrong.

But the AI Observatory was showing something very different.

One recognised AI crawler — ClaudeBot — was reaching the website but was succeeding on only a small proportion of its business retrieval attempts.

In the Observatory evidence window shown below, ClaudeBot had made 14 autonomous accesses, but only 1 of 7 business retrieval attempts succeeded. Its measured business success rate was therefore just 14.3%.

AI Observatory showing ClaudeBot with a low business retrieval success rate on a client website
AI Observatory evidence showing normal activity from several recognised machine systems, while ClaudeBot recorded only 1 successful business retrieval from 7 attempts.

This is the kind of failure that is easy to miss. The website itself was not down. There was no obvious WordPress error, no broken page, and no visible problem for ordinary users.

The issue existed specifically in the machine retrieval path.

A website can appear perfectly healthy to people while an AI crawler is being rejected somewhere deeper in the delivery stack.
The First Clue

ClaudeBot Was Being Rejected with HTTP 418

The Observatory told us that ClaudeBot was reaching the site, but its retrieval success rate was abnormally low. The next step was to reproduce the behaviour outside the Observatory and see exactly what the server returned.

We sent controlled requests to the website using the ClaudeBot user agent.

curl.exe -i -A "ClaudeBot" "https://example.com/"

Instead of the expected successful response, the request was rejected with:

HTTP 418 ClaudeBot was being actively refused somewhere in the website delivery path.

HTTP 418 is not a normal response for a public webpage. In this context it was effectively a signal that a security rule was intercepting the request before the page could be delivered normally.

That narrowed the problem considerably, but it still did not tell us which layer was responsible.

What we knew

ClaudeBot could reach the infrastructure, but some of its retrieval attempts were being rejected before normal page delivery.

What we did not yet know

Whether the rejection was coming from Cloudflare, the web server, ModSecurity, WordPress, or another security component further down the stack.

So the problem stopped being simply “Claude cannot retrieve the site”.

It became a much more useful engineering question: where, exactly, is the request being stopped?

Isolating the Fault

Was Cloudflare Blocking ClaudeBot — or Was the Problem at the Origin?

The next job was to separate the delivery layers.

The site sat behind Cloudflare, so an HTTP 418 response did not automatically tell us whether the rejection was happening at the edge or on the hosting server behind it.

The key diagnostic question was simple: does ClaudeBot still fail when we test the origin directly?

We compared normal public requests through Cloudflare with direct-origin testing using the same ClaudeBot user agent.

Through Cloudflare

ClaudeBot requests were returning the abnormal HTTP 418 response.

Direct to the Origin

The same behaviour was reproduced at the hosting server itself.

That was the decisive result.

If the same failure could be reproduced at the origin, then Cloudflare was not creating the block. The request was being rejected further downstream.

The edge was doing its job. The origin was refusing the crawler.

From there, the investigation moved into the hosting security layer, where the behaviour was traced to ModSecurity.

This distinction mattered. Without isolating the layers first, it would have been easy to start changing Cloudflare rules, crawler permissions or WordPress configuration unnecessarily.

Instead, the evidence pointed directly at the correct part of the stack.

The Fix

Once the Fault Was Isolated, the Host Could Fix the Right Thing

With the failure reproduced at the origin and the security layer identified, the problem could be handed to the hosting provider with useful evidence rather than a vague report that “ClaudeBot seems to be blocked”.

The hosting environment was managed by Synergy Wholesale, and once the issue had been narrowed to ModSecurity they were able to correct the behaviour at the server side.

This is where evidence matters: instead of asking the host to inspect an entire WordPress stack, CDN configuration and security chain, we could point directly to the origin-layer block.

After the change, the same controlled ClaudeBot request was repeated.

Before the Fix

HTTP 418 ClaudeBot was rejected before the requested page could be delivered normally.

After the Fix

HTTP 200 OK The request passed through the delivery stack and the page was returned successfully.

The retest confirmed that Cloudflare was serving the request normally and that the origin was no longer rejecting the ClaudeBot user agent.

The important result was not simply that the error disappeared. It was that we could reproduce the fault, isolate its location, apply the correction and then verify the recovery.

That turns a suspected AI visibility problem into a measurable infrastructure incident with a clear before-and-after result.

Why This Matters

A Healthy Website Can Still Have an AI Retrieval Problem

This incident is a useful reminder that AI visibility is not only a content, schema or search problem.

A website can look completely normal to its owner, its customers and even conventional monitoring systems while an AI crawler is encountering a very different experience.

Everything can appear normal

Pages load, forms work, search engines index the site, Cloudflare serves traffic and human visitors see no obvious error.

But machine retrieval can still fail

A crawler can be blocked by a firewall rule, origin security layer, server configuration or another component that human browsing does not expose.

That matters because no amount of good content can help if the system trying to retrieve it cannot get the page.

Good content, valid schema and strong entity signals only become useful to a retrieval system if the underlying infrastructure actually allows that system to reach them.

This is why we treat AI visibility as a systems problem rather than a single SEO task. Content, structured data, identity, corroboration and infrastructure all have to work together.

In this case, the content was not the problem. The crawler permissions were not the problem. Cloudflare was not the problem.

The failure existed deeper in the stack, at the origin security layer.

If the machine cannot retrieve the page, everything above that point becomes irrelevant.

The value of the Observatory was not that it produced another score. It exposed a real retrieval failure that could be investigated, reproduced and fixed.

What the Observatory Added

From “Something Looks Wrong” to a Reproducible Technical Diagnosis

Without the Observatory, this problem could easily have remained invisible.

The website was online, the pages loaded normally, and there was no obvious customer-facing failure to investigate. The first useful signal came from the crawler evidence itself.

Observation

ClaudeBot showed an unusually low retrieval success rate compared with other recognised machine systems.

Verification

Controlled requests using the ClaudeBot user agent reproduced the failure outside the monitoring interface.

Isolation

Direct-origin testing showed that the block persisted behind Cloudflare, narrowing the fault to the hosting environment.

Confirmation

After the ModSecurity issue was corrected, the same test returned HTTP 200 OK and normal page delivery was restored.

That sequence is the important part.

Observe → reproduce → isolate → correct → verify.

That is a much stronger basis for technical decision-making than relying on assumptions about whether an AI crawler should be able to access a site.

It also means changes can be kept focused. We did not need to alter unrelated Cloudflare rules, rewrite content, change schema, or make speculative WordPress changes.

The problem was not guessed at. It was measured, traced and verified.
What This Demonstrates

AI Visibility Needs Retrieval Evidence, Not Assumptions

This case was not about a dramatic website outage. It was about something much quieter: one recognised AI crawler was being obstructed while the site continued to look healthy to almost everyone else.

That is exactly why retrieval evidence matters.

The Observatory showed that ClaudeBot was struggling. Controlled testing reproduced the failure. Direct-origin checks isolated the fault. The host corrected the ModSecurity behaviour. The same request then returned normally.

The result was a complete evidence chain from detection through to verification.

For us, that is the practical role of the AI Observatory: not to predict whether an AI system will recommend a business, but to establish whether recognised machine systems are actually reaching the site, retrieving its pages successfully, and encountering infrastructure problems that might otherwise remain invisible.

Retrieval is not the whole of AI visibility — but without retrieval, the rest of the visibility stack has nothing to work with.

The Observatory forms one part of Sydney Business Web's wider AI visibility methodology, alongside entity structure, technical diagnostics and external corroboration.

You can read more about the AI Observatory here .

References & Further Reading

Technical Background and Related Sydney Business Web Resources

This case sits within a wider body of work on AI crawler access, retrieval evidence, entity structure and infrastructure reliability.

Sydney Business Web

External Technical References

  • Anthropic: Web Crawlers and ClaudeBot — Anthropic's documentation describing ClaudeBot, Claude-User and Claude-SearchBot and how site owners can control access.
  • OWASP ModSecurity — technical documentation for the ModSecurity web application firewall and its role in inspecting and controlling HTTP traffic.

External references are provided for technical context. The crawler behaviour, origin testing and remediation described in this article are based on direct observations from the client incident.

FAQ

Frequently Asked Questions

What does HTTP 418 mean in this case?

HTTP 418 is not a normal success response for a public webpage. In this incident it indicated that the ClaudeBot request was being intercepted and rejected by a security rule before the page could be delivered normally.

Was Cloudflare blocking ClaudeBot?

No. Direct-origin testing reproduced the same failure behind Cloudflare, which showed that the rejection was occurring at the hosting origin rather than at the Cloudflare edge.

What actually caused the block?

The investigation isolated the behaviour to ModSecurity at the origin server. Once the hosting provider corrected the relevant behaviour, the same ClaudeBot test returned HTTP 200 OK.

Could a website owner notice this without monitoring?

Not necessarily. The website can continue to load normally for human visitors while a particular machine crawler is being rejected. That is why crawler-specific retrieval evidence can reveal problems that ordinary uptime monitoring does not.

Does fixing crawler access guarantee AI visibility?

No. Successful retrieval is only one part of AI visibility. Content quality, entity clarity, structured data, corroboration and relevance still matter. But if the crawler cannot retrieve the site, those other signals cannot be used effectively.

What does the AI Observatory actually measure?

The AI Observatory records recognised AI crawler activity at the edge and helps show whether those systems are reaching the site successfully, failing requests, or encountering abnormal infrastructure behaviour.

Why is direct-origin testing important?

It helps separate CDN or edge behaviour from hosting-server behaviour. If the same failure occurs when the origin is tested directly, the investigation can move past the edge layer and focus on the server-side stack.

Can ModSecurity block AI crawlers without blocking normal visitors?

Yes. Security rules can react differently to user agents, request patterns, headers or other characteristics. A site may therefore appear completely healthy to normal visitors while a specific crawler is being refused.

AI Retrieval Diagnostics

Do You Know Whether AI Crawlers Can Actually Reach Your Website?

A website can look healthy while recognised AI systems are being blocked, rejected or failing somewhere in the delivery path.

Sydney Business Web uses the AI Observatory to measure real crawler activity, identify abnormal retrieval behaviour and provide evidence that can be investigated at the correct technical layer.

If you want to know whether AI systems are actually reaching your website successfully — rather than simply assuming they can — we can check it.

Observe

Measure recognised AI crawler activity and retrieval success.

Diagnose

Trace unusual failures through Cloudflare, hosting, security and application layers.

The AI Observatory is available as part of Sydney Business Web's AI visibility diagnostic and monitoring work.

Learn more about the AI Observatory  or  talk to us about your website's AI visibility .