LIVE SEARCH & AI CRAWLER EVIDENCE
See Which Search & AI Crawlers
Are Reading Your Website
A live cloud-based measurement system is running right now, watching search and AI-related crawlers as they reach this website and measuring whether they successfully retrieve useful public information.
This is not a mock-up. The figures below are being generated from real production traffic observed during the last 24 hours.
And this is not just about Sydney Business Web.
This page is a live demonstration of a measurement system that can be applied to suitable business websites. It allows a business to see, at a glance, which identifiable search and AI-related crawlers have actually reached the site, what kinds of information they retrieved, and whether that retrieval succeeded.
LIVE — ROLLING 24-HOUR EVIDENCE
Who Has Actually Been Here?
These are real search and AI-related machine systems observed reaching the website during the current rolling 24-hour window. The figures below show qualifying machine activity that passed our public evidence filters.
Distinct autonomous crawler systems whose identities met the public corroboration threshold during this rolling 24-hour period. This is an observed count — not a system limit.
Successful qualifying retrievals of substantive public business information from the website.
The proportion of qualifying business-information retrieval attempts that succeeded.
Successful retrievals of deliberate machine-facing discovery resources such as sitemaps and robots.txt.
This is the interesting bit: the totals above are not estimates. Below are the individual search and AI-related crawler systems that produced the qualifying evidence during this particular 24-hour window.
MACHINES OBSERVED
Which Search & AI Crawlers Produced the Evidence?
Every qualifying system observed during the current rolling 24-hour window is shown below. There is no fixed list and no five-system limit. The table expands and contracts automatically as real crawler activity changes.
| Machine | Role |
Autonomous Accesses |
Business Retrievals |
Success Rate |
Discovery Retrievals |
|---|---|---|---|---|---|
| Loading current machine evidence… | |||||
A system appears when qualifying activity from it is observed inside the rolling 24-hour window, and can disappear again when that activity falls outside the window. At different times this dashboard may therefore show three systems, five systems, eight systems or more. It simply reports what the current evidence shows.
That is deliberate. A bigger number would be easy to produce; a defensible number is more useful.
Synthetic requests created by Sydney Business Web to test the observatory are excluded.
Owner-triggered tools such as Google Inspection activity are retained separately but do not count as autonomous headline evidence.
A crawler name in a User-Agent is not enough. If the identity cannot be sufficiently corroborated, it does not enter the public headline figures.
Support assets, security/configuration probes, HEAD-only checks and unclassified resources do not inflate business-information retrieval figures.
Our private cloud-based observatory software continuously monitors more machine activity than appears in this public dashboard.. The public figures are intentionally smaller because only observations that satisfy the measurement rules are promoted into evidence.
NOW IMAGINE THIS ON YOUR WEBSITE
What Could This Tell You About Your Website?
The live evidence above happens to be from Sydney Business Web because this is where we built and commissioned the system. The principle is not limited to our website. A suitable business website can be instrumented to observe its own search and AI-related machine retrieval activity.
“Are Google, Bing, ChatGPT, Claude and other machine systems actually reaching my website — and can they retrieve what I want them to see?”
Instead of assuming the answer is yes, the purpose of the observatory is to give you evidence of what is actually happening.
Who is reaching the site?
See which sufficiently corroborated search and AI-related crawler systems have actually appeared during the observation period — rather than relying on a theoretical list of bots that might visit someday.
Are they retrieving business information?
Separate meaningful retrieval of public business content from requests for images, scripts, fonts, diagnostic resources and other background website traffic.
Is retrieval succeeding?
Measure whether qualifying requests for business information actually succeed. A crawler arriving at the door is not the same thing as successfully retrieving the information inside.
Has something changed?
Because the evidence is time-based, changes become visible. A machine system can appear, disappear, increase its activity or begin encountering retrieval failures.
Website analytics usually asks:
“What are my human visitors doing?”
Sessions, page views, conversions, traffic sources and visitor behaviour remain extremely useful business measures.
This observatory asks:
“What are machine systems actually retrieving?”
It adds a different layer of visibility: observable retrieval activity from search and AI-related machine systems.
Suppose an OpenAI-related crawler appears in your evidence today, successfully retrieves several business pages, and then disappears from tomorrow's rolling window.
That does not mean ChatGPT has suddenly forgotten your business. It means exactly what the instrument says: qualifying activity from that machine system was observed during one measurement window and was not observed during the other.
The dashboard reports evidence — not a story invented around the evidence.
WHY THIS MATTERS
AI visibility should not begin and end with asking an AI system what it thinks about your business.
Those tests can be useful, but they examine the result from the outside. Retrieval evidence gives us another view entirely: whether identifiable machine systems are actually accessing the website and successfully retrieving its information.
That is where the engineering comes in. We deliberately reject a substantial amount of raw machine traffic before anything is allowed into the public evidence.
Now Show Me How It Works ↓HOW THE MEASUREMENT WORKS
How Do We Know These Numbers Are Real?
The public dashboard is not simply counting anything that calls itself a crawler. Our cloud-based observatory software continuously records machine activity reaching the website, then applies a series of filters before an observation is allowed into the public evidence.
Where identity, purpose or relevance is uncertain, the observation can remain in our private observatory but does not enter the headline evidence.
Observe Real Traffic
The system observes requests as they reach the production website. It records enough technical information to determine what requested the resource, what was requested and what response the website returned.
REAL PRODUCTION ACTIVITYCorroborate Identity
A request saying “Googlebot”, “ClaudeBot” or another crawler name is not automatically trusted. The system looks for additional technical evidence supporting the claimed provider.
IDENTITY MUST BE SUPPORTEDSeparate Autonomous Activity
Requests we trigger ourselves for testing or inspection are not treated as independent crawler behaviour. Diagnostic activity is retained separately rather than inflating the headline figures.
AUTONOMOUS ACTIVITY ONLYClassify What Was Retrieved
Business pages are not treated the same as stylesheets, images, fonts, security probes or machine-discovery files. The requested resource is classified before it can contribute to a public metric.
MEANINGFUL RESOURCES COUNTCount Only the Evidence
Once an observation passes the required filters, it can contribute to the rolling public measurements: machine systems observed, retrieval attempts, successful retrievals and success rates.
PUBLIC EVIDENCEWHAT GETS FILTERED OUT?
Quite a Lot — Deliberately.
Raw machine traffic contains many requests that are useful operationally but would be misleading if they were all presented as evidence of meaningful AI or search retrieval.
Our Own Tests
Synthetic requests generated to commission, test or diagnose the observatory are excluded from production evidence.
Owner-Triggered Diagnostics
Genuine tools such as Google Inspection activity may be visible to the observatory, but they are not autonomous crawling and therefore do not enter autonomous headline figures.
Unsupported Identity Claims
A User-Agent can claim almost anything. A crawler-labelled request that does not satisfy the public corroboration threshold is not promoted into public evidence.
Support Assets
Images, JavaScript, CSS, fonts, video and similar supporting files can show that a machine reached the site, but they do not count as substantive business-information retrieval.
Security & Configuration Probes
Requests looking for configuration files, credentials or other sensitive paths are not evidence that a system retrieved useful public business information.
HEAD-Only Requests
A HEAD request can check whether a resource exists without retrieving its body. We do not treat that as equivalent to retrieving the information itself.
Did the machine retrieve useful public content?
This includes qualifying retrievals of substantive pages and other explicitly approved public information resources carrying information about the business, its services or its published work.
Could the machine discover and navigate the site?
Files such as robots.txt and XML sitemaps are important machine-facing resources, but retrieving them is not the same thing as retrieving substantive business information.
WHAT THE DASHBOARD OUTPUTS MEAN
Four Numbers. Four Different Questions.
Qualifying Machine Systems
How many distinct autonomous machine systems met the public evidence threshold during the rolling window?
Business Retrievals
How many qualifying requests successfully retrieved substantive public business information?
Retrieval Success Rate
Of the qualifying attempts to retrieve business information, what proportion actually succeeded?
Discovery Retrievals
How often were deliberate machine-discovery resources successfully retrieved?
FOR THOSE WHO WANT THE ENGINE ROOM
The Technical Architecture
The observatory operates at the website edge using Cloudflare. A production Worker records selected request and response characteristics into Cloudflare Analytics Engine. A separate, read-only reporting Worker queries those measurements, applies the public measurement rules and supplies the dashboard through a dedicated evidence endpoint.
Raw telemetry is evidence material. It is not yet a conclusion.
The purpose of the filtering system is to turn raw machine activity into a smaller set of measurements we are prepared to defend publicly.Retrieval proves retrieval — not understanding or recommendation.
The next section explains exactly what this evidence does, and does not, allow us to claim.
What Does This Actually Prove? ↓THE EVIDENCE BOUNDARY
What This Proves — and What It Does Not
Good measurement becomes less useful the moment we claim more than the evidence actually shows. This observatory has a deliberately narrow job: to measure observable machine access and retrieval.
Retrieval proves retrieval. It does not automatically prove what happened inside an AI system after the information was retrieved.
What We Can Say
Observable production traffic shows that the website received qualifying machine requests during the measured period.
The request was not accepted simply because it carried a familiar crawler name.
The observatory can distinguish requests for substantive business information from discovery files, support assets and other traffic.
The website response allows us to measure whether retrieval attempts were successfully completed.
Retrieval of resources such as robots.txt and sitemaps can be measured separately from business-content retrieval.
What We Cannot Say
Successful retrieval tells us the information was accessible, not how an AI model interpreted it internally.
Retrieval is a necessary precursor to many forms of representation, but it is not itself evidence of citation.
A crawler retrieving a page does not prove that an AI assistant later selected the business for a recommendation.
Machine access is technical activity. It must never be represented as endorsement or approval by the provider.
Leads, enquiries, sales and other business outcomes are separate downstream measurements.
WHERE THIS MEASUREMENT SITS
AI Visibility Is a Chain, Not a Single Event
Retrieval is important because a machine cannot reliably work with information it cannot first access. But retrieval is only one part of the larger path from being discoverable to producing a business outcome.
Highlighted stages: the principal part of the chain measured by this retrieval observatory.
Retrieval evidence shows observable machine access to this website. It does not by itself prove AI understanding, citation, recommendation or endorsement.
WHY DRAW THE LINE SO CLEARLY?
Because a measurement is only useful if you know what it measures.
We could make the story sound more impressive by blending crawler activity, AI answers, rankings, citations and recommendations into one vague “AI visibility score”.
We deliberately do not.
Access is access. Retrieval is retrieval. Representation is representation. Outcomes are outcomes. Each deserves its own evidence.
The methodology and evidence are published.
The final part of this page links to the supporting technical material, related AI Visibility Verification work and the next step for businesses interested in measuring their own machine retrieval.
See the Evidence & References ↓SUPPORTING WORK & VERIFICATION
Don’t Just Take the Dashboard at Face Value.
The live display is the visible end of a larger measurement system. We have published the reasoning, measurement boundaries and real-world retrieval evidence behind it so that the method can be examined rather than treated as a black-box marketing claim.
Measuring Machine Retrieval
The first part of our AI Visibility Verification work explains why crawler activity must be measured carefully and why a crawler name alone is not sufficient evidence of identity.
Read Part One → PART 02Evidence from Real Machine Access
The second part examines actual production retrieval evidence, including what the observations can establish and where their evidential limits lie.
Read Part Two → REFERENCEAI Visibility Glossary
A practical reference for the terminology used across AI search, machine retrieval, entities, structured data, verification and Sydney Business Web’s AI Visibility framework.
Open the Glossary →The figures displayed above are not hard-coded into this page. The page requests its current measurements from the Sydney Business Web evidence reporting endpoint, which aggregates qualifying observations from the underlying cloud observatory.
View the Raw Public Evidence Feed ↗NOW BRING IT BACK TO YOUR BUSINESS
Could We Measure This on Your Website?
Potentially, yes.
This Sydney Business Web implementation is our working production example. The same measurement principles can be applied to suitable business websites where the required access and technical environment allow reliable observation of machine traffic.
The aim is simple: instead of wondering whether search and AI-related systems are reaching your website, give the business a way to observe and measure what is actually happening.
Which qualifying search and AI-related crawler systems were observed
Whether they successfully retrieved substantive business information
Whether machine-discovery resources were successfully accessible
How machine activity changes across a rolling observation period
Which observations were strong enough to enter defensible public evidence
NOT EVERY WEBSITE IS IDENTICAL
We Would Check the Environment First.
This is instrumentation, not a plugin badge. Before promising a dashboard, we need to know that the website and its delivery stack give us the technical access required to measure the right things.
Technical Suitability
We examine the website, hosting, DNS/CDN arrangement and the available observation points.
Measurement Design
We determine which resources and machine behaviours matter for the business and define what should — and should not — count.
Observatory Deployment
The measurement layer is implemented, tested and separated from synthetic commissioning traffic.
Baseline & Reporting
Real production observations establish the baseline from which machine retrieval can be monitored over time.
AI VISIBILITY SHOULD BE OBSERVABLE
Want to Know What the Machines Are Actually Doing on Your Website?
Talk to Sydney Business Web about whether a machine-retrieval observatory is technically appropriate for your website and how it could fit into a wider AI Visibility measurement program.
No claims of guaranteed AI citation, recommendation or ranking. We measure what the available evidence allows us to measure.Engineer → Corroborate → Verify Retrieval → Test Representation
Sydney Business Web — AI Visibility VerificationPRACTICAL QUESTIONS
Live AI Retrieval Evidence FAQ
The measurement is deliberately narrow. These are the practical questions we would expect a business owner to ask after seeing the live evidence above.
01 Can this system be installed on any website? +
Not necessarily. The website, hosting, DNS/CDN arrangement and available technical access need to be assessed first. The measurement system requires a reliable observation point where relevant machine requests and website responses can be recorded without interfering with normal website operation.
02 Does seeing an AI crawler mean my business is being recommended by AI? +
No. It can provide evidence that a qualifying machine system reached the website and successfully retrieved information. It does not by itself prove that an AI system understood, cited, represented, recommended or endorsed the business.
03 Why do the crawler names change from day to day? +
Because the public dashboard uses a rolling 24-hour observation window. A crawler system appears when qualifying activity from it is observed inside that window and can disappear again when that activity falls outside it.
The displayed population is therefore evidence of what was actually observed during the current period — not a fixed list of crawler brands.
04 Why are some machine requests excluded from the public figures? +
Because counting everything would produce a larger number, but a less useful one.
Synthetic tests, owner-triggered diagnostics, insufficiently corroborated crawler identities, support assets, security and configuration probes, HEAD-only requests and other non-qualifying activity are prevented from inflating the headline retrieval evidence.
05 Can the dashboard show more than five crawler systems? +
Yes. There is no five-system limit and no fixed crawler list. The dashboard is generated from the qualifying systems actually observed during the rolling measurement window.
It may therefore display three systems, five systems, eight systems or more as real machine activity changes.
06 Can this tell me whether ChatGPT, Claude or another AI system has actually read my website? +
It can show qualifying retrieval activity attributable to relevant machine systems when that activity is actually observed and meets the required corroboration and measurement rules.
That is strong evidence of machine access and retrieval. It does not reveal what an AI model subsequently understood, retained, cited or did with the retrieved information.
We measure what happened at the website. We do not pretend that this reveals everything that happened afterwards inside somebody else's AI system.
EVIDENCE & TECHNICAL SOURCES
Internal Research & External References
The live retrieval evidence above sits on top of published Sydney Business Web research and first-party technical documentation from the companies operating the infrastructure, search engines and machine systems being measured.
Internal Research & Evidence
Our published work explaining the retrieval methodology, evidence boundaries, terminology and wider AI Visibility framework.
AI Visibility Verification — Measuring Machine Retrieval
Why machine retrieval should be measured separately, why crawler activity needs careful interpretation and why a crawler name alone is not sufficient evidence of identity.
AI Retrieval Evidence from Real Machine Access
Real production retrieval evidence from search and AI-related machine systems, together with the limits on what those observations legitimately allow us to conclude.
AI Visibility Glossary
Practical definitions covering AI Visibility, retrieval, entities, structured data, verification, GEO, AEO and other machine-search concepts used throughout this work.
AI Visibility Services & Pricing
How entity engineering, corroboration, technical accessibility, retrieval measurement and representation testing fit within our wider AI Visibility work.
External Technical Sources
Wherever practical, we reference the organisations actually operating the infrastructure and crawler systems rather than secondary marketing or SEO commentary.
Workers Analytics Engine
Cloudflare documentation for the analytics platform used to record and aggregate measurements from the production retrieval observatory.
Analytics Engine SQL API
Documentation for querying Analytics Engine datasets through the SQL API used by the separate reporting layer behind the public evidence dashboard.
Verified Bots
Cloudflare's documentation covering verified automated systems and the distinction between an asserted crawler identity and independently verified bot traffic.
Googlebot & Crawler Verification
Google's first-party documentation describing Googlebot, crawler identification and methods for verifying whether a request genuinely originates from Google.
Verify Bingbot
Microsoft's guidance for determining whether traffic claiming to be Bingbot actually originates from Bing infrastructure.
OpenAI Publishers & Developers
OpenAI guidance covering web discovery, crawler controls and OAI-SearchBot access for content that may be surfaced through ChatGPT search.
Anthropic Web Crawlers & Site Controls
Anthropic's first-party documentation describing its web crawler activity and the controls available to website owners.
About Applebot
Apple's documentation describing Applebot, its uses, crawler identification, verification methods and website access controls.
We deliberately favour documentation published by the organisations operating the infrastructure and machine systems being discussed.
NOW PUT THE INSTRUMENT ON YOUR WEBSITE
Want to Know Which Search & AI Crawlers Are Actually Reaching Your Business?
Sydney Business Web can assess whether your website is suitable for machine-retrieval monitoring and, where technically appropriate, build an observatory around your own production traffic.
Stop relying entirely on crawler lists, assumptions and occasional AI searches. Measure what is actually happening at your website.
We assess technical suitability first. Retrieval monitoring measures observable machine access and retrieval — it does not guarantee AI citation, recommendation, ranking or endorsement.
