Crawler

SiteTransformerBot

SiteTransformerBot is the web crawler of Site Transformer. It only ever visits a site whose owner asked us to crawl it — it does not roam the open web and it builds no public index.

Why it visited your site

Because someone who proved they control this site started a crawl of it from their Site Transformer dashboard.

Site Transformer is a reverse proxy for technical SEO. Before a site can be crawled, our customer has to add it to their account and verify ownership by publishing a DNS record we ask them for. Only then does the Start crawl button in their dashboard do anything at all. The crawl maps the site's URL structure and page templates so we can propose a configuration for that customer to review.

The requests in your access log are that crawl. We do not crawl the open web, we do not resell or republish what we fetch, and we never visit a site nobody registered with us. If you believe this crawl was not authorised by the site's owner, write to [email protected] and we will stop it.

How to identify it

Every request the crawler makes — including the /robots.txt fetch that opens a run — carries the same three identifying fields:

  • User-Agent — always exactly SiteTransformerBot/1.0 (+https://www.sitetransformer.com/bot). The version moves only when crawl behaviour changes materially, never on an ordinary release.
  • From — [email protected], the mailbox that reaches a person here.
  • X-SiteTransformer-Crawl — the id of the crawl run. Every request in one run carries the same id, so you can group our lines in your log and quote the id if you write to us.

Match on the product token SiteTransformerBot rather than on the whole string: that is the token to use in robots.txt, and it survives a version bump.

request · headers
# every request, /robots.txt included
User-Agent: SiteTransformerBot/1.0 (+https://www.sitetransformer.com/bot)
From: [email protected]
X-SiteTransformer-Crawl: rC7m2Kd9x1QpVb3TfLzYs

# the method is always GET

Where it connects from

Our crawler connects from a small, fixed set of addresses. Checking the source IP against that set is the only reliable way to tell the real crawler from something imitating it — and those are the addresses to allow in a firewall or WAF if you want the crawl to succeed.

That set changes as our infrastructure grows, so we publish the current addresses where they cannot go stale: on the Start crawl panel in the dashboard, next to the User-Agent. Ask the site owner who started the crawl to read them to you, or write to [email protected] and we will send them.

How it behaves

The crawler is deliberately slow. It maps one site for its owner; it is not racing to index the web.

  • It reads your robots.txt first. Before anything else it fetches /robots.txt from the origin server and obeys it, including Crawl-delay.
  • It reads your sitemap. URLs listed in the sitemaps named in robots.txt, or in /sitemap.xml, are crawled too.
  • About one request per second to start. A run starts at roughly one request per second and, on a server that answers quickly, may grow to at most ten. It never exceeds your robots.txt Crawl-delay.
  • Four connections at most. No more than four requests are ever in flight at the same time.
  • It slows down when you push back. A 429 or 503 makes it back off, and a Retry-After header is honoured. Very long delays are capped, so if you want the crawler gone rather than slow, block it — see below.
  • GET only. It never posts, never submits a form and never signs in, and it skips URLs that look like they would change something (cart, checkout, log out, unsubscribe).

How to slow it down or block it

The robots.txt of the origin server is the control surface, and it is the only one you need: write a group for the token SiteTransformerBot and the crawler picks it up on its next run. You do not have to ask us for anything.

Allow it, but keep your rules

robots.txt · allow
User-agent: SiteTransformerBot
Disallow: /admin/
Disallow: /cart/
Disallow: /search

A named group replaces the User-agent: * group for us — it does not add to it. If you write one, repeat every Disallow line you still want us to respect.

Slow it down

robots.txt · slow down
User-agent: SiteTransformerBot
Crawl-delay: 5

Crawl-delay: 5 gives you one request every five seconds. The crawl simply takes longer; nothing about it fails.

Block it entirely

robots.txt · block
User-agent: SiteTransformerBot
Disallow: /

The crawler stops before it fetches a single page. The customer who started the run sees it finish with a blocked by robots.txt result telling them your robots.txt does not allow us — so they know to talk to you instead of retrying.

A change takes effect on the next run rather than the one already in flight: a crawl reads robots.txt when it starts. If you need a crawl stopped right now, mail us — the address is below.

Verifying it is really us

Check the source IP address of the request against the addresses above. That is the only check worth anything: a User-Agent is plain text, so anyone can copy ours, and scrapers do.

Traffic that claims to be SiteTransformerBot from any other address is not us — block it however you like, and no crawl of your site is affected. And because a User-Agent rule catches impostors and the real crawler alike, prefer an IP-based rule when you are rate-limiting rather than blocking.

Contact

Abuse reports, questions, or a crawl you want stopped now: [email protected].

Send the hostname and, if you have it, the X-SiteTransformer-Crawl id from your log — with that id we can find the exact run and stop it. We answer in English and Turkish.

Change log

The version in the User-Agent moves only when crawl behaviour changes materially — a new fetch policy, a different rate profile — not on every release of the product.

  • 1.0 · 7 September 2026 — initial release.