About the StumbleBot crawler
If you run a website and saw StumbleBot/1.0 (+https://stumble-bot.com/bot) in your logs, this page explains what it is.
What it does
Stumble-Bot is a small, invite-only web discovery service. Its indexer reads RSS and Atom feeds, public APIs (YouTube feeds, Flickr public feeds, the Wikipedia API, the Hacker News Algolia API, the Marginalia search API) and a curated list of websites. For each new link it fetches the page once to read its title, description, preview image and frame headers, and re-checks it roughly once a month to make sure the link is still alive. It does not crawl links inside pages, download media, or index page text.
How polite it is
- At most one request every two seconds to any single host, and a few concurrent fetches in total.
- Reads at most 1.5 MB of a page and stops.
- Honors HTTP caching headers (ETag / Last-Modified) on feeds and backs off exponentially on errors.
- Never fetches private, local or non-web addresses.
How your site is shown
Members see your page inside a sandboxed frame with the Stumble-Bot bar above it — exactly as a visitor would see it in a new tab, with your own URL, scripts and ads intact — or, if your site sends X-Frame-Options or a Content-Security-Policy: frame-ancestors header that forbids framing, as a preview card with a button that opens your page in a new tab. We send no referrer to your site.
Opting out
Email [email protected] with your domain and we will remove it from the index and stop fetching it. A Disallow rule for StumbleBot in your robots.txt is honored for new fetches as well.