# gritt.io — robots.txt # # ⚠️ This file REPLACES the Cloudflare-managed robots.txt the moment gritt-web takes the # /robots.txt route. The managed version currently live disallows GPTBot, ClaudeBot, # Google-Extended, CCBot, Bytespider, Amazonbot, Applebot-Extended and meta-externalagent. # That block is REMOVED here, deliberately — AI answer engines are a discovery channel for # founders searching "who invests in seed-stage fintech", and Gritt cannot be cited in an # answer it is not allowed to read. # # ⚠️ robots.txt is not the only control. If Cloudflare's zone-level AI crawler blocking is # enabled (Security → Bots → "Block AI bots" / AI Scrapers & Crawlers), those crawlers are # refused at the edge regardless of what this file says. That setting must be OFF for the # policy below to take effect. # # Note the crawl surface is genuinely public: profiles carry bio, portfolio and links, while # the investor's email address is gated server-side by plan. Being read is the product. # ── Default: everything indexable except private app surfaces ────────────────────────────── User-agent: * Allow: / # Auth pages — noindexed anyway, but kept out of crawl budget Disallow: /login/ Disallow: /register/ Disallow: /forgot-password/ Disallow: /auth/ Disallow: /campaign/ Disallow: /billing/ Disallow: /settings/ Disallow: /payment-success/ # Server routes — never useful to a crawler Disallow: /api/ # ── AI answer engines — explicitly welcomed ──────────────────────────────────────────────── # Listed individually rather than relying on the wildcard above, so the intent is unambiguous # and survives anyone re-tightening the default. User-agent: GPTBot Allow: / Disallow: /api/ Disallow: /campaign/ Disallow: /billing/ Disallow: /settings/ User-agent: OAI-SearchBot Allow: / Disallow: /api/ User-agent: ChatGPT-User Allow: / Disallow: /api/ User-agent: ClaudeBot Allow: / Disallow: /api/ Disallow: /campaign/ Disallow: /billing/ Disallow: /settings/ User-agent: Claude-Web Allow: / Disallow: /api/ User-agent: anthropic-ai Allow: / Disallow: /api/ User-agent: PerplexityBot Allow: / Disallow: /api/ User-agent: Google-Extended Allow: / Disallow: /api/ User-agent: Applebot-Extended Allow: / Disallow: /api/ User-agent: meta-externalagent Allow: / Disallow: /api/ User-agent: Bytespider Allow: / Disallow: /api/ User-agent: CCBot Allow: / Disallow: /api/ User-agent: Amazonbot Allow: / Disallow: /api/ # ── Sitemaps ─────────────────────────────────────────────────────────────────────────────── # /sitemap.xml is a sitemap INDEX pointing at the three child sitemaps (all live, served by gritt-web): # /sitemap/sitemap-pages.xml — search / SEO / marketing pages, A–G tier # /sitemap/sitemap-profile.xml — investor profiles, by investment count + # /sitemap/sitemap-company.xml — companies with >=5 investors, + (its own ?p= index if >45k) # Declaring the index alone is enough — Google discovers the children through it. Only ever declare a # sitemap that returns 200; advertising a 404 wastes crawl budget and is flagged in Search Console. Sitemap: https://www.gritt.io/sitemap.xml