# # robots.txt # # This file is to prevent the crawling and indexing of certain parts # of your site by web crawlers and spiders run by sites like Yahoo! # and Google. By telling these "robots" where not to go on your site, # you save bandwidth and server resources. # # This file will be ignored unless it is at the root of your host: # Used: http://example.com/robots.txt # Ignored: http://example.com/site/robots.txt # # For more information about the robots.txt standard, see: # http://www.robotstxt.org/robotstxt.html # User-triggered AI agents: someone asked their assistant to fetch/browse this # specific page in real time (not bulk crawling/training). Deliberately kept # separate from the AI crawler opt-out block below — matches the same # distinction already made in scripts/ai-robots/bglog-nginx-block-ai-crawlers.conf. User-agent: ChatGPT Agent User-agent: ChatGPT-User User-agent: Claude-User User-agent: Manus-User User-agent: MistralAI-User User-agent: MistralAI-User/1.0 User-agent: NotebookLM User-agent: NovaAct User-agent: Operator User-agent: Perplexity-User Allow: / # Social sharing/link preview crawlers need access to Open Graph metadata. User-agent: facebookexternalhit Allow: / User-agent: FacebookBot Allow: / User-agent: Facebot Allow: / User-agent: meta-externalagent Allow: / User-agent: Meta-ExternalAgent Allow: / User-agent: meta-externalfetcher Allow: / User-agent: Meta-ExternalFetcher Allow: / User-agent: meta-webindexer Disallow: /search/ Disallow: /user/login Allow: / User-agent: * # CSS, JS, Images Allow: /core/*.css$ Allow: /core/*.css? Allow: /core/*.js$ Allow: /core/*.js? Allow: /core/*.gif Allow: /core/*.jpg Allow: /core/*.jpeg Allow: /core/*.png Allow: /core/*.svg Allow: /profiles/*.css$ Allow: /profiles/*.css? Allow: /profiles/*.js$ Allow: /profiles/*.js? Allow: /profiles/*.gif Allow: /profiles/*.jpg Allow: /profiles/*.jpeg Allow: /profiles/*.png Allow: /profiles/*.svg # Directories Disallow: /core/ Disallow: /profiles/ # Files Disallow: /README.md Disallow: /composer/Metapackage/README.txt Disallow: /composer/Plugin/ProjectMessage/README.md Disallow: /composer/Plugin/Scaffold/README.md Disallow: /composer/Plugin/VendorHardening/README.txt Disallow: /composer/Template/README.txt Disallow: /modules/README.txt Disallow: /sites/README.txt Disallow: /themes/README.txt # Paths (clean URLs) Disallow: /admin/ Disallow: /comment/reply/ Disallow: /filter/tips Disallow: /node/add/ Disallow: /search/ Disallow: /user/register Disallow: /user/password Disallow: /user/login Disallow: /user/logout Disallow: /media/oembed Disallow: /*/media/oembed # Legacy uploaded documents. These URLs are migrated file attachments, not # landing pages, and should not be indexed as standalone search results. Disallow: /ClientFiles/ Disallow: /sites/default/files/legacy-clientfiles/ClientFiles/ # Paths (no clean URLs) Disallow: /index.php/admin/ Disallow: /index.php/comment/reply/ Disallow: /index.php/filter/tips Disallow: /index.php/node/add/ Disallow: /index.php/search/ Disallow: /index.php/user/password Disallow: /index.php/user/register Disallow: /index.php/user/login Disallow: /index.php/user/logout Disallow: /index.php/media/oembed Disallow: /index.php/*/media/oembed # SEO backlink/research crawlers: pure resource drain, no benefit to # bglog.net (they feed their own paid tools, not Google/Bing search # results). SemrushBot is deliberately left crawlable - kept for the # site's own SEO monitoring. User-agent: DotBot User-agent: AhrefsBot User-agent: MJ12bot Disallow: / # AI crawler opt-out. Based on https://github.com/ai-robots-txt/ai.robots.txt # Added as a polite robots.txt layer; nginx can hard-block selected crawlers. User-agent: AddSearchBot User-agent: AI2Bot User-agent: AI2Bot-DeepResearchEval User-agent: Ai2Bot-Dolma User-agent: aiHitBot User-agent: amazon-kendra User-agent: Amazonbot User-agent: AmazonBuyForMe User-agent: Amzn-SearchBot User-agent: Amzn-User User-agent: Andibot User-agent: Anomura User-agent: anthropic-ai User-agent: ApifyBot User-agent: ApifyWebsiteContentCrawler User-agent: Applebot User-agent: Applebot-Extended User-agent: Aranet-SearchBot User-agent: atlassian-bot User-agent: Awario User-agent: AzureAI-SearchBot User-agent: bedrockbot User-agent: bigsur.ai User-agent: Bravebot User-agent: Brightbot User-agent: Brightbot 1.0 User-agent: BuddyBot User-agent: Bytespider User-agent: CCBot User-agent: Channel3Bot User-agent: ChatGLM-Spider User-agent: Claude-Code User-agent: Claude-SearchBot User-agent: Claude-Web User-agent: ClaudeBot User-agent: Cloudflare-AutoRAG User-agent: CloudVertexBot User-agent: Code User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Cotoyogi User-agent: Crawl4AI User-agent: Crawlspace User-agent: Datenbank Crawler User-agent: DeepSeekBot User-agent: Devin User-agent: Diffbot User-agent: DuckAssistBot User-agent: Echobot Bot User-agent: EchoboxBot User-agent: ExaBot User-agent: Factset_spyderbot User-agent: FirecrawlAgent User-agent: FriendlyCrawler User-agent: Gemini-Deep-Research User-agent: GPTBot User-agent: HenkBot User-agent: iAskBot User-agent: iaskspider User-agent: iaskspider/2.0 User-agent: IbouBot User-agent: ICC-Crawler User-agent: ImagesiftBot User-agent: imageSpider User-agent: img2dataset User-agent: ISSCyberRiskCrawler User-agent: kagi-fetcher User-agent: Kangaroo Bot User-agent: KlaviyoAIBot User-agent: KunatoCrawler User-agent: laion-huggingface-processor User-agent: LAIONDownloader User-agent: LCC User-agent: LinerBot User-agent: Linguee Bot User-agent: LinkupBot User-agent: MyCentralAIScraperBot User-agent: NagetBot User-agent: netEstate Imprint Crawler User-agent: newsai User-agent: OAI-SearchBot User-agent: omgili User-agent: omgilibot User-agent: OpenAI User-agent: opencode User-agent: PanguBot User-agent: Panscient User-agent: panscient.com User-agent: PerplexityBot User-agent: PetalBot User-agent: PhindBot User-agent: Poggio-Citations User-agent: Poseidon Research Crawler User-agent: QualifiedBot User-agent: QuillBot User-agent: quillbot.com User-agent: SBIntuitionsBot User-agent: Scrapy User-agent: SemrushBot-OCOB User-agent: SemrushBot-SWA User-agent: ShapBot User-agent: Sidetrade indexer bot User-agent: Spider User-agent: TavilyBot User-agent: Terra Cotta User-agent: TerraCotta User-agent: Thinkbot User-agent: TikTokSpider User-agent: Timpibot User-agent: Trae User-agent: TwinAgent User-agent: VelenPublicWebCrawler User-agent: WARDBot User-agent: Webzio-Extended User-agent: webzio-extended User-agent: wpbot User-agent: WRTNBot User-agent: YaK User-agent: YandexAdditional User-agent: YandexAdditionalBot User-agent: YouBot User-agent: ZanistaBot Disallow: / Sitemap: https://bglog.net/sitemap.xml