# Banana Media Network — robots.txt # # Lives in the theme root. Ghost serves a theme's robots.txt in place of its own # default, so the Ghost disallows below are copied from that default and must # stay: dropping them exposes /ghost/ and the member comment endpoints. # # POSITION OF THIS FILE: nothing is blocked except Ghost admin and a handful of # internal API paths. Every crawler is allowed through — search, AI search, and # the ones that collect training data alike. # # That last group was refused until 16 August 2026. It is allowed now on the # owner's decision, and the same decision is written into /terms/ so the two # cannot drift apart. See the Content signals section below for what moved and # what deliberately did not. Sitemap: https://bananamedianetwork.com/sitemap.xml # ── Content signals ───────────────────────────────────────────────────────── # # Cloudflare's Content Signals Policy. Three independent signals: # # search = yes — index us and link to us. # ai-input = yes — fetch us to answer a reader's question, and cite us. # ai-train = yes — you may use this to train or fine-tune a model. # # ai-train was no until 16 August 2026 and is now yes, on the owner's decision. # This is a licensing position, not an oversight, and it was changed everywhere # at once rather than here alone: # # /terms/ the "Text and data mining" section, and the plain-language # answer above it, both rewritten to permit it. # # default.hbs the pair removed. Those two tags # were the W3C machine-readable form of a reservation that no # longer exists, and leaving them would have contradicted this # line on every page of the site. # # The three had to move together. A signal saying yes while the terms say no is # not a small inconsistency — the signal is what an AI company reads for # permission and the terms are what a court would read, so disagreeing is worse # than either answer on its own. # # What this does NOT change: nothing above about search. Training tokens never # controlled indexing, and Googlebot, bingbot and Applebot are unaffected. # # What it does NOT permit: republishing articles in full, or scraping at a rate # that degrades the site. Both are still refused in /terms/, and permission to # mine is not permission to reprint. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=yes Disallow: /ghost/ Disallow: /email/ Disallow: /members/api/comments/counts/ Disallow: /r/ Disallow: /webmentions/receive/ Disallow: /.ghost/analytics/api/ # ── No crawler is turned away ─────────────────────────────────────────────── # # Every bot is allowed, on the owner's instruction, and that includes the ones # that collect training data. There are no per-agent Disallow groups in this # file any more. Google-Extended and Applebot-Extended were the last two and # they are gone with the rest. # # What that changed, precisely, so nobody has to guess later: # # Search visibility nothing. Blocking GPTBot or Google-Extended never cost # a single position — those tokens control training, not # indexing. Googlebot, bingbot and Applebot were allowed # before this change and are allowed after it. # # Training GPTBot, ClaudeBot, CCBot, Bytespider, Google-Extended # and Applebot-Extended all honour a Disallow here, and # there is no longer one to honour. What remains is the # declaration, not a block. # # THERE IS NO LONGER A RESERVATION TO ENFORCE, AND THAT IS THE POINT. # # Two decisions were made here, and they are separate ones: # # Access every crawler may fetch every public page. Decided first. # Licensing the content may be used for training. Decided after, and # deliberately — it is not a side effect of the first. # # Before this, the site allowed access while reserving the training right, which # is a coherent position and was the position for months. It is simply no longer # the one held. # # What did not change: republishing articles in full is still refused, and so is # scraping at a rate that degrades the site for readers. Both are in /terms/, # and neither depends on the training question. # ── Notes on what is NOT here, and why ────────────────────────────────────── # # GPTBot, CCBot, ClaudeBot, Bytespider and the rest of the training crawlers # have no Disallow group. Until 16 August 2026 they were refused at Cloudflare's # edge instead; that block is gone, and it was removed on purpose rather than # reinstated here. # # Worth recording how that block disappeared, because nothing announced it. It # lived in a Cloudflare dashboard setting, this file assumed it was working, and # a re-test on 16 August found GPTBot, CCBot, ClaudeBot and Bytespider all # receiving 200 — while this file still described them as refused. # # The lesson outlived the policy: anything the site claims about itself belongs # in a file that ships with a deploy, not in a dashboard toggle that can be # changed by someone else and reported by nobody. Re-test the agents named here # before trusting any sentence in this file about what reaches the site. # # OAI-SearchBot, ChatGPT-User, PerplexityBot and DuckAssistBot are AI search # agents. They fetch a page because a reader asked something, then cite and link # back, which is referral traffic and exactly what we want. They are allowed by # the * group above. # # They were blocked at the edge until recently, which meant this site was absent # from ChatGPT search, Perplexity and DuckDuckGo AI answers no matter what this # file said. Re-tested on 16 August 2026: all four now receive 200. Nothing here # needs changing to keep it that way — but if reach outside the Philippines ever # drops off a cliff again, test those four user agents first. # # Googlebot, bingbot, Applebot and facebookexternalhit were tested the same day # and all received 200. # # Mediapartners-Google and AdsBot-Google are allowed by the * group and were # confirmed reaching the site. Blocking them would not reduce advertising, it # would make the advertising irrelevant and fail an AdSense review.