# Search engines: unrestricted. # # This group stays FIRST on purpose. A conforming crawler obeys only its most # specific matching User-agent group, so the archiver groups below do not apply # to Googlebot, Bingbot or any other search crawler. Order does not matter to a # conforming parser, but a buggy one that takes the first matching group gets # the permissive rules here rather than a Disallow. Failing open for an archiver # is a small miss; failing closed for Googlebot would be a disaster. User-agent: * Allow: / Disallow: /api/ # Web archiving crawlers. # # archive.org_bot is the Internet Archive's current crawler and the only entry # here that matters in practice. ia_archiver was Alexa Internet's crawler, which # seeded the early Internet Archive but shut down in 2022; it is listed because # some mirrors still send it, not because it is active. # # Known limit, stated rather than implied: the Internet Archive has largely # stopped honoring robots.txt for archiving since 2017, so treat this as a # request rather than a block. Removing pages already in the Wayback Machine # requires emailing info@archive.org. Enforcing a block would require rejecting # these user agents at the edge, which this file cannot do. User-agent: archive.org_bot Disallow: / User-agent: ia_archiver Disallow: / User-agent: ia_archiver-web.archive.org Disallow: / User-agent: Wayback Disallow: / User-agent: heritrix Disallow: / Sitemap: https://www.uncutgems.com/sitemap.xml