Sitemap: https://colincogle.name/sitemap.xml ######################################################################## # For the record, I have no issues with AI. By all means, ask Siri or # Cortana to read my content. However, I'm not going to help some rich # asshole, business owner, or capitalist pig profit off of my content # without any kind of kickback (or at least attribution). # # All content on this site is made available under Creative Commons 4.0 # ShareAlike-Attribution (CC-BY-SA) and that's the terms of licensing. # Don't like it? Don't index me. ######################################################################## # Webz.io - Multi-purpose, commercial uses; including LLMs. User-Agent: Omgili Disallow: / User-Agent: Omgilibot Disallow: / # 80legs - some crawler User-Agent: 008 Disallow: / User-Agent: voltron Disallow: / # You're okay for now, Ahrefsbot. #User-Agent: AhrefsBot #Disallow: / # Who asks Alexa about me? User-Agent: AmazonBot Disallow: / # Common Crawl Bot. It was used to train ChatGPT 3, but I do like its # mission statement, so I will allow it. #User-Agent: CCBot #Disallow: / # OpenAI's bot. This is the one used for training data. User-Agent: GPTBot Disallow: / # OpenAI's ChatGPT bot, used to answer user questions. # This is fine by me. #User-Agent: ChatGPT-User #Disallow: / # Anthropic's AI training bot for Claude. # Blocking for now. Might reconsider later. User-Agent: anthropic-ai Disallow: / User-Agent: cohere-ai Disallow: / User-Agent: MJ12bot Disallow: / # I can't figure out what this company is supposed to do. User-Agent: Piplbot Disallow: / # Google Bard training, not to be confused with Googlebot. User-Agent: Google-Extended Disallow: / # Contextual advertising. User-Agent: peer39_crawler Disallow: / # Some other AI thing. User-Agent: PerplexityBot Disallow: / # I don't use them for marketing, so they have no reason to collect my # data. User-Agent: SemrushBot Disallow: / # Screw Elongated Muskrat. If you're using X, switch to Mastodon. User-Agent: Twitterbot Disallow: /