{"id":13899882,"url":"https://github.com/healsdata/ai-training-opt-out","last_synced_at":"2025-07-17T17:32:28.151Z","repository":{"id":68601613,"uuid":"597809896","full_name":"healsdata/ai-training-opt-out","owner":"healsdata","description":"Known tags and settings suggested to opt out of having your content used for AI training.","archived":false,"fork":false,"pushed_at":"2024-06-21T00:53:54.000Z","size":41,"stargazers_count":116,"open_issues_count":0,"forks_count":3,"subscribers_count":6,"default_branch":"main","last_synced_at":"2024-08-07T19:52:41.465Z","etag":null,"topics":["ai","meta","opt-out","robots-txt"],"latest_commit_sha":null,"homepage":"","language":"HTML","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"unlicense","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/healsdata.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-02-05T17:52:15.000Z","updated_at":"2024-08-07T02:50:28.000Z","dependencies_parsed_at":"2023-02-21T09:46:05.489Z","dependency_job_id":"2d07a847-0f0f-4d6a-8f47-6e0c9af531c7","html_url":"https://github.com/healsdata/ai-training-opt-out","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/healsdata%2Fai-training-opt-out","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/healsdata%2Fai-training-opt-out/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/healsdata%2Fai-training-opt-out/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/healsdata%2Fai-training-opt-out/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/healsdata","download_url":"https://codeload.github.com/healsdata/ai-training-opt-out/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":226286978,"owners_count":17600701,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","meta","opt-out","robots-txt"],"created_at":"2024-08-06T19:01:13.446Z","updated_at":"2024-11-25T06:31:09.018Z","avatar_url":"https://github.com/healsdata.png","language":"HTML","funding_links":[],"categories":["Open-Source Repos \u0026 Tools","HTML"],"sub_categories":["AI Crawler Management"],"readme":"# AI Training Opt Out\nKnown tags and settings suggested to opt out of having your content used for AI training.\n\n# Contents\n\n* [**robots.txt**](/robots.txt) A copy-and-paste collection of tags to add to your own robots.txt. (You can automate generation of this file with [darkvisitors.com](https://darkvisitors.com/))\n* [**meta-tags.html**](/meta-tags.html) A copy-and-paste collection of tags to add to your own `\u003chead\u003e`\n* [**headers.txt**](/headers.txt) HTTP headers you can add to your responses. This is more more involved and installation is outside the scope of this document.\n* [**ai.txt**](/ai.txt) An alternative to robots.txt created by Spawning, the company behind [haveibeentrained.com](https://haveibeentrained.com/).\n* [**ip-ranges.txt**](/ip-ranges.txt) Known IP ranges for AI crawlers. These will change over time, so links to the canonical source is included.\n* [**tdmrep.json**](/.well-known/tdmrep.json) A Web protocol, capable of expressing the reservation of rights relative to text \u0026 data mining (TDM)\n\n# Other Opt-Outs\n\n* **OpenAI** (Includes ChaGPT and DALL·E): You can opt-out of having your input and output to their services used to train by emailing your organization ID to [support@openai.com](mailto:support@openai.com). *Note: This doesn't include any data they scraped to train their model.*\n* **StabilityAI**: Stable Diffusion 3 will honor opt-out requests on [haveibeentrained.com](https://haveibeentrained.com/).\n* **AWS**: \"AWS may be using your data to train its AI models, and you may have unwittingly consented to it. Prepare to jump through a series of complex hoops to stop it.\" -- [How to Stop Feeding AWS’s AI With Your Data](https://www.lastweekinaws.com/blog/How-to-Stop-Feeding-AWSs-AI-With-Your-Data/)\n* **Substack** \"If you do NOT want your publication to be used to train AI, open your publication, go to Settings \u003e Publication details and switch it on.\"\n* **[Wordpress](https://wordpress.com/support/privacy-settings/#prevent-third-party-sharing)** and **[Tumblr](https://help.tumblr.com/hc/en-us/articles/115011611747-Privacy-options#01H692KHGF5N3SVHDV02P5W34P)** are both opt-out for your post content.\n* **The Stack** Find your repo(s) on [Am I in The Stack?](https://huggingface.co/spaces/bigcode/in-the-stack) and then click Opt-Out at the bottom to open a request.\n\n# References\n\n* [How to Block ChatGPT From Using Your Website Content](https://www.searchenginejournal.com/how-to-block-chatgpt-from-using-your-website-content/478384/)\n* [All Deviations Are Opted Out of AI Datasets](https://www.deviantart.com/team/journal/UPDATE-All-Deviations-Are-Opted-Out-of-AI-Datasets-934500371)\n* [OpenAI Terms of Use](https://openai.com/terms/)\n* [Stability AI plans to let artists opt out of Stable Diffusion 3 image training](https://arstechnica.com/information-technology/2022/12/stability-ai-plans-to-let-artists-opt-out-of-stable-diffusion-3-image-training/)\n* [Stop AI Data Mining in its Tracks with AI.txt](https://site.spawning.ai/spawning-ai-txt)\n* [Sites scramble to block ChatGPT web crawler after instructions emerge](https://arstechnica.com/information-technology/2023/08/openai-details-how-to-keep-chatgpt-from-gobbling-up-website-data/)\n* [An update on web publisher controls](https://blog.google/technology/ai/an-update-on-web-publisher-controls/) -- Google's VP of Trust\n* [Dark Visitors: A List of Known AI Agents on the Internet](https://darkvisitors.com/) \n* [TDM Reservation Protocol (TDMRep)](https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20240202/)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhealsdata%2Fai-training-opt-out","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhealsdata%2Fai-training-opt-out","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhealsdata%2Fai-training-opt-out/lists"}