{"id":22116147,"url":"https://github.com/blacksujit/youtube-analysis","last_synced_at":"2025-03-24T05:30:36.693Z","repository":{"id":257005901,"uuid":"857066312","full_name":"Blacksujit/Youtube-Analysis","owner":"Blacksujit","description":"This project is designed to classify YouTube comments as  toxic  or  non-toxic  using  BERT (Bidirectional Encoder Representations from Transformers). By fine-tuning a pre-trained BERT model, we leverage state-of-the-art NLP capabilities to identify harmful content in online conversations. ","archived":false,"fork":false,"pushed_at":"2024-10-10T20:46:49.000Z","size":749,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-01-29T11:22:43.157Z","etag":null,"topics":["bert","bert-model","bertmodel","embeddings","fine","generator","tuning","wordcloud","words"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Blacksujit.png","metadata":{"files":{"readme":"Readme.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-09-13T18:26:52.000Z","updated_at":"2024-11-14T15:24:43.000Z","dependencies_parsed_at":null,"dependency_job_id":"3c43d5e7-cfc1-4185-9606-883fec7e88c9","html_url":"https://github.com/Blacksujit/Youtube-Analysis","commit_stats":null,"previous_names":["blacksujit/youtube_toxic_comments_classification_model","blacksujit/youtube-analysis"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Blacksujit%2FYoutube-Analysis","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Blacksujit%2FYoutube-Analysis/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Blacksujit%2FYoutube-Analysis/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Blacksujit%2FYoutube-Analysis/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Blacksujit","download_url":"https://codeload.github.com/Blacksujit/Youtube-Analysis/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":245216297,"owners_count":20579136,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["bert","bert-model","bertmodel","embeddings","fine","generator","tuning","wordcloud","words"],"created_at":"2024-12-01T12:19:34.388Z","updated_at":"2025-03-24T05:30:36.664Z","avatar_url":"https://github.com/Blacksujit.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# 🎯 YouTube Toxic Comment Classification Using BERT\n\n![image](https://github.com/user-attachments/assets/b5fd4135-0f9c-4b6c-9e20-93a18bdcc2b9)\n\n\n## 🚀 Overview\n\nThis project is designed to classify YouTube comments as **toxic** or **non-toxic** using **BERT** (Bidirectional Encoder Representations from Transformers). By fine-tuning a pre-trained BERT model, we leverage state-of-the-art NLP capabilities to identify harmful content in online conversations. \n\nThe model is trained to identify toxicity, which is crucial for creating safer and more respectful online platforms. This project can be extended for different toxicity levels or used in moderation tools.\n\n---\n\n## 🔎 Demo\n\nHere's a quick demo of the toxicity classification tool:\n\n![Toxicity Prediction Demo](https://user-images.githubusercontent.com/12345678/toxicity-demo.gif) \u003c!-- Add a GIF demo or link to a video here --\u003e\n\n---\n\n## 🗂️ Table of Contents\n\n1. [Dataset](#dataset)\n2. [Installation](#installation)\n3. [Data Preprocessing](#data-preprocessing)\n4. [Model Architecture](#model-architecture)\n5. [Training](#training)\n6. [Evaluation](#evaluation)\n7. [Results](#results)\n8. [Usage](#usage)\n9. [Future Improvements](#future-improvements)\n10. [License](#license)\n\n---\n\n## 📊 Dataset\n\nThe dataset consists of YouTube comments with corresponding labels indicating whether the comment is **toxic** (`1`) or **non-toxic** (`0`). This binary classification problem is aimed at improving content moderation in online discussions.\n\n- **Columns**:\n  - `comment_id`: Unique identifier for each comment.\n  - `content`: Text of the comment.\n  - `label`: `1` for toxic comments, `0` for non-toxic comments.\n---\n\n## 💻 Installation\n\nTo get started with this project, follow these instructions:\n\n### Prerequisites\n\n- Python 3.7+\n- PyTorch 1.6+\n- Hugging Face Transformers library\n- CUDA-enabled GPU (optional but recommended)\n\n### Setup\n\n1. **Clone the Repository**:\n    ```bash\n    git clone https://github.com/your-repository/youtube-toxic-comment-classification.git\n    cd youtube-toxic-comment-classification\n    ```\n\n\n---\n\n## 🛠️ Data Preprocessing\n\nThe raw comments are preprocessed before being fed into the BERT model. This includes:\n\n- Removing URLs, special characters, and extra spaces.\n- Converting all text to lowercase.\n- Tokenizing using BERT tokenizer, which converts the text into input tokens compatible with BERT.\n\n\n\n## 🧠 Model Architecture:\n\nBERT (Bidirectional Encoder Representations from Transformers)\nThis project uses BERT to handle the NLP task of toxic comment detection. BERT is a transformer-based model that understands the context of words in sentences, making it highly effective for text classification tasks.\n\n## Key Components:\n\nTokenizer: Converts sentences into token IDs.\n\nBERT for Sequence Classification: Pre-trained BERT model fine-tuned for binary classification.\n\n\n## 📊 Results:\n\nAfter fine-tuning BERT, we achieved the following performance metrics:\n\n1.) Accuracy: ~90%\n\n2.) F1 Score: ~0.85\n\n3.) Precision: ~0.88\n\n4.) Recall: ~0.83\n\n\n## 📈 Future Improvements\n\n1.) Multi-label Classification: Extend the model to classify different types of toxicity (e.g., hate speech, threats, etc.).\n\n2.) Data Augmentation: Generate synthetic examples to address class imbalance.\n\n3.) Model Optimization: Experiment with models like DistilBERT for faster performance.\n\n4.) Multi-language Support: Expand to detect toxicity in different languages.\n\n## 📜 License:\n\n\nThis project is licensed under the MIT License. See the LICENSE file for more details.\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fblacksujit%2Fyoutube-analysis","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fblacksujit%2Fyoutube-analysis","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fblacksujit%2Fyoutube-analysis/lists"}