{"id":21898607,"url":"https://github.com/zzarif/multilabel-scifi-tags-classifier","last_synced_at":"2026-04-29T22:39:41.229Z","repository":{"id":242588778,"uuid":"809820779","full_name":"zzarif/Multilabel-Scifi-Tags-Classifier","owner":"zzarif","description":"Classify 160 different tags from scifi and fantasy questions.","archived":false,"fork":false,"pushed_at":"2024-08-07T08:15:48.000Z","size":82981,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-10-06T20:47:58.736Z","etag":null,"topics":["blurr","f1-score","fantasy","fastai","flask","huggingface","multilabel-text-classification","nlp","onnx-inference","questions","scifi","scraping","selenium","stackapi","stackexchange","tags","vercel"],"latest_commit_sha":null,"homepage":"https://multilabel-scifi-tags-classifier.vercel.app","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/zzarif.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-06-03T14:07:56.000Z","updated_at":"2024-08-07T08:15:52.000Z","dependencies_parsed_at":"2024-06-03T23:20:11.177Z","dependency_job_id":"ff0fa99a-9a16-40ad-bd0b-5201a7cdb5c7","html_url":"https://github.com/zzarif/Multilabel-Scifi-Tags-Classifier","commit_stats":null,"previous_names":["zzarif/stackexchange-scifi-tags-classifier"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/zzarif/Multilabel-Scifi-Tags-Classifier","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zzarif%2FMultilabel-Scifi-Tags-Classifier","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zzarif%2FMultilabel-Scifi-Tags-Classifier/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zzarif%2FMultilabel-Scifi-Tags-Classifier/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zzarif%2FMultilabel-Scifi-Tags-Classifier/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/zzarif","download_url":"https://codeload.github.com/zzarif/Multilabel-Scifi-Tags-Classifier/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zzarif%2FMultilabel-Scifi-Tags-Classifier/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":32447292,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-29T22:27:22.272Z","status":"ssl_error","status_checked_at":"2026-04-29T22:10:49.234Z","response_time":110,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["blurr","f1-score","fantasy","fastai","flask","huggingface","multilabel-text-classification","nlp","onnx-inference","questions","scifi","scraping","selenium","stackapi","stackexchange","tags","vercel"],"created_at":"2024-11-28T14:33:23.962Z","updated_at":"2026-04-29T22:39:41.198Z","avatar_url":"https://github.com/zzarif.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003ch1 align=\"center\"\u003e\n  \u003cbr\u003e\n  \u003c!-- \u003ca href=\"http://www.amitmerchant.com/electron-markdownify\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/amitmerchant1990/electron-markdownify/master/app/img/markdownify.png\" alt=\"Markdownify\" width=\"200\"\u003e\u003c/a\u003e\n  \u003cbr\u003e --\u003e\n  Multilabel Scifi Tags Classifier\n  \u003cbr\u003e\n\u003c/h1\u003e\n\n\u003ch4 align=\"center\"\u003eClassify 160 different tags from scifi and fantasy questions.\u003c/h4\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003c!-- \u003ca href=\"https://badge.fury.io/js/electron-markdownify\"\u003e\n    \u003cimg src=\"https://badge.fury.io/js/electron-markdownify.svg\"\n         alt=\"Gitter\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://gitter.im/amitmerchant1990/electron-markdownify\"\u003e\u003cimg src=\"https://badges.gitter.im/amitmerchant1990/electron-markdownify.svg\"\u003e\u003c/a\u003e --\u003e\n  \u003c!-- \u003ca href=\"\"\u003e\n      \u003cimg src=\"https://img.shields.io/badge/website-online-blue.svg\"\u003e\n  \u003c/a\u003e --\u003e\n  \u003ca href=\"https://github.com/zzarif/Multilabel-Scifi-Tags-Classifier\"\u003e\n    \u003cimg src=\"https://img.shields.io/github/last-commit/zzarif/Multilabel-Scifi-Tags-Classifier\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://www.kaggle.com/datasets/zibranzarif/multilabel-scifi-tags-classifier-dataset\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/dataset-kaggle-blue.svg\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://multilabel-scifi-tags-classifier.vercel.app\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/website-online-red.svg\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://opensource.org/licenses/MIT\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/license-MIT-yellow.svg\"\u003e\n  \u003c/a\u003e\n\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"#-overview\"\u003eOverview\u003c/a\u003e •\n  \u003ca href=\"#%EF%B8%8F-data-collection\"\u003eData Collection\u003c/a\u003e •\n  \u003ca href=\"#-data-preprocessing\"\u003eData Preprocessing\u003c/a\u003e •\n  \u003ca href=\"#-model-training\"\u003eModel Training\u003c/a\u003e •\n  \u003ca href=\"#-model-compression-and-onnx-inference\"\u003eCompression\u003c/a\u003e •\n  \u003ca href=\"#-model-deployment\"\u003eModel Deployment\u003c/a\u003e •\n  \u003ca href=\"#-web-deployment\"\u003eWeb Deployment\u003c/a\u003e •\n  \u003ca href=\"#%EF%B8%8F-build-from-source\"\u003eBuild from Source\u003c/a\u003e •\n  \u003ca href=\"#%EF%B8%8F-contact\"\u003eContact\u003c/a\u003e\n\u003c/p\u003e\n\n## 📋 Overview\n\nA multi-label text classification model from data collection, model training and deployment. The model can classify **160** different types of question tags from https://scifi.stackexchange.com. The keys of `tag_types_encoded.json` shows the list of question tags.\n\n## 🗂️ Data Collection\n\nData was collected from https://scifi.stackexchange.com/questions, a part of the Stack Exchange network of Q\u0026A sites dedicated to science fiction and fantasy topics. Some popular question tags include Star Wars, Harry Potter, Marvel, DC Comics, Star Trek, Lord of the Rings, and Game of Thrones. The data collection was divided into two steps:\n\n1. **Question URL Scraping:** The scifi question URLs were scraped with `question_url_scraper.py` and the URLs are stored along with the question titles in `question_urls.csv` file. Scroll to [this](#run-the-selenium-scraper) section for details.\n2. **Fetching Question Details:** For each of the question URL in `question_urls.csv`, the question details (title, URL, description, tags) were fetched with `fetch_question_detail.py`. The question details are stored in `question_details.csv` file. Alternatively, `question_detail_scraper.py` could be used to scrape the question details. Scroll to [this](#fetch-question-details) section for details.\n\nIn total, **30,000** scifi question details were collected.\n\n## 🔄 Data Preprocessing\n\nInitially, there were 2095 different question tags in the dataset. After analyzing, it was found that 1935 of them were rare tags (tags that appeared in less than 0.2% of the questions). So, the rare tags were removed. As a result, a very small portion of the questions were void of any tag at all. So those question rows were removed as well. Finally, the dataset had **160** different tags across **27,493** questions.\n\n## 💪 Model Training\n\nThree different models from HuggingFace Transformers were fine-tuned using Fastai and Blurr. All of the models achieved **99%+** accuracy. Following are the list of models:\n\n1. [distilroberta-base](https://huggingface.co/distilbert/distilroberta-base)\n2. [roberta-base](https://huggingface.co/FacebookAI/roberta-base)\n3. [bert-base-uncased](https://huggingface.co/google-bert/bert-base-uncased)\n\nThe model training notebooks can be viewed [here](notebooks/).\n\n## 📦 Model Compression and ONNX Inference\n\nThe trained models required a storage space between **300-500MB**. So, the models were compressed using ONNX quantization which reduced its size between **80-120MB**. Following are the key performance metrics for each of the models and their compressed version respectively:\n\n\u003ctable\u003e\n\u003ctr\u003e\n    \u003cth\u003e\u003c/th\u003e\n    \u003cth\u003edistilroberta-base\u003c/th\u003e\n    \u003cth\u003edistilroberta-base (quantized)\u003c/th\u003e\n    \u003cth\u003eroberta-base\u003c/th\u003e\n    \u003cth\u003eroberta-base (quantized)\u003c/th\u003e\n    \u003cth\u003ebert-base-uncased\u003c/th\u003e\n    \u003cth\u003ebert-base-uncased (quantized)\u003c/th\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n    \u003cth\u003eF1 Score (Micro)\u003c/th\u003e\n    \u003ctd\u003e0.745\u003c/td\u003e\n    \u003ctd\u003e0.745\u003c/td\u003e\n    \u003ctd\u003e0.783\u003c/td\u003e\n    \u003ctd\u003e0.782\u003c/td\u003e\n    \u003ctd\u003e0.715\u003c/td\u003e\n    \u003ctd\u003e0.717\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n    \u003cth\u003eF1 Score (Macro)\u003c/th\u003e\n    \u003ctd\u003e0.278\u003c/td\u003e\n    \u003ctd\u003e0.277\u003c/td\u003e\n    \u003ctd\u003e0.521\u003c/td\u003e\n    \u003ctd\u003e0.518\u003c/td\u003e\n    \u003ctd\u003e0.146\u003c/td\u003e\n    \u003ctd\u003e0.15\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n    \u003cth\u003eSize\u003c/th\u003e\n    \u003ctd\u003e314 MB\u003c/td\u003e\n    \u003ctd\u003e79 MB\u003c/td\u003e\n    \u003ctd\u003e476 MB\u003c/td\u003e\n    \u003ctd\u003e120 MB\u003c/td\u003e\n    \u003ctd\u003e419 MB\u003c/td\u003e\n    \u003ctd\u003e105 MB\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n## 🤗 Model Deployment\n\n`distilroberta-base` (quantized) with **99.37%** accuracy was the final compressed model that was deployed to HuggingFace Spaces Gradio App. The implementation can be found in [deployment](deployment) folder or [here](https://huggingface.co/spaces/zzarif/Multilabel-Scifi-Tags-Classifier).\n\n![SE Scifi Tags Classifier](deployment/hf_model_deployed.png)\n\n## 🌐 Web Deployment\n\nDeveloped a Flask Webapp and deployed to Vercel. It takes scifi and fantasy questions as input and classifies the relevant tags associated with the question via HuggingFace API. The webapp is live [here](https://multilabel-scifi-tags-classifier.vercel.app/).\n\n1. The webapp takes scifi and fantasy questions as input:\n\n![Flask App Scifi Tags Classifier](deployment/web_deployed_model0.png)\n\n2. It utilizes HuggingFace API to classify the relevant tags:\n\n![Flask App Scifi Tags Classifier](deployment/web_deployed_model1.png)\n\n## ⚙️ Build from Source\n\n1. Clone the repo\n\n```bash\ngit clone https://github.com/zzarif/Multilabel-Scifi-Tags-Classifier.git\ncd Multilabel-Scifi-Tags-Classifier/\n```\n\n2. Initialize and activate virtual environment\n\n```bash\nvirtualenv --no-site-packages venv\nsource venv/Scripts/activate\n```\n\n3. Install dependencies\n\n```bash\npip install -r requirements.txt\n```\n\n_Note: Select virtual environment interpreter from_ `Ctrl`+`Shift`+`P`\n\n## Run the Selenium Scraper\n\n```bash\npython scraper/question_url_scraper.py\n```\n\nWait for the script to finish (might take a few hours depending on your network bandwidth). When complete, this will generate [question_urls.csv](data/question_urls.csv) file. This file has **30,000** StackExchange Scifi question titles and URLs. Now, we need to fetch the question details (description, tags, etc.) for each of these question URLs.\n\n## Fetch Question Details\n\nFetching details from **30,000** question URLs is a rather resource intensive task. To efficiently fetch the details we can either request the question details via Stack API (recommended) or we can scrape the details with selenium scraper. (There are other ways too as mentioned [here](https://stackoverflow.com/a/40017359/23817375).)\n\n### Method 1: Request question details via [Stack API](https://api.stackexchange.com/)\n\nThis method explains how we can utilize StackExchange REST APIs to request **30,000** question details. To do so:\n\n1. Register your v2.0 application at [Stackapps](https://stackapps.com/apps/oauth/register) to get an API key.\n2. `deactivate` your active virtual environment.\n3. Open your `venv/Scripts/activate` file and add this line at the end of file (replace `\u003cyour_api_key\u003e` with the API key from Stackapps):\n\n```bash\nexport STACK_API_KEY=\"\u003cyour_api_key\u003e\"\n```\n\n4. Activate the virtual environment again:\n\n```bash\nsource venv/Scripts/activate\n```\n\n5. Now, fetch the question details via Stack API:\n\n```bash\npython stackapi/fetch_question_detail.py\n```\n\nWait for the script to finish (might take a few hours depending on your network bandwidth). The script might get interrupted midway, because the Stack API is [throttled](https://api.stackexchange.com/docs/throttle) to max 10,000 calls per day for registered apps and it only allows for only a limited number of calls within a timeframe. In that case, simply wait and re-run the script the next day and it will resume from where it got interrupted. When complete, this will generate [question_details.csv](data/question_details.csv) file. It has the details (title, url, description, tags) of **30,000** scifi and fantasy questions from StackExchange.\n\n### Method 2: Scrape question details via Selenium Scraper\n\n```bash\npython scraper/question_detail_scraper.py\n```\n\nThis method scrapes question details using `selenium` and `multiprocessing`. This is how the scraping is done:\n\n1. The script reads the [question_urls.csv](data/question_urls.csv) file containing the question URLs.\n2. It divides the URLs into chunks based on the number of CPU cores available.\n3. For each chunk, it creates a separate process to scrape the question details concurrently.\n4. Each process uses `selenium` to navigate to each question URL, scrape the title, description, and tags, and store the data in a list.\n5. After scraping all the questions in a chunk, the process saves the data as a CSV file specific to that chunk.\n6. The script waits for all processes to finish before terminating.\n\nWhen complete, we have to merge all the chunk specific CSV files into one [question_details.csv](data/question_details.csv) file. By utilizing `multiprocessing`, the script can scrape multiple question details simultaneously, improving the overall efficiency of the scraping process. However, this method is, often times, not reliable due to SE's screen-scraping guidelines as mentioned [here](https://meta.stackexchange.com/a/446) and poses the potential risk of IP range ban.\n\n\n## ✉️ Contact\n\n[![LinkedIn](https://img.shields.io/badge/LinkedIn-0077B5?logo=linkedin\u0026logoColor=white)](https://www.linkedin.com/in/zibran-zarif-amio-b82717263/) [![Mail](https://img.shields.io/badge/Gmail-EA4335?logo=gmail\u0026logoColor=fff)](mailto:zibran.zarif.amio@gmail.com)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzzarif%2Fmultilabel-scifi-tags-classifier","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fzzarif%2Fmultilabel-scifi-tags-classifier","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzzarif%2Fmultilabel-scifi-tags-classifier/lists"}