{"id":15158028,"url":"https://github.com/kavayk29/quora-duplicate-question-pair","last_synced_at":"2026-01-21T02:04:11.362Z","repository":{"id":252278042,"uuid":"839963279","full_name":"Kavayk29/Quora-Duplicate-Question-Pair","owner":"Kavayk29","description":"This project improves information retrieval by detecting duplicate question pairs in the Quora dataset using data exploration, text preprocessing, feature engineering, and models like Random Forest and LSTM, aiming to streamline question-answering.","archived":false,"fork":false,"pushed_at":"2024-11-27T14:58:40.000Z","size":29399,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-04-07T14:48:13.386Z","etag":null,"topics":["beautifulsoup4","bilstm","gensim","keras","lstm","matplotlib","numpy","pandas","pytorch","random-forest","seaborn","sklearn","tensorflow","xgboost"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Kavayk29.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-08-08T17:19:04.000Z","updated_at":"2024-11-27T14:58:59.000Z","dependencies_parsed_at":"2024-08-08T20:05:15.501Z","dependency_job_id":"41ee7fe3-6a54-48ca-9ba9-4dc8af136369","html_url":"https://github.com/Kavayk29/Quora-Duplicate-Question-Pair","commit_stats":{"total_commits":17,"total_committers":1,"mean_commits":17.0,"dds":0.0,"last_synced_commit":"6bb3d3ca2af7a4a08315d307b18c827cebface90"},"previous_names":["kavayk29/quora-duplicate-question-pair"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Kavayk29%2FQuora-Duplicate-Question-Pair","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Kavayk29%2FQuora-Duplicate-Question-Pair/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Kavayk29%2FQuora-Duplicate-Question-Pair/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Kavayk29%2FQuora-Duplicate-Question-Pair/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Kavayk29","download_url":"https://codeload.github.com/Kavayk29/Quora-Duplicate-Question-Pair/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247675627,"owners_count":20977376,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["beautifulsoup4","bilstm","gensim","keras","lstm","matplotlib","numpy","pandas","pytorch","random-forest","seaborn","sklearn","tensorflow","xgboost"],"created_at":"2024-09-26T20:21:59.409Z","updated_at":"2026-01-21T02:04:11.337Z","avatar_url":"https://github.com/Kavayk29.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"Project Overview\nThis project aims to detect duplicate question pairs in the Quora dataset. By identifying similar questions, the system can help streamline the question-answering process and improve the efficiency of information retrieval on the platform.\n\nKey Features:\nData Exploration: Load and explore the Quora dataset to understand its structure and characteristics.\nText Preprocessing: Implement various techniques to clean and preprocess the text data, including the removal of HTML tags and special characters.\nFeature Engineering: Extract meaningful features from the text data to improve the model’s ability to detect duplicate questions.\nModeling: Apply machine learning models such as Random Forest and XGBoost, as well as deep learning models like LSTM and BiLSTM, to predict duplicate question pairs.\nEvaluation: Assess model performance using metrics like accuracy, precision, and recall.\nTechnologies Used:\nPython: Core language for data processing and modeling.\nPandas: For handling and manipulating data structures.\nNumpy: For numerical operations and array management.\nSeaborn \u0026 Matplotlib: For data visualization and analysis.\nBeautifulSoup: For text cleaning and preprocessing.\nHow to Use:\nLoad the Dataset: Begin by loading the Quora dataset using the provided code.\nPreprocess the Data: Clean and prepare the text data for modeling.\nTrain the Models: Utilize the provided scripts to train and evaluate different models on the dataset.\nAnalyze Results: Review the model performance metrics and visualizations to understand the results.\nConclusion:\nThis project provides a comprehensive approach to detecting duplicate questions on Quora. By combining data preprocessing, feature engineering, and advanced modeling techniques, it delivers a robust solution for improving information retrieval on the platform.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkavayk29%2Fquora-duplicate-question-pair","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkavayk29%2Fquora-duplicate-question-pair","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkavayk29%2Fquora-duplicate-question-pair/lists"}