{"id":19383777,"url":"https://github.com/mariaorabi/data-mining-disease-tweets-analysis","last_synced_at":"2026-06-20T14:31:48.164Z","repository":{"id":208309525,"uuid":"721313012","full_name":"Mariaorabi/Data-Mining-Disease-Tweets-Analysis","owner":"Mariaorabi","description":"Analyze tweets related to diseases using data mining techniques to derive insights and patterns.","archived":false,"fork":false,"pushed_at":"2023-11-20T20:32:51.000Z","size":22780,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-02-24T17:18:03.157Z","etag":null,"topics":["data-analysis-python","data-anlysis","data-mining","data-processing","disease","juypter-notebook","nlp","python","tweet-analysis"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Mariaorabi.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2023-11-20T19:54:17.000Z","updated_at":"2023-11-20T20:45:06.000Z","dependencies_parsed_at":null,"dependency_job_id":"2a25cb4e-8324-4d40-86f4-e5d4105d2dd9","html_url":"https://github.com/Mariaorabi/Data-Mining-Disease-Tweets-Analysis","commit_stats":{"total_commits":5,"total_committers":1,"mean_commits":5.0,"dds":0.0,"last_synced_commit":"f3283e46b71cd3e74eaeaaedc974f82b6f671a57"},"previous_names":["mariaorabi/data-mining-disease-tweets-analysis"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/Mariaorabi/Data-Mining-Disease-Tweets-Analysis","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Mariaorabi%2FData-Mining-Disease-Tweets-Analysis","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Mariaorabi%2FData-Mining-Disease-Tweets-Analysis/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Mariaorabi%2FData-Mining-Disease-Tweets-Analysis/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Mariaorabi%2FData-Mining-Disease-Tweets-Analysis/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Mariaorabi","download_url":"https://codeload.github.com/Mariaorabi/Data-Mining-Disease-Tweets-Analysis/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Mariaorabi%2FData-Mining-Disease-Tweets-Analysis/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34573729,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-20T02:00:06.407Z","response_time":98,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["data-analysis-python","data-anlysis","data-mining","data-processing","disease","juypter-notebook","nlp","python","tweet-analysis"],"created_at":"2024-11-10T09:27:50.227Z","updated_at":"2026-06-20T14:31:48.146Z","avatar_url":"https://github.com/Mariaorabi.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# Disease Tweets Analysis\n\n## Introduction\nIn this project, I aim to analyze tweets related to four diseases: AIDS/HIV, cancer, Corona (COVID-19), and diabetes. The analysis involves preprocessing the tweets, extracting key information, and deriving insights through natural language processing (NLP) techniques.\n\n## Data Files\n- Four files containing tweets related to the diseases, each categorized by specific keywords.\n\n## Preprocessing\n- Remove user mentions and the word \"LINK@\" from the tweets.\n- Preserve social media symbols as individual words to avoid separation into meaningless signs.\n\n## Parts of Speech Analysis\n### a. Grammatical Analysis using Spacy\n- Utilize the Spacy package for grammatical analysis.\n- Ignore stop words, perform lemmatization, and filter words based on the English dictionary.\n- \n### b. Most Common Words \n- Report the 20 most common words for each disease.\n- Discuss the relevance and value of the results.\n\n### c. Most Common Adjective Words \n- Identify the 20 most common adjective words for each disease.\n- Evaluate the logical and meaningful aspects of the results.\n\n### d. Most Common Verbs\n- Extract the 20 most common verbs for each disease.\n- Provide insights into the logical and meaningful implications of the results.\n\n### e. Most Common Noun Words \n- Analyze the 20 most common noun words for each disease.\n- Discuss the sense and meaningfulness of the results.\n\n## Parsing \n### a. Dependency Parsing Function\n- Implement a function for checking words directly related to the disease names in tweets.\n- Convert verbs to their base forms (lemmas) and summarize the 10 common verbs along with their relative frequency for each disease.\n- Discuss differences between diseases and compare with results from section d of question 2.\n\n### b. Most Common Adjective Heads for Disease Names \n- Find the adjectives for which the disease name is the most common head.\n- Discuss differences between diseases and compare with results from section c of question 2.\n\n## Conclusion\n- Summarize key findings from the analysis.\n- Reflect on any variations between diseases and the reasons behind them.\n\n## How to Submit\n- The code must be submitted in a Jupyter notebook file.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmariaorabi%2Fdata-mining-disease-tweets-analysis","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmariaorabi%2Fdata-mining-disease-tweets-analysis","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmariaorabi%2Fdata-mining-disease-tweets-analysis/lists"}