{"id":31476749,"url":"https://github.com/parthapray/disease-symptom-knowledge-database_flattening","last_synced_at":"2025-10-02T01:41:16.680Z","repository":{"id":306907500,"uuid":"1027632546","full_name":"ParthaPRay/Disease-Symptom-Knowledge-Database_Flattening","owner":"ParthaPRay","description":"This repo shows the coding organizing symptoms in comma separated rows flattening for given disease from Disease-Symptom Knowledge Database based github link","archived":false,"fork":false,"pushed_at":"2025-07-28T10:09:51.000Z","size":110,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2025-07-28T11:39:46.650Z","etag":null,"topics":["disease-symptom","knowledge-database"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ParthaPRay.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-07-28T09:51:46.000Z","updated_at":"2025-07-28T10:09:55.000Z","dependencies_parsed_at":"2025-07-28T11:39:48.714Z","dependency_job_id":"dff9dbd8-d937-4a2d-b079-a33cf0940976","html_url":"https://github.com/ParthaPRay/Disease-Symptom-Knowledge-Database_Flattening","commit_stats":null,"previous_names":["parthapray/disease-symptom-knowledge-database"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/ParthaPRay/Disease-Symptom-Knowledge-Database_Flattening","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ParthaPRay%2FDisease-Symptom-Knowledge-Database_Flattening","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ParthaPRay%2FDisease-Symptom-Knowledge-Database_Flattening/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ParthaPRay%2FDisease-Symptom-Knowledge-Database_Flattening/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ParthaPRay%2FDisease-Symptom-Knowledge-Database_Flattening/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ParthaPRay","download_url":"https://codeload.github.com/ParthaPRay/Disease-Symptom-Knowledge-Database_Flattening/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ParthaPRay%2FDisease-Symptom-Knowledge-Database_Flattening/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":277942696,"owners_count":25903112,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-01T02:00:09.286Z","response_time":88,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["disease-symptom","knowledge-database"],"created_at":"2025-10-02T01:41:15.429Z","updated_at":"2025-10-02T01:41:16.672Z","avatar_url":"https://github.com/ParthaPRay.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Disease-Symptom Data Cleaner and Flattener\n\nThis repository provides a Python script for cleaning, normalizing, and flattening disease-symptom datasets [Unified Medical Language System (UMLS)](https://www.nlm.nih.gov/research/umls/index.html) specific [https://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/index.html](https://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/index.html).\nIt is designed to process medical data from Excel files (especially those with composite codes separated by `^`), remove unwanted junk characters, and output a clean, analysis-ready CSV.\n\n## Features\n\n* **Directly loads medical data** from a GitHub-hosted Excel file\n* **Fills missing disease names** using last valid entry (forward fill)\n* **Removes unnecessary columns** (e.g., occurrence counts)\n* **Eliminates rows with missing symptoms**\n* **Handles composite disease codes**: splits codes joined by `^` into separate records\n* **Removes junk characters** (e.g., `Â`) from disease and symptom columns\n* **Exports a clean, flattened CSV** ready for ML or analytics\n\n## Input\n\n\"raw_data.xlsx\" The script loads raw data which is obtained in .xlsx format from [https://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/index.html](https://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/index.html) from below github link:\n\n```\nhttps://raw.githubusercontent.com/anujdutt9/Disease-Prediction-from-Symptoms/master/notebook/dataset/raw_data.xlsx\n```\n\n## Output\n\nA cleaned and flattened CSV named:\n\n```\nflattened_url.csv\n```\n\nEach row contains a single disease code and its associated symptoms.\n\n## Usage\n\n1. **Clone this repository** or copy the script into your project.\n\n2. Ensure you have [Python 3.9+](https://www.python.org/downloads/) and the following packages installed:\n\n   ```bash\n   pip install pandas\n   ```\n\n3. **Run the script**:\n\n   ```bash\n   python flatten.py\n   ```\n\n\n\n4. **Result:**\n   You’ll get a `flattened_url.csv` file in your working directory.\n\n## Example\n\n**Input disease string:**\n\n```\nUMLS:C0376358_malignant neoplasm of prostate^UMLS:C0600139_carcinoma prostate^UMLS:C0600159_carcinoma digitailis\n```\n\n**Becomes three rows:**\n\n```\nUMLS:C0376358_malignant neoplasm of prostate, symptom_1, symptom_2, ...\nUMLS:C0600139_carcinoma prostate, symptom_1, symptom_2, ...\nUMLS:C0600159_carcinoma digitailis, symptom_1, symptom_2, ...\n```\n\n## Full Script\n\n```python\nimport pandas as pd\nimport re\n\n# Download Excel file directly from GitHub RAW link\nfile_url = 'https://raw.githubusercontent.com/anujdutt9/Disease-Prediction-from-Symptoms/master/notebook/dataset/raw_data.xlsx'\ndf = pd.read_excel(file_url)\n\n# Fill down Disease names\ndf['Disease'] = df['Disease'].ffill()\n\n# Drop 'Count of Disease Occurrence' if exists\nif 'Count of Disease Occurrence' in df.columns:\n    df = df.drop(columns=['Count of Disease Occurrence'])\n\n# Remove rows where Symptom is NaN\ndf = df[df['Symptom'].notna()]\n\n# Replace '^' with ',' in symptoms\ndf['Symptom'] = df['Symptom'].str.replace('^', ',', regex=False)\n\n# --- Remove junk symbols like Â etc. from Disease and Symptom ---\ndef clean_text(text):\n    # Remove any character that is not printable ASCII or common punctuation/space\n    return re.sub(r'[^\\x20-\\x7E]', '', str(text))\n\ndf['Disease'] = df['Disease'].apply(clean_text)\ndf['Symptom'] = df['Symptom'].apply(clean_text)\n\n# Prepare new rows for disease splits\nexpanded_rows = []\n\nfor disease, group in df.groupby('Disease', sort=False):\n    symptoms = ','.join(group['Symptom'])\n    disease_codes = [d.strip() for d in disease.split('^')]\n    for code in disease_codes:\n        expanded_rows.append({'Disease': code, 'Symptom': symptoms})\n\n# Create result DataFrame preserving order\nresult = pd.DataFrame(expanded_rows)\n\n# (Optional) Clean junk symbols from the final result again, just in case\nresult['Disease'] = result['Disease'].apply(clean_text)\nresult['Symptom'] = result['Symptom'].apply(clean_text)\n\n# Save to CSV\noutput_csv = 'flattened_url.csv'\nresult.to_csv(output_csv, index=False)\n\nresult.head()\n```\n\n### Raw data to flattended data\n\nBefore Raw data\n\n\u003cimg width=\"1076\" height=\"234\" alt=\"image\" src=\"https://github.com/user-attachments/assets/ba927d36-f878-47be-9c2e-2882bc5a7152\" /\u003e\n\nAfter\n\nFlattened data (Comma separated)\n\n\u003cimg width=\"1796\" height=\"23\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5747fff8-c3e3-411d-8ada-554ac35a8d7e\" /\u003e\n\n\n   ```bash\nUMLS:C0008031_pain chest,UMLS:C0392680_shortness of breath,UMLS:C0012833_dizziness,UMLS:C0004093_asthenia,UMLS:C0085639_fall,UMLS:C0039070_syncope,UMLS:C0042571_vertigo,UMLS:C0038990_sweat,UMLS:C0700590_sweating increased,UMLS:C0030252_palpitation,UMLS:C0027497_nausea,UMLS:C0002962_angina pectoris,UMLS:C0438716_pressure chest\n   ```\n\n\n\n## License\n\nMIT\n\n\n## References\n1. https://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/index.html\n2. https://github.com/anujdutt9/Disease-Prediction-from-Symptoms\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fparthapray%2Fdisease-symptom-knowledge-database_flattening","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fparthapray%2Fdisease-symptom-knowledge-database_flattening","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fparthapray%2Fdisease-symptom-knowledge-database_flattening/lists"}