{"id":30305543,"url":"https://github.com/kammitama5/data_science_001","last_synced_at":"2026-07-13T20:31:45.319Z","repository":{"id":84573696,"uuid":"73207440","full_name":"kammitama5/Data_Science_001","owner":"kammitama5","description":"A collection of Python Scripts used for DataScience, using Jupyter, etc","archived":false,"fork":false,"pushed_at":"2016-12-13T19:18:00.000Z","size":279,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"master","last_synced_at":"2025-10-20T04:53:16.657Z","etag":null,"topics":["data-science","jupyter-notebook"],"latest_commit_sha":null,"homepage":null,"language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kammitama5.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2016-11-08T16:54:52.000Z","updated_at":"2017-07-08T03:46:40.000Z","dependencies_parsed_at":"2023-03-11T08:01:19.255Z","dependency_job_id":null,"html_url":"https://github.com/kammitama5/Data_Science_001","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/kammitama5/Data_Science_001","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kammitama5%2FData_Science_001","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kammitama5%2FData_Science_001/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kammitama5%2FData_Science_001/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kammitama5%2FData_Science_001/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kammitama5","download_url":"https://codeload.github.com/kammitama5/Data_Science_001/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kammitama5%2FData_Science_001/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35436278,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-13T02:00:06.543Z","response_time":119,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["data-science","jupyter-notebook"],"created_at":"2025-08-17T08:09:51.626Z","updated_at":"2026-07-13T20:31:45.313Z","avatar_url":"https://github.com/kammitama5.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"----------------------\n\n# Data_Science_001\n\n----------------------\n\n1. A collection of Python Scripts used for DataScience, using Jupyter, etc\n2. I am part of a team working on doing data analysis on SEER Cancer data, learning on a weekly basis.\n   I may post some notes and files based my learning for this group.\n3. Also included is some of the exercises from DataCamp's \"Intro to Data Science(using Python)\" and \"Intro to Data Science using R. \n\n\n\n#In Data Science:\n\n1. You usually start with a question to answer -\u003e a problem to solve\n2. You wrangle data -\u003e data acquisition/ data cleaning\n3. You explore data -\u003e build intuition, find patterns\n4. Make conclusions and predicitions (using machine learning, etc)\n5. You communicate your findings (whether through a blog or through visualization).\n  1. Data visualization is useful for communication\n  \n\nYou may need to return to your original question and refine it -\u003e scientific process\n\n1. create hypothesis, use empiricism to observe, make a conclusion\n2. Sometimes not necessarily in order -\u003e sometimes data acquisition happens before question to ask.\n\n\n=======================\n#Categorical Variables:\n\nThere are two types of categorical variables:\n\n\n1. Nominal:\n  -\u003e Categorical variable without an implied order (reminds me of unsorted in Clojure)\n2. Ordinal:\n  -\u003e Have a natural ordering (eg. Small, Medium, Large, Extra-Large)\n  \n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkammitama5%2Fdata_science_001","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkammitama5%2Fdata_science_001","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkammitama5%2Fdata_science_001/lists"}