{"id":28025422,"url":"https://github.com/johannesvc/data-science-portfolio","last_synced_at":"2026-05-11T07:04:38.072Z","repository":{"id":291252020,"uuid":"977036806","full_name":"JohannesVC/data-science-portfolio","owner":"JohannesVC","description":"A curated portfolio of applied data science projects focused on machine learning, NLP, and social impact.","archived":false,"fork":false,"pushed_at":"2025-05-03T11:11:15.000Z","size":17438,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-05-11T04:25:19.698Z","etag":null,"topics":["academic-portfolio","data-science","deep-learning","keras","machine-learning","media-bias","nlp","pandas","scikit-learn"],"latest_commit_sha":null,"homepage":"https://johannes.vc","language":"HTML","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/JohannesVC.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-05-03T09:28:03.000Z","updated_at":"2025-05-03T11:11:18.000Z","dependencies_parsed_at":"2025-05-03T12:32:23.816Z","dependency_job_id":null,"html_url":"https://github.com/JohannesVC/data-science-portfolio","commit_stats":null,"previous_names":["johannesvc/data-science-portfolio"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/JohannesVC/data-science-portfolio","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JohannesVC%2Fdata-science-portfolio","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JohannesVC%2Fdata-science-portfolio/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JohannesVC%2Fdata-science-portfolio/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JohannesVC%2Fdata-science-portfolio/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/JohannesVC","download_url":"https://codeload.github.com/JohannesVC/data-science-portfolio/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JohannesVC%2Fdata-science-portfolio/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":278551498,"owners_count":26005386,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-06T02:00:05.630Z","response_time":65,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["academic-portfolio","data-science","deep-learning","keras","machine-learning","media-bias","nlp","pandas","scikit-learn"],"created_at":"2025-05-11T04:22:34.691Z","updated_at":"2025-10-06T02:44:01.245Z","avatar_url":"https://github.com/JohannesVC.png","language":"HTML","funding_links":[],"categories":[],"sub_categories":[],"readme":"## Data Science Portfolio  \n**Johannes Van Cauwenberghe**  \nLondon, UK — MSc Data Science (University of London)\n\nThis portfolio presents a selection of applied data science projects across domains including natural language processing, causal inference, urban analytics, and ethical AI. Each project folder contains the codebase (`.ipynb` or `.qmd`), dependencies (`requirements.txt`), and a corresponding PDF or HTML report. The goal is not only to demonstrate technical proficiency, but to highlight *how data can inform public-interest questions*---from sustainability and urban policy to media bias and ethics.\n\n---\n\n## Featured Projects\n\n### Optimising News Recommenders Beyond Accuracy (final project)\n**[Optimising Beyond Accuracy: Tuning for Diversity and Novelty in Attention-based News Recommenders](https://github.com/JohannesVC/OptimisingBeyondAccuracy)** \nA deep learning–based recommender system for news articles that balances accuracy with diversity and novelty, using the MIND dataset and attention-based neural architectures. Regularisation terms for Intra-List Diversity and Surprisal are embedded directly in the model’s loss function.\n📄 [Read report (ResearchGate)](https://www.researchgate.net/publication/391382172_Optimising_Beyond_Accuracy_Tuning_for_Diversity_and_Novelty_in_Attention-based_News_Recommenders)\n\n---\n\n### Visualising Urban Change  \n**[Visualising Urban Change: The Impact of Traffic Calming Measures on Symbolic Capital and Socio-Economic Dynamics in London](./visualising-symbolic-capital-London)**  \nA spatial analysis of London’s Low-Traffic Neighbourhoods (LTNs) and their relationship to \"symbolic capital\" retail (boutiques, bakeries, etc.), using Companies House, OpenStreetMap, income/demographic data from statistical agencies, and Voronoi diagrams.\n📄 [Read report (PDF)](./visualising-symbolic-capital-London/report-no-code.pdf)\n\n---\n\n### Sustainable Finance under Uncertainty  \n**[Bayesian Investment Analysis](./bayesian-investment-analysis)**  \nA decision-support system for green investors using a custom Bayesian Network in `pgmpy`, predicting environmental and financial viability from policy and market factors.  \n📄 [Read report (PDF)](./bayesian-investment-analysis/report.pdf)\n\n---\n\n### Political Bias Detection in News Media  \n**[Political Bias in News: A Feature-Weighted Classifier](./nlp-political-bias-in-news)**  \nClassified news articles as left- or right-leaning using the Log-Odds Ratio with Informative Dirichlet Prior (LOR-IDP). Combined AllSides and NewsCatcher data for a bias-aware pipeline.  \n📄 [Read report (PDF)](./nlp-political-bias-in-news/report.pdf)\n\n---\n\n### NLP at Scale with Common Crawl  \n**[Mapping Corporate Ethics: Clustering UK Companies’ Values Using NLP](./nlp-common-crawl-corporate-ethics)**  \nClustered mission statements from UK firms scraped from the Common Crawl using PySpark, TF-IDF, and topic modelling.  \n📄 [Read report (PDF)](./nlp-common-crawl-corporate-ethics/report.pdf)\n\n---\n\n### Systematic Deep Learning for News Classification  \n**[A Systematic Exploration of News Classification Using Neural Networks](./deep-learning-for-news-classification)**  \nBenchmarked various deep learning models across different news categories. Emphasised reproducibility and interpretability, with structured hyperparameter tuning and model comparison using Keras 3.8.  \n🌐 [Read report (HTML)](https://storage.googleapis.com/data-science-portfolio/deep-learning-for-news-classification/report.html)\n\n---\n\n### Multimodal Bias Classification  \n**[Multimodal Architectures for Bias Classification in News](./deep-learning-for-multimodal-bias-classification)**  \nIntegrated textual and metadata features (e.g. source, region, headline) using hybrid deep learning models. Demonstrated that combining non-textual features improves political bias prediction, especially for under-represented classes.  \n🌐 [Read report (HTML)](https://storage.googleapis.com/data-science-portfolio/deep-learning-for-multimodal-bias-classification/report.html)\n\n---\n\n### Quantifying Indifference in UK Media  \n**[Quantifying Indifference Towards Palestinian Suffering Across UK News Sources](./quantifying-indifference-towards-palestinian-suffering)**  \nAnalysed over 4,000 articles from UK news outlets (including 744 scraped from BBC News) using NER, sentiment scoring, and date-level casualty records to uncover disparities in how Palestinian vs. Israeli suffering is reported. Offers an interactive data-driven critique of media framing and moral distance, particularly of the BBC.  \n🌐 [Read report (HTML)](https://storage.cloud.google.com/data-science-portfolio/quantifying-indifference-towards-palestinian-suffering/report.html)\n\n---\n\n## Navigation and Reproducibility\n\n- Each folder contains its own `README.md` for full context.\n- Environments are reproducible via `requirements.txt`.\n- Reports are provided as PDFs or self-contained HTMLs.\n- Please note that some datasets are not uploaded due to size or licensing.\n\n---\n\n## 📬 Contact  \n\nFeel free to reach out via [LinkedIn](https://www.linkedin.com/in/johannesvc/) or email (johannes.vc@hotmail.com) for collaborations, freelance data work, or research inquiries.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjohannesvc%2Fdata-science-portfolio","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjohannesvc%2Fdata-science-portfolio","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjohannesvc%2Fdata-science-portfolio/lists"}