{"id":15288211,"url":"https://github.com/scikit-adaptation/skada","last_synced_at":"2025-04-12T10:50:59.245Z","repository":{"id":208606628,"uuid":"721572834","full_name":"scikit-adaptation/skada","owner":"scikit-adaptation","description":"Domain adaptation toolbox compatible with scikit-learn and pytorch","archived":false,"fork":false,"pushed_at":"2025-02-21T13:55:07.000Z","size":1678,"stargazers_count":112,"open_issues_count":58,"forks_count":21,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-04-12T02:07:01.681Z","etag":null,"topics":["data-shift","domain-adaptation","sklearn"],"latest_commit_sha":null,"homepage":"https://scikit-adaptation.github.io/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"bsd-3-clause","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/scikit-adaptation.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":".github/CONTRIBUTING.md","funding":null,"license":"COPYING","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2023-11-21T10:42:27.000Z","updated_at":"2025-04-10T09:28:53.000Z","dependencies_parsed_at":"2025-04-12T02:17:07.641Z","dependency_job_id":null,"html_url":"https://github.com/scikit-adaptation/skada","commit_stats":null,"previous_names":["scikit-adaptation/skada"],"tags_count":6,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scikit-adaptation%2Fskada","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scikit-adaptation%2Fskada/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scikit-adaptation%2Fskada/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scikit-adaptation%2Fskada/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/scikit-adaptation","download_url":"https://codeload.github.com/scikit-adaptation/skada/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248557844,"owners_count":21124165,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["data-shift","domain-adaptation","sklearn"],"created_at":"2024-09-30T15:44:44.930Z","updated_at":"2025-04-12T10:50:59.237Z","avatar_url":"https://github.com/scikit-adaptation.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# SKADA - Domain Adaptation with scikit-learn and PyTorch\n\n[![PyPI version](https://badge.fury.io/py/skada.svg)](https://badge.fury.io/py/skada)\n[![Build Status](https://github.com/scikit-adaptation/skada/actions/workflows/testing.yml/badge.svg)](https://github.com/scikit-adaptation/skada/actions)\n[![Codecov Status](https://codecov.io/gh/scikit-adaptation/skada/branch/main/graph/badge.svg)](https://codecov.io/gh/scikit-adaptation/skada)\n[![License](https://img.shields.io/badge/License-BSD_3--Clause-blue.svg)](https://opensource.org/licenses/BSD-3-Clause)\n[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.12666838.svg)](https://doi.org/10.5281/zenodo.12666838)\n\n\u003e [!WARNING]\n\u003e This library is currently in a phase of active development. All features are subject to change without prior notice. If you are interested in collaborating, please feel free to reach out by opening an issue or starting a discussion.\n\nSKADA is a library for domain adaptation (DA) with a scikit-learn and PyTorch/skorch\ncompatible API with the following features:\n\n- DA estimators and transformers with a scikit-learn compatible API (fit, transform, predict).\n- PyTorch/skorch API for deep learning DA algorithms.\n- Classifier/Regressor and data Adapter DA algorithms compatible with scikit-learn pipelines.\n- Compatible with scikit-learn validation loops (cross_val_score, GridSearchCV, etc).\n\n**Citation**: If you use this library in your research, please cite the following reference:\n\n```\nGnassounou T., Kachaiev O., Flamary R., Collas A., Lalou Y., de Mathelin A., Gramfort A., Bueno R., Michel F., Mellot A.,  Loison V., Odonnat A., Moreau T. (2024). SKADA : Scikit Adaptation (version 0.3.0). URL: https://scikit-adaptation.github.io/\n```\n\nor in Bibtex format :\n\n```bibtex\n@misc{gnassounou2024skada,\nauthor = {Gnassounou, Théo and Kachaiev, Oleksii and Flamary, Rémi and Collas, Antoine and Lalou, Yanis and de Mathelin, Antoine and Gramfort, Alexandre and Bueno, Ruben and Michel, Florent and Mellot, Apolline and  Loison, Virginie and Odonnat, Ambroise and Moreau, Thomas},\nmonth = {7},\ntitle = {SKADA : Scikit Adaptation},\nurl = {https://scikit-adaptation.github.io/},\nyear = {2024}\n}\n```\n\n\n## Implemented algorithms\n\nThe following algorithms are currently implemented.\n\n### Domain adaptation algorithms\n\n- Sample reweighting methods (Gaussian [1], Discriminant [2], KLIEPReweight [3],\n  DensityRatio [4], TarS [21], KMMReweight [23])\n- Sample mapping methods (CORAL [5], Optimal Transport DA OTDA [6], LinearMonge [7], LS-ConS [21])\n- Subspace methods (SubspaceAlignment [8], TCA [9], Transfer Subspace Learning [27])\n- Other methods (JDOT [10], DASVM [11], OT Label Propagation [28])\n\nAny methods that can be cast as an adaptation of the input data can be used in one of two ways:\n- a scikit-learn transformer (Adapter) which provides both a full Classifier/Regressor estimator\n - or an `Adapter` that can be used in a DA pipeline with `make_da_pipeline`.\n Refer to the examples below and visit [the gallery](https://scikit-adaptation.github.io/auto_examples/index.html)for more details.\n\n### Deep learning domain adaptation algorithms\n\n- Deep Correlation alignment (DeepCORAL [12])\n- Deep joint distribution optimal (DeepJDOT [13])\n- Divergence minimization (MMD/DAN [14])\n- Adversarial/discriminator based DA (DANN [15], CDAN [16])\n\n### DA metrics\n\n- Importance Weighted [17]\n- Prediction entropy [18]\n- Soft neighborhood density [19]\n- Deep Embedded Validation (DEV) [20]\n- Circular Validation [11]\n\n\n## Installation\n\nThe library is not yet available on PyPI. You can install it from the source code.\n```python\npip install git+https://github.com/scikit-adaptation/skada\n```\n\n## Short examples\n\nWe provide here a few examples to illustrate the use of the library. For more\ndetails, please refer to this [example](https://scikit-adaptation.github.io/auto_examples/plot_how_to_use_skada.html), the [quick start guide](https://scikit-adaptation.github.io/quickstart.html) and the [gallery](https://scikit-adaptation.github.io/auto_examples/index.html).\n\nFirst, the DA data in the SKADA API is stored in the following format:\n\n```python\nX, y, sample_domain\n```\n\nWhere `X` is the input data, `y` is the target labels and `sample_domain` is the\ndomain labels (positive for source and negative for target domains). We provide\nbelow an example ho how to fit a DA estimator:\n\n```python\nfrom skada import CORAL\n\nda = CORAL()\nda.fit(X, y, sample_domain=sample_domain) # sample_domain passed by name\n\nypred = da.predict(Xt) # predict on test data\n```\n\nOne can also use `Adapter` classes to create a full pipeline with DA:\n\n```python\nfrom skada import CORALAdapter, make_da_pipeline\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LogisticRegression\n\npipe = make_da_pipeline(StandardScaler(), CORALAdapter(), LogisticRegression())\n\npipe.fit(X, y, sample_domain=sample_domain) # sample_domain passed by name\n```\n\nPlease note that for `Adapter` classes that implement sample reweighting, the\nsubsequent classifier/regressor must require sample_weights as input. This is\ndone with the `set_fit_requires` method. For instance, with `LogisticRegression`, you\nwould use `LogisticRegression().set_fit_requires('sample_weight')`:\n\n```python\nfrom skada import GaussianReweightAdapter, make_da_pipeline\npipe = make_da_pipeline(GaussianReweightAdapter(),\n                        LogisticRegression().set_fit_request(sample_weight=True))\n```\n\nFinally SKADA can be used for cross validation scores estimation and hyperparameter\nselection :\n\n```python\nfrom sklearn.model_selection import cross_val_score, GridSearchCV\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LogisticRegression\n\nfrom skada import CORALAdapter, make_da_pipeline\nfrom skada.model_selection import SourceTargetShuffleSplit\nfrom skada.metrics import PredictionEntropyScorer\n\n# make pipeline\npipe = make_da_pipeline(StandardScaler(), CORALAdapter(), LogisticRegression())\n\n# split and score\ncv = SourceTargetShuffleSplit()\nscorer = PredictionEntropyScorer()\n\n# cross val score\nscores = cross_val_score(pipe, X, y, params={'sample_domain': sample_domain},\n                         cv=cv, scoring=scorer)\n\n# grid search\nparam_grid = {'coraladapter__reg': [0.1, 0.5, 0.9]}\ngrid_search = GridSearchCV(estimator=pipe,\n                           param_grid=param_grid,\n                           cv=cv, scoring=scorer)\n\ngrid_search.fit(X, y, sample_domain=sample_domain)\n```\n\n## Acknowledgements\n\nThis toolbox has been created and is maintained by the SKADA team that includes the following members:\n\n* [Théo Gnassounou](https://tgnassou.github.io/)\n* [Oleksii Kachaiev](https://kachayev.github.io/talks/)\n* [Rémi Flamary](https://remi.flamary.com/)\n* [Antoine Collas](https://www.antoinecollas.fr/)\n* [Yanis Lalou](https://github.com/YanisLalou)\n* [Antoine de Mathelin](https://scholar.google.com/citations?user=h79bffAAAAAJ\u0026hl=fr)\n* [Ruben Bueno]()\n\n## License\n\nThe library is distributed under the 3-Clause BSD license.\n\n## References\n\n[1] Shimodaira Hidetoshi. [\"Improving predictive inference under covariate shift by weighting the log-likelihood function.\"](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=235723a15c86c369c99a42e7b666dfe156ad2cba) Journal of statistical planning and inference 90, no. 2 (2000): 227-244.\n\n[2] Sugiyama Masashi, Taiji Suzuki, and Takafumi Kanamori. [\"Density-ratio matching under the Bregman divergence: a unified framework of density-ratio estimation.\"](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=f1467208a75def8b2e52a447ab83644db66445ea) Annals of the Institute of Statistical Mathematics 64 (2012): 1009-1044.\n\n[3] Sugiyama Masashi, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul Von Bünau, and Motoaki Kawanabe. [\"Direct importance estimation for covariate shift adaptation.\"](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=af14e09a9f829b9f0952eac244b0ac0c8bda2ca8) Annals of the Institute of Statistical Mathematics 60 (2008): 699-746.\n\n[4] Sugiyama Masashi, and Klaus-Robert Müller. [\"Input-dependent estimation of generalization error under covariate shift.\"](https://web.archive.org/web/20070221112234id_/http://sugiyama-www.cs.titech.ac.jp:80/~sugi/2005/IWSIC.pdf) (2005): 249-279.\n\n[5] Sun Baochen, Jiashi Feng, and Kate Saenko. [\"Correlation alignment for unsupervised domain adaptation.\"](https://arxiv.org/pdf/1612.01939.pdf) Domain adaptation in computer vision applications (2017): 153-171.\n\n[6] Courty Nicolas, Flamary Rémi, Tuia Devis, and Alain Rakotomamonjy. [\"Optimal transport for domain adaptation.\"](https://arxiv.org/pdf/1507.00504.pdf) IEEE Trans. Pattern Anal. Mach. Intell 1, no. 1-40 (2016): 2.\n\n[7] Flamary, R., Lounici, K., \u0026 Ferrari, A. (2019). [Concentration bounds for linear monge mapping estimation and optimal transport domain adaptation](https://arxiv.org/pdf/1905.10155.pdf). arXiv preprint arXiv:1905.10155.\n\n[8] Fernando, B., Habrard, A., Sebban, M., \u0026 Tuytelaars, T. (2013). [Unsupervised visual domain adaptation using subspace alignment](https://openaccess.thecvf.com/content_iccv_2013/papers/Fernando_Unsupervised_Visual_Domain_2013_ICCV_paper.pdf). In Proceedings of the IEEE international conference on computer vision (pp. 2960-2967).\n\n[9] Pan, S. J., Tsang, I. W., Kwok, J. T., \u0026 Yang, Q. (2010). [Domain adaptation via transfer component analysis](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=4823e52161ec339d4d3526099a5477321f6a9a0f). IEEE transactions on neural networks, 22(2), 199-210.\n\n[10] Courty, N., Flamary, R., Habrard, A., \u0026 Rakotomamonjy, A. (2017). [Joint distribution optimal transportation for domain adaptation](https://proceedings.neurips.cc/paper_files/paper/2017/file/0070d23b06b1486a538c0eaa45dd167a-Paper.pdf). Advances in neural information processing systems, 30.\n\n[11] Bruzzone, L., \u0026 Marconcini, M. (2009). [Domain adaptation problems: A DASVM classification technique and a circular validation strategy.](https://rslab.disi.unitn.it/papers/R82-PAMI.pdf) IEEE transactions on pattern analysis and machine intelligence, 32(5), 770-787.\n\n[12] Sun, B., \u0026 Saenko, K. (2016). [Deep coral: Correlation alignment for deep domain adaptation](https://arxiv.org/pdf/1607.01719.pdf). In Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part III 14 (pp. 443-450). Springer International Publishing.\n\n[13] Damodaran, B. B., Kellenberger, B., Flamary, R., Tuia, D., \u0026 Courty, N. (2018). [Deepjdot: Deep joint distribution optimal transport for unsupervised domain adaptation](https://openaccess.thecvf.com/content_ECCV_2018/papers/Bharath_Bhushan_Damodaran_DeepJDOT_Deep_Joint_ECCV_2018_paper.pdf). In Proceedings of the European conference on computer vision (ECCV) (pp. 447-463).\n\n[14] Long, M., Cao, Y., Wang, J., \u0026 Jordan, M. (2015, June). [Learning transferable features with deep adaptation networks](https://proceedings.mlr.press/v37/long15.pdf). In International conference on machine learning (pp. 97-105). PMLR.\n\n[15] Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., ... \u0026 Lempitsky, V. (2016). [Domain-adversarial training of neural networks](https://www.jmlr.org/papers/volume17/15-239/15-239.pdf). Journal of machine learning research, 17(59), 1-35.\n\n[16] Long, M., Cao, Z., Wang, J., \u0026 Jordan, M. I. (2018). [Conditional adversarial domain adaptation](https://proceedings.neurips.cc/paper_files/paper/2018/file/ab88b15733f543179858600245108dd8-Paper.pdf). Advances in neural information processing systems, 31.\n\n[17] Sugiyama, M., Krauledat, M., \u0026 Müller, K. R. (2007). [Covariate shift adaptation by importance weighted cross validation](https://www.jmlr.org/papers/volume8/sugiyama07a/sugiyama07a.pdf). Journal of Machine Learning Research, 8(5).\n\n[18] Morerio, P., Cavazza, J., \u0026 Murino, V. (2017).[ Minimal-entropy correlation alignment for unsupervised deep domain adaptation](https://arxiv.org/pdf/1711.10288.pdf). arXiv preprint arXiv:1711.10288.\n\n[19] Saito, K., Kim, D., Teterwak, P., Sclaroff, S., Darrell, T., \u0026 Saenko, K. (2021). [Tune it the right way: Unsupervised validation of domain adaptation via soft neighborhood density](https://openaccess.thecvf.com/content/ICCV2021/papers/Saito_Tune_It_the_Right_Way_Unsupervised_Validation_of_Domain_Adaptation_ICCV_2021_paper.pdf). In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 9184-9193).\n\n[20] You, K., Wang, X., Long, M., \u0026 Jordan, M. (2019, May). [Towards accurate model selection in deep unsupervised domain adaptation](https://proceedings.mlr.press/v97/you19a/you19a.pdf). In International Conference on Machine Learning (pp. 7124-7133). PMLR.\n\n[21] Zhang, K., Schölkopf, B., Muandet, K., Wang, Z. (2013). [Domain Adaptation under Target and Conditional Shift](http://proceedings.mlr.press/v28/zhang13d.pdf). In International Conference on Machine Learning (pp. 819-827). PMLR.\n\n[22] Loog, M. (2012). Nearest neighbor-based importance weighting. In 2012 IEEE International Workshop on Machine Learning for Signal Processing, pages 1–6. IEEE (https://arxiv.org/pdf/2102.02291.pdf)\n\n[23] Domain Adaptation Problems: A DASVM ClassificationTechnique and a Circular Validation StrategyLorenzo Bruzzone, Fellow, IEEE, and Mattia Marconcini, Member, IEEE (https://rslab.disi.unitn.it/papers/R82-PAMI.pdf)\n\n[24] Loog, M. (2012). Nearest neighbor-based importance weighting. In 2012 IEEE International Workshop on Machine Learning for Signal Processing, pages 1–6. IEEE (https://arxiv.org/pdf/2102.02291.pdf)\n\n[25] J. Huang, A. Gretton, K. Borgwardt, B. Schölkopf and A. J. Smola. Correcting sample selection bias by unlabeled data. In NIPS, 2007. (https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=07117994f0971b2fc2df95adb373c31c3d313442)\n\n[26] Long, M., Wang, J., Ding, G., Sun, J., and Yu, P. (2014). [Transfer joint matching for unsupervised domain adaptation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1410–1417](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=a279f53f386ac78345b67e13c1808880c718efdf)\n\n[27] S. Si, D. Tao and B. Geng. In IEEE Transactions on Knowledge and Data Engineering, (2010) [Bregman Divergence-Based Regularization for Transfer Subspace Learning](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=4118b4fc7d61068b9b448fd499876d139baeec81)\n\n[28] Solomon, J., Rustamov, R., Guibas, L., \u0026 Butscher, A. (2014, January). [Wasserstein propagation for semi-supervised learning](https://proceedings.mlr.press/v32/solomon14.pdf). In International Conference on Machine Learning (pp. 306-314). PMLR.\n\n[29] Montesuma, Eduardo Fernandes, and Fred Maurice Ngole Mboula. [\"Wasserstein barycenter for multi-source domain adaptation.\"](https://openaccess.thecvf.com/content/CVPR2021/papers/Montesuma_Wasserstein_Barycenter_for_Multi-Source_Domain_Adaptation_CVPR_2021_paper.pdf) In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16785-16793. 2021.\n\n[30] Gnassounou, Theo, Rémi Flamary, and Alexandre Gramfort. [\"Convolution Monge Mapping Normalization for learning on sleep data.\"](https://proceedings.neurips.cc/paper_files/paper/2023/file/21718991f6acf19a42376b5c7a8668c5-Paper-Conference.pdf) Advances in Neural Information Processing Systems 36 (2024).\n\n[31] Redko, Ievgen, Nicolas Courty, Rémi Flamary, and Devis Tuia.[ \"Optimal transport for multi-source domain adaptation under target shift.\"](https://proceedings.mlr.press/v89/redko19a/redko19a.pdf) In The 22nd International Conference on artificial intelligence and statistics, pp. 849-858. PMLR, 2019.\n\n[32] Hu, D., Liang, J., Liew, J. H., Xue, C., Bai, S., \u0026 Wang, X. (2023). [Mixed Samples as Probes for Unsupervised Model Selection in Domain Adaptation](https://proceedings.neurips.cc/paper_files/paper/2023/file/7721f1fea280e9ffae528dc78c732576-Paper-Conference.pdf). Advances in Neural Information Processing Systems 36 (2024).\n\n[33] Kang, G., Jiang, L., Yang, Y., \u0026 Hauptmann, A. G. (2019). [Contrastive Adaptation Network for Unsupervised Domain Adaptation](https://arxiv.org/abs/1901.00976). In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 4893-4902).\n\n[34] Jin, Ying, Wang, Ximei, Long, Mingsheng, Wang, Jianmin. [Minimum Class Confusion for Versatile Domain Adaptation](https://arxiv.org/pdf/1912.03699). ECCV, 2020.\n\n[35] Zhang, Y., Liu, T., Long, M., \u0026 Jordan, M. I. (2019). [Bridging Theory and Algorithm for Domain Adaptation](https://arxiv.org/abs/1904.05801). In Proceedings of the 36th International Conference on Machine Learning, (pp. 7404-7413).\n\n[36] Xiao, Zhiqing, Wang, Haobo, Jin, Ying, Feng, Lei, Chen, Gang, Huang, Fei, Zhao, Junbo.[SPA: A Graph Spectral Alignment Perspective for Domain Adaptation](https://arxiv.org/pdf/2310.17594). In Neurips, 2023.\n\n[37] Xie, Renchunzi, Odonnat, Ambroise, Feofanov, Vasilii, Deng, Weijian, Zhang, Jianfeng and An, Bo. [MaNo: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution Shifts](https://arxiv.org/pdf/2405.18979). In NeurIPS, 2024.","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fscikit-adaptation%2Fskada","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fscikit-adaptation%2Fskada","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fscikit-adaptation%2Fskada/lists"}