{"id":20305551,"url":"https://github.com/vitae-transformer/vitae-transformer-matting","last_synced_at":"2026-03-05T23:01:58.860Z","repository":{"id":38349531,"uuid":"476163140","full_name":"ViTAE-Transformer/ViTAE-Transformer-Matting","owner":"ViTAE-Transformer","description":"A comprehensive list [AIM@IJCAI'21, P3M@MM'21, GFM@IJCV'22, RIM@CVPR'23, P3MNet@IJCV'23] of our research works related to image matting, including papers, codes, datasets, demos, and citations. Note: The repo for [IJCV'23] \"Rethinking Portrait Matting with Privacy Preserving\" has been moved to: https://github.com/ViTAE-Transformer/P3M-Net","archived":false,"fork":false,"pushed_at":"2023-04-11T05:48:56.000Z","size":5,"stargazers_count":231,"open_issues_count":1,"forks_count":22,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-12-04T17:53:15.271Z","etag":null,"topics":["computer-vision","deep-learning","image-matting","privacy-preserving","survey","vision-transformer"],"latest_commit_sha":null,"homepage":"","language":"TeX","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ViTAE-Transformer.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-03-31T05:27:24.000Z","updated_at":"2025-07-09T19:03:33.000Z","dependencies_parsed_at":"2025-01-14T11:22:59.578Z","dependency_job_id":"315fe75e-8634-4907-85b9-5c9f154ea847","html_url":"https://github.com/ViTAE-Transformer/ViTAE-Transformer-Matting","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/ViTAE-Transformer/ViTAE-Transformer-Matting","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ViTAE-Transformer%2FViTAE-Transformer-Matting","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ViTAE-Transformer%2FViTAE-Transformer-Matting/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ViTAE-Transformer%2FViTAE-Transformer-Matting/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ViTAE-Transformer%2FViTAE-Transformer-Matting/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ViTAE-Transformer","download_url":"https://codeload.github.com/ViTAE-Transformer/ViTAE-Transformer-Matting/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ViTAE-Transformer%2FViTAE-Transformer-Matting/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":30154272,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-05T22:39:40.138Z","status":"ssl_error","status_checked_at":"2026-03-05T22:39:24.771Z","response_time":93,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["computer-vision","deep-learning","image-matting","privacy-preserving","survey","vision-transformer"],"created_at":"2024-11-14T17:08:55.918Z","updated_at":"2026-03-05T23:01:58.829Z","avatar_url":"https://github.com/ViTAE-Transformer.png","language":"TeX","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003ch1 align=\"center\"\u003eImage Matting\u003c/h1\u003e\n\n\nThis repo contains a comprehensive list of our research works related to **image matting**, including papers, codes, datasets, demos, and citations. For any related questions, please contact \u003cstrong\u003e\u003ca href=\"https://github.com/jizhiziLi/\"\u003eJizhizi Li\u003c/a\u003e\u003c/strong\u003e at [jili8515@uni.sydney.edu.au](mailto:jili8515@uni.sydney.edu.au) and \u003cstrong\u003e\u003ca href=\"https://github.com/xymsh\"\u003eSihan Ma\u003c/a\u003e\u003c/strong\u003e at [sima7436@uni.sydney.edu.au](mailto:sima7436@uni.sydney.edu.au).\n\n***\n\u003e\n\u003e\u003ch3\u003e\u003cstrong\u003e\u003ci\u003e🚀 News\u003c/i\u003e\u003c/strong\u003e\u003c/h3\u003e\n\u003e\n\u003e [2023-04-10]: Publish the paper [Deep Image Matting: A Comprehensive Survey](https://arxiv.org/abs/2304.04672) on arXiv.\n\u003e \n\u003e [2023-03-28]: The paper [Rethinking Portrait Matting with Privacy Preserving](https://arxiv.org/abs/2203.16828) has been accepted by the International Journal of Computer Vision ([IJCV](https://www.springer.com/journal/11263)) 🎉\n\u003e \n\u003e [2023-02-28]: The paper [Referring Image Matting](https://arxiv.org/abs/2206.05149) has been accepted by the Computer Vision and Pattern Recognition Conference ([CVPR](https://cvpr2023.thecvf.com/)) 🎉\n***\n\n\n## Overview\n\n[1. Deep Image Matting: A Comprehensive Survey, arXiv, 2023](#survey) \u0026emsp; \u003ca href=\"https://arxiv.org/abs/2304.04672\"\u003e\u003cimg  src=\"https://img.shields.io/badge/arxiv-Paper-brightgreen\" \u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/matting-survey\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/JizhiziLi/matting-survey.svg?logo=github\u0026label=Stars\"\u003e\u003c/a\u003e\n\n\n[2. Rethinking Portrait Matting with Privacy Preserving, IJCV, 2023](#rethink_p3m) \u0026emsp; \u003ca href=\"https://arxiv.org/abs/2203.16828\"\u003e\u003cimg  src=\"https://img.shields.io/badge/arxiv-Paper-brightgreen\" \u003e\u003c/a\u003e \u003ca href=\"https://github.com/ViTAE-Transformer/P3M-Net\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/ViTAE-Transformer/P3M-Net.svg?logo=github\u0026label=Stars\"\u003e\u003c/a\u003e \u003ca href=\"https://github.com/ViTAE-Transformer/P3M-Net#ppt-setting-and-p3m-10k-dataset\"\u003e\u003cimg src=\"https://img.shields.io/badge/dataset-P3M--10k-orange\"\u003e\n\n[3. Referring Image Matting, CVPR, 2023](#rim) \u0026emsp; \u003ca href=\"https://arxiv.org/abs/2206.05149\"\u003e\u003cimg  src=\"https://img.shields.io/badge/arxiv-Paper-brightgreen\" \u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/rim/\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/JizhiziLi/rim.svg?logo=github\u0026label=Stars\"\u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/rim/#refmatte\"\u003e\u003cimg src=\"https://img.shields.io/badge/dataset-RefMatte-orange\"\u003e\n\n[4. Bridging Composite and Real: Towards End-to-end Deep Image Matting, IJCV, 2022](#gfm)  \u0026emsp; \u003ca href=\"https://arxiv.org/abs/2010.16188\"\u003e\u003cimg  src=\"https://img.shields.io/badge/arxiv-Paper-brightgreen\" \u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/gfm/\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/JizhiziLi/gfm.svg?logo=github\u0026label=Stars\"\u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/gfm/#am-2k\"\u003e\u003cimg src=\"https://img.shields.io/badge/dataset-AM--2k-orange\"\u003e \u003ca href=\"https://github.com/JizhiziLi/gfm/#bg-20k\"\u003e\u003cimg src=\"https://img.shields.io/badge/dataset-BG--20k-orange\"\u003e\n\n[5. Privacy-preserving Portrait Matting, ACM MM, 2021](#p3m) \u0026emsp; \u003ca href=\"https://arxiv.org/abs/2104.14222\"\u003e\u003cimg  src=\"https://img.shields.io/badge/arxiv-Paper-brightgreen\" \u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/p3m/\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/JizhiziLi/p3m.svg?logo=github\u0026label=Stars\"\u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/p3m/#ppt-setting-and-p3m-10k-dataset\"\u003e\u003cimg src=\"https://img.shields.io/badge/dataset-P3M--10k-orange\"\u003e\n\n[6. Deep Automatic Natural Image Matting, IJCAI, 2021](#aim) \u0026emsp; \u003ca href=\"https://arxiv.org/abs/2107.07235\"\u003e\u003cimg  src=\"https://img.shields.io/badge/arxiv-Paper-brightgreen\" \u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/aim/\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/JizhiziLi/aim.svg?logo=github\u0026label=Stars\"\u003e\u003c/a\u003e \u003ca href=\"https://github.com/JizhiziLi/aim/#aim-500\"\u003e\u003cimg src=\"https://img.shields.io/badge/dataset-AIM--500-orange\"\u003e\u003c/a\u003e\n\n\n## Projects\n\n### \u003cspan id=\"survey\"\u003e📘 Deep Image Matting: A Comprehensive Survey [arXiv-2023]\u003c/span\u003e\n\n\u003cem\u003eJizhizi Li, Jing Zhang, and Dacheng Tao\u003c/em\u003e\n\n[Paper](https://arxiv.org/abs/2304.04672) | [Github Code](https://github.com/JizhiziLi/matting-survey) | [BibTex](./assets/arXiv_2023_Survey/matting_survey.bib)\n\nImage matting refers to extracting precise alpha matte from natural images, and it plays a critical role in various downstream applications, such as image editing. The emergence of deep learning has revolutionized the field of image matting and given birth to multiple new techniques, including automatic, interactive, and referring image matting. Here we present a comprehensive review of recent advancements in image matting in the era of deep learning.\n\n\u003cimg src=\"https://github.com/JizhiziLi/matting-survey/raw/master/src/timeline.jpg\" width=\"100%\"\u003e\n\n***\n\n### \u003cspan id=\"rethink_p3m\"\u003e📘 Rethinking Portrait Matting with Privacy Preserving [IJCV-2023]\u003c/span\u003e\n\n\u003cimg src=\"https://github.com/ViTAE-Transformer/P3M-Net/raw/main/demo/gif/p_3cf7997c.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/ViTAE-Transformer/P3M-Net/raw/main/demo/gif/2.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/ViTAE-Transformer/P3M-Net/raw/main/demo/gif/3.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/ViTAE-Transformer/P3M-Net/raw/main/demo/gif/4.gif\" width=\"25%\"\u003e\n\n\u003cem\u003eSihan Ma\u003csup\u003e\u0026#8727;\u003c/sup\u003e, Jizhizi Li\u003csup\u003e\u0026#8727;\u003c/sup\u003e, Jing Zhang, He Zhang, and Dacheng Tao. (*equal contribution)\u003c/em\u003e\n\n[Paper](https://arxiv.org/abs/2203.16828) | [Github Code](https://github.com/ViTAE-Transformer/P3M-Net) | [Dataset](https://github.com/ViTAE-Transformer/P3M-Net#ppt-setting-and-p3m-10k-dataset) | [Demo](https://colab.research.google.com/drive/1pD_XKx31Lgd7zwq46dRpz2jGdsH1ZIay?usp=sharing) | [BibTex](./assets/IJCV_2023_Rethink_P3M/rethink_p3m.bib)\n\nThis paper introduces three variants of P3M-Net based on both transformer and CNN backbones to solve the portrait matting problem with privacy preserving. Also a simple yet effective Copy and Paste strategy (P3M-CP) is devised to enable the matting model to process both face-blurred and normal images without extra effort during inference.\n\n\u003cimg src=\"https://github.com/ViTAE-Transformer/P3M-Net/raw/main/demo/p3m-net-variants.png\" width=\"100%\"\u003e\n\n---\n\n### \u003cspan id=\"rim\"\u003e📘 Referring Image Matting [CVPR-2023]\u003c/span\u003e\n\n\u003cimg src=\"https://github.com/JizhiziLi/RIM/raw/master/demo/src/more_k.jpg\" width=\"50%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/RIM/raw/master/demo/src/more_rw.jpg\" width=\"50%\"\u003e\n\n\u003cem\u003eJizhizi Li, Jing Zhang, and Dacheng Tao\u003c/em\u003e\n\n[Paper](https://arxiv.org/abs/2206.05149) |  [Github Code](https://github.com/JizhiziLi/rim/) | [Dataset](https://github.com/JizhiziLi/rim/#refmatte) | [BibTex](./assets/CVPR_2023_RIM/rim.bib)\n\nImage matting refers to extracting the accurate foregrounds in the image. Current automatic methods tend to extract all the salient objects in the image indiscriminately. In this paper, we propose a new task named Referring Image Matting (RIM), referring to extracting the meticulous alpha matte of the specific object that can best match the given natural language description. We then propose a large-scale dataset RefMatte and a carefully designed method CLIPMat to serve as a baseline suite for RIM. We believe the new task RIM along with the RefMatte dataset and the method CLIPMat will open new research directions in this area and facilitate future studies.\n\n\u003cimg src=\"https://github.com/JizhiziLi/RIM/raw/master/demo/src/clipmat.png\" width=\"100%\"\u003e\n\n---\n\n### \u003cspan id=\"gfm\"\u003e📘 Bridging Composite and Real: Towards End-to-end Deep Image Matting [IJCV-2022]\u003c/span\u003e\n\n\n\u003cimg src=\"https://github.com/JizhiziLi/GFM/raw/master/demo/src/homepage/spring.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/GFM/raw/master/demo/src/homepage/summer.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/GFM/raw/master/demo/src/homepage/autumn.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/GFM/raw/master/demo/src/homepage/winter.gif\" width=\"25%\"\u003e\n\n\n\u003cem\u003eJizhizi Li\u003csup\u003e1\u0026#8727;\u003c/sup\u003e, Jing Zhang\u003csup\u003e1\u0026#8727;\u003c/sup\u003e, Stephen J. Maybank, and Dacheng Tao\u003c/em\u003e. (*equal contribution)\n\n\n[Paper](https://arxiv.org/abs/2010.16188) |  [Github Code](https://github.com/JizhiziLi/gfm/) | [Dataset](https://github.com/JizhiziLi/gfm/#am-2k) | [Demo](https://colab.research.google.com/drive/1EaQ5h4u9Q_MmDSFTDmFG0ZOeSsFuRTsJ?usp=sharing) | [BibTex](./assets/IJCV_2022_GFM/gfm.bib)\n\nWe propose a novel Glance and Focus Matting network (GFM), which employs a shared encoder and two separate decoders to learn both tasks in a collaborative manner for end-to-end image matting. We also establish a novel Animal Matting dataset (AM-2k) to serve for end-to-end matting task. Furthermore, we investigate the domain gap issue between composition images and natural images systematically, propose a carefully designed composite route RSSN and a large-scale high-resolution background dataset (BG-20k) to serve as better candidates for composition.\n\n\u003cimg src=\"https://github.com/JizhiziLi/GFM/raw/master/demo/src/homepage/gfm.png\" width=\"100%\"\u003e\n\n---\n\n### \u003cspan id=\"p3m\"\u003e📘 Privacy-Preserving Portrait Matting [ACM MM-21]\u003c/span\u003e\n\n\u003cimg src=\"https://github.com/JizhiziLi/P3M/raw/master/demo/gif/p_2c2e4470.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/P3M/raw/master/demo/gif/p_4dfffce8.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/P3M/raw/master/demo/gif/p_d4fd9815.gif\" width=\"25%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/P3M/raw/master/demo/gif/p_64da52e3.gif\" width=\"25%\"\u003e\n\n\u003cem\u003eJizhizi Li\u003csup\u003e\u0026#8727;\u003c/sup\u003e, Sihan Ma\u003csup\u003e\u0026#8727;\u003c/sup\u003e, Jing Zhang, and Dacheng Tao. (*equal contribution)\u003c/em\u003e\n\n[Paper](https://arxiv.org/abs/2104.14222) | [Github Code](https://github.com/JizhiziLi/P3M) | [Dataset](https://github.com/JizhiziLi/p3m/#ppt-setting-and-p3m-10k-dataset) | [BibTex](./assets/ACM_MM_2021_P3M/p3m.bib)\n\nThis work presents P3M-10k, which is the first large-scale anonymized benchmark for Privacy-Preserving Portrait Matting, to solve the increasing concerns about the privacy in image matting. They also propose P3M-Net, which leverages the power of a unified framework for both semantic perception and detail matting, and specifically emphasizes the interaction between them and the encoder to facilitate the matting process. \n\n\u003cimg src=\"https://github.com/JizhiziLi/P3M/raw/master/demo/network.png\" width=\"100%\"\u003e\n\n---\n\n### \u003cspan id=\"aim\"\u003e📘 Deep Automatic Natural Image Matting [IJCAI-21]\u003c/span\u003e\n\n\u003cimg src=\"https://github.com/JizhiziLi/AIM/raw/master/demo/o_e92b90fc.gif\" width=\"33%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/AIM/raw/master/demo/o_2749e288.gif\" width=\"33%\"\u003e\u003cimg src=\"https://github.com/JizhiziLi/AIM/raw/master/demo/o_87e77b32.gif\" width=\"33%\"\u003e\n\n\u003cem\u003eJizhizi Li, Jing Zhang, and Dacheng Tao\u003c/em\u003e\n\n[Paper](https://arxiv.org/abs/2107.07235) | [Github Code](https://github.com/JizhiziLi/aim/) | [Dataset](https://github.com/jizhiziLi/aim#aim-500) | [BibTex](./assets/IJCAI_2021_AIM/aim.bib)\n\nWe investigate the difficulties when extending the automatic matting methods to natural images with salient transparent/meticulous foregrounds or non-salient foregrounds by proposing a novel end-to-end matting network, which can predict a generalized trimap for any image of the above types as a unified semantic representation and simultaneously guide the matting network to focus on the transition areas via an attention mechanism. We also construct a test set AIM-500 that contains 500 diverse natural images covering all types along with manually labeled alpha mattes, making it feasible to benchmark the generalization ability of AIM models.\n\n\u003cimg src=\"https://github.com/JizhiziLi/AIM/raw/master/demo/network.png\" width=\"100%\"\u003e\n\n\n\n\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvitae-transformer%2Fvitae-transformer-matting","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fvitae-transformer%2Fvitae-transformer-matting","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvitae-transformer%2Fvitae-transformer-matting/lists"}