{"id":15628254,"url":"https://github.com/skalskip/top-cvpr-2024-papers","last_synced_at":"2025-04-04T10:08:53.559Z","repository":{"id":232618071,"uuid":"784797264","full_name":"SkalskiP/top-cvpr-2024-papers","owner":"SkalskiP","description":"This repository is a curated collection of the most exciting and influential CVPR 2024 papers. 🔥 [Paper + Code + Demo]","archived":false,"fork":false,"pushed_at":"2024-06-24T08:39:46.000Z","size":60,"stargazers_count":708,"open_issues_count":3,"forks_count":59,"subscribers_count":15,"default_branch":"master","last_synced_at":"2025-04-03T14:26:10.558Z","etag":null,"topics":["computer-vision","cvpr","cvpr2024","image-segmentation","object-detection","paper","transformers","vision-and-language"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc0-1.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/SkalskiP.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-04-10T15:20:30.000Z","updated_at":"2025-03-26T13:11:19.000Z","dependencies_parsed_at":null,"dependency_job_id":"7af75908-ffc7-4a6f-8d8f-2135f2e4043f","html_url":"https://github.com/SkalskiP/top-cvpr-2024-papers","commit_stats":null,"previous_names":["skalskip/top-cvpr-2024-papers"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Ftop-cvpr-2024-papers","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Ftop-cvpr-2024-papers/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Ftop-cvpr-2024-papers/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Ftop-cvpr-2024-papers/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/SkalskiP","download_url":"https://codeload.github.com/SkalskiP/top-cvpr-2024-papers/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247157283,"owners_count":20893220,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["computer-vision","cvpr","cvpr2024","image-segmentation","object-detection","paper","transformers","vision-and-language"],"created_at":"2024-10-03T10:21:40.021Z","updated_at":"2025-04-04T10:08:53.536Z","avatar_url":"https://github.com/SkalskiP.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"![visitor badge](https://visitor-badge.laobi.icu/badge?page_id=SkalskiP.top-cvpr-2024-papers)\n\n\u003cdiv align=\"center\"\u003e\n  \u003ch1 align=\"center\"\u003etop CVPR 2024 papers\u003c/h1\u003e\n  \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2023-papers\"\u003e2023\u003c/a\u003e\n\u003c/div\u003e\n\n\u003cbr\u003e\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg width=\"600\" src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/347853f9-9e93-4ca0-858b-a7c3f6bba073\" alt=\"vancouver\"\u003e\n\u003c/div\u003e\n\n## 👋 hello\n\nComputer Vision and Pattern Recognition is a massive conference. In **2024** alone,\n**11,532** papers were submitted, and **2,719** were accepted. I created this repository\nto help you search for crème de la crème of CVPR publications. If the paper you are\nlooking for is not on my short list, take a peek at the full\n[list](https://cvpr.thecvf.com/Conferences/2024/AcceptedPapers) of accepted papers.\n\n## 🗞️ papers and posters\n\n*🔥 - highlighted papers*\n\n\u003c!--- AUTOGENERATED_PAPERS_LIST --\u003e\n\u003c!---\n   WARNING: DO NOT EDIT THIS LIST MANUALLY. IT IS AUTOMATICALLY GENERATED.\n   HEAD OVER TO https://github.com/SkalskiP/top-cvpr-2024-papers/blob/master/CONTRIBUTING.md FOR MORE DETAILS ON HOW TO MAKE CHANGES PROPERLY.\n--\u003e\n### 3d from multi-view and sensors\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/31668.png?t=1717417393.7589533\" title=\"SpatialTracker: Tracking Any 2D Pixels in 3D Space\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/56498f78-2ca0-46ee-9231-6aa1806b6ebc\" alt=\"SpatialTracker: Tracking Any 2D Pixels in 3D Space\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2404.04319\" title=\"SpatialTracker: Tracking Any 2D Pixels in 3D Space\"\u003e\n        \u003cstrong\u003e🔥 SpatialTracker: Tracking Any 2D Pixels in 3D Space\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, Xiaowei Zhou\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2404.04319\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/henry123-boy/SpaTracker\"\u003ecode\u003c/a\u003e]   \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e 3D from multi-view and sensors\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Fri 21 Jun 1:30 p.m. EDT — 3 p.m. EDT #84\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/31616.png?t=1716470830.0209699\" title=\"ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/0453bf88-9d54-4ecf-8a45-01af0f604faf\" alt=\"ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2403.01807\" title=\"ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models\"\u003e\n        \u003cstrong\u003eViewDiff: 3D-Consistent Image Generation with Text-to-Image Models\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Lukas Höllein, Aljaž Božič, Norman Müller, David Novotny, Hung-Yu Tseng, Christian Richardt, Michael Zollhöfer, Matthias Nießner\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2403.01807\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/facebookresearch/ViewDiff\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/SdjoCqHzMMk\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e 3D from multi-view and sensors\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 8 p.m. EDT — 9:30 p.m. EDT #20\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2405.12979\" title=\"OmniGlue: Generalizable Feature Matching with Foundation Model Guidance\"\u003e\n        \u003cstrong\u003eOmniGlue: Generalizable Feature Matching with Foundation Model Guidance\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Hanwen Jiang, Arjun Karpur, Bingyi Cao, Qixing Huang, Andre Araujo\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2405.12979\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/google-research/omniglue\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"https://huggingface.co/spaces/qubvel-hf/omniglue\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e 3D from multi-view and sensors\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Fri 21 Jun 1:30 p.m. EDT — 3 p.m. EDT #32\n\u003c/p\u003e\n\u003cbr/\u003e\n\n### deep learning architectures and techniques\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30529.png?t=1717455193.7819567\" title=\"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/4aaf3f87-cc62-4fa3-af99-c8c1c83c0069\" alt=\"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/pdf/2311.06242\" title=\"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks\"\u003e\n        \u003cstrong\u003e🔥 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, Lu Yuan\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/pdf/2311.06242\"\u003epaper\u003c/a\u003e]  [\u003ca href=\"https://youtu.be/cOlyA00K1ec\"\u003evideo\u003c/a\u003e] [\u003ca href=\"https://huggingface.co/spaces/gokaygokay/Florence-2\"\u003edemo\u003c/a\u003e] [\u003ca href=\"https://youtu.be/cOlyA00K1ec\"\u003ecolab\u003c/a\u003e]\n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Deep learning architectures and techniques\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 8 p.m. EDT — 9:30 p.m. EDT #102\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### document analysis and understanding\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2405.04408\" title=\"DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks\"\u003e\n        \u003cstrong\u003eDocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Jiaxin Zhang, Dezhi Peng, Chongyu Liu, Peirong Zhang, Lianwen Jin\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2405.04408\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/ZZZHANG-jx/DocRes\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"https://huggingface.co/spaces/qubvel-hf/documents-restoration\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Document analysis and understanding\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #101\n\u003c/p\u003e\n\u003cbr/\u003e\n\n### efficient and scalable vision\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/e95eac04-5a45-402c-885d-14395879abd3\" title=\"EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/e95eac04-5a45-402c-885d-14395879abd3\" alt=\"EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.00863\" title=\"EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything\"\u003e\n        \u003cstrong\u003e🔥 EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Yunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang, Fanyi Xiao, Chenchen Zhu, Xiaoliang Dai, Dilin Wang, Fei Sun, Forrest Iandola, Raghuraman Krishnamoorthi, Vikas Chandra\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.00863\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/yformer/EfficientSAM\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"https://huggingface.co/spaces/SkalskiP/EfficientSAM\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Efficient and scalable vision\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #144\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30022.png?t=1718402790.003817\" title=\"MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training\"\u003e\n        \u003cimg src=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30022.png?t=1718402790.003817\" alt=\"MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2311.17049\" title=\"MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training\"\u003e\n        \u003cstrong\u003eMobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, Oncel Tuzel\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2311.17049\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/apple/ml-mobileclip\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"https://huggingface.co/spaces/Xenova/webgpu-mobileclip\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Efficient and scalable vision\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #130\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### explainable computer vision\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/6d87318b-57c1-40c7-9de6-5cb47145e119\" title=\"Describing Differences in Image Sets with Natural Language\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/6d87318b-57c1-40c7-9de6-5cb47145e119\" alt=\"Describing Differences in Image Sets with Natural Language\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.02974\" title=\"Describing Differences in Image Sets with Natural Language\"\u003e\n        \u003cstrong\u003e🔥 Describing Differences in Image Sets with Natural Language\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez, Serena Yeung-Levy\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.02974\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/Understanding-Visual-Datasets/VisDiff\"\u003ecode\u003c/a\u003e]   \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Explainable computer vision\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Fri 21 Jun 8 p.m. EDT — 9:30 p.m. EDT #115\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### image and video synthesis and generation\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2311.16973\" title=\"DemoFusion: Democratising High-Resolution Image Generation With No $$$\"\u003e\n        \u003cstrong\u003eDemoFusion: Democratising High-Resolution Image Generation With No $$$\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Ruoyi Du, Dongliang Chang, Timothy Hospedales, Yi-Zhe Song, Zhanyu Ma\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2311.16973\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/PRIS-CV/DemoFusion\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"https://huggingface.co/spaces/radames/Enhance-This-DemoFusion-SDXL\"\u003edemo\u003c/a\u003e] [\u003ca href=\"https://colab.research.google.com/github/camenduru/DemoFusion-colab/blob/main/DemoFusion_colab.ipynb\"\u003ecolab\u003c/a\u003e]\n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Image and video synthesis and generation\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 8 p.m. EDT — 9:30 p.m. EDT #132\n\u003c/p\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/b0833f6b-6924-4f28-b409-ae85aaaa4dd6\" title=\"DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/2a0219f5-9f1e-47e1-a968-d4d98154feb2\" alt=\"DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2306.14435\" title=\"DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing\"\u003e\n        \u003cstrong\u003e🔥 DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Yujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan, Hanshu Yan, Wenqing Zhang, Vincent Y. F. Tan, Song Bai\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2306.14435\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/Yujun-Shi/DragDiffusion\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/rysOFTpDBhc\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Image and video synthesis and generation\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 8 p.m. EDT — 9:30 p.m. EDT #392\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30657.png?t=1717473392.6694562\" title=\"Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/709e3619-25d9-409e-b6ad-ca082611fe09\" alt=\"Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2311.17919\" title=\"Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models\"\u003e\n        \u003cstrong\u003e🔥 Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Daniel Geng, Inbum Park, Andrew Owens\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2311.17919\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/dangeng/visual_anagrams\"\u003ecode\u003c/a\u003e]   [\u003ca href=\"https://colab.research.google.com/github/dangeng/visual_anagrams/blob/main/notebooks/colab_demo_free_tier.ipynb\"\u003ecolab\u003c/a\u003e]\n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Image and video synthesis and generation\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Fri 21 Jun 8 p.m. EDT — 9:30 p.m. EDT #118\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### low-level vision\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/8eb6b4f0-4ae6-4615-9921-f73fa2aa3766\" title=\"XFeat: Accelerated Features for Lightweight Image Matching\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/50b6d16f-c2d8-49a4-8c15-a31d6f9a3c44\" alt=\"XFeat: Accelerated Features for Lightweight Image Matching\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2404.19174\" title=\"XFeat: Accelerated Features for Lightweight Image Matching\"\u003e\n        \u003cstrong\u003eXFeat: Accelerated Features for Lightweight Image Matching\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Guilherme Potje, Felipe Cadar, Andre Araujo, Renato Martins, Erickson R. Nascimento\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2404.19174\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/verlab/accelerated_features\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/RamC70IkZuI\"\u003evideo\u003c/a\u003e] [\u003ca href=\"https://huggingface.co/spaces/qubvel-hf/xfeat\"\u003edemo\u003c/a\u003e] [\u003ca href=\"https://colab.research.google.com/github/verlab/accelerated_features/blob/main/notebooks/xfeat_matching.ipynb\"\u003ecolab\u003c/a\u003e]\n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Low-level vision\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #245\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/038bef8f-a6df-440d-9ebc-b58f69beb338\" title=\"Robust Image Denoising through Adversarial Frequency Mixup\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/03cc753c-f875-479e-bca2-e0375e9929a6\" alt=\"Robust Image Denoising through Adversarial Frequency Mixup\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://openaccess.thecvf.com/content/CVPR2024/html/Ryou_Robust_Image_Denoising_through_Adversarial_Frequency_Mixup_CVPR_2024_paper.html\" title=\"Robust Image Denoising through Adversarial Frequency Mixup\"\u003e\n        \u003cstrong\u003eRobust Image Denoising through Adversarial Frequency Mixup\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Donghun Ryou, Inju Ha, Hyewon Yoo, Dongwan Kim, Bohyung Han\n    \u003cbr/\u003e\n    [\u003ca href=\"https://openaccess.thecvf.com/content/CVPR2024/html/Ryou_Robust_Image_Denoising_through_Adversarial_Frequency_Mixup_CVPR_2024_paper.html\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/dhryougit/AFM\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/zQ0pwFSk7uo\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Low-level vision\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #250\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### multi-modal learning\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2310.03744\" title=\"Improved Baselines with Visual Instruction Tuning\"\u003e\n        \u003cstrong\u003e🔥 Improved Baselines with Visual Instruction Tuning\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Haotian Liu, Chunyuan Li, Yuheng Li, Yong Jae Lee\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2310.03744\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/LLaVA-VL/LLaVA-NeXT\"\u003ecode\u003c/a\u003e]   \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Multi-modal learning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Fri 21 Jun 8 p.m. EDT — 9:30 p.m. EDT #209\n\u003c/p\u003e\n\u003cbr/\u003e\n\n### recognition: categorization, detection, retrieval\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/31301.png?t=1717420504.9897285\" title=\"DETRs Beat YOLOs on Real-time Object Detection\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/3732bfdd-4be4-45cd-8353-e056094f9fec\" alt=\"DETRs Beat YOLOs on Real-time Object Detection\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2304.08069\" title=\"DETRs Beat YOLOs on Real-time Object Detection\"\u003e\n        \u003cstrong\u003eDETRs Beat YOLOs on Real-time Object Detection\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Yian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei, Guanzhong Wang, Qingqing Dang, Yi Liu, Jie Chen\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2304.08069\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/lyuwenyu/RT-DETR\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://www.youtube.com/watch?v=UOc0qMSX4Ac\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Recognition: Categorization, detection, retrieval\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #229\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/f9023a28-aca5-4965-a194-984c62348dc0\" title=\"YOLO-World: Real-Time Open-Vocabulary Object Detection\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/b9f0bb1e-91d4-4ea3-83c6-ee0817afc1bf\" alt=\"YOLO-World: Real-Time Open-Vocabulary Object Detection\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2401.17270\" title=\"YOLO-World: Real-Time Open-Vocabulary Object Detection\"\u003e\n        \u003cstrong\u003eYOLO-World: Real-Time Open-Vocabulary Object Detection\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, Ying Shan\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2401.17270\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/AILab-CVC/YOLO-World\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/X7gKBGVz4vs\"\u003evideo\u003c/a\u003e] [\u003ca href=\"https://huggingface.co/spaces/SkalskiP/YOLO-World\"\u003edemo\u003c/a\u003e] [\u003ca href=\"https://colab.research.google.com/github/roboflow-ai/notebooks/blob/main/notebooks/zero-shot-object-detection-with-yolo-world.ipynb\"\u003ecolab\u003c/a\u003e]\n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Recognition: Categorization, detection, retrieval\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #223\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/31732.png?t=1717298372.5822952\" title=\"Object Recognition as Next Token Prediction\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/bcdc1aba-8ecb-4e63-a8a7-d287ca728bbb\" alt=\"Object Recognition as Next Token Prediction\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.02142\" title=\"Object Recognition as Next Token Prediction\"\u003e\n        \u003cstrong\u003e🔥 Object Recognition as Next Token Prediction\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Kaiyu Yue, Bor-Chun Chen, Jonas Geiping, Hengduo Li, Tom Goldstein, Ser-Nam Lim\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.02142\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/kaiyuyue/nxtp\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/xeI8dZIpoco\"\u003evideo\u003c/a\u003e]  [\u003ca href=\"https://colab.research.google.com/drive/1pJX37LP5xGLDzD3H7ztTmpq1RrIBeWX3?usp=sharing\"\u003ecolab\u003c/a\u003e]\n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Recognition: Categorization, detection, retrieval\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #199\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### segmentation, grouping and shape analysis\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/62d34981-73d6-49b2-8058-46ec99bac94d\" title=\"RobustSAM: Segment Anything Robustly on Degraded Images\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/ee15d3bc-c391-44f9-b35b-24af714ef119\" alt=\"RobustSAM: Segment Anything Robustly on Degraded Images\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://openaccess.thecvf.com/content/CVPR2024/html/Chen_RobustSAM_Segment_Anything_Robustly_on_Degraded_Images_CVPR_2024_paper.html\" title=\"RobustSAM: Segment Anything Robustly on Degraded Images\"\u003e\n        \u003cstrong\u003e🔥 RobustSAM: Segment Anything Robustly on Degraded Images\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Wei-Ting Chen, Yu-Jiet Vong, Sy-Yen Kuo, Sizhou Ma, Jian Wang\n    \u003cbr/\u003e\n    [\u003ca href=\"https://openaccess.thecvf.com/content/CVPR2024/html/Chen_RobustSAM_Segment_Anything_Robustly_on_Degraded_Images_CVPR_2024_paper.html\"\u003epaper\u003c/a\u003e]  [\u003ca href=\"https://www.youtube.com/watch?v=Awukqkbs6zM\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Segmentation, grouping and shape analysis\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #378\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30253.png?t=1716781257.513028\" title=\"Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/0c43b789-f2e8-4ff9-ae46-b5a87de1b921\" alt=\"Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://openaccess.thecvf.com/content/CVPR2024/html/Zhang_Frozen_CLIP_A_Strong_Backbone_for_Weakly_Supervised_Semantic_Segmentation_CVPR_2024_paper.html\" title=\"Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation\"\u003e\n        \u003cstrong\u003e🔥 Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Bingfeng Zhang, Siyue Yu, Yunchao Wei, Yao Zhao, Jimin Xiao\n    \u003cbr/\u003e\n    [\u003ca href=\"https://openaccess.thecvf.com/content/CVPR2024/html/Zhang_Frozen_CLIP_A_Strong_Backbone_for_Weakly_Supervised_Semantic_Segmentation_CVPR_2024_paper.html\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/zbf1991/WeCLIP\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/Lh489nTm_M0\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Segmentation, grouping and shape analysis\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #351\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/2f2bf794-3981-48c8-992d-04dd32ee9ced\" title=\"Semantic-aware SAM for Point-Prompted Instance Segmentation\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/f1ed2755-1df1-45fe-810b-5fc98b4b52e1\" alt=\"Semantic-aware SAM for Point-Prompted Instance Segmentation\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.15895\" title=\"Semantic-aware SAM for Point-Prompted Instance Segmentation\"\u003e\n        \u003cstrong\u003e🔥 Semantic-aware SAM for Point-Prompted Instance Segmentation\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Zhaoyang Wei, Pengfei Chen, Xuehui Yu, Guorong Li, Jianbin Jiao, Zhenjun Han\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.15895\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/zhaoyangwei123/SAPNet\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/42-tJFmT7Ao\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Segmentation, grouping and shape analysis\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #331\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2403.15789\" title=\"In-Context Matting\"\u003e\n        \u003cstrong\u003e🔥 In-Context Matting\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    He Guo, Zixuan Ye, Zhiguo Cao, Hao Lu\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2403.15789\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/tiny-smart/in-context-matting\"\u003ecode\u003c/a\u003e]   \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Segmentation, grouping and shape analysis\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #343\n\u003c/p\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/bfe79038-706d-491b-ac99-083f421dc5ec\" title=\"General Object Foundation Model for Images and Videos at Scale\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/4f0ed38d-28aa-4766-b290-940cbc6711d6\" alt=\"General Object Foundation Model for Images and Videos at Scale\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.09158\" title=\"General Object Foundation Model for Images and Videos at Scale\"\u003e\n        \u003cstrong\u003e🔥 General Object Foundation Model for Images and Videos at Scale\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Junfeng Wu, Yi Jiang, Qihao Liu, Zehuan Yuan, Xiang Bai, Song Bai\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.09158\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/FoundationVision/GLEE\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://www.youtube.com/watch?v=PSVhfTPx0GQ\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Segmentation, grouping and shape analysis\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Wed 19 Jun 1:30 p.m. EDT — 3 p.m. EDT #350\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### self-supervised or unsupervised representation learning\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30014.png?t=1717339970.9614518\" title=\"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/9a03d726-0459-48f1-9f1e-5f12c7382084\" alt=\"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.14238\" title=\"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks\"\u003e\n        \u003cstrong\u003e🔥 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, Jifeng Dai\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.14238\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/OpenGVLab/InternVL\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"https://huggingface.co/spaces/OpenGVLab/InternVL\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Self-supervised or unsupervised representation learning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Fri 21 Jun 8 p.m. EDT — 9:30 p.m. EDT #412\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### video: low-level analysis, motion, and tracking\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/29590.png?t=1717456006.3308516\" title=\"Matching Anything by Segmenting Anything\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/bb451f47-ba3e-4e34-a7c0-3410b64d9339\" alt=\"Matching Anything by Segmenting Anything\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2406.04221\" title=\"Matching Anything by Segmenting Anything\"\u003e\n        \u003cstrong\u003e🔥 Matching Anything by Segmenting Anything\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Siyuan Li, Lei Ke, Martin Danelljan, Luigi Piccinelli, Mattia Segu, Luc Van Gool, Fisher Yu\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2406.04221\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/siyuanliii/masa\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/KDQVujKAWFQ\"\u003evideo\u003c/a\u003e]  \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Video: Low-level analysis, motion, and tracking\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #421\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/9711186c-b05b-472d-b095-d98dbe386171\" title=\"DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/18caf2db-5dab-4251-9eeb-e2397c67eb3f\" alt=\"DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2403.02075\" title=\"DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction\"\u003e\n        \u003cstrong\u003eDiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Weiyi Lv, Yuhang Huang, Ning Zhang, Ruei-Sung Lin, Mei Han, Dan Zeng\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2403.02075\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/Kroery/DiffMOT\"\u003ecode\u003c/a\u003e]   \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Video: Low-level analysis, motion, and tracking\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 8 p.m. EDT — 9:30 p.m. EDT #455\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n### vision, language, and reasoning\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/31492.png?t=1717327133.6073072\" title=\"Alpha-CLIP: A CLIP Model Focusing on Wherever You Want\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/4480d88a-7f8f-48c2-bcb0-bde3b694dfd8\" alt=\"Alpha-CLIP: A CLIP Model Focusing on Wherever You Want\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.03818\" title=\"Alpha-CLIP: A CLIP Model Focusing on Wherever You Want\"\u003e\n        \u003cstrong\u003eAlpha-CLIP: A CLIP Model Focusing on Wherever You Want\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Zeyi Sun, Ye Fang, Tong Wu, Pan Zhang, Yuhang Zang, Shu Kong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.03818\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/SunzeY/AlphaCLIP\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/QCEIKPZpZz0\"\u003evideo\u003c/a\u003e] [\u003ca href=\"https://huggingface.co/spaces/Zery/Alpha-CLIP_LLaVA-1.5\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Vision, language, and reasoning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 1:30 p.m. EDT — 3 p.m. EDT #327\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2401.06209\" title=\"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs\"\u003e\n        \u003cstrong\u003e🔥 Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, Saining Xie\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2401.06209\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/tsb0601/MMVP\"\u003ecode\u003c/a\u003e]   \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Vision, language, and reasoning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 1:30 p.m. EDT — 3 p.m. EDT #390\n\u003c/p\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/30109.png?t=1717509456.89997\" title=\"LISA: Reasoning Segmentation via Large Language Model\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/fc2699d9-7bd2-4c3a-8e6c-4961505cc802\" alt=\"LISA: Reasoning Segmentation via Large Language Model\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2308.00692\" title=\"LISA: Reasoning Segmentation via Large Language Model\"\u003e\n        \u003cstrong\u003e🔥 LISA: Reasoning Segmentation via Large Language Model\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, Jiaya Jia\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2308.00692\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/dvlab-research/LISA\"\u003ecode\u003c/a\u003e]  [\u003ca href=\"http://103.170.5.190:7870/\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Vision, language, and reasoning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 1:30 p.m. EDT — 3 p.m. EDT #413\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/53e03a08-4dd9-451a-975e-e3654fa5bc71\" title=\"ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/6d1536ae-3f96-49d9-a05f-9648b925cdb5\" alt=\"ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2312.00784\" title=\"ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts\"\u003e\n        \u003cstrong\u003eViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Mu Cai, Haotian Liu, Dennis Park, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Yong Jae Lee\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2312.00784\"\u003epaper\u003c/a\u003e] [\u003ca href=\"https://github.com/WisconsinAIVision/ViP-LLaVA\"\u003ecode\u003c/a\u003e] [\u003ca href=\"https://youtu.be/j_l1bRQouzc\"\u003evideo\u003c/a\u003e] [\u003ca href=\"https://pages.cs.wisc.edu/~mucai/vip-llava.html\"\u003edemo\u003c/a\u003e] \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Vision, language, and reasoning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 1:30 p.m. EDT — 3 p.m. EDT #317\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\n\u003cp align=\"left\"\u003e\n    \u003ca href=\"https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202024/31040.png?t=1718300473.5736258\" title=\"MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI\"\u003e\n        \u003cimg src=\"https://github.com/SkalskiP/top-cvpr-2024-papers/assets/26109316/8b9f69b7-3384-40e6-828f-90bf7b43e345\" alt=\"MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI\" width=\"400px\" align=\"left\" /\u003e\n    \u003c/a\u003e\n    \u003ca href=\"https://arxiv.org/abs/2311.16502\" title=\"MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI\"\u003e\n        \u003cstrong\u003e🔥 MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI\u003c/strong\u003e\n    \u003c/a\u003e\n    \u003cbr/\u003e\n    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, Wenhu Chen\n    \u003cbr/\u003e\n    [\u003ca href=\"https://arxiv.org/abs/2311.16502\"\u003epaper\u003c/a\u003e]    \n    \u003cbr/\u003e\n    \u003cstrong\u003eTopic:\u003c/strong\u003e Vision, language, and reasoning\n    \u003cbr/\u003e\n    \u003cstrong\u003eSession:\u003c/strong\u003e Thu 20 Jun 1:30 p.m. EDT — 3 p.m. EDT #382\n\u003c/p\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\u003c!--- AUTOGENERATED_PAPERS_LIST --\u003e\n\n## 🦸 contribution\n\nWe would love your help in making this repository even better! If you know of an amazing\npaper that isn't listed here, or if you have any suggestions for improvement, feel free\nto open an\n[issue](https://github.com/SkalskiP/top-cvpr-2024-papers/issues)\nor submit a\n[pull request](https://github.com/SkalskiP/top-cvpr-2024-papers/pulls).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fskalskip%2Ftop-cvpr-2024-papers","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fskalskip%2Ftop-cvpr-2024-papers","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fskalskip%2Ftop-cvpr-2024-papers/lists"}