{"id":27604152,"url":"https://github.com/skalskip/vlms-zero-to-hero","last_synced_at":"2025-10-06T11:52:17.350Z","repository":{"id":269458192,"uuid":"906122697","full_name":"SkalskiP/vlms-zero-to-hero","owner":"SkalskiP","description":"This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.","archived":false,"fork":false,"pushed_at":"2025-01-23T11:23:09.000Z","size":346,"stargazers_count":1063,"open_issues_count":1,"forks_count":97,"subscribers_count":46,"default_branch":"master","last_synced_at":"2025-05-08T22:44:02.116Z","etag":null,"topics":["bert-model","clip","computer-vision","embeddings","gpt","gpt-2","lora","natural-language-processing","seq2seq","vision-language-model","word2vec"],"latest_commit_sha":null,"homepage":"https://www.youtube.com/@SkalskiP","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/SkalskiP.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-12-20T08:02:26.000Z","updated_at":"2025-05-08T22:10:02.000Z","dependencies_parsed_at":"2025-04-22T19:42:02.576Z","dependency_job_id":"97ac21b5-bb90-43c8-826a-12ca10736a2e","html_url":"https://github.com/SkalskiP/vlms-zero-to-hero","commit_stats":null,"previous_names":["skalskip/vlms-zero-to-hero"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Fvlms-zero-to-hero","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Fvlms-zero-to-hero/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Fvlms-zero-to-hero/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkalskiP%2Fvlms-zero-to-hero/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/SkalskiP","download_url":"https://codeload.github.com/SkalskiP/vlms-zero-to-hero/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254264765,"owners_count":22041793,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["bert-model","clip","computer-vision","embeddings","gpt","gpt-2","lora","natural-language-processing","seq2seq","vision-language-model","word2vec"],"created_at":"2025-04-22T19:25:00.936Z","updated_at":"2025-10-06T11:52:12.315Z","avatar_url":"https://github.com/SkalskiP.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n  \u003ch1 align=\"center\"\u003eVLMs zero-to-hero\u003c/h1\u003e\n\n  \u003cp\u003ecoming: january 2025...\u003c/p\u003e\n\n\u003c/div\u003e\n\n# hello\n\nWelcome to VLMs Zero to Hero! This series will take you on a journey from the \nfundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.\n\n# tutorials\n\n|                                                                                                                                 **notebook**                                                                                                                                  |                                                                                              **open in colab**                                                                                              | **video** |                **paper**                |\n|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------:|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------:|:---------:|:---------------------------------------:|\n| 01.01. [Word2Veq: Distributed Representations of Words and Phrases and their Compositionality](https://github.com/SkalskiP/vlms-zero-to-hero/blob/master/01_natural_language_processing_fundamentals/01_01_word2vec_with_sub_sampling_and_negative_sampling_in_pytorch.ipynb) | [link](https://colab.research.google.com/github/SkalskiP/vlms-zero-to-hero/blob/master/01_natural_language_processing_fundamentals/01_01_word2vec_with_sub_sampling_and_negative_sampling_in_pytorch.ipynb) |   soon    | [link](https://arxiv.org/abs/1310.4546) |\n\n# roadmap\n\n## natural language processing (NLP) fundamentals\n\n- Word2Veq: [Efficient Estimation of Word Representations in Vector Space](https://arxiv.org/pdf/1301.3781) (2013) and [Distributed Representations of Words and Phrases and their Compositionality](https://arxiv.org/pdf/1310.4546) (2013) \n- Seq2Seq: [Sequence to Sequence Learning with Neural Networks](https://arxiv.org/pdf/1409.3215) (2014)\n- [Attention Is All You Need](https://arxiv.org/pdf/1706.03762) (2017)\n- BERT: [Pre-training of Deep Bidirectional Transformers for Language Understanding](https://arxiv.org/pdf/1810.04805) (2018)\n- GPT: [Improving Language Understanding by Generative Pre-Training](https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf) (2018)\n\n## computer vision (CV) fundamentals\n\n- AlexNet: [ImageNet Classification with Deep Convolutional Neural Networks](https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf) (2012)\n- VGG: [Very Deep Convolutional Networks for Large-Scale Image Recognition](https://arxiv.org/pdf/1409.1556) (2014)\n- ResNet: [Deep Residual Learning for Image Recognition](https://arxiv.org/pdf/1512.03385) (2015)\n\n## early vision-language models\n\n- [Show and Tell: A Neural Image Caption Generator](https://arxiv.org/pdf/1411.4555) (2014) and [Show, Attend and Tell: Neural Image Caption Generation with Visual Attention](https://arxiv.org/pdf/1502.03044) (2015)\n- [A Picture is Worth 16x16 Words: Transformers for Image Recognition at Scale](https://arxiv.org/pdf/2010.11929) (2020)\n- CLIP: [Learning Transferable Visual Models from Natural Language Supervision](https://arxiv.org/pdf/2103.00020) (2021)\n\n## scale and efficiency\n\n- [Scaling Laws for Neural Language Models](https://arxiv.org/pdf/2001.08361) (2020)\n- LoRA: [Low-Rank Adaptation of Large Language Models](https://arxiv.org/pdf/2106.09685) (2021)\n- QLoRA: [Efficient Fine-tuning of Quantized LLMs](https://arxiv.org/pdf/2305.14314) (2023)\n\n## modern vision-language models\n\n- Flamingo: [A Visual Language Model for Few-Shot Learning](https://arxiv.org/pdf/2204.14198) (2022)\n- LLaVA: [Visual Instruction Tuning](https://arxiv.org/pdf/2304.08485) (2023)\n- BLIP-2: [Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models](https://arxiv.org/pdf/2301.12597) (2023)\n- PaliGemma: [A versatile 3B VLM for transfer](https://arxiv.org/pdf/2407.07726) (2024)\n\n## extra\n\n- BLEU: [a Method for Automatic Evaluation of Machine Translation](https://aclanthology.org/P02-1040.pdf) (2002)\n\n# contribute and suggest more papers\n\nAre there important papers, models, or techniques we missed? Do you have a favorite \nbreakthrough in vision-language research that isn't listed here? We’d love to hear \nyour suggestions!","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fskalskip%2Fvlms-zero-to-hero","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fskalskip%2Fvlms-zero-to-hero","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fskalskip%2Fvlms-zero-to-hero/lists"}