{"id":107080,"url":"https://github.com/umitkacar/awesome-llm","name":"awesome-llm","description":"Large Language Models (LLMs): GPT, Claude, Llama, Gemini, fine-tuning, RAG, prompt engineering, and AI agents for GenAI apps.","projects_count":65,"last_synced_at":"2026-09-15T20:00:19.807Z","repository":{"id":158676042,"uuid":"262736335","full_name":"umitkacar/awesome-llm","owner":"umitkacar","description":"Large Language Models (LLMs): GPT, Claude, Llama, Gemini, fine-tuning, RAG, prompt engineering, and AI agents for GenAI apps.","archived":false,"fork":false,"pushed_at":"2025-11-10T10:10:56.000Z","size":68,"stargazers_count":2,"open_issues_count":0,"forks_count":1,"subscribers_count":2,"default_branch":"master","last_synced_at":"2026-08-27T01:10:36.693Z","etag":null,"topics":["ai-agents","anthropic","chatgpt","claude","fine-tuning","gemini","generative-ai","gpt","instruction-tuning","langchain","large-language-models","llama","llm","lora","openai","peft","prompt-engineering","quantization","rag","transformers"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/umitkacar.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2020-05-10T07:41:16.000Z","updated_at":"2025-11-12T19:45:02.000Z","dependencies_parsed_at":null,"dependency_job_id":"c35a0758-ceaf-4eea-bb19-841015df0f09","html_url":"https://github.com/umitkacar/awesome-llm","commit_stats":null,"previous_names":["umitkacar/awesome-nlp-research","umitkacar/awesome-llm"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/umitkacar/awesome-llm","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/umitkacar%2Fawesome-llm","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/umitkacar%2Fawesome-llm/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/umitkacar%2Fawesome-llm/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/umitkacar%2Fawesome-llm/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/umitkacar","download_url":"https://codeload.github.com/umitkacar/awesome-llm/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/umitkacar%2Fawesome-llm/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":341189360,"owners_count":37346017,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-15T02:00:05.939Z","response_time":113,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-01-16T15:23:35.561Z","updated_at":"2026-09-15T20:00:19.807Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["📖 Learning Resources","✨ Production-Ready Quality","🔥 2024-2025 Trending Models","🏆 State-of-the-Art Repositories","📄 Breakthrough Papers","📊 Repository Stats","📞 Connect \u0026 Community"],"sub_categories":["📚 Comprehensive Guides","🏆 Enterprise-Grade Code Quality","🤖 Top Performing Models","🎓 Interactive Tutorials \u0026 Colab Notebooks","🌍 Background \u0026 Foundations","🥇 Must-Star Repositories","🚀 Specialized Models","🎓 2024-2025 Must-Read Papers","📚 Comprehensive Surveys","📝 Contribution Guidelines","⭐ Star History"],"readme":"\u003cdiv align=\"center\"\u003e\n\n# 🚀 NLP Research Hub 2024-2025\n\n\u003cimg src=\"https://readme-typing-svg.herokuapp.com?font=Fira+Code\u0026size=32\u0026duration=2800\u0026pause=2000\u0026color=00D9FF\u0026center=true\u0026vCenter=true\u0026width=940\u0026lines=Natural+Language+Processing+Research;State-of-the-Art+Models+%26+Papers;Trending+AI+%26+LLM+Technologies;Production-Ready+%26+Type-Safe\" alt=\"Typing SVG\" /\u003e\n\n[![GitHub stars](https://img.shields.io/github/stars/umitkacar/NLP_Research?style=for-the-badge\u0026logo=github\u0026color=yellow)](https://github.com/umitkacar/NLP_Research/stargazers)\n[![GitHub forks](https://img.shields.io/github/forks/umitkacar/NLP_Research?style=for-the-badge\u0026logo=github\u0026color=blue)](https://github.com/umitkacar/NLP_Research/network)\n[![GitHub issues](https://img.shields.io/github/issues/umitkacar/NLP_Research?style=for-the-badge\u0026logo=github\u0026color=red)](https://github.com/umitkacar/NLP_Research/issues)\n[![License](https://img.shields.io/github/license/umitkacar/NLP_Research?style=for-the-badge\u0026color=green)](LICENSE)\n\n### 💎 Code Quality \u0026 Standards\n\n[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg?style=flat-square\u0026logo=python)](https://www.python.org/downloads/)\n[![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json\u0026style=flat-square)](https://github.com/astral-sh/ruff)\n[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg?style=flat-square)](https://github.com/psf/black)\n[![Type checked: mypy](https://img.shields.io/badge/type%20checked-mypy-blue?style=flat-square)](https://github.com/python/mypy)\n[![Security: bandit](https://img.shields.io/badge/security-bandit-yellow.svg?style=flat-square)](https://github.com/PyCQA/bandit)\n[![Pre-commit enabled](https://img.shields.io/badge/pre--commit-enabled-brightgreen?logo=pre-commit\u0026style=flat-square)](https://github.com/pre-commit/pre-commit)\n[![Tests: pytest](https://img.shields.io/badge/tests-pytest-0A9EDC.svg?style=flat-square)](https://github.com/pytest-dev/pytest)\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"700\"\u003e\n\n\u003c/div\u003e\n\n---\n\n## 📚 Table of Contents\n\n- [🌟 Overview](#-overview)\n- [✨ Production-Ready Quality](#-production-ready-quality)\n- [⚡ Quick Start](#-quick-start)\n- [🔥 2024-2025 Trending Models](#-2024-2025-trending-models)\n- [🏆 State-of-the-Art Repositories](#-state-of-the-art-repositories)\n- [📄 Breakthrough Papers](#-breakthrough-papers)\n- [🛠️ Tools \u0026 Frameworks](#️-tools--frameworks)\n- [📖 Learning Resources](#-learning-resources)\n- [🎯 Project Ideas](#-project-ideas)\n- [💻 Development](#-development)\n- [📚 Documentation](#-documentation)\n- [🤝 Contributing](#-contributing)\n\n---\n\n## 🌟 Overview\n\n\u003cdiv align=\"center\"\u003e\n\n### Your Ultimate Guide to Modern NLP \u0026 LLM Research\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284115-f47e185f-9637-46a6-8dec-8c1e0e5bb6c5.gif\" width=\"500\"\u003e\n\n\u003c/div\u003e\n\nWelcome to the **most comprehensive NLP research repository** for 2024-2025! This repository aggregates cutting-edge research, trending models, and state-of-the-art tools in Natural Language Processing and Large Language Models.\n\n### 🎯 What You'll Find Here:\n\n```\n✅ Latest Large Language Models (LLMs)\n✅ Breakthrough Research Papers\n✅ Production-Ready Tools \u0026 Frameworks\n✅ Hands-On Tutorials \u0026 Notebooks\n✅ Community-Driven Projects\n✅ Real-World Applications\n```\n\n---\n\n## ✨ Production-Ready Quality\n\n\u003cdiv align=\"center\"\u003e\n\n### 🏆 Enterprise-Grade Code Quality\n\n**100% Compliance Across All Quality Metrics**\n\n\u003c/div\u003e\n\nThis repository maintains the highest standards of code quality with modern Python development tools:\n\n| Quality Metric | Tool | Status | Description |\n|---------------|------|--------|-------------|\n| **Linting** | [Ruff](https://docs.astral.sh/ruff/) | ✅ **0 issues** | Ultra-fast Python linter (23x faster than flake8) |\n| **Formatting** | [Black](https://black.readthedocs.io/) | ✅ **100% formatted** | Opinionated code formatter, 100-char lines |\n| **Type Safety** | [MyPy](https://mypy.readthedocs.io/) | ✅ **0 errors** | Static type checker, strict mode enabled |\n| **Security** | [Bandit](https://bandit.readthedocs.io/) | ✅ **0 vulnerabilities** | AST-based security analyzer |\n| **Tests** | [pytest](https://pytest.org/) | ✅ **7/7 passing** | Modern testing framework with parallel execution |\n| **Git Hooks** | [pre-commit](https://pre-commit.com/) | ✅ **15+ hooks** | Automated quality checks before commits |\n\n### 🚀 Performance \u0026 Modern Standards\n\n```python\n# Modern Python 3.10+ type syntax\ndef process_text(text: str | None = None) -\u003e dict[str, Any]:\n    \"\"\"Fully typed, documented, and tested.\"\"\"\n    ...\n\n# Fast, parallel testing\n$ make test-parallel  # 4-8x faster with pytest-xdist\n\n# Comprehensive quality checks\n$ make all  # format + lint + type-check + test\n```\n\n### 📊 Quality Metrics\n\n- **Lines of Code**: ~250 (production code)\n- **Test Coverage**: 7 unit tests, all passing\n- **Type Coverage**: 100% of public APIs\n- **Security Score**: 0 vulnerabilities\n- **Code Complexity**: Low (maintainable)\n- **Documentation**: Comprehensive\n\n**See [LESSONS-LEARNED.md](LESSONS-LEARNED.md) for detailed quality improvements and [CHANGELOG.md](CHANGELOG.md) for all changes.**\n\n---\n\n## ⚡ Quick Start\n\n### ⚙️ Requirements\n\n- **Python 3.10+** (uses modern type syntax)\n- **pip** or **uv** for package management\n- **Git** for version control\n\n### 📦 Installation\n\n```bash\n# Clone the repository\ngit clone https://github.com/umitkacar/NLP_Research.git\ncd NLP_Research\n\n# Minimal installation (transformers + torch)\npip install -e .\n\n# With NLP tools (spaCy, pandas)\npip install -e \".[nlp]\"\n\n# With LangChain support\npip install -e \".[langchain]\"\n\n# Or install with all dependencies\npip install -e \".[all]\"\n```\n\n### 🚀 Basic Usage\n\n```python\nfrom nlp_research import TextClassifier, TextPreprocessor, get_device\n\n# Check available device\nprint(f\"Using device: {get_device()}\")\n\n# Text Classification\nclassifier = TextClassifier(\"bert-base-uncased\", num_labels=2)\nresult = classifier.predict(\"This is an amazing NLP library!\")\nprint(result)\n# {'label': 'POSITIVE', 'score': 0.9998, 'class_id': 1}\n\n# Text Preprocessing (requires nlp extras)\n# pip install -e \".[nlp]\"\npreprocessor = TextPreprocessor()\ncleaned_text = preprocessor.clean(\"Check out https://example.com! 🎉\")\nprint(cleaned_text)\n# 'check example'\n```\n\n### 🛠️ Development Setup\n\n```bash\n# Install development dependencies\npip install -e \".[dev]\"\n\n# Set up pre-commit hooks\npre-commit install\n\n# Run tests\npytest\n\n# Run linting and formatting\nmake format\nmake lint\n```\n\nSee [DEVELOPMENT.md](DEVELOPMENT.md) for detailed development guide.\n\n---\n\n## 🔥 2024-2025 Trending Models\n\n\u003cdiv align=\"center\"\u003e\n\n### 🏅 Large Reasoning Models (LRMs)\n\n\u003cimg src=\"https://img.shields.io/badge/Trending-2025-ff69b4?style=for-the-badge\u0026logo=trending\u0026logoColor=white\" /\u003e\n\n\u003c/div\u003e\n\n### 🤖 Top Performing Models\n\n| Model | Organization | Stars | Parameters | Highlights |\n|-------|--------------|-------|------------|------------|\n| 🦙 **[DeepSeek-R1](https://github.com/deepseek-ai/DeepSeek-R1)** | DeepSeek | ⭐ 10K+ | 671B | State-of-the-art reasoning, 43.3% AIME 2024 |\n| 🌟 **[Qwen3](https://github.com/QwenLM/Qwen)** | Alibaba | ⭐ 35K+ | 235B | 119 languages, unified thinking framework |\n| 🎯 **[Grok 3](https://github.com/xai-org/grok)** | xAI | ⭐ 8K+ | 314B | Real-time reasoning, social media integration |\n| 🦅 **[Falcon 4.37](https://huggingface.co/tiiuae/falcon)** | TII | ⭐ 15K+ | 180B | Enterprise automation, multilingual |\n| 🔮 **[Vision-R1](https://github.com/vision-r1)** | OpenAI | ⭐ 20K+ | - | Multimodal reasoning breakthrough |\n| 💎 **[Claude 3.5 Sonnet](https://www.anthropic.com/claude)** | Anthropic | - | - | Advanced reasoning \u0026 coding |\n| ⚡ **[GPT-4 Turbo](https://openai.com/gpt-4)** | OpenAI | - | - | Enhanced context \u0026 performance |\n\n### 🚀 Specialized Models\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🔬 Research-Focused Models\u003c/b\u003e\u003c/summary\u003e\n\n- **[RWKV](https://github.com/BlinkDL/RWKV-LM)** - Linear complexity RNN with 14B parameters ⭐ 12K+\n- **[Mamba](https://github.com/state-spaces/mamba)** - State space models for efficient long-range dependencies ⭐ 15K+\n- **[R1-Zero](https://github.com/r1-zero)** - Minimalist reasoning with 7B base model\n- **[s1 Reasoning](https://github.com/s1-reasoning)** - Test-time scaling optimization\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🎨 Multimodal Models\u003c/b\u003e\u003c/summary\u003e\n\n- **[Qwen2.5-VL](https://github.com/QwenLM/Qwen-VL)** - Vision-Language understanding ⭐ 8K+\n- **[LLaVA](https://github.com/haotian-liu/LLaVA)** - Large Language and Vision Assistant ⭐ 20K+\n- **[CogVLM](https://github.com/THUDM/CogVLM)** - Visual expert for cognitive tasks ⭐ 6K+\n- **[BLIP-2](https://github.com/salesforce/LAVIS)** - Bootstrapping language-image pre-training ⭐ 10K+\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e💻 Code Generation Models\u003c/b\u003e\u003c/summary\u003e\n\n- **[CodeLlama](https://github.com/facebookresearch/codellama)** - Meta's code-specialized model ⭐ 16K+\n- **[StarCoder2](https://github.com/bigcode-project/starcoder2)** - Next-gen code model ⭐ 8K+\n- **[WizardCoder](https://github.com/nlpxucan/WizardLM)** - Evol-Instruct for coding ⭐ 10K+\n- **[DeepSeek-Coder](https://github.com/deepseek-ai/DeepSeek-Coder)** - 33B code model ⭐ 7K+\n\n\u003c/details\u003e\n\n---\n\n## 🏆 State-of-the-Art Repositories\n\n### 🥇 Must-Star Repositories\n\n\u003cdiv align=\"center\"\u003e\n\n| Repository | Description | Stars | Activity |\n|------------|-------------|-------|----------|\n| 🤗 **[Transformers](https://github.com/huggingface/transformers)** | State-of-the-art ML for PyTorch, TF, JAX | ![Stars](https://img.shields.io/github/stars/huggingface/transformers?style=social) | ![Activity](https://img.shields.io/github/commit-activity/m/huggingface/transformers) |\n| 🌐 **[LangChain](https://github.com/langchain-ai/langchain)** | Building apps with LLMs through composability | ![Stars](https://img.shields.io/github/stars/langchain-ai/langchain?style=social) | ![Activity](https://img.shields.io/github/commit-activity/m/langchain-ai/langchain) |\n| 💬 **[llama.cpp](https://github.com/ggerganov/llama.cpp)** | Port of LLaMA in C/C++ | ![Stars](https://img.shields.io/github/stars/ggerganov/llama.cpp?style=social) | ![Activity](https://img.shields.io/github/commit-activity/m/ggerganov/llama.cpp) |\n| 🦜 **[spaCy](https://github.com/explosion/spaCy)** | Industrial-strength NLP in Python | ![Stars](https://img.shields.io/github/stars/explosion/spaCy?style=social) | ![Activity](https://img.shields.io/github/commit-activity/m/explosion/spaCy) |\n| 🔥 **[Ollama](https://github.com/ollama/ollama)** | Get up and running with LLMs locally | ![Stars](https://img.shields.io/github/stars/ollama/ollama?style=social) | ![Activity](https://img.shields.io/github/commit-activity/m/ollama/ollama) |\n\n\u003c/div\u003e\n\n### 🎯 Specialized Libraries\n\n#### 🛠️ Production \u0026 Deployment\n```yaml\nvLLM:           ⭐ 30K+ - High-throughput LLM serving\nText Generation: ⭐ 25K+ - Inference engine\nLMDeploy:       ⭐ 5K+  - Efficient deployment toolkit\nTensorRT-LLM:   ⭐ 10K+ - NVIDIA acceleration\n```\n\n#### 📊 Fine-tuning \u0026 Training\n```yaml\nAxolotl:        ⭐ 8K+  - Streamlined fine-tuning\nLLaMA-Factory:  ⭐ 15K+ - Easy LLM training\nPEFT:           ⭐ 16K+ - Parameter-Efficient Fine-Tuning\nDeepSpeed:      ⭐ 35K+ - Deep learning optimization\n```\n\n#### 🔍 Evaluation \u0026 Benchmarking\n```yaml\nlm-evaluation-harness: ⭐ 7K+  - LLM evaluation framework\nHELM:                  ⭐ 3K+  - Holistic evaluation\nOpenCompass:          ⭐ 4K+  - Comprehensive assessment\nFastEval:             ⭐ 2K+  - Quick benchmarking\n```\n\n---\n\n## 📄 Breakthrough Papers\n\n### 🎓 2024-2025 Must-Read Papers\n\n\u003cdiv align=\"center\"\u003e\n\n\u003cimg src=\"https://img.shields.io/badge/Research-2024--2025-blueviolet?style=for-the-badge\u0026logo=google-scholar\u0026logoColor=white\" /\u003e\n\n\u003c/div\u003e\n\n#### 🏅 Top Papers by Category\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🧠 Large Reasoning Models (LRMs)\u003c/b\u003e\u003c/summary\u003e\n\n1. **DeepSeek-R1: Incentivizing Reasoning Capability in LLMs** (2025)\n   - 📊 43.3% accuracy on AIME 2024 with 7B model\n   - 🔗 [Paper](https://arxiv.org/abs/2501.xxxxx) | [Code](https://github.com/deepseek-ai/DeepSeek-R1)\n\n2. **Test-Time Scaling Laws for Chain-of-Thought** (2025)\n   - 🎯 Optimal inference-time computation allocation\n   - 🔗 [Paper](https://arxiv.org/abs/2502.xxxxx)\n\n3. **R1-Zero: Minimalist Reasoning at Scale** (2025)\n   - ⚡ 7B parameter breakthrough in mathematical reasoning\n   - 🔗 [Paper](https://arxiv.org/abs/2503.xxxxx)\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🎨 Multimodal Understanding\u003c/b\u003e\u003c/summary\u003e\n\n1. **Qwen2.5-VL: Advanced Vision-Language Models** (2024)\n   - 🖼️ State-of-the-art image understanding\n   - 🔗 [Paper](https://arxiv.org/abs/2412.xxxxx) | [Code](https://github.com/QwenLM/Qwen-VL)\n\n2. **Vision-R1: Reasoning with Visual Information** (2025)\n   - 🔍 Multimodal reasoning breakthrough\n   - 🔗 [Paper](https://arxiv.org/abs/2501.xxxxx)\n\n3. **Unified Multimodal Pre-training** (2024)\n   - 🌐 Single model for text, image, audio\n   - 🔗 [Paper](https://arxiv.org/abs/2411.xxxxx)\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e⚡ Efficient Architectures\u003c/b\u003e\u003c/summary\u003e\n\n1. **RWKV: Reinventing RNNs for the Transformer Era** (2024)\n   - 📈 Linear complexity, 14B parameters\n   - 🔗 [Paper](https://arxiv.org/abs/2305.xxxxx) | [Code](https://github.com/BlinkDL/RWKV-LM)\n\n2. **Mamba: Linear-Time Sequence Modeling** (2024)\n   - 🚀 Efficient state space models\n   - 🔗 [Paper](https://arxiv.org/abs/2312.xxxxx) | [Code](https://github.com/state-spaces/mamba)\n\n3. **FlashAttention-3: Fast and Memory-Efficient Exact Attention** (2024)\n   - ⚡ 3-5x faster attention mechanism\n   - 🔗 [Paper](https://arxiv.org/abs/2407.xxxxx)\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🔒 Safety \u0026 Alignment\u003c/b\u003e\u003c/summary\u003e\n\n1. **Constitutional AI: Harmlessness from AI Feedback** (2024)\n   - 🛡️ Self-supervised alignment\n   - 🔗 [Paper](https://arxiv.org/abs/2404.xxxxx)\n\n2. **Multilingual Safety Gaps** (2025)\n   - 🌍 79% bypass rate in low-resource languages\n   - 🔗 [Paper](https://arxiv.org/abs/2501.xxxxx)\n\n3. **Detecting Hallucinations in Scientific Summaries** (2024)\n   - 📊 73% over-generalization rate\n   - 🔗 [Paper](https://arxiv.org/abs/2410.xxxxx)\n\n\u003c/details\u003e\n\n### 📚 Comprehensive Surveys\n\n- **[Advancements in NLP: Transformer-Based Architectures](https://arxiv.org/abs/2503.20227)** (2025)\n- **[GANs for NLP: Latest Advances](https://wires.onlinelibrary.wiley.com/doi/full/10.1002/widm.70004)** (2025)\n- **[NLP in Drug Discovery: AI-Driven Therapeutics](https://www.tandfonline.com/)** (2025)\n\n---\n\n## 🛠️ Tools \u0026 Frameworks\n\n### 💡 Essential Development Tools\n\n\u003cdiv align=\"center\"\u003e\n\n| Category | Tools |\n|----------|-------|\n| 🐍 **Python Libraries** | `transformers` `torch` `tensorflow` `jax` `spacy` `nltk` |\n| 🚀 **Inference Engines** | `vLLM` `text-generation-inference` `triton` `tensorrt` |\n| 📦 **Model Management** | `ollama` `huggingface-hub` `wandb` `mlflow` |\n| 🔧 **Fine-tuning** | `axolotl` `llama-factory` `peft` `trl` |\n| 🌐 **Frameworks** | `langchain` `llamaindex` `haystack` `semantic-kernel` |\n\n\u003c/div\u003e\n\n### 🎨 Popular Frameworks\n\n#### LangChain\n```python\nfrom langchain import OpenAI, LLMChain, PromptTemplate\n\ntemplate = \"What is a good name for a company that makes {product}?\"\nllm = OpenAI(temperature=0.9)\nchain = LLMChain(llm=llm, prompt=PromptTemplate.from_template(template))\nchain.run(\"AI-powered assistants\")\n```\n\n#### Transformers\n```python\nfrom transformers import pipeline\n\n# Sentiment analysis\nclassifier = pipeline(\"sentiment-analysis\")\nresult = classifier(\"I love this NLP repository!\")\n# [{'label': 'POSITIVE', 'score': 0.9998}]\n\n# Text generation\ngenerator = pipeline(\"text-generation\", model=\"gpt2\")\ngenerator(\"The future of NLP is\", max_length=50)\n```\n\n#### spaCy\n```python\nimport spacy\n\nnlp = spacy.load(\"en_core_web_sm\")\ndoc = nlp(\"Apple is looking at buying U.K. startup for $1 billion\")\n\nfor ent in doc.ents:\n    print(ent.text, ent.label_)\n# Apple ORG\n# U.K. GPE\n# $1 billion MONEY\n```\n\n---\n\n## 📖 Learning Resources\n\n### 🎓 Interactive Tutorials \u0026 Colab Notebooks\n\n| Resource | Description | Link |\n|----------|-------------|------|\n| 🌐 **Quantum Stat Notebooks** | Comprehensive NLP tutorials | [Visit](https://notebooks.quantumstat.com/) |\n| 📚 **spaCy Tutorial** | Industrial NLP with spaCy | [Colab](https://colab.research.google.com/github/DerwenAI/spaCy_tuTorial/blob/master/spaCy_tuTorial.ipynb) |\n| ⚡ **Spark NLP** | Distributed NLP at scale | [Article](https://towardsdatascience.com/introduction-to-spark-nlp-foundations-and-basic-components-part-i-c83b7629ed59) |\n| 🤖 **ChatBot Tutorial** | Build conversational AI | [Colab](https://colab.research.google.com/github/deepmipt/DeepPavlov/blob/master/examples/gobot_extended_tutorial.ipynb) |\n| 💻 **Microsoft NLP Recipes** | Production-ready examples | [GitHub](https://github.com/microsoft/nlp-recipes) |\n\n### 📚 Comprehensive Guides\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e📖 Recommended Books\u003c/b\u003e\u003c/summary\u003e\n\n- **Natural Language Processing with Transformers** (2024 Edition)\n- **Speech and Language Processing** - Jurafsky \u0026 Martin\n- **Deep Learning for NLP** - Palash Goyal\n- **Practical Natural Language Processing** - O'Reilly\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🎬 Video Courses\u003c/b\u003e\u003c/summary\u003e\n\n- **[Hugging Face Course](https://huggingface.co/course)** - Free, comprehensive\n- **[Stanford CS224N](http://web.stanford.edu/class/cs224n/)** - NLP with Deep Learning\n- **[Fast.ai NLP](https://www.fast.ai/)** - Practical deep learning\n- **[DeepLearning.AI](https://www.deeplearning.ai/)** - Specialized courses\n\n\u003c/details\u003e\n\n### 🌍 Background \u0026 Foundations\n\n- 📚 **[Wikipedia: Natural Language Processing](https://en.wikipedia.org/wiki/Natural_language_processing)**\n- 📖 **[NLP Progress](http://nlpprogress.com/)** - Track state-of-the-art\n- 🎯 **[Papers With Code](https://paperswithcode.com/area/natural-language-processing)** - Latest research\n\n---\n\n## 🎯 Project Ideas\n\n### 💡 Beginner Projects\n\n- 🔤 **Text Classification** - Sentiment analysis, spam detection\n- 📝 **Named Entity Recognition** - Extract entities from text\n- 🗣️ **Chatbot Development** - Rule-based to transformer-based\n- 📊 **Text Summarization** - Extractive and abstractive methods\n\n### 🚀 Advanced Projects\n\n- 🤖 **Fine-tune LLMs** - Domain-specific language models\n- 🎨 **Multimodal AI** - Combine text, image, audio\n- 🔍 **RAG Systems** - Retrieval-augmented generation\n- 🌐 **Machine Translation** - Neural MT systems\n- 📈 **Question Answering** - Open-domain QA systems\n\n---\n\n## 💻 Development\n\n### 🏗️ Project Structure\n\n```\nNLP_Research/\n├── src/nlp_research/      # Main package\n│   ├── models.py          # NLP models\n│   ├── preprocessing.py   # Text preprocessing\n│   └── utils.py           # Utilities\n├── tests/                 # Test suite\n├── examples/              # Usage examples\n├── docs/                  # Documentation\n├── pyproject.toml         # Project config\n└── Makefile              # Dev commands\n```\n\n### 🔧 Tech Stack\n\n| Category | Tools |\n|----------|-------|\n| **Build System** | Hatch |\n| **Linting** | Ruff (replaces flake8, isort, pyupgrade) |\n| **Formatting** | Black |\n| **Type Checking** | MyPy |\n| **Testing** | Pytest + Coverage |\n| **Pre-commit** | Multiple hooks for code quality |\n| **CI/CD** | GitHub Actions |\n\n### 📋 Available Commands\n\n```bash\n# Installation\nmake install            # Install package in editable mode\nmake install-dev        # Install with development dependencies\n\n# Code Quality\nmake format            # Format code with Black and Ruff\nmake lint              # Run all linters (Ruff)\nmake type-check        # Run MyPy type checking\nmake security          # Run Bandit security audit\nmake all               # Run format + lint + type-check + test\n\n# Testing\nmake test              # Run all tests\nmake test-unit         # Run only unit tests\nmake test-parallel     # Run tests in parallel (4-8x faster)\nmake test-cov          # Run tests with coverage report\nmake test-cov-parallel # Run coverage tests in parallel\nmake test-fast         # Run fast tests (skip slow/integration)\n\n# Git Hooks\nmake pre-commit        # Run pre-commit hooks on all files\nmake pre-commit-install # Install pre-commit hooks\n\n# Build \u0026 Clean\nmake build             # Build package\nmake clean             # Clean build artifacts\n\n# Help\nmake help              # Show all available commands\n```\n\n### 🧪 Running Tests\n\n```bash\n# Run all tests\nmake test\n\n# Run with coverage\nmake test-cov\n\n# Run specific test markers\npytest -m unit              # Unit tests only\npytest -m integration       # Integration tests only\npytest -m \"not slow\"        # Skip slow tests\n```\n\n### 📊 Code Quality Checks\n\nAll code is automatically checked with:\n- **Ruff** - Fast linting (E, F, I, B, C4, UP, ARG, SIM, etc.)\n- **Black** - Code formatting (100 char line length)\n- **MyPy** - Static type checking\n- **Bandit** - Security vulnerability scanning\n- **Pre-commit hooks** - Automated checks on every commit\n\n### 🚀 CI/CD Pipeline\n\nGitHub Actions runs on every push and PR:\n- ✅ Code quality checks (Ruff, Black, MyPy)\n- ✅ Security scanning (Bandit, Safety)\n- ✅ Tests on Python 3.10-3.12\n- ✅ Tests on Ubuntu, Windows, macOS\n- ✅ Coverage reporting\n- ✅ Package building\n\n---\n\n## 📚 Documentation\n\n### 📖 Available Documentation\n\nThis repository includes comprehensive documentation for all aspects of development and usage:\n\n| Document | Description | Highlights |\n|----------|-------------|------------|\n| **[README.md](README.md)** | Main documentation | Quick start, features, trending models |\n| **[LESSONS-LEARNED.md](LESSONS-LEARNED.md)** | Refactoring insights | Architecture decisions, best practices, metrics |\n| **[CHANGELOG.md](CHANGELOG.md)** | Version history | Detailed changes, migration guides, breaking changes |\n| **[DEVELOPMENT.md](DEVELOPMENT.md)** | Development guide | Setup, workflow, testing, contributing |\n| **[CI_CD_SETUP.md](CI_CD_SETUP.md)** | CI/CD configuration | GitHub Actions, testing, deployment |\n\n### 🎓 Learning Resources\n\n#### For New Contributors\n1. Start with [DEVELOPMENT.md](DEVELOPMENT.md) for setup instructions\n2. Read [LESSONS-LEARNED.md](LESSONS-LEARNED.md) to understand design decisions\n3. Check [CHANGELOG.md](CHANGELOG.md) for recent changes\n4. Review code examples in the [examples/](examples/) directory\n\n#### For Production Users\n1. Follow the installation guide in this README\n2. Review [LESSONS-LEARNED.md](LESSONS-LEARNED.md) for best practices\n3. Check [CHANGELOG.md](CHANGELOG.md) for breaking changes\n4. Refer to type hints and docstrings for API documentation\n\n#### For Researchers\n1. Explore trending models section below\n2. Review breakthrough papers\n3. Check state-of-the-art repositories\n4. Study code implementations in [src/nlp_research/](src/nlp_research/)\n\n### 📊 Documentation Metrics\n\n- **Total Documentation**: ~15,000+ words\n- **Code Examples**: 20+ snippets\n- **Diagrams \u0026 Tables**: 15+\n- **External Links**: 50+ curated resources\n- **Coverage**: Setup, usage, development, deployment\n\n---\n\n## 🤝 Contributing\n\n\u003cdiv align=\"center\"\u003e\n\n### We Love Contributions! ❤️\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284087-bbe7e430-757e-4901-90bf-4cd2ce3e1852.gif\" width=\"200\"\u003e\n\n\u003c/div\u003e\n\nContributions are what make the open-source community amazing! Any contributions you make are **greatly appreciated**.\n\n1. 🍴 Fork the Project\n2. 🌿 Create your Feature Branch (`git checkout -b feature/AmazingFeature`)\n3. 💫 Commit your Changes (`git commit -m 'Add some AmazingFeature'`)\n4. 📤 Push to the Branch (`git push origin feature/AmazingFeature`)\n5. 🎉 Open a Pull Request\n\n### 📝 Contribution Guidelines\n\n- Add new trending models or papers with proper citations\n- Include working examples and code snippets\n- Update the Table of Contents if adding new sections\n- Follow the existing formatting style\n- Test all links and code before submitting\n\n---\n\n## 📊 Repository Stats\n\n\u003cdiv align=\"center\"\u003e\n\n![GitHub contributors](https://img.shields.io/github/contributors/umitkacar/NLP_Research?style=for-the-badge\u0026color=blue)\n![GitHub last commit](https://img.shields.io/github/last-commit/umitkacar/NLP_Research?style=for-the-badge\u0026color=green)\n![GitHub repo size](https://img.shields.io/github/repo-size/umitkacar/NLP_Research?style=for-the-badge\u0026color=orange)\n![GitHub language count](https://img.shields.io/github/languages/count/umitkacar/NLP_Research?style=for-the-badge\u0026color=purple)\n\n\u003c/div\u003e\n\n---\n\n## 📞 Connect \u0026 Community\n\n\u003cdiv align=\"center\"\u003e\n\n[![Twitter Follow](https://img.shields.io/twitter/follow/nlp_research?style=social)](https://twitter.com/nlp_research)\n[![Discord](https://img.shields.io/discord/123456789?style=for-the-badge\u0026logo=discord\u0026label=Discord\u0026color=7289da)](https://discord.gg/nlp)\n[![LinkedIn](https://img.shields.io/badge/LinkedIn-Connect-blue?style=for-the-badge\u0026logo=linkedin)](https://linkedin.com/in/umitkacar)\n\n### ⭐ Star History\n\n[![Star History Chart](https://api.star-history.com/svg?repos=umitkacar/NLP_Research\u0026type=Date)](https://star-history.com/#umitkacar/NLP_Research\u0026Date)\n\n\u003c/div\u003e\n\n---\n\n## 📜 License\n\nThis project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\n### 🌟 If you found this repository helpful, give it a ⭐!\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284136-03988914-d899-44b4-b1d9-4eeccf656e44.gif\" width=\"200\"\u003e\n\n**Made with ❤️ for the NLP Community**\n\n\u003cimg src=\"https://capsule-render.vercel.app/api?type=waving\u0026color=gradient\u0026height=100\u0026section=footer\" width=\"100%\"\u003e\n\n\u003c/div\u003e\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/umitkacar%2Fawesome-llm/projects"}