awesome-knowledge-distillation
Awesome Knowledge Distillation
https://github.com/dkozlov/awesome-knowledge-distillation
Last synced: about 1 hour ago
JSON representation
-
Caffe
- Face Model Compression by Distilling Knowledge from Neurons
- KnowledgeDistillation Layer (Caffe implementation)
- Knowledge distillation, realized in caffe
- Cross Modal Distillation for Supervision Transfer
- Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised Detection
- Knowledge Distillation via Instance Relationship Graph
-
Keras
-
Lasagne + Theano
-
Lua
-
MXNet
-
PyTorch
- Transformer model distillation
- TinyBERT
- Attention Transfer
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Interpreting Deep Classifier by Visual Distillation of Dark Knowledge
- Mean teachers are better role models
- Relational Knowledge Distillation
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Fast Human Pose Estimation Pytorch
- MEAL: Multi-Model Ensemble via Adversarial Learning
- MEAL-V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks
- Using Teacher Assistants to Improve Knowledge Distillation
- A Comprehensive Overhaul of Feature Distillation
- Contrastive Representation Distillation
- Channel Distillation
- Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
- MGD: Matching Guided Distillation
- torchdistill: A Modular, Configuration-Driven Framework for Knowledge Distillation
- distiller: A large scale study of Knowledge Distillation
- Knowledge-Distillation-Zoo: Pytorch implementation of various Knowledge Distillation (KD) methods
- Neural Network Distiller by Intel AI Lab: a Python package for neural network compression research.
- KD_Lib : A Pytorch Knowledge Distillation library for benchmarking and extending works in the domains of Knowledge Distillation, Pruning, and Quantization.
- Vision Transformer Distillation
- Cross-Layer Distillation with Semantic Calibration
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge Distillation
- Distilling Knowledge via Knowledge Review
- Hierarchical Self-supervised Augmented Knowledge Distillation
- Causal Distillation for Language Models
- UniversalNER
- MobileSAM
- Logit-Standardization-KD
- Delayed Eps-Shrinking for Faster Once-For-All Training
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge Distillation
- EchoDFKD
- Autoregressive Distillation of Diffusion Transformers (ARD)
- Simple Unsupervised Knowledge Distillation With Space Similarity
-
Tensorflow
- Deep Reinforcement Learning, knowledge transfer
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- Distillation
- An example application of neural network distillation to MNIST
- Data-free Knowledge Distillation for Deep Neural Networks
- Inspired by net2net, network distillation
- Knowledge Distillation using Tensorflow
- Knowledge Distillation Methods with Tensorflow
- Zero-Shot Knowledge Distillation in Deep Networks in ICML2019
- Knowledge_distillation_benchmark via Tensorflow2.0
-
Theano
-
Torch
-
Uncategorized
-
Uncategorized
- Neural Network Ensembles
- Neural Network Ensembles, Cross Validation, and Active Learning
- Combining labeled and unlabeled data with co-training
- Ensemble Methods in Machine Learning
- Model Compression
- Learning with Pseudo-Ensembles
- Cross Modal Distillation for Supervision Transfer
- Distilling Model Knowledge
- Learning Using Privileged Information: Similarity Control and Knowledge Transfer
- Distillation as a Defense to Adversarial Perturbations against Deep Neural Networks
- Do deep convolutional nets really need to be deep and convolutional?
- MobileID: Face Model Compression by Distilling Knowledge from Neurons
- Recurrent Neural Network Training with Dark Knowledge Transfer
- Adapting Models to Signal Degradation using Distillation - Chyi Su, Subhransu Maji, 2016
- Data-Free Knowledge Distillation For Deep Neural Networks
- Local Affine Approximators for Improving Knowledge Transfer
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Revisiting knowledge transfer for training object class detectors
- A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning
- Rocket Launching: A Universal and Efficient Framework for Training Well-performing Light Net
- Data Distillation: Towards Omni-Supervised Learning
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- Learning from Noisy Labels with Distillation - Jia Li, ICCV 2017
- Deep Mutual Learning
- Distilling a Neural Network Into a Soft Decision Tree
- Multimodal Recurrent Neural Networks with Information Transfer Layers for Indoor Scene Labeling - Pui Chau, Gang Wang, 2018
- Born Again Neural Networks
- YASENN: Explaining Neural Networks via Partitioning Activation Sequences
- Knowledge Distillation with Adversarial Samples Supporting Decision Boundary
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Self-supervised knowledge distillation using singular value decomposition
- Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised Detection
- Learning to Steer by Mimicking Features from Heterogeneous Auxiliary Networks
- A Generalized Meta-loss function for regression and classification using privileged information
- Large scale distributed neural network training through online distillation
- KDGAN: Knowledge Distillation with Generative Adversarial Networks
- Deep Face Recognition Model Compression via Knowledge Transfer and Distillation
- Relational Knowledge Distillation
- Graph-based Knowledge Distillation by Multi-head Attention Network
- Knowledge Adaptation for Efficient Semantic Segmentation
- Structured Knowledge Distillation for Semantic Segmentation
- Fast Human Pose Estimation
- MEAL: Multi-Model Ensemble via Adversarial Learning
- Learning Lightweight Lane Detection CNNs by Self Attention Distillation
- Improved Knowledge Distillation via Teacher Assistant: Bridging the Gap Between Student and Teacher - Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Hassan Ghasemzadeh, AAAI 2020
- A Comprehensive Overhaul of Feature Distillation
- Contrastive Representation Distillation
- Distillation-Based Training for Multi-Exit Architectures
- Learning Metrics from Teachers: Compact Networks for Image Embedding
- On the Efficacy of Knowledge Distillation
- Revisit Knowledge Distillation: a Teacher-free Framework
- Ensemble Distribution Distillation
- Improving Generalization and Robustness with Noisy Collaboration in Knowledge Distillation
- Self-training with Noisy Student improves ImageNet classification - Thang Luong, Quoc V. Le, CVPR 2020
- Variational Student: Learning Compact and Sparser Networks in Knowledge Distillation Framework
- Preparing Lessons: Improve Knowledge Distillation with Better Supervision
- Positive-Unlabeled Compression on the Cloud
- Variational Information Distillation for Knowledge Transfer
- Knowledge Distillation via Instance Relationship Graph
- Knowledge Distillation via Route Constrained Optimization
- Similarity-Preserving Knowledge Distillation
- Distilling Object Detectors with Fine-grained Feature Imitation
- Knowledge Squeezed Adversarial Network Compression
- Stagewise Knowledge Distillation
- Knowledge Distillation from Internal Representations
- Knowledge Flow: Improve Upon Your Teachers - Jen Liu, Jian Peng, Alexander G. Schwing, 2019
- Graph Representation Learning via Multi-task Knowledge Distillation
- Deep geometric knowledge distillation with graphs
- Correlation Congruence for Knowledge Distillation
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation
- BAM! Born-Again Multi-Task Networks for Natural Language Understanding - Thang Luong, Urvashi Khandelwal, Christopher D. Manning, Quoc V. Le, ACL 2019
- Self-Knowledge Distillation in Natural Language Processing
- Rethinking Data Augmentation: Self-Supervision and Self-Distillation
- MSD: Multi-Self-Distillation Learning via Multi-classifiers within Deep Neural Networks
- Efficient Video Classification Using Fewer Frames
- Retaining Privileged Information for Multi-Task Learning - Wei Lehman
- Data-Free Learning of Student Networks
- Positive-Unlabeled Compression on the Cloud
- When Does Label Smoothing Help?
- The State of Knowledge Distillation for Classification
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Channel Distillation: Channel-Wise Attention for Knowledge Distillation
- Residual Knowledge Distillation
- ResKD: Residual-Guided Knowledge Distillation
- Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
- MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks
- MGD: Matching Guided Distillation
- Reducing the Teacher-Student Gap via Spherical Knowledge Distillation
- Regularizing Class-wise Predictions via Self-knowledge Distillation
- Training data-efficient image transformers & distillation through attention (DeiT)
- Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks - Jin Yoon, 2020
- Cross-Layer Distillation with Semantic Calibration - Ping Mei, Yuan Zhang, Can Wang, Yan Feng, Chun Chen, AAAI 2021
- Subclass Distillation
- MobileStyleGAN: A Lightweight Convolutional Neural Network for High-Fidelity Image Synthesis
- Knowledge Distillation: A Survey
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge Distillation - Chul Moon, CVPR 2021
- Complementary Relation Contrastive Distillation
- Distilling Knowledge via Knowledge Review
- Hierarchical Self-supervised Augmented Knowledge Distillation
- Causal Distillation for Language Models
- How many Observations are Enough? Knowledge Distillation for Trajectory Forecasting
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
- MobileSAMv2: Faster Segment Anything to Everything - Ho Kim, Choong Seon Hong, 2023
- Dark knowledge
- Model Compression
- Heterogeneous Knowledge Transfer in Video Emotion Recognition, Attribution and Summarization - Gang Jiang, Boyang Li, Leonid Sigal, 2015
- Cross Modal Distillation for Supervision Transfer
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Defensive Collaborative Multi-task Training - Defending against Adversarial Attack towards Deep Neural Networks
- Unifying distillation and privileged information - Paz, Léon Bottou, Bernhard Schölkopf, Vladimir Vapnik, ICLR 2016
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- FitNets: Hints for Thin Deep Nets
- Knowledge Distillation for Small-footprint Highway Networks
- Sequence-Level Knowledge Distillation - papernotes](https://github.com/dennybritz/deeplearning-papernotes/blob/master/notes/seq-knowledge-distillation.md), Yoon Kim, Alexander M. Rush, EMNLP 2016
- Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Learning Loss for Knowledge Distillation with Conditional Adversarial Networks - Chang Hsu, Jiawei Huang, 2017
- Knowledge Projection for Deep Neural Networks
- Moonshine: Distilling with Cheap Convolutions
- Efficient Neural Architecture Search via Parameters Sharing
- Deep Co-Training for Semi-Supervised Image Recognition
- Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples
- Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation
- Contrastive Representation Distillation
- Positive-Unlabeled Compression on the Cloud
- Precision Shaking and DORPO: Conceptual Foundations of LLM Knowledge Distillation Methods
- Neural Network Ensembles, Cross Validation, and Active Learning
- Distilling the Knowledge in a Neural Network
- Learning Using Privileged Information: Similarity Control and Knowledge Transfer
- Learning Transferable Architectures for Scalable Image Recognition
- Retaining Privileged Information for Multi-Task Learning - Wei Lehman, KDD 2019
- When Does Label Smoothing Help?
- On Distillation of Guided Diffusion Models
- Progressive Distillation for Fast Sampling of Diffusion Models
- TRACT: Denoising Diffusion Models with Transitive Closure Time-Distillation
- Adversarial Diffusion Distillation
-
Programming Languages
Categories
Sub Categories
Keywords
pytorch
5
knowledge-distillation
4
image-classification
2
deep-learning
2
network-compression
2
attention
1
truncated-svd
1
regularization
1
quantization
1
pruning-structures
1
pruning
1
onnx
1
jupyter-notebook
1
group-lasso
1
early-exit
1
distillation
1
deep-neural-networks
1
automl-for-compression
1
tensorflow
1
lua
1
lane-detection
1
cnn
1
transformers
1
computer-vision
1
attention-mechanism
1
teacher-student
1
knowledge-transfer
1
iccv2019
1
vision-api
1
tensorflow-basics
1
keras-distillation
1
keras
1
stacked-hourglass-networks
1
pose-estimation
1
transformer
1
semantic-segmentation
1
pytorch-ecosystem
1
pascal-voc
1
object-detection
1
nlp
1
natural-language-processing
1
imagenet
1
google-colab
1
glue
1
colab-notebook
1
coco
1
cifar100
1
cifar10
1
amazon-sagemaker-lab
1
artificial-intelligence
1