{"id":22047144,"url":"https://github.com/zycv/speaker-recognition-based-on-deep-learning-an-overview","last_synced_at":"2026-01-05T14:12:30.235Z","repository":{"id":163941592,"uuid":"380305258","full_name":"zycv/Speaker-Recognition-Based-on-Deep-Learning-An-Overview","owner":"zycv","description":"This repo is to list the references papers of 《Speaker Recognition Based on Deep Learning: An Overview》","archived":false,"fork":false,"pushed_at":"2021-06-26T18:12:13.000Z","size":10,"stargazers_count":38,"open_issues_count":0,"forks_count":5,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-01-28T20:44:03.501Z","etag":null,"topics":["deep-learning","speaker-recognition","speaker-verification"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/zycv.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2021-06-25T17:06:35.000Z","updated_at":"2024-12-12T03:06:42.000Z","dependencies_parsed_at":"2023-07-20T07:31:27.712Z","dependency_job_id":null,"html_url":"https://github.com/zycv/Speaker-Recognition-Based-on-Deep-Learning-An-Overview","commit_stats":null,"previous_names":[],"tags_count":2,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zycv%2FSpeaker-Recognition-Based-on-Deep-Learning-An-Overview","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zycv%2FSpeaker-Recognition-Based-on-Deep-Learning-An-Overview/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zycv%2FSpeaker-Recognition-Based-on-Deep-Learning-An-Overview/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zycv%2FSpeaker-Recognition-Based-on-Deep-Learning-An-Overview/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/zycv","download_url":"https://codeload.github.com/zycv/Speaker-Recognition-Based-on-Deep-Learning-An-Overview/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":245115782,"owners_count":20563234,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["deep-learning","speaker-recognition","speaker-verification"],"created_at":"2024-11-30T13:25:26.403Z","updated_at":"2026-01-05T14:12:30.191Z","avatar_url":"https://github.com/zycv.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# Speaker-Recognition-Based-on-Deep-Learning-An-Overview\nThis repo is to list the references papers of [Speaker Recognition Based on Deep Learning: An Overview](https://arxiv.org/abs/2012.00931)\n\nThe analysis of the paper can refer to here:[基于深度学习的说话人识别概述](https://zhuanlan.zhihu.com/p/381526743)\n\nAll references are sorted in the order of the original paper, and the pdf document can be downloaded in the `Release` on the right. \n\n# References list\n\n\n1 An overview of text-independent speaker recognition From features to supervectors .pdf\n\n2 Forensic speaker recognition.pdf\n\n3 The inference of identity in forensic speaker recognition.pdf\n\n4 an overview of speaker identification accuracy and robustness issues.pdf\n\n5 Talker‐Recognition Procedure Based on Analysis of Variance.pdf\n\n6 Speaker Verification Using Adapted Gaussain Mixture Models.pdf\n\n7 Support vector machines using GMM supervectors for speaker verification.pdf\n\n8 Joint factor analysis versus eigenchannels in speaker recognition.pdf\n\n9 Front-end factor analysis for speaker verification.pdf\n\n10 Bayesian speaker verification with heavy tailed priors.pdf\n\n11 Analysis of i-vector length normalization in speaker recognition systems.pdf\n\n12 A novel scheme for speaker recognition using a phonetically-aware deep neural network.pdf\n\n13 Deep neural networks for small footprint text-dependent speaker verification.pdf\n\n14 Xvectors Robust DNN embeddings for speaker recognition.pdf\n\n15 Voxceleb  a large-scale speaker identification dataset.pdf\n\n17 Speaker recognition by machines and humans  A tutorial review.pdf\n\n18 An overview of automatic speaker recognition technology.pdf\n\n19 An overview of statistical pattern recognition techniques for speaker verification.pdf\n\n20 Spoofing and countermeasures for speaker verification A survey.pdf\n\n21 Speaker diarization A review of recent research.pdf\n\n22 The Attacker's Perspective on Automatic Speaker Verification An Overview.pdf\n\n23 Speaker verification using deep neural networks A review.pdf\n\n24 Learning Discriminative Features for Speaker Identification and Verification.pdf\n\n25 Centroid-based deep metric learning for speaker recognition.pdf\n\n26 An End-to-End Text-Independent Speaker Identification System on Short Utterances.pdf\n\n27 Combining deep embeddings of acoustic and articulatory features for speaker identification.pdf\n\n28 A Memory Augmented Architecture for Continuous Speaker Identification in Meetings.pdf\n\n29 X-vectors Robust neural embeddings for speaker recognition.pdf\n\n30 A study of interspeaker variability in speaker verification.pdf\n\n31 Deep neural network based posteriors for text-dependent speaker verification.pdf\n\n32 Deep Neural Networks and Hidden Markov Models in i-vector-based Text-Dependent Speaker Verification.pdf\n\n33 A deep neural network speaker verification system targeting microphone speech.pdf\n\n34 Exploiting sequence information for text-dependent speaker verification.pdf\n\n35 Text-dependent speaker verification based on i-vectors, neural networks and hidden Markov models.pdf\n\n36 Deep neural network approaches to speaker and language recognition.pdf\n\n37 Deep Neural Networks for extracting Baum-Welch statistics for Speaker Recognition.pdf\n\n38 Time delay deep neural network-based universal background models for speaker recognition.pdf\n\n39 Advances in deep neural network approaches.pdf\n\n40 Phone-centric local variability vector for text-constrained speaker verification.pdf\n\n41 A unified deep neural network for speaker and language recognition.pdf\n\n42 The IBM 2016 speaker recognition system.pdf\n\n43 Application of convolutional neural networks to speaker recognition in noisy conditions.pdf\n\n44 Insights into deep neural networks for speaker recognition.pdf\n\n45 Exploring robustness of DNN RNN for extracting speaker Baum-Welch statistics in mismatched conditions.pdf\n\n47 Exploring the role of phonetic bottleneck features for speaker and language recognition.pdf\n\n48 Analysis and Optimization of Bottleneck Features for Speaker RecognitionK.pdf\n\n49 Augmenting short-term cepstral features with long-term discriminative features for speaker verification of telephone data.pdf\n\n50 Combination of cepstral and phonetically discriminative features for speaker verification.pdf\n\n51 Deep Neural Network Embeddings for Text-Independent Speaker Verification.pdf\n\n52 Towards directly modeling raw speech signal for speaker verification using CNNs.pdf\n\n53 Speaker recognition from raw waveform with sincnet.pdf\n\n54 Avoiding speaker overfitting in end-to-end dnns using raw waveform for text-independent speaker verification.pdf\n\n55 A complete end-to-end speaker verification system using deep neural networks  From raw signals to verification result.pdf\n\n56 Rawnet  Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification.pdf\n\n57 Short utterance compensation in speaker verification via cosine-based teacher-student learning of speaker embeddings.pdf\n\n58 Ensemble additive margin softmax for speaker verification.pdf\n\n59 Voxceleb2 Deep speaker recognition.pdf\n\n60 Utterance-level aggregation for speaker recognition in the wild.pdf\n\n61 Frequency and temporal convolutional attention for text-independent speaker recognition.pdf\n\n62 End-to-end text-independent speaker verification with triplet loss on short utterances..pdf\n\n63 Text-independent speaker verification based on triplet convolutional neural network embeddings.pdf\n\n64 Seq2seq attentional siamese neural networks for text-dependent speaker verification.pdf\n\n65 Orthogonal training for text-independent speaker verification.pdf\n\n66 Orthogonality regularizations for end-to-end speaker.pdf\n\n67 JHU-HLTCOE system for the VoxSRC speaker recognition challenge.pdf\n\n68 Multi-resolution multi-head attention in deep speaker embedding.pdf\n\n69 Deep speaker  an end-to-end neural speaker embedding system.pdf\n\n70 Deep speaker representation using orthogonal decomposition and recombination for speaker verification.pdf\n\n71 Magneto X-vector magnitude estimation network plus offset for improved speaker recognition.pdf\n\n72 Deep Speaker Embeddings for Short-Duration Speaker Verification.pdf\n\n73 Deep Discriminative Embeddings for Duration Robust Speaker Verification.pdf\n\n74 Boundary discriminative large margin cosine loss for text-independent speaker verification.pdf\n\n75 Text-independent speaker verification using 3d convolutional neural networks.pdf\n\n76 Deep speaker feature learning for text-independent speaker verification.pdf\n\n77 Attention-based models for text-dependent speaker verification.pdf\n\n78 Generalized end-to-end loss for speaker verification.pdf\n\n79 End-to-end text-dependent speaker verification.pdf\n\n80 Improving deep CNN networks with long temporal context for text-independent speaker verification.pdf\n\n81 Deep speaker embedding learning with multi-level pooling for text-independent speaker verification.pdf\n\n82 Speaker recognition for multi-speaker conversations using x-vectors.pdf\n\n83 Speaker embedding extraction with phonetic information.pdf\n\n84 Self-Attentive Speaker Embeddings for Text-Independent Speaker Verification..pdf\n\n85 Attentive statistics pooling for deep speaker embedding.pdf\n\n86 Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification.pdf\n\n87 Gaussian-constrained training for speaker verification.pdf\n\n88 Margin matters  Towards more discriminative deep neural network embeddings for speaker recognition.pdf\n\n89 State-of-the-Art Speaker Recognition for Telephone and Video Speech The JHU-MIT Submission for NIST SRE18.pdf\n\n90 Statistics pooling time delay neural network based on X-vector for speaker verification.pdf\n\n91 Bayesian x-vector  Bayesian Neural Network based x-vector System for Speaker Verification.pdf\n\n92 Deep Speaker Embedding Extraction with Channel-Wise Feature Responses and Additive Supervision Softmax Loss Function..pdf\n\n93 An improved deep embedding learning method for short duration speaker verification.pdf\n\n94 An Effective Deep Embedding Learning Architecture for Speaker Verification..pdf\n\n95 Speaker characterization using tdnn-lstm based speaker embedding.pdf\n\n96 A time delay neural network architecture for efficient modeling of long temporal contexts.pdf\n\n97 Self-supervised speaker embeddings.pdf\n\n98 JHU-HLTCOE system for the VoxSRC speaker recognition challenge.pdf\n\n99 Semi-Orthogonal Low-Rank Matrix Factorization for Deep Neural Networks..pdf\n\n100 The JHU Speaker Recognition System for the VOiCES 2019 Challenge.pdf\n\n101 Densely Connected Time Delay Neural Network for Speaker Verification.pdf\n\n102 Compact Speaker Embedding  lrx-vector.pdf\n\n103 Deep residual learning for image recognition.pdf\n\n104 Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms.pdf\n\n105 Wav2Spk  A Simple DNN Architecture for Learning Speaker Embeddings from Waveforms..pdf\n\n106 BERTPHONE Phonetically-aware Encoder Representations for Utterance-level Speaker and Language Recognition.pdf\n\n107 Self-attention encoding and pooling for speaker recognition.pdf\n\n108 Autospeech Neural architecture search for speaker recognition.pdf\n\n109 Evolutionary Algorithm Enhanced Neural Architecture Search for Text-Independent Speaker Verification.pdf\n\n110 A Comparative Re-Assessment of Feature Extractors for Deep Speaker Embeddings.pdf\n\n111 State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and speakers in the wild evaluations.pdf\n\n112 Am-mobilenet1d  A portable model for speaker recognition.pdf\n\n113 Attention is all you need.pdf\n\n114 A structured self-attentive sentence embedding.pdf\n\n115 Deeply Fused Speaker Embeddings for Text-Independent Speaker Verification..pdf\n\n116 An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales.pdf\n\n117 Cnn with phonetic attention for text-independent speaker verification.pdf\n\n118 Self multi-head attention for speaker recognition.pdf\n\n119 Vector-Based Attentive Pooling for Text-Independent Speaker Verification.pdf\n\n120 NetVLAD CNN architecture for weakly supervised place recognition.pdf\n\n121 Ghostvlad for set-based face recognition.pdf\n\n122 Exploring the encoding layer and loss function in end-to-end speaker and language recognition system.pdf\n\n123 Spatial pyramid encoding with convex length normalization for text-independent speaker verification.pdf\n\n124 Total variability layer in deep neural network embeddings for speaker verification.pdf\n\n125 Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System.pdf\n\n126 Shortcut Connections Based Deep Speaker Embeddings for End-to-End Speaker Verification System.pdf\n\n127 A deep neural network for short-segment speaker recognition.pdf\n\n128 Improving multi-scale aggregation using feature pyramid module for robust speaker verification of variable-duration utterances.pdf\n\n129 Sphereface Deep hypersphere embedding for face recognition.pdf\n\n130 Angular Softmax for Short-Duration Text-independent Speaker Verification.pdf\n\n131 On deep speaker embeddings for text-independent speaker recognition.pdf\n\n132 Unified hypersphere embedding for speaker recognition.pdf\n\n133 Dynamic Margin Softmax Loss for Speaker Verification.pdf\n\n134 Large margin softmax loss for speaker verification.pdf\n\n135 On Parameter Adaptation in Softmax-based Cross-Entropy Loss for Improved Convergence Speed and Accuracy in DNN-based Speaker Recognition.pdf\n\n136 Additive margin softmax for face verification.pdf\n\n137 Cosface Large margin cosine loss for deep face recognition.pdf\n\n138 Arcface Additive angular margin loss for deep face recognition.pdf\n\n139 Discriminative neural embedding learning for short-duration text-independent speaker verification.pdf\n\n140 A discriminative feature learning approach for deep face recognition.pdf\n\n141 Multi-Task Discriminative Training of Hybrid DNN-TVM Model for Speaker Verification with Noisy and Far-Field Speech.pdf\n\n142 Multi-task learning for text-dependent speaker verification.pdf\n\n143 Deep feature for text-dependent speaker verification.pdf\n\n144 DNN based speaker embedding using content information for text-dependent speaker verification.pdf\n\n145 On the Usage of Phonetic Information for Text-Independent Speaker Embedding Extraction..pdf\n\n146 Collaborative joint training with multitask recurrent model for speech and speaker recognition.pdf\n\n147 SNR-invariant multitask deep neural networks for robust speaker verification.pdf\n\n148 Multi-task network for noise-robust keyword spotting and speaker verification using CTC-based soft VAD and global query attention.pdf\n\n149 End-to-end attention based text-dependent speaker verification.pdf\n\n150 Deep neural network-based speaker embeddings for end-to-end speaker verification.pdf\n\n151 End-to-end DNN based speaker recognition inspired by i-vector and PLDA.pdf\n\n153 Optimization of False Acceptance Rejection Rates and Decision Threshold for End-to-End Text-Dependent Speaker Verification Systems.pdf\n\n154 Joint i-vector with end-to-end system for short duration text-independent speaker verification.pdf\n\n155 Tristounet  triplet loss for speaker turn embedding.pdf\n\n156 Speaker verification by partial AUC optimization with mahalanobis distance metric learning.pdf\n\n157 End-to-end Text-dependent Speaker Verification Using Novel Distance Measures..pdf\n\n158 Facenet A unified embedding for face recognition and clustering.pdf\n\n159 Prototypical networks for few-shot learning.pdf\n\n160 Few shot speaker recognition using deep neural networks.pdf\n\n161 In defence of metric learning for speaker recognition.pdf\n\n162 Meta-learning for short utterance speaker recognition with imbalance length pairs.pdf\n\n163 Angular Margin Centroid Loss for Text-independent Speaker Recognition.pdf\n\n164 End-to-end losses based on speaker basis vectors and all-speaker hard negative mining for speaker verification.pdf\n\n165 DIHARD II is still hard Experimental results and discussions from the DKU-LENOVO team.pdf\n\n166 Speaker diarization using deep neural network embeddings.pdf\n\n167 Diarization is Hard  Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge..pdf\n\n168 Speaker diarization using latent space clustering in generative adversarial network.pdf\n\n169 Speaker diarization with lstm.pdf\n\n170 Speaker segmentation using deep speaker vectors for fast speaker change scenarios.pdf\n\n171 Context and Uncertainty Modeling for Online Speaker Change Detection.pdf\n\n172 Pre-training of speaker embeddings for low-latency speaker change detection in broadcast news.pdf\n\n173 Speaker, environment and channel change detection and clustering via the bayesian information criterion.pdf\n\n174 Automatic segmentation, classification and clustering of broadcast news audio.pdf\n\n176 An Unsupervised Neural Prediction Framework for Learning Speaker Embeddings Using Recurrent Neural Networks.pdf\n\n178 Convolutional neural network for speaker change detection in telephone speaker diarization system.pdf\n\n179 Speaker diarization with PLDA i-vector scoring and unsupervised calibration.pdf\n\n180 Speaker diarization with i-vectors from DNN senone posteriors.pdf\n\n181 Bayesian HMM Based x-Vector Clustering for Speaker Diarization.pdf\n\n182 But system for the second dihard speech diarization challenge.pdf\n\n183 Optimizing Bayesian HMM based x-vector clustering for the second DIHARD speech diarization challenge.pdf\n\n184 A comparison of neural network feature transforms for speaker diarization.pdf\n\n185 Speaker diarisation using 2D self-attentive combination of embeddings.pdf\n\n186 Speaker embeddings incorporating acoustic conditions for diarization.pdf\n\n187 Speaker diarization with session-level speaker embedding refinement using graph neural networks.pdf\n\n188 LSTM based similarity measurement with spectral clustering for speaker diarization.pdf\n\n189 Self-attentive similarity measurement strategies in speaker diarization.pdf\n\n190 Multilayer bootstrap networks.pdf\n\n191 Universal Background Sparse Coding and Multilayer Bootstrap Network for Speaker Clustering.pdf\n\n192 An Investigation of Speaker Clustering Algorithms in Adverse Acoustic Environments.pdf\n\n193 Fully supervised speaker diarization.pdf\n\n194 DNN-based speaker clustering for speaker diarisation.pdf\n\n195 Active learning based constrained clustering for speaker diarization.pdf\n\n196 Discriminative neural clustering for speaker diarisation.pdf\n\n197 Supervised online diarization with sample mean loss for multi-domain data.pdf\n\n198 Diarization resegmentation in the factor analysis subspace.pdf\n\n199 Speaker Diarization based on Bayesian HMM with Eigenvoice Priors.pdf\n\n200 Neural speech turn segmentation and affinity propagation for speaker diarization.pdf\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzycv%2Fspeaker-recognition-based-on-deep-learning-an-overview","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fzycv%2Fspeaker-recognition-based-on-deep-learning-an-overview","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzycv%2Fspeaker-recognition-based-on-deep-learning-an-overview/lists"}