SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2005.07143 · 2020

ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

arXiv · PDF · Open in the Atlas

Code that ran

We lifted 27 functions out of this paper's own repositories and ran 26 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.

RepositoryRoleRan
gzhu06/tdspkr-mismatch-study pwc_unofficial 10 of 11
ol-mega/ppca pwc_unofficial 7 of 7
PunkMale/ECAPA-TDNN-CNCeleb pwc_unofficial 4 of 4
LKLQQ/ecapa_tdnn pwc_unofficial 3 of 3
taoruijie/ecapa-tdnn pwc_unofficial 2 of 2
FunctionStatusWhere it lives
ComputeErrorRates Ran PunkMale/ECAPA-TDNN-CNCeleb/tools.py
code served (permissive licence) · get_code("2f8f70723c6fc25c")
SE_Res2Block Ran gzhu06/tdspkr-mismatch-study/backbones/aggregator/ECAPA-TDNN.py
code served (permissive licence) · get_code("bb83d95621d7219f")
calculate_eer Ran gzhu06/tdspkr-mismatch-study/evaluate/eer_monitor.py
code served (permissive licence) · get_code("0c85a144d4c6f68a")
compute_fa_miss Ran LKLQQ/ecapa_tdnn/src/metrics.py
code served (permissive licence) · get_code("8a1f2f7df5f4b6fe")
cosine_similarity Ran gzhu06/tdspkr-mismatch-study/evaluate/eer_monitor.py
code served (permissive licence) · get_code("5bdb987f035ec4ec")
data_catalog Ran gzhu06/tdspkr-mismatch-study/preprocessing/wavform_extract.py
code served (permissive licence) · get_code("83e2c902cae568f3")
findAllUtt Ran PunkMale/ECAPA-TDNN-CNCeleb/dataset.py
code served (permissive licence) · get_code("18d1c1af085a8476")
find_files Ran gzhu06/tdspkr-mismatch-study/preprocessing/wavform_extract.py
code served (permissive licence) · get_code("9dbd7f6108f96407")
get_EER Ran LKLQQ/ecapa_tdnn/src/metrics.py
code served (permissive licence) · get_code("009c08f278c64a6a")
get_EER_from_scores Ran LKLQQ/ecapa_tdnn/src/metrics.py
code served (permissive licence) · get_code("ad04c8794c939048")
get_last_checkpoint_if_any Ran gzhu06/tdspkr-mismatch-study/utils.py
code served (permissive licence) · get_code("d15bcb24b6237cdb")
get_padding_elem Ran ol-mega/ppca/speechbrain/nnet/CNN.py
code served (permissive licence) · get_code("8feeecfbeb765eab")
get_padding_elem_transposed Ran ol-mega/ppca/speechbrain/nnet/CNN.py
code served (permissive licence) · get_code("26b6efc761f6e025")
init_args Ran PunkMale/ECAPA-TDNN-CNCeleb/tools.py
code served (permissive licence) · get_code("fcc13d262968ef5f")
init_args Ran taoruijie/ecapa-tdnn/tools.py
code served (permissive licence) · get_code("9f9dfffa0e162e11")
latest_checkpoint_path Ran gzhu06/tdspkr-mismatch-study/trainer/train_utils.py
code served (permissive licence) · get_code("1d6005e9f2222eee")
natural_sort Ran gzhu06/tdspkr-mismatch-study/utils.py
code served (permissive licence) · get_code("1677ad306233c7b7")
pack_padded_sequence Ran ol-mega/ppca/speechbrain/nnet/RNN.py
code served (permissive licence) · get_code("4d856b195967f759")
pad_packed_sequence Ran ol-mega/ppca/speechbrain/nnet/RNN.py
code served (permissive licence) · get_code("7c9d9bad06c577fb")
parse_arguments Ran ol-mega/ppca/speechbrain/core.py
code served (permissive licence) · get_code("bd8ed1453a55a2df")
pickle2array Ran gzhu06/tdspkr-mismatch-study/utils.py
code served (permissive licence) · get_code("69748fb6d2228cc0")
quantize_midrise Ran ol-mega/ppca/anonymization_methods/quantize_lpc_poles.py
code served (permissive licence) · get_code("298f364b3f304c9f")
quantize_midtread Ran ol-mega/ppca/anonymization_methods/quantize_lpc_poles.py
code served (permissive licence) · get_code("8b5629f75ceda907")
split_data Ran gzhu06/tdspkr-mismatch-study/trainer/train_utils.py
code served (permissive licence) · get_code("92b93a196c3d8b18")
tuneThresholdfromScore Ran PunkMale/ECAPA-TDNN-CNCeleb/tools.py
code served (permissive licence) · get_code("4b2feff2caedf01e")
tuneThresholdfromScore Ran taoruijie/ecapa-tdnn/tools.py
code served (permissive licence) · get_code("e40eb8a3ee8e21ed")
trial_eer Not yet run gzhu06/tdspkr-mismatch-study/evaluate/eer_monitor.py
code served (permissive licence) · get_code("72855f03c7884551")

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

Current speaker verification techniques rely on a neural network to extract speaker representations. The successful x-vector architecture is a Time Delay Neural Network (TDNN) that applies statistics pooling to project variable-length utterances into fixed-length speaker characterizing embeddings. In this paper, we propose multiple enhancements to this architecture based on recent trends in the related fields of face verification and computer vision. Firstly, the initial frame layers can be restructured into 1-dimensional Res2Net modules with impactful skip connections. Similarly to SE-ResNet, we introduce Squeeze-and-Excitation blocks in these modules to explicitly model channel interdependencies. The SE block expands the temporal context of the frame layer by rescaling the channels according to global properties of the recording. Secondly, neural networks are known to learn hierarchical features, with each layer operating on a different level of complexity. To leverage this complementary information, we aggregate and propagate features of different hierarchical levels. Finally, we improve the statistics pooling module with channel-dependent frame attention. This enables the network to focus on different subsets of frames during each of the channel's statistics estimation. The proposed ECAPA-TDNN architecture significantly outperforms state-of-the-art TDNN based systems on the VoxCeleb test sets and the 2019 VoxCeleb Speaker Recognition Challenge.

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2005.07143")
get_code_for_paper("2005.07143")
have("2005.07143")

Connect an agent — have() is free.