Xing Niu, Prashant Mathur, Marcello Federico, Brian Thompson, Juan Zuluaga-Gomez, Zhaocheng Huang, Rohit Paturi, Sundararajan Srinavasan
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end and multi-task training model, named Speaker-Turn Aware Conversational Speech Translation, that combines automatic speech recognition, speech translation and speaker turn detection using special tokens in a serialized labeling format. We run experiments on the Fisher-CALLHOME corpus, which we adapted by merging the two single-speaker channels into one multi-speaker channel, thus representing the more realistic and challenging scenario with multi-speaker turns and cross-talk. Experimental results across single-and multi-speaker conditions and against conventional ST systems, show that our model outperforms the reference systems on the multi-speaker condition, while attaining comparable performance on the single-speaker condition. We release scripts for data processing and model training. 1 * Work conducted during an internship at Amazon.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2311.00697")
get_code_for_paper("2311.00697")
have("2311.00697")
Connect an agent — have() is free.