SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2305.14838 · NeurIPS · 2023

ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text Translation

Michael Zeng, Xuedong Huang, Shujie Liu, Chenyang Le, Yao Qian, Long Zhou, Yanmin Qian

arXiv · PDF · Open in the Atlas

Code that ran

We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite architecture of public pretrained speech-only and language-only models and optimized data-efficiently for spoken language tasks. Particularly, we propose to incorporate cross-modality learning into transfer learning and conduct them simultaneously for downstream tasks in a multi-task learning manner. Our approach has demonstrated effectiveness in end-to-end speech-to-text translation tasks, achieving a new state-of-the-art average BLEU score of 31.5 on the multilingual speech to English text translation task for 21 languages, as measured on the public CoVoST2 evaluation set. 2

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2305.14838")
get_code_for_paper("2305.14838")
have("2305.14838")

Connect an agent — have() is free.