SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2410.01036 · EMNLP · 2024

950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages

Marco Gaido, Luisa Bentivogli, Sara Papi, Bruno Kessler, Alessio Brutti, Mauro Cettolo, Matteo Fondazione, Roberto Gretter, Marco Matassoni, Mohamed Nabih

arXiv · PDF · Open in the Atlas

Code that ran

We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing speech FMs (SFMs) fall short of full compliance with the open-source principles, even if claimed otherwise, as no existing SFM has model weights, code, and training data publicly available under open-source terms. In this work, we take the first step toward filling this gap by focusing on the 24 official languages of the European Union (EU). We collect suitable training data by surveying automatic speech recognition datasets and unlabeled speech corpora under open-source compliant licenses, for a total of 950k hours. Additionally, we release automatic transcripts for 441k hours of unlabeled data under the permissive CC-BY license, thereby facilitating the creation of open-source SFMs for the EU languages. github.com/hlt-mt/mosel hf.co/datasets/FBK-MT/mosel Name License Hours Languages Label CommonVoice (Ardila et al., 2020) CC-0 6,732

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2410.01036")
get_code_for_paper("2410.01036")
have("2410.01036")

Connect an agent — have() is free.