Shinji Watanabe, Alkis Koudounas, Hayato Futami, Emiru Tsunoo, Quentin Jodelet, Osamu Take
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report nonoverlapping metric subsets, preventing direct comparisons. We introduce COMPASS, a unified and reproducible benchmarking framework integrating 46 metrics across eight dimensions, and deploy it on 1,248 model-language configurations from FLEURS and CVSS, spanning cascaded and end-to-end architectures over ten language pairs. Architectures exhibit complementary strengths: best-vs-worst gaps exceed 30% on naturalness and speaker preservation but remain within a few points on translation quality, so single-metric rankings systematically misrepresent system quality. Correlation filtering reduces 46 metrics to 10 per direction, with three axes requiring different metrics across X→EN and EN→X (e.g., TER/UTMOS vs. ChrF++/NISQA-MOS); these subsets preserve rankings (Spearman's ρ > 0.80) while cutting evaluation time by ≈ 2.5×. Human validation across dubbing, podcasts, and medical domains shows standalone MOS predictors fail to predict listener preference, while top domainspecific metrics correlate with human judgment (ρ ≥ 0.90). We release COMPASS as a foundation for domain-aware S2ST evaluation.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2606.03241")
get_code_for_paper("2606.03241")
have("2606.03241")
Connect an agent — have() is free.