S2Cap: A Benchmark and a Baseline for Singing Style Captioning

cs.AI updates on arXiv.org 13小时前

S2Cap: A Benchmark and a Baseline for Singing Style Captioning

本文提出构建唱歌风格字幕数据集S2Cap，包含丰富唱歌声音属性，填补现有数据集在唱歌风格字幕任务中的不足，并开发高效算法实现风格字幕。

arXiv:2409.09866v3 Announce Type: replace-cross Abstract: Singing voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voices capture only a narrow range of attributes and lack acoustic features, leading to limited utility towards downstream tasks, such as style captioning. To fill this gap, we formally define the singing style captioning task and present S2Cap, a dataset of singing voices with detailed descriptions covering diverse vocal, acoustic, and demographic characteristics. Using this dataset, we develop an efficient and straightforward baseline algorithm for singing style captioning. The dataset is available at https://zenodo.org/records/15673764.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

唱歌风格字幕数据集 S2Cap

相关文章

MS MARCO Web Search: A Large-Scale Information-Rich Web Dataset Featuring Millions of Real Clicked Query-Document Labels

This Week In Machine Learning & AI - 5/27/16: The White House on AI & Aggressive Self-Driving Cars

CinePile: A Novel Dataset and Benchmark Specifically Designed for Authentic Long-Form Video Understanding

‘RAG Me Up’: A Generic AI Framework (Server + UIs) that Enables You to Do RAG on Your Own Dataset Easily

HuggingFace Releases ? FineWeb: A New Large-Scale (15-Trillion Tokens, 44TB Disk Space) Dataset for LLM Pretraining

Unlocking the Language of Proteins: How Large Language Models Are Revolutionizing Protein Sequence Understanding

MAGPIE: A Self-Synthesis Method for Generating Large-Scale Alignment Data by Prompting Aligned LLMs with Nothing

Midjourney: ↩️ @kortizart To the best of our knowledge; you are not in our dataset. Here's a "portrait by Karla Ortiz" vs a "portrait by artist". FW...

Hugging Face: We're excited to welcome @argilla_io to the Hugging Face team! ? Time to democratise good Machine Learning, one dataset at a time!...

Hugging Face: Hugging Face is hosting a demo site for @iclr_conf authors to find and claim their papers and discuss those papers on dedicated pages Th...