TechCrunch News 04月22日
Two undergrads built an AI speech model to rival NotebookLM
index_new5.html
../../../zaker_core/zaker_tpl_static/wap/tpl_guoji1.html

 

Nari Labs推出开源AI语音生成模型Dia,两位本科生在三个月内开发完成。该模型可在普通PC上运行,支持生成类似播客风格的对话,并提供声音克隆功能。Dia在技术上具有竞争力,但在安全性和数据来源方面存在挑战。尽管如此,Nari Labs计划在其基础上构建具有社交功能的合成语音平台,并拓展对英语以外语言的支持。

🎙️ **模型特性**: Dia是一款开源的AI语音生成模型,由Nari Labs开发,可在普通PC上运行。它能够根据脚本生成对话,用户可以自定义说话者的语调,并添加非语言提示,如咳嗽和笑声。

💻 **技术细节**: Dia拥有16亿个参数,可在Hugging Face和GitHub上获取。它支持声音克隆功能,并且在TechCrunch的测试中表现良好,生成的声音质量与其他工具相当。

⚠️ **潜在风险**: Dia缺乏安全措施,容易被用于制作虚假信息或诈骗录音。Nari Labs虽然声明不负责任滥用行为,但未披露训练数据来源,可能涉及版权问题。

🚀 **未来计划**: Nari Labs计划基于Dia构建一个具有社交功能的合成语音平台,并发布技术报告,同时扩展对英语以外语言的支持。

A pair of undergrads, neither with extensive AI expertise, say that they’ve created an openly available AI model that can generate podcast-style clips similar to Google’s NotebookLM.

The market for synthetic speech tools is vast and growing. ElevenLabs is one of the largest players, but there’s no shortage of challengers (see PlayAI, Sesame, and so on). Investors believe that these tools have immense potential. According to PitchBook, startups developing voice AI tech raised over $398 million in VC funding last year.

Toby Kim, one of the Korea-based co-founders of Nari Labs, the group behind the newly released model, said that he and his fellow co-founder started learning about speech AI three months ago. Inspired by NotebookLM, they wanted to create a model that offered more control over generated voices and “freedom in the script.”

Kim says they used Google’s TPU Research Cloud program, which provides researchers with free access to the company’s TPU AI chips, to train Nari’s model, Dia. Weighing in at 1.6 billion parameters, Dia can generate dialogue from a script, letting users customize speakers’ tones and insert disfluencies, coughs, laughs, and other nonverbal cues.

Parameters are the internal variables models use to make predictions. Generally, models with more parameters perform better.

Available from the AI dev platform Hugging Face and GitHub, Dia can run on most modern PCs with at least 10GB of VRAM. It generates a random voice unless prompted with a description of an intended style, but it can also clone a person’s voice.

In TechCrunch’s brief testing of Dia through Nari’s web demo, Dia worked quite well, uncomplaining generating two-way chats about any subject. The quality of the voices seems competitive with other tools out there, and the voice cloning function is among the easiest this reporter has tried.

Here’s a sample:

Like many voice generators, Dia offers little in the way of safeguards, however. It’d be trivially easy to craft disinformation or a scammy recording. On Dia’s project pages, Nari discourages abuse of the model to impersonate, deceive, or otherwise engage in illicit campaigns, but the group says it “isn’t responsible” for misuse.

Nari also hasn’t disclosed which data it scraped to train Dia. It’s possible Dia was developed using copyrighted content — a commenter on Hacker News notes that one sample sounds like the hosts of NPR’s “Planet Money” podcast. Training models on copyrighted content is a widespread but legally dubious practice. Some AI companies claim that fair use shields them from liability, while rights holders assert that fair use doesn’t apply to training.

In any event, Kim says Nari’s plan is to create a synthetic voice platform with a “social aspect” on top of Dia and larger, future models. Nari also intends to release a technical report for Dia, and to expand the model’s support to languages beyond English.

Fish AI Reader

Fish AI Reader

AI辅助创作,多种专业模板,深度分析,高质量内容生成。从观点提取到深度思考,FishAI为您提供全方位的创作支持。新版本引入自定义参数,让您的创作更加个性化和精准。

FishAI

FishAI

鱼阅,AI 时代的下一个智能信息助手,助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

AI 语音生成 开源 Dia Nari Labs
相关文章