NVIDIA aims to solve AI’s issues with many languages

AI News 3小时前

../../../zaker_core/zaker_tpl_static/wap/tpl_guoji1.html

尽管人工智能（AI）看似无处不在，但目前仅覆盖极少数语言，将大部分全球人口排除在外。NVIDIA正致力于弥合这一差距，尤其是在欧洲。该公司发布了一套强大的开源工具，使开发者能够为25种欧洲语言构建高质量的语音AI。这不仅包括主要语言，更重要的是，为克罗地亚语、爱沙尼亚语和马耳他语等被科技巨头忽视的语言提供了支持。其核心是名为Granary的庞大语音库，包含约一百万小时的音频数据，并辅以Canary-1b-v2和Parakeet-tdt-0.6b-v3两个新AI模型，分别侧重于高精度转录翻译和实时应用。通过自动化数据处理流程，NVIDIA大幅降低了AI语言数据获取的成本和时间，加速了数字包容性，让更多地区的用户能够享受AI语音技术带来的便利。

🌍 **AI语言覆盖不均，NVIDIA致力于缩小数字鸿沟：** 文章指出，当前AI技术主要服务于少数主流语言，忽视了全球大部分语言使用者。NVIDIA通过发布开源工具，旨在为25种欧洲语言（包括克罗地亚语、爱沙尼亚语、马耳他语等）开发高质量语音AI，解决AI语言覆盖不均的问题，促进数字包容性。

📚 **Granary语音库赋能AI学习：** NVIDIA推出了名为Granary的庞大语音库，收录了约一百万小时的精选人类语音数据。这些数据经过精心策划，用于训练AI识别语音细微差别和进行语言翻译，为开发者提供了高质量的训练资源。

🚀 **两大AI模型提升语音AI能力：** NVIDIA发布了两个新的AI模型：Canary-1b-v2，一个专为复杂转录和翻译任务设计的高精度模型；以及Parakeet-tdt-0.6b-v3，一个专注于实时应用、速度极快的模型。这两个模型都支持标点、大小写处理和词级时间戳，能够满足专业级应用的需求。

💡 **自动化流程降低数据获取成本：** NVIDIA利用其NeMo工具包，开发了一个自动化的数据处理流程，能够将原始、未标记的音频数据转化为AI可学习的高质量结构化数据。这一创新显著缩短了数据标注的时间和成本，提高了效率，并发现其Granary数据仅需一半的量就能达到目标准确率。

🌐 **推动全球开发者创新与应用：** NVIDIA将这些强大的工具和方法开放给全球开发者社区，旨在激发新一轮的创新浪潮。通过赋能开发者构建能够准确理解本地语言的语音AI工具，NVIDIA期望创造一个AI能够真正理解并服务于所有人的语言的世界。

While AI might feel ubiquitous, it primarily operates in a tiny fraction of the world’s 7,000 languages, leaving a huge portion of the global population behind. NVIDIA aims to fix this glaring blind spot, particularly within Europe.

The company has just released a powerful new set of open-source tools aimed at giving developers the power to build high-quality speech AI for 25 different European languages. This includes major languages, but more importantly, it offers a lifeline to those often overlooked by big tech, such as Croatian, Estonian, and Maltese.

The goal is to let developers create the kind of voice-powered tools many of us take for granted, from multilingual chatbots that actually understand you to customer service bots and translation services that work in the blink of an eye.

The centrepiece of this initiative is Granary, an enormous library of human speech. It contains around a million hours of audio, all curated to help teach AI the nuances of speech recognition and translation.

To make use of this speech data, NVIDIA is also providing two new AI models designed for language tasks:

Canary-1b-v2

Parakeet-tdt-0.6b-v3

If you’re keen to dive into the science behind it, the paper on Granary will be presented at the Interspeech conference in the Netherlands this month. For the developers eager to get their hands dirty, the dataset and both models are already available on Hugging Face.

The real magic, however, lies in how this data was created. We all know that training AI requires vast amounts of data, but getting it is usually a slow, expensive, and frankly tedious process of human annotation.

To get around this, NVIDIA’s speech AI team – working with researchers from Carnegie Mellon University and Fondazione Bruno Kessler – built an automated pipeline. Using their own NeMo toolkit, they were able to take raw, unlabelled audio and whip it into high-quality, structured data that an AI can learn from.

This isn’t just a technical achievement; it’s a huge leap for digital inclusivity. It means a developer in Riga or Zagreb can finally build voice-powered AI tools that properly understand their local languages. And they can do it more efficiently. The research team found that their Granary data is so effective that it takes about half the amount of it to reach a target accuracy level compared to other popular datasets.

The two new models demonstrate this power. Canary is frankly a beast, offering translation and transcription quality that rivals models three times its size, but with up to ten times the speed. Parakeet, meanwhile, can chew through a 24-minute meeting recording in one go, automatically figuring out what language is being spoken. Both models are smart enough to handle punctuation, capitalisation, and provide word-level timestamps, which is required for building professional-grade applications.

By putting these powerful tools and the methods behind them into the hands of the global developer community, NVIDIA isn’t just releasing a product. It’s kickstarting a new wave of innovation, hoping to create a world where AI speaks your language, no matter where you’re from.

(Photo by Aedrian Salazar)

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is co-located with other leading events including Intelligent Automation Conference, BlockX, Digital Transformation Week, and Cyber Security & Cloud Expo.

Explore other upcoming enterprise technology events and webinars powered by TechForge here.

The post NVIDIA aims to solve AI’s issues with many languages appeared first on AI News.

Fish AI Reader

FishAI

联系邮箱 441953276@qq.com

相关标签