<!-- canonical: https://www.trendhunter.com/trends/valle -->
<!-- robots: noindex -->

# AI-Powered Language Models
Microsft Develops a New Text-to-Speech AI Dubbed the VALL-E

By Amy Duong | Published 2023-01-10 | Tech
Source: Trend Hunter, https://www.trendhunter.com/trends/valle
References: [engadget](https://www.engadget.com/microsofts-vall-e-ai-can-simulate-any-persons-voice-from-a-short-audio-sample-112520213.html) & [valle-demo.github.io](https://valle-demo.github.io/)

![AI-Powered Language Models](https://cdn.trendhunterstatic.com/thumbs/496/valle.jpeg)

[Microsoft](https://www.trendhunter.com/trends/vasa1) develops its own [text-to-speech artificial intelligence](https://www.trendhunter.com/trends/bookfab-audiobook) model dubbed the VALL-E. It is able to simulate a user's voice from only a three-second audio sample. As reported by Ars Technica, the speech can only match the timbre and the emotional tone of the specific speaker.

It goes the extra mile and also imitates the acoustics of a room as well. The VALL-E is what the company notes as a 'neural codec language model.' It stems from [Meta's](https://www.trendhunter.com/trends/voicebox-ai) AI-powered compression neural net Encodec. This generates audio with text input and other short audio samples from a speaker. The VALL-E technology is trained with over 60,000 hours of English language speakers and from 7,000+ speakers on the Meta LibriLight audio library.

Image Credit: Getty Images

## Trend Insights (Trend Hunter)

- Score: 3.5/10
- Popularity: 22% | Activity: 68% | Freshness: 14%
- Audience gender: 50% men, 50% women
- Primary generations: Gen Z, Millennial, Gen X
- Top markets: North America, South America, Europe, Asia, Africa

## Categories

[Trend Hunter](https://www.trendhunter.com/trends) > [Tech](https://www.trendhunter.com/tech) > [AI](https://www.trendhunter.com/ai)

## Key Themes

### Key Themes Behind This Trend

- **Audio Deepfakes:** The development of AI-powered language models like VALL-E creates opportunities for malicious actors to create convincing audio deepfakes for nefarious purposes.
- **Personalized Text-to-speech:** VALL-E's ability to simulate a user's voice from a brief audio sample represents an opportunity for businesses to offer personalized text-to-speech services.
- **Realistic Voice Assistants:** AI-powered language models like VALL-E contribute to the development of more realistic and human-like voice assistants.

### Where This Applies

- **Media and Entertainment:** The entertainment industry can leverage VALL-E's capabilities to create realistic voiceovers and more immersive audio experiences.
- **Digital Assistants:** Developers of digital assistants and voice-enabled devices can integrate AI-powered language models like VALL-E to create more human-like interactions with users.
- **Cybersecurity:** As the use of AI-powered language models like VALL-E becomes more widespread, the cybersecurity industry has an opportunity to develop new solutions to detect and prevent audio deepfakes.

## Related on Trend Hunter

- [Video Creation AI Tools](https://www.trendhunter.com/trends/vasa1.md)
- [Voice Cloning Technological Tools](https://www.trendhunter.com/trends/voice-engine.md)
- [AI Voice Replicators](https://www.trendhunter.com/trends/voicebox-ai.md)
- [AI Voice Generators](https://www.trendhunter.com/trends/elevenlabs1.md)
- [AI-Generated Language Models](https://www.trendhunter.com/trends/massively-multilingual-speech.md)
- [Emotion-Adjustable AI Voices](https://www.trendhunter.com/trends/hume-octave.md)
- [AI-Powered Voice Producers](https://www.trendhunter.com/trends/audiobox.md)
- [AI Voice Platforms](https://www.trendhunter.com/trends/ai-voicer.md)
- [Commercial AI Language Models](https://www.trendhunter.com/trends/llama-2.md)
