rotating globe
26 Sep 2026


Google expands Gemini with voice, AI

New Gemini updates bring voice cloning, expressive speech and animated faces to AI interactions

Google is adding a more human touch to Gemini with new voice and video features that allow its artificial intelligence models to speak more naturally and appear as animated digital characters.

The company has rolled out new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, alongside Gemini 3.8 Live with Live Avatar. The updates are designed to make AI conversations more expressive, interactive and closer to the way people communicate.

The latest additions expand Gemini beyond text-based responses. Users and developers can now create AI-generated speech with different voices and styles, while Live Avatar adds a face that can speak and respond during conversations.

Gemini gets voice cloning

One of the key features in the new Gemini text-to-speech models is the ability to create customised voices. Developers can provide instructions describing how a voice should sound, including its style, tone and delivery.

Gemini 3.8 Flash TTS can also reproduce a voice using a short audio sample. According to reports on the rollout, the system can clone a voice from around 30 seconds of recorded audio.

The feature could be useful for applications that need personalised or character-based voices. Content creators could use it for narration, while developers could build virtual assistants, games and interactive characters with distinctive voices.

Google has also focused on making AI-generated speech more expressive. Rather than simply reading text aloud, the model can adjust aspects such as tone, pacing and delivery. This is intended to make conversations sound less mechanical.

The model supports a wide range of languages and dialects, making the technology potentially useful for global applications.

Multiple voices for conversations

The new Gemini TTS technology can also generate conversations involving multiple speakers.

This means developers can create dialogue in which different AI voices interact rather than relying on a single synthetic voice. The system is designed to handle changes in speakers and produce speech that sounds more like a performed conversation.

That could be useful for podcasts, audiobooks, video production and virtual characters. Instead of recording every line separately, creators could generate complete dialogue scenes through AI.

Google is positioning Gemini 3.8 Flash-Lite TTS as a more efficient option for applications that need to generate large amounts of speech. This could be particularly relevant for developers building voice-based AI agents and other high-volume services.

The new models are available through Google’s developer ecosystem, including Google AI Studio and the Gemini API.

Gemini Live gets an animated face

Google is also giving its conversational AI a visual identity with Gemini 3.8 Live with Live Avatar.

The feature allows users to interact with an AI avatar that can listen, respond and speak while displaying facial expressions and synchronised lip movements.

Instead of receiving an audio response from an AI assistant, users can see a digital character communicating with them. Google says the system is designed for near real-time interactions, allowing conversations to feel more dynamic.

The technology combines Gemini’s conversational capabilities with AI-generated video. The avatar’s mouth moves with its speech, while facial expressions add another layer to the interaction.

Google is initially targeting businesses with Live Avatar. Potential uses include customer service, interactive product demonstrations, training programmes and digital guides.

Companies can use predefined avatars and can also create characters with their own visual appearance and voice. The feature can be used across different experiences, including websites, mobile applications and interactive kiosks.

Focus shifts to natural AI interaction

The latest Gemini features reflect the wider direction of generative AI development. AI companies are increasingly working on how systems communicate, rather than focusing only on their ability to generate text.

Voice has become an important part of that shift. More natural speech can make AI assistants easier to use, particularly in situations where typing is inconvenient.

The addition of AI avatars takes the experience another step further. A digital face can provide visual feedback during a conversation and make interactions feel more immediate.

However, voice cloning also brings questions around digital identity and misuse. A system that can reproduce someone’s voice from a short recording can potentially be used to create convincing synthetic audio.

Google says it has safeguards around its AI-generated content. Its Live Avatar technology also uses SynthID, Google’s system for identifying AI-generated media.

The company is therefore expanding Gemini across several forms of communication at the same time. Gemini can already generate text and images, while its latest updates strengthen its capabilities in speech, dialogue and AI-generated video.

For developers, the new tools offer more control over the personality and presentation of AI systems. For users, they could make future interactions with digital assistants feel less like communicating with software and more like having a conversation.

Google’s latest Gemini rollout shows where conversational AI is heading: not just systems that can answer questions, but AI that can speak with expression, reproduce customised voices and communicate through a digital face.