The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestIndexTeam2.5
IndexTeamText → Speech

Bilibili's IndexTTS-2.5 Refines Zero-Shot Voice Cloning

The latest text-to-speech model from bilibili's Index team pairs voice cloning with emotion control across Chinese, English, and Japanese.

Aug 10, 2026
NotableOther
IndexTTS-2.5

Bilibili's Index team has published IndexTTS-2.5, the newest iteration of its text-to-speech system on Hugging Face. The model targets zero-shot voice cloning, meaning it can reproduce a speaker's voice from a short reference sample without task-specific fine-tuning.

Beyond straightforward cloning, IndexTTS-2.5 offers emotion control and cross-lingual synthesis. According to the release, it supports Chinese, English, and Japanese, positioning it for use cases that span dubbing, localization, and expressive narration.

Why it matters

Zero-shot cloning with emotion steering has become a competitive frontier in open speech models, and multilingual coverage is what separates a demo from a production tool. A few practical implications:

  • Creators can generate voiceovers in three major languages from a single reference clip.
  • Emotion control gives finer editorial command over tone without re-recording.
  • Coming from bilibili, a large video platform, the model reflects real content-production demands.

The model ships under a custom license, so teams evaluating it for commercial work should review the terms on the model page before deploying. As a point release in the IndexTTS line, version 2.5 signals steady, incremental progress rather than a ground-up redesign.

Sources

  • IndexTeam/IndexTTS-2.5

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Languageszh, en, ja
LicenseOTHER
Downloads14.1K
Likes229

Modalities

Text → Speech

0 comments

No comments yet. Be the first to weigh in.

More in Text → Speech

Unknown/Text → Speech

Nari Labs Ships Qwen3-Based TTS and ASR Models

The startup pairs speech synthesis and recognition built on Qwen3, pitching accuracy, low latency, and lower cost.

Sep 14, 2026
StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
StepFun/Text → Speech

StepFun's StepAudio 3 Gen Unifies TTS and Music

A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.

Sep 10, 2026