The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • Hardware estimates
  • RSS feed
  • llms.txt
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestKRAFTON1.0
KRAFTONText → Speech

KRAFTON releases 1B zero-shot voice-cloning TTS

Raon-OpenTTS-1B is a flow-matching diffusion transformer that clones English voices from short reference clips.

May 21, 2026
UpdateOther
Raon-OpenTTS-1B

KRAFTON has published Raon-OpenTTS-1B, a one-billion-parameter text-to-speech model that generates English speech and can mimic a target voice from a short reference sample. The release is available on Hugging Face and marks the first version of the Raon open TTS line.

The model is a flow-matching diffusion transformer, an architecture that has become popular for high-quality audio and image synthesis because it can produce natural output in relatively few sampling steps. Positioned as a zero-shot voice-cloning system, it aims to reproduce a speaker's timbre without per-voice fine-tuning, relying instead on a reference clip at inference time.

What it offers

  • Roughly 1B parameters, a size that stays practical to run on a single modern GPU
  • Zero-shot voice cloning from a short reference sample
  • English-language output
  • A flow-matching diffusion-transformer backbone

Why it matters

Open zero-shot TTS remains a competitive space, and a compact 1B model that clones voices lowers the barrier for developers building narration, accessibility, and assistant features without training custom voices. The tradeoff is scope: this initial release targets English only, and the repository carries a non-standard "other" license, so teams should check the terms before commercial use. As a version 1.0, it establishes a baseline the family can iterate on.

Sources

  • KRAFTON/Raon-OpenTTS-1B

    Hugging Face

    Visit
OlderMisoLabs Debuts MisoTTS, an Open Voice ModelMisoLabs · Text → Speech · 5 months agoNewerOpenBMB's MiniCPM5-1B targets on-device AIOpenBMB · Text / LLM · 5 months ago

Get the model

Hugging Face

Specs

Parameters1B
LanguagesEnglish
LicenseOTHER
Downloads61
Likes81

Can you run it?

Runs on just about anything.

  • BF16 (as published)

    Any modern laptop · 8 GB RAM

    3 GB
  • FP8

    Any modern laptop · 8 GB RAM

    2 GB
Your machine
  • Precision not published; assuming 16-bit (BF16) weights.
  • Parameter count is an editorial estimate — no weights metadata on Hugging Face.

Estimated from the parameter count. Assumes an 8K context and ~1 GB runtime overhead; actual needs vary. How we estimate


Modalities

Text → Speech

The Weekly Weights

Every open release that mattered, one email a week.

0 comments

No comments yet. Be the first to weigh in.

More from KRAFTON

All KRAFTON releases →
A.X-K2 Raon Speech 21B-A3B
KRAFTON/Any-to-AnyRuns on a laptop

KRAFTON releases A.X-K2 Raon speech MoE model

The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Jul 27, 2026
Raon-Speech-9B
KRAFTON/Any-to-AnyRuns on a laptop

KRAFTON Releases 9B Bilingual Speech Model

The gaming giant behind 'PUBG' has released Raon-Speech-9B, a multimodal model for English and Korean speech recognition and synthesis.

Mar 30, 2026

More in Text → Speech

All Text → Speech →
Unknown/Text → Speech

Nari Labs Ships Qwen3-Based TTS and ASR Models

The startup pairs speech synthesis and recognition built on Qwen3, pitching accuracy, low latency, and lower cost.

Sep 14, 2026
StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
StepFun/Text → Speech

StepFun's StepAudio 3 Gen Unifies TTS and Music

A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.

Sep 10, 2026