OpenMOSS Releases Transcribe-Diarize ASR Model
The open-weights team behind MOSS turns to long-form speech recognition with built-in speaker diarization and timestamps.

OpenMOSS has published MOSS Transcribe-Diarize, an open-weights speech recognition model aimed squarely at the messy realities of real-world audio: long recordings, multiple speakers, and the need to know who said what and when.
Unlike a plain transcription model that spits out an undifferentiated wall of text, this release combines three tasks that usually require stitching together separate tools. It transcribes long-form audio, applies speaker diarization to separate distinct voices, and produces timestamped output so passages can be aligned back to the original recording.
Why it matters
Diarization is one of the most persistent pain points in applied speech work. Teams building meeting summarizers, podcast tools, call-center analytics, or interview transcription pipelines typically bolt a diarization system onto a separate ASR engine, then fight to reconcile the two. Packaging these capabilities in a single open model lowers that friction considerably.
- Long-form transcription for extended recordings
- Speaker diarization to attribute speech to individual voices
- Timestamped output for alignment and navigation
The model ships under a custom license, so teams evaluating it for production should read the terms on the model page closely. As a 1.0 initial release, it also arrives without published benchmarks in the record, meaning the practical proof will come from real deployments. Still, for a space long dominated by a handful of proprietary APIs, another capable open option is worth attention.
Sources
- Visit
OpenMOSS-Team/MOSS-Transcribe-Diarize
Hugging Face
More in Speech → Text

KRAFTON releases A.X-K2 Raon speech MoE model
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Microsoft's VibeVoice ASR Goes BitNet for CPU Speech
A BitNet-quantized speech recognition model trades GPU dependence for efficient CPU inference in English and Chinese.
CrisperWhisper 2.0 Large targets verbatim transcription
A Whisper-based ASR model that keeps every filler word and stamps timestamps to the individual word, now covering English and German.
0 comments
No comments yet. Be the first to weigh in.