OpenMOSS Releases Transcribe-Diarize ASR Model
The open-weights team behind MOSS turns to long-form speech recognition with built-in speaker diarization and timestamps.

OpenMOSS has published MOSS Transcribe-Diarize, an open-weights speech recognition model aimed squarely at the messy realities of real-world audio: long recordings, multiple speakers, and the need to know who said what and when.
Unlike a plain transcription model that spits out an undifferentiated wall of text, this release combines three tasks that usually require stitching together separate tools. It transcribes long-form audio, applies speaker diarization to separate distinct voices, and produces timestamped output so passages can be aligned back to the original recording.
Why it matters
Diarization is one of the most persistent pain points in applied speech work. Teams building meeting summarizers, podcast tools, call-center analytics, or interview transcription pipelines typically bolt a diarization system onto a separate ASR engine, then fight to reconcile the two. Packaging these capabilities in a single open model lowers that friction considerably.
- Long-form transcription for extended recordings
- Speaker diarization to attribute speech to individual voices
- Timestamped output for alignment and navigation
The model ships under a custom license, so teams evaluating it for production should read the terms on the model page closely. As a 1.0 initial release, it also arrives without published benchmarks in the record, meaning the practical proof will come from real deployments. Still, for a space long dominated by a handful of proprietary APIs, another capable open option is worth attention.
Sources
- Visit
OpenMOSS-Team/MOSS-Transcribe-Diarize
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.