Our Models for speech
Arabic-first models for text-to-speech and speech-to-text — built and fine-tuned on Gulf Arabic, each with its own codename and benchmark.
MModelSayvors · v2.1
MModelSayvors · v2
TTurtleSayvors Voice · v1
MModelSayvors · v2
MModelSayvors · v1
TTS models
Voice generation models that turn text into natural Gulf-Arabic speech. Each has its own codename and benchmark.
M
ModelSayvors · v2.1
CurrentOur fastest, most natural Gulf-Arabic voice — tuned for live phone conversations.
| MOS (naturalness) | 4.8 |
|---|---|
| Real-time factor (RTF) | 0.18 |
| Arabic EPW — Gulf | < 5% |
| Arabic EPC | < 4.5% |
| Gulf dialect coverage | High |
M
ModelSayvors · v2
StableBalanced quality and speed for high-volume outbound calling.
| MOS (naturalness) | 4.6 |
|---|---|
| Real-time factor (RTF) | 0.21 |
| Arabic EPW — Gulf | < 6% |
| Arabic EPC | < 5.5% |
| Gulf dialect coverage | High |
T
TurtleSayvors Voice · v1
LegacyThe original voice model — slower, with simpler prosody.
| MOS (naturalness) | 4.3 |
|---|---|
| Real-time factor (RTF) | 0.27 |
| Arabic EPW — Gulf | < 8% |
| Arabic EPC | < 7% |
| Gulf dialect coverage | High |
STT models
Transcription models that turn spoken Arabic into text in real time. Each has its own codename and benchmark.
M
ModelSayvors · v2
CurrentLatest transcription model with strong Gulf-dialect accuracy.
| WER — English | 2.4% |
|---|---|
| WER — Arabic Gulf | 4.8% |
| EPC — Arabic | 3.9% |
| Inference latency | 180 ms |
| Gulf dialect | High |
M
ModelSayvors · v1
LegacyFirst-generation transcription baseline.
| WER — English | 3.6% |
|---|---|
| WER — Arabic Gulf | 8.2% |
| EPC — Arabic | 6.7% |
| Inference latency | 340 ms |
| Gulf dialect | High |