At SJM Labs, we've always believed that AI built for the Gulf should actually sound like the Gulf. That's the idea behind Namat (نمط), a Saudi-first text-to-speech platform designed to capture the way people here actually speak, not a generic, flattened version of Arabic.
If you've spent any time with existing Arabic TTS tools, you've probably run into the same wall we did. Type in a script meant for a Saudi audience, and what comes out sounds like it was written for a news broadcast in Cairo or Amman, or worse, like a machine reading Arabic phonetically without understanding how it's actually spoken. Modern Standard Arabic has its place, but it isn't how people in Riyadh, Jeddah, or the wider Gulf talk to each other, and it definitely isn't how a Saudi brand should sound in its ads, its audiobooks, or its content. Most TTS systems treat Arabic as a single, uniform language and default to MSA regardless of who's supposed to be listening. That works for formal announcements. It falls apart the moment you need something that sounds local.
Namat starts from a different premise: dialect isn't a detail to smooth over on the way to "good enough" Arabic. It's the whole point. A Najdi audiobook should sound Najdi. A Hejazi voiceover should carry Hejazi rhythm and pronunciation. A Gulf ad campaign should sound like it was made for the Gulf, because it was. That's the gap we set out to close.
What's actually different under the hood
Powering Namat is SJM-Voice-1, our voice model built specifically around Saudi and Gulf speech patterns. We didn't want a system that just translates text into sound. We wanted one that understands dialect, accent, the way certain words get pronounced differently depending on region, tone, and the natural rhythm of how sentences actually flow when they're spoken out loud rather than read off a page. That's a much harder problem than generic speech synthesis, and it's the one we chose to focus on.
At launch, Namat will support six language and dialect options: English, MSA, Saudi White dialect, Gulf White dialect, Najdi, and Hejazi. We're not treating these as an afterthought or a future roadmap item. All six are part of the plan from day one, because a platform that claims to be Saudi-first has to actually cover the range of how Saudis and Gulf Arabic speakers sound, not just one flavor of it.
Two other capabilities are core to the platform rather than bolted on later. The first is zero-shot voice cloning, which lets a user capture a specific voice from a short reference sample and generate new speech in that same voice, without needing hours of training data. The second is emotion tagging, so generated speech doesn't come out flat. Content that's meant to sound excited, calm, serious, or warm should actually carry that tone, not just get the words right.
Who this is actually for
We built Namat with three groups of people in mind, because we kept hearing the same frustration from all three.
Content creators need voiceovers that sound like their actual audience, not a translated stand-in that technically says the right words but feels foreign the moment it's spoken. Whether it's a YouTube video, a podcast intro, or short-form content aimed at a Saudi or Gulf audience, the voice needs to match the culture, not just the language.
Ad producers face a narrower version of the same problem, but with higher stakes. Localizing a campaign across different Gulf markets traditionally means booking a new voice actor for every region, coordinating studio time, and managing revisions across multiple people and schedules. Namat is designed to collapse that process. One script, multiple dialects, generated directly, with the accent handled correctly instead of approximated.
Audiobook authors and publishers are the third group, and arguably the one most underserved by existing tools. Long-form narration is unforgiving. A flat, mispronounced, or tonally wrong voice becomes exhausting over an hour, let alone an entire book. Namat is built to handle long-form text properly, not just short marketing clips, so that narration holds up over the length of a real audiobook rather than just a demo reel.
Why we're building it self-hosted
Namat is being built as a web app at namat.cc, with an iOS app planned to ship alongside it. Both are designed to run on infrastructure we control end to end, rather than wrapping a third-party TTS API and reselling access to it. That's a deliberate choice, not a default one.
Self-hosting means we control quality directly. If a dialect needs more work, if a voice sounds off, if pronunciation needs adjusting, we can act on that ourselves instead of waiting on someone else's roadmap or being limited by what a third-party API happens to expose. It also means we control cost, which matters for a product meant to be genuinely usable by individual creators and small teams, not just enterprises with large budgets. And it means Namat's direction is shaped by what our users in this region actually need, not by the priorities of a general-purpose model built for a global, dialect-agnostic audience.
What comes next
Namat is still in development, and we're building it deliberately rather than rushing to ship something half finished. Dialect accuracy, natural tone, and voice cloning that actually sounds like the reference voice are hard to get right, and we'd rather take the time to do that properly than launch early and disappoint the people we built this for.
Looking further ahead, we're already thinking about where SJM-Voice-1 goes next, including real-time generation for AI voice agents, which opens up an entirely different set of use cases beyond static content generation. That's a future chapter, not part of this one, but it's part of why we're building the underlying model with room to grow rather than optimizing narrowly for a single launch feature set.
We started SJM Labs because we kept running into the same gap: AI tools that technically support Arabic but don't actually understand how it's spoken in this region. Namat is our answer to that gap for voice. More details, and Namat itself, are on the way.