From the community's voice and words, we build models that understand and speak Tamazight — speech recognition, speech synthesis, and translation — across every dialect. Because a language thousands of years old deserves its place in the age of AI.
Tens of millions speak Amazigh across North Africa, yet it remains a “low-resource” language for AI: very little speech and text data, scattered across dialects with no unified base, and the ⵜⵉⴼⵉⵏⴰⵖ Tifinagh script underrepresented in global models.
There is no model without data. So we start where everything should begin: with the community itself — voice, word, and story.
dialects across 6 countries — from Tashelhit to Tuareg to Zenaga
supported scripts: Tifinagh · Latin · Arabic
all data in the public domain, owned by everyone
contribution types: voice · lexicon · translation · heritage
Our goal is an open suite of models that understand and produce Tamazight. We start from the data, and build one capability at a time.
Turn spoken Amazigh into text — the foundation of every voice application.
In developmentNatural Amazigh speech generated from text — to voice what is written in Tifinagh.
In developmentTranslate between Arabic/French and Amazigh in Tifinagh — in both directions.
In developmentOur first product: a conversational AI like ChatGPT or Gemini — but it understands and replies in Tamazight. Try it now.
● LiveA “data-first” approach: a community collects, native speakers validate, open data is published, and models are trained.
Through the awal app: record your voice in your dialect, add words and meanings, translate, and document heritage.
Every contribution is reviewed by validators from the same dialect (approve / reject / request a fix) — quality guarded by native speakers.
What's approved is released anonymized under CC0 as an open dataset in the Mozilla Common Voice format.
On top of this clean data we train open Amazigh models — recognition, synthesis, and translation.
The community app (PWA) that collects and validates voices, words, translations, and heritage — with a four-language interface.
● Live nowHuman-validated Amazigh data in the public domain (CC0), open to all — the fuel for every model.
● GrowingThe first product built on this data: a conversational AI assistant — like ChatGPT or Gemini — that chats in Tamazight.
● LiveEvery recording and every word you contribute today brings the first Amazigh AI model a step closer to reality. One task is enough to start.