ⵜⴰⵡⵎⴰⵜ · An Amazigh AI initiative
ⵜⴰⵡⵎⴰⵜ ⴰⵢ

Building the first Amazigh AI model.

From the community's voice and words, we build models that understand and speak Tamazight — speech recognition, speech synthesis, and translation — across every dialect. Because a language thousands of years old deserves its place in the age of AI.

Open data · CC0 Community-driven · collected & validated by native speakers Sovereign · owned by the community Every dialect
Why now

A language of millions, almost absent from AI.

Tens of millions speak Amazigh across North Africa, yet it remains a “low-resource” language for AI: very little speech and text data, scattered across dialects with no unified base, and the ⵜⵉⴼⵉⵏⴰⵖ Tifinagh script underrepresented in global models.

There is no model without data. So we start where everything should begin: with the community itself — voice, word, and story.

10

dialects across 6 countries — from Tashelhit to Tuareg to Zenaga

3

supported scripts: Tifinagh · Latin · Arabic

CC0

all data in the public domain, owned by everyone

4

contribution types: voice · lexicon · translation · heritage

What we build

Open Amazigh models — for every language task.

Our goal is an open suite of models that understand and produce Tamazight. We start from the data, and build one capability at a time.

Speech recognition

Turn spoken Amazigh into text — the foundation of every voice application.

In development

Speech synthesis

Natural Amazigh speech generated from text — to voice what is written in Tifinagh.

In development

Machine translation

Translate between Arabic/French and Amazigh in Tifinagh — in both directions.

In development

yaz — the Amazigh chat

Our first product: a conversational AI like ChatGPT or Gemini — but it understands and replies in Tamazight. Try it now.

● Live
How we build it

From your voice… to a model.

A “data-first” approach: a community collects, native speakers validate, open data is published, and models are trained.

Contribute

Through the awal app: record your voice in your dialect, add words and meanings, translate, and document heritage.

Human validation

Every contribution is reviewed by validators from the same dialect (approve / reject / request a fix) — quality guarded by native speakers.

Published openly

What's approved is released anonymized under CC0 as an open dataset in the Mozilla Common Voice format.

We train the models

On top of this clean data we train open Amazigh models — recognition, synthesis, and translation.

The ecosystem

One piece of a bigger picture.

Digital sovereignty for Amazigh

Our language deserves its model. And your voice is what builds it.

Every recording and every word you contribute today brings the first Amazigh AI model a step closer to reality. One task is enough to start.