Jais Brings Arabic-Focused Language Models Into the Open

Jais and Jais-chat are open language models trained on Arabic, English and code, developed by researchers from the United Arab Emirates with Cerebras. The team says they outperform existing freely available Arabic models, while larger commercial systems remain ahead on average in benchmarks.

WTF Index NEUTRAL
◄ Terminator 0 Idiocracy 0 ►

The article describes an open-model release focused on Arabic language access and benchmark performance, without a clear lean toward either risk.

Jais Brings Arabic-Focused Language Models Into the Open

Jais and Jais-chat are open language models designed to work with Arabic as well as English and code. Developed by researchers from the United Arab Emirates in collaboration with Cerebras, the models aim to address a gap in open language technology for Arabic.

Their release offers developers and researchers models they can access through Hugging Face, with Jais-chat also available to try through Arabic-GPT. The team reports strong results on Arabic benchmarks, while noting that leading commercial systems still perform better on average.

Built with Arabic at the center

Jais is a 13 billion parameter model pre-trained on 395 billion tokens. Of those, 116 billion are Arabic tokens. Jais-chat is an instruction-tuned version, trained with an additional 10 million instruction/response pairs to make it suitable for conversational use.

The models draw on Arabic websites, books, news and Wikipedia. The team says it filtered all data before training. The project also uses 232 billion tokens of English data from The Pile by EleutherAI, along with 46 billion code tokens.

That mix reflects a practical challenge for Arabic language models: the team says the amount of Arabic data available is limited. English and code data supplement the Arabic material, while the models remain focused on improving Arabic capabilities.

How the team describes the results

In benchmarks, Jais and Jais-chat outperform existing freely available Arabic models by 11 to 15 points in accuracy, according to the team. The models are also described as competitive with Meta's LLaMa2 on English.

That does not mean they lead every comparison. The source says commercial models such as OpenAI's ChatGPT and Anthropic's Claude remain ahead on average in the benchmarks, though those systems are significantly larger. For some tasks, including writing, the team says Jais and Jais-chat are on par with ChatGPT.

These distinctions matter when interpreting a benchmark claim. Results can vary by task, and the reported strength in Arabic does not establish that Jais is a general replacement for larger commercial chatbots. The clearest stated advantage is performance over other freely available Arabic models.

Open access and safeguards

Jais and Jais-chat are available on Hugging Face, and users can try them on Arabic-GPT. That availability gives people a way to explore the models directly and gives developers access to Arabic-focused systems they can build on.

For Jais-chat, the team also provides security mechanisms, including filters and classifiers for unwanted requests and outputs. These tools are part of the model offering described in the source; no further details about their operation or effectiveness are provided there.

A different route to training

The models were trained on Cerebras CS-2 systems rather than Nvidia GPUs. Cerebras makes a wafer-sized AI chip that is installed in those systems, making the hardware a notable part of how the project was carried out.

Jais is presented as the first Arabic-centric open model at this scale. Its significance rests on combining substantial Arabic training data, an instruction-tuned chat version and public availability. The team’s benchmark results suggest a stronger open option for Arabic, while comparisons with commercial models remain qualified by their larger size and their average lead across benchmarks.