A major new family of open language models aims to make local language AI simple and fast. The models were shown at an industry summit in India this week. India hosted the event where the research arm presented the work. Cohere Labs built the models with an eye toward real world use and local language needs. The company says the models run on normal laptops and do not need a cloud connection.

Model Family Details
The base Tiny Aya model has about three point three five billion parameters. The family also includes a version tuned for global use that follows directions better. The company made regional variants that focus on different language groups. One of them is focused on South Asian languages of Bengali Hindi Punjabi Urdu Gujarati Tamil Telugu and Marathi. Another variant targets African languages and one covers Europe West Asia and much of Asia Pacific. The aim is to give each region a model that understands local words and cultural context while still staying broadly multilingual.
On Device Use
The models can run on everyday hardware. The company trained them on a single cluster of high end GPUs from a major chip maker. Nvidia made the H100 chips used in training. The research team then optimized the models to use far less compute when they run on a laptop. That makes it possible to do translation and chat with low latency and no network. Offline use opens doors in places with poor internet service and in apps that must protect user privacy.
Developer Access Points
The models and the training data are available on common model hubs and local deployment tools. The release page notes that developers can download the models from a popular model host. HuggingFace hosts the models and sample data. The company also supports downloads on Kaggle and on local runtime tools such as Ollama. The team will share a full technical report on training and evaluation methods so researchers can study the results and adapt them.

Many countries use many languages and dialects. Large models that focus on English do not always serve diverse users well. An open family of models that is tuned for local speech and that runs offline lets app makers build products for people who speak native languages. This can power translation for field workers voice assistants in low bandwidth areas and tools that work in private without sending text to cloud servers.