Glossary · Prometheus

What is a mixture of experts (MoE)?

What is a mixture of experts (MoE)? A clear explanation for Azerbaijani business — and how Prometheus applies it.

Understanding Mixture of Experts (MoE) Architecture

A Mixture of Experts (MoE) is an advanced AI architecture where a massive model is partitioned into numerous specialized sub-networks, known as 'experts.' Instead of activating the entire model for every query, a lightweight routing mechanism—the gating network—dynamically selects only the most relevant experts for each specific token or task. This sparse activation allows the model to scale to an immense total parameter count while only utilizing a fraction of its capacity during inference, effectively delivering high-tier performance without the prohibitive compute costs associated with traditional dense models. Prometheus, developed by Allmaz, is the first large language model built natively for the Azerbaijani language utilizing this MoE framework. By combining a specialized routing system with a native tokenizer, Prometheus provides organizations across Azerbaijan with a high-capacity AI solution that remains computationally efficient. This architecture enables the model to handle complex linguistic nuances and domain-specific knowledge while being deployed fully on-premise, ensuring that sensitive organizational data never leaves the local network.

Capabilities

Business Advantages of MoE Architecture

Efficient Scaling: Access massive model capacities, such as the 587B parameter tier, without activating every parameter per request, keeping operational inference costs manageable.

Domain Specialization: Expert sub-networks naturally develop proficiency in specific linguistic patterns and reasoning tasks, significantly enhancing the quality and accuracy of outputs.

Hardware Flexibility: With available sizes in 39B, 99B, and 587B parameters, organizations can precisely align the model tier with their existing on-premise compute resources.

Linguistic Optimization: When paired with a native tokenizer, MoE routing concentrates expert capacity on the unique agglutinative morphology of the Azerbaijani language.

Practical On-Premise Deployment: Lower active parameter counts per request make it feasible to run high-capacity models within private data centers without requiring supercomputing clusters.

Future-Proof Scalability: As an active area of AI research, the MoE design pattern allows for the integration of new capabilities and continuous performance improvements over time.

Prometheus: Native Azerbaijani MoE Features

Three Deployment Tiers

Prometheus is available in 39B, 99B, and 587B parameter configurations. Organizations can choose the tier that best balances accuracy requirements with available on-premise hardware, all within the same MoE framework.

Native Azerbaijani Tokenizer

The model ships with a tokenizer purpose-built for Azerbaijani, correctly handling the ə character and the language's agglutinative morphology. This means the MoE routing operates on linguistically meaningful tokens rather than fragmented approximations.

4.6× Efficiency on Azerbaijani Text

Because the tokenizer and expert routing are co-designed for Azerbaijani, Prometheus processes Azerbaijani text 4.6× more efficiently than architectures built for other languages and adapted afterward.

Trained on 651M+ Curated Azerbaijani Words

The experts inside Prometheus were trained on more than 651 million curated Azerbaijani words, giving each specialist sub-network deep exposure to authentic language patterns across a wide range of topics.

TUMLU Benchmark Validation

Model quality is verified against TUMLU, a benchmark comprising 38,139 native Azerbaijani questions spanning 11 disciplines, providing an objective, domain-diverse measure of real-world performance.

Fully On-Premise Deployment

Prometheus runs entirely within your own network infrastructure. No data is transmitted to external servers, satisfying strict data-sovereignty and compliance requirements common in regulated industries.

The Prometheus MoE Workflow

1Input tokenization: incoming text is broken into tokens. In Prometheus, the native tokenizer converts Azerbaijani text into tokens that respect the language's morphological structure, including characters like ə.
2Gating decision: a lightweight routing network scores each token against all available experts and selects the most relevant subset to activate, leaving the remaining experts idle for that request.
3Expert processing: the selected experts process their assigned tokens in parallel, each applying the specialized knowledge encoded during training on the 651M+ word Azerbaijani corpus.
4Output aggregation: the results from active experts are combined and passed through the model's remaining layers to produce a coherent, contextually accurate response.
5On-premise inference: the entire process runs inside your organization's own infrastructure, so sensitive data never leaves your network at any stage.
6Benchmark-verified output: responses are grounded in a model validated on the TUMLU benchmark, giving teams confidence that quality has been measured against a rigorous, domain-diverse standard.

Frequently Asked Questions

Does a 587B-parameter MoE model require 587B parameters worth of compute on every request?

No. The defining property of MoE is sparse activation: only a subset of experts is engaged per token. The total parameter count reflects the model's overall knowledge capacity, not the actual compute consumed per individual inference call.

Why is a native tokenizer necessary for the Azerbaijani language?

Azerbaijani is agglutinative, meaning grammatical information is added via suffixes to root words, and it uses unique characters like ə. Generic tokenizers often fragment these into meaningless pieces, wasting tokens and degrading understanding. Prometheus's native tokenizer enables 4.6× greater efficiency on Azerbaijani text.

What is the TUMLU benchmark and how does it validate Prometheus?

TUMLU is a rigorous evaluation benchmark containing 38,139 native Azerbaijani questions across 11 different disciplines. It provides an objective, multi-domain measure of how the model understands the language in real-world contexts, rather than relying on translated proxy benchmarks.

Is Prometheus suitable for industries with strict data-sovereignty requirements?

Yes. Prometheus is designed for fully on-premise deployment. All inference and processing occur within your own network infrastructure, ensuring that no data is ever transmitted to external servers or third-party vendors.

How should I choose between the 39B, 99B, and 587B parameter versions?

The choice depends on your specific balance of accuracy, complexity, and hardware. Smaller tiers (39B) offer faster responses and lower infrastructure requirements, while larger tiers (587B) provide deeper reasoning and broader language coverage. Allmaz can assist in evaluating your specific workload to determine the ideal fit.

Deploy Prometheus in Your Organization

Prometheus is the first large language model built natively for Azerbaijani — combining a Mixture of Experts architecture, a purpose-built tokenizer, and fully on-premise deployment. Contact the Allmaz team to discuss which parameter tier fits your infrastructure and how Prometheus can support your business goals.

Request a demo