MEGA โ€” reference MoE language model

A small, fully reproducible Mixture-of-Experts causal language model trained from scratch as part of the MEGA reference implementation. This is a research artifact, not a general-purpose assistant: it was trained on a tiny structured corpus (below) and its competence is limited to that corpus.

Model details

architecture decoder-only Transformer + MoE
layers 6
d_model 256
heads / KV heads 8 / 2 (GQA)
attention gqa
MoE 16 routed experts, top-2, 2 shared
context 256 tokens
vocab 800
total params 31.53M
active params / token 9.33M
precision fp32 (CPU-trained)

Intended use

Demonstrating and testing the MEGA training pipeline: tokenizer, data packing, Muon+AdamW optimisation, MoE routing and auxiliary-loss-free balancing, checkpointing and Hub export. Do not use it where correctness matters.

Training data

A deterministic, structure-rich synthetic corpus (Frankenstein-Labs/mega-corpus): bilingual (fr/en) templated sentences, arithmetic, word problems and small Python snippets, generated from grammar templates. No scraped data and no third-party licence obligations.

Evaluation

metric value
perplexity 3.068
next_token_accuracy 0.6671
arithmetic_accuracy 0.35

How to use

import torch
from mega.hub.serialization import from_pretrained
from mega.tokenizer.tokenizer import MegaTokenizer

model = from_pretrained(".")            # or "<user>/<repo>"
tok = MegaTokenizer.from_pretrained(".")

ids = tok.encode("Le chat noir mange une pomme .", add_bos=True)
out = model.generate(torch.tensor([ids]), max_new_tokens=16, temperature=0.0)
print(tok.decode(out[0].tolist()))

Limitations

  • Trained on a few thousand tokens of templated text; expect plausible but arbitrary output outside that domain.
  • French/English only.
  • No safety tuning. Outputs are unmoderated.
  • CPU-trained at small scale; no distributed parallelism was used.

Citation

Part of the MEGA reference project. Licensed Apache-2.0.

Downloads last month
13
Safetensors
Model size
31.5M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Frankenstein-Labs/mega-small 1