tinymixtral-it β€” instruction-tuned MoE

tinymixtral-it (a.k.a. v3.0-it) is the instruction-tuned version of the MoE flagship mikecovlee/tinymixtral: 477.5M total / 276.1M active parameters (top-2 of 4 experts), 2048 context, 32k SentencePiece vocabulary.

It is the MoE counterpart of the dense mikecovlee/tinymistral-276m-it: the exact same 3M-tier SFT recipe (same data, hyper-parameters and schedule) is applied to both the MoE and the dense base, so the two instruction-tuned models can be compared at the same active-parameter count and the same SFT budget.

Model details

Parameter Value
hidden_size 1024
num_layers 16
Attention Grouped Query Attention (16 heads / 4 KV heads)
Head dim 64
RoPE theta 1,000,000
Norm RMSNorm + per-head QK-Norm (pre-RoPE)
Experts 4 routed (top-2), aux loss 1e-3
Expert FFN SwiGLU, intermediate = 2048
Vocab size 32,000 (tied embeddings)
Max position 2,048
Total params ~477.5M
Active params ~276.1M

Training

  • Initialization: the public base mikecovlee/tinymixtral (v3.0).
  • Data: 2,168,835 deduplicated and eval-decontaminated English instruction conversations, blended from 10 public sources (Tulu3, OpenHermes, SlimOrca, OpenOrca, UltraChat, MetaMath, OrcaMath, OpenMathInstruct-2, SQuAD2, TriviaQA).
  • Recipe: 1 epoch, sequence packing to 1024 tokens, batch 24, bf16, AdamW, lr 2e-5 (cosine, 100-step warmup), weight decay 0.1, seed 42.

Evaluation

Measured with lm-evaluation-harness v0.4.12, 0-shot (cuda, bf16).

Harness = mean of the 7 primary metrics (hellaswag acc_norm, piqa acc, winogrande acc, arc_easy acc, arc_challenge acc_norm, openbookqa acc_norm, lambada acc).

Model 7-task harness 7-task all-acc GSM8K strict / flex IFEval prompt / inst
tinymixtral-it (MoE + 3M SFT) 0.3994 0.3691 0.0182 / 0.0205 0.1756 / 0.2782
tinymixtral (MoE base) 0.3992 β€” 0.0000 / 0.0159 β€”
tinymistral-276m-it (dense + 3M SFT) 0.3892 0.3631 0.0174 / 0.0205 0.1460 / 0.2602
tinymistral-276m (dense base) 0.3904 0.3890 0.0000 / 0.0136 β€”

Harness note. Means quoted here use the 7-task harness (excludes BoolQ); the base v1.0/v3.0 cards report an 8-task mean (includes BoolQ). The two are not directly comparable.

Per-task (tinymixtral-it): hellaswag 0.340, piqa 0.634, winogrande 0.528, arc_easy 0.448, arc_challenge 0.260, openbookqa 0.302, lambada 0.289 β€” plus BoolQ 0.426, reported separately (BoolQ is excluded from the 7-task mean above, to match the dense iso-active comparison).

Additional metric:

Metric Score
Open-ended answer quality (LLM rubric, 4,955 held-out prompts, 0–100) 15.0 Β± 0.3

Note: relative to the base model, instruction following and open-ended answer quality improve substantially while multiple-choice common-sense accuracy drops (BoolQ is the most sensitive task). Arithmetic remains far below practical use. The dense counterpart receives the same SFT recipe and lands close on the 7-task harness (0.3892 vs 0.3994) but is less steerable (IFEval 0.1460 / 0.2602).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "mikecovlee/tinymixtral-it"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "What is 12% of 250?"}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))

Conversational use via the bundled chat template (tokenizer.apply_chat_template); greedy decoding works well at this scale.

Limitations

  • Absolute numbers are bounded by the 477.5M MoE total / 276.1M active scale and the 8.05B-token pretraining budget; arithmetic remains far below practical use.
  • English-centric, no safety alignment.
  • Multiple-choice common-sense accuracy drops slightly after SFT (BoolQ most sensitive).

Family

Naming. The MoE family is published under tinymixtral; the dense 276M iso-active ablation companions use the tinymistral spelling. Both belong to the same project.

Citation

@misc{tinymixtralit2026,
  title  = {TinyMixtral: a small Mixture-of-Experts language-model family},
  author = {Michael Lee},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/mikecovlee/tinymixtral-it}}
}

License

MIT (Copyright (C) 2026 Michael Lee).

Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mikecovlee/tinymixtral-it

Finetuned
(1)
this model

Datasets used to train mikecovlee/tinymixtral-it

Collection including mikecovlee/tinymixtral-it