🥖 Baguettotron-MoE

Pleias

Paper (NeurIPS 2026) · SYNTH dataset · Baguettotron-350M · Baguettotron-600M

Baguettotron-MoE is a 13B-total / 1B-active Mixture-of-Experts reasoning model trained from scratch on 50B tokens of SYNTH, a fully synthetic corpus amplified from 58,698 Wikipedia articles. It is the sparse member of the Baguettotron suite in our NeurIPS 2026 paper, It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs.

  • Single training stage: no separate mid-training, SFT or RLHF. The model follows instructions, reasons in <think> traces and recalls facts because SYNTH trains those capabilities directly.
  • Less than one pass over SYNTH: reaches the same final loss as the dense Baguettotron-600M (1.19) on ~32% of its tokens.
  • Highest FActScore in the paper: 46.3% macro, ahead of OLMoE-1B-7B-Instruct (33.4%) and DeepSeek-MoE-16B-Chat (38.3%), trained on 40–100× fewer tokens.
  • Knows what it doesn't know: abstains on 67% of held-out Wikipedia entities outside its seed corpus, against 7% for OLMoE.

Model design

A Mixtral-style sparse decoder with top-1 routing over 16 experts and no shared expert.

Baguettotron-MoE architecture

Total parameters 13.3B
Active parameters per token 1.08B (excl. the 84M input-embedding lookup)
Layers 47
Hidden size 1280
Attention GQA, 20 query / 4 KV heads, head dim 64
Experts 16 routed, top-1, no shared expert
Expert MLP SwiGLU, intermediate size 4480
Router Softmax, no top-k renormalization
Position encoding RoPE, θ = 10,000
Norm RMSNorm (pre-norm), ε = 1e-6
Embeddings Untied
Vocabulary 65,536 (Pleias tokenizer)
Context length 2,048 tokens

Why trust_remote_code is needed. Stock Mixtral renormalizes the selected router scores to sum to 1, so with top-1 routing every expert output gets weight 1.0. This model was trained with the raw softmax probability of the chosen expert as the gate (typically 0.16–0.5). Loading it as plain Mixtral inflates expert outputs 2–6×. modeling_baguettotron_moe.py overrides only the router and keeps everything else identical to Mixtral.

Training

  • Data: SYNTH, ~50B tokens (less than one pass over the ~75B-token corpus)
  • Steps: 31,847 at global batch 768 × 2,048 tokens (~1.57M tokens/step)
  • Optimizer: AdamW, peak LR 3e-3, 1k-step warmup, linear decay over the final 16.6% of steps to 0.2% of peak; weight decay 0.01, grad clip 1.0
  • Hardware: 16× H100 (4 nodes × 4 GPUs) on MareNostrum 5 (BSC), torchtitan with FSDP (experts sharded through FSDP, no expert parallelism), ~17.5% MFU

Evaluation

All scores are from the paper. Each model is prompted in its native format; Baguettotron models get no system prompt and have <think>\n seeded after the assistant turn.

Accuracy vs training tokens

General benchmarks

Model Training tokens MCQ (22 tasks) Open-ended (8 tasks)
Baguettotron-MoE 50B 39.6 23.8
Baguettotron-600M 158B 42.2 24.3
Baguettotron-350M 200B 36.3 16.0
Qwen3-0.6B ~36T 46.9 25.0
LFM2.5-350M 28T 42.3 20.8
SmolLM2-360M-Instruct 4T 22.8 21.6
Gemma-3-270M-IT 2T 22.5 14.9

Factual precision (FActScore, 500 Wikipedia entities)

Model Training tokens S/(S+C) Macro
Baguettotron-MoE (13B / 1B active) 50B 81.7 46.3
DeepSeek-MoE-16B-Chat (16B / 2.8B active) ~2T 82.9 38.3
OLMoE-1B-7B-Instruct (7B / 1B active) ~5.1T 82.6 33.4
  • S/(S+C): share of atomic facts supported by Wikipedia among supported + contradicted
  • Macro: per-entity supported / all facts, with inconclusive facts counted in the denominator
  • Held-out entities: on 600 Wikipedia Good Articles outside the seed corpus, the model abstains on 67% (vs. 20% in-seed), and precision on attempted answers drops from 82% to 62%. OLMoE abstains on 7% and keeps 81%, since web-scale data covers more of these entities.

Prompt format

The model was trained on ChatML with a <think> block and no system prompt. The bundled chat template opens the assistant turn with <think>\n:

<|im_start|>user
What do you know about the Treaty of Westphalia?<|im_end|>
<|im_start|>assistant
<think>

The model writes its reasoning, closes it with </think>, answers, and ends the turn with <|im_end|>.

RAG. Pass sources inside the user turn. The answer then cites them with <ref>[quote]</ref>:

<|im_start|>user
{question}

<source_1>[…]</source_1>
<source_2>[…]</source_2><|im_end|>
<|im_start|>assistant
<think>

Reasoning notation. Traces use SYNTH's compact stenographic style (→ derivation, ↺ backtracking, ∴ conclusion, …). The confidence markers ● (high), ◐ (partial) and ○ (low) are informative: on traces dominated by uncertain markers, FActScore drops from 0.50 to 0.41, and the model commits to ~15–20% fewer facts. See the Baguettotron-350M card for the full notation.

Inference

The bf16 weights are 26.6 GB and fit on a single 80 GB GPU.

vLLM

vllm serve PleIAs/baguettotron-moe \
  --trust-remote-code \
  --model-impl transformers \
  --reasoning-parser deepseek_r1 \
  --enforce-eager
  • --trust-remote-code --model-impl transformers loads the custom router through vLLM's Transformers backend, which still swaps in fused MoE kernels for the experts.
  • --reasoning-parser deepseek_r1 moves the <think> trace into a separate reasoning field.
  • Tested with vLLM 0.24.0 and transformers 5.13.0.
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
    model="PleIAs/baguettotron-moe",
    messages=[{"role": "user", "content": "What do you know about the Treaty of Westphalia?"}],
    temperature=0.1,
    top_p=0.95,
    presence_penalty=0.1,
    max_tokens=1536,
)
print(resp.choices[0].message.reasoning)  # <think> trace
print(resp.choices[0].message.content)    # answer

Recommended sampling: temperature=0.1, top_p=0.95, presence_penalty=0.1. Prompt and generation share the 2,048-token context, so keep max_tokens below 2,048 minus the prompt length.

Transformers

Requires transformers>=5 (the modeling code subclasses MixtralTopKRouter).

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "PleIAs/baguettotron-moe"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, trust_remote_code=True, dtype="bfloat16", device_map="auto"
)

messages = [{"role": "user", "content": "What do you know about the Treaty of Westphalia?"}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=1536, do_sample=True, temperature=0.1, top_p=0.95)
print(tokenizer.decode(out[0, inputs.shape[1]:]))

Limitations

  • 2,048-token context shared by prompt, reasoning and answer
  • Knowledge is bounded by the seed corpus (~58.7k Wikipedia articles): recall is precise on seed topics, and the model mostly abstains outside them
  • Weaker on broad web-knowledge benchmarks (e.g. ARC-Challenge, GeoBench) than models trained on trillions of web tokens
  • Reasoning traces are English-only, even when the prompt and answer are in another language
  • No preference tuning or safety alignment; not intended for high-stakes use without further evaluation

Citation

@inproceedings{langlais2026itsalltraining,
  title         = {It's All Training: A Fully Synthetic Single-Stage Recipe for {LLMs}},
  author        = {Langlais, Pierre-Carl and Delobelle, Pieter and Detrois, Yannick and Chizhov, Pavel and Rosas-Hinostroza, Carlos and Si Smail, Neil and Burtin, Benjamin and Shcharbakova, Hanna and Yamshchikov, Ivan and Stasenko, Anastasia},
  booktitle     = {Advances in Neural Information Processing Systems},
  year          = {2026},
  eprint        = {2609.37891},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.37891}
}

Acknowledgements

Trained on MareNostrum 5 (BSC) through the EuroHPC Extreme Scale Access call (EHPC-EXT-2025E01-092, JULIP: Powerful LLMs made in Europe). Parts of this research received funding from SPRIN-D, the German Federal Agency for Breakthrough Innovation.

Downloads last month
17
Safetensors
Model size
13B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train PleIAs/Baguettotron-MoE

Paper for PleIAs/Baguettotron-MoE