Instructions to use PleIAs/Baguettotron-MoE with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PleIAs/Baguettotron-MoE with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PleIAs/Baguettotron-MoE", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("PleIAs/Baguettotron-MoE", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PleIAs/Baguettotron-MoE with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PleIAs/Baguettotron-MoE" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/Baguettotron-MoE", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PleIAs/Baguettotron-MoE
- SGLang
How to use PleIAs/Baguettotron-MoE with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PleIAs/Baguettotron-MoE" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/Baguettotron-MoE", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PleIAs/Baguettotron-MoE" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/Baguettotron-MoE", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PleIAs/Baguettotron-MoE with Docker Model Runner:
docker model run hf.co/PleIAs/Baguettotron-MoE
🥖 Baguettotron-MoE
Paper (NeurIPS 2026) · SYNTH dataset · Baguettotron-350M · Baguettotron-600M
Baguettotron-MoE is a 13B-total / 1B-active Mixture-of-Experts reasoning model trained from scratch on 50B tokens of SYNTH, a fully synthetic corpus amplified from 58,698 Wikipedia articles. It is the sparse member of the Baguettotron suite in our NeurIPS 2026 paper, It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs.
- Single training stage: no separate mid-training, SFT or RLHF. The model follows instructions, reasons in
<think>traces and recalls facts because SYNTH trains those capabilities directly. - Less than one pass over SYNTH: reaches the same final loss as the dense Baguettotron-600M (1.19) on ~32% of its tokens.
- Highest FActScore in the paper: 46.3% macro, ahead of OLMoE-1B-7B-Instruct (33.4%) and DeepSeek-MoE-16B-Chat (38.3%), trained on 40–100× fewer tokens.
- Knows what it doesn't know: abstains on 67% of held-out Wikipedia entities outside its seed corpus, against 7% for OLMoE.
Model design
A Mixtral-style sparse decoder with top-1 routing over 16 experts and no shared expert.
| Total parameters | 13.3B |
| Active parameters per token | 1.08B (excl. the 84M input-embedding lookup) |
| Layers | 47 |
| Hidden size | 1280 |
| Attention | GQA, 20 query / 4 KV heads, head dim 64 |
| Experts | 16 routed, top-1, no shared expert |
| Expert MLP | SwiGLU, intermediate size 4480 |
| Router | Softmax, no top-k renormalization |
| Position encoding | RoPE, θ = 10,000 |
| Norm | RMSNorm (pre-norm), ε = 1e-6 |
| Embeddings | Untied |
| Vocabulary | 65,536 (Pleias tokenizer) |
| Context length | 2,048 tokens |
Why trust_remote_code is needed. Stock Mixtral renormalizes the selected router scores to sum to 1, so with top-1 routing every expert output gets weight 1.0. This model was trained with the raw softmax probability of the chosen expert as the gate (typically 0.16–0.5). Loading it as plain Mixtral inflates expert outputs 2–6×. modeling_baguettotron_moe.py overrides only the router and keeps everything else identical to Mixtral.
Training
- Data: SYNTH, ~50B tokens (less than one pass over the ~75B-token corpus)
- Steps: 31,847 at global batch 768 × 2,048 tokens (~1.57M tokens/step)
- Optimizer: AdamW, peak LR 3e-3, 1k-step warmup, linear decay over the final 16.6% of steps to 0.2% of peak; weight decay 0.01, grad clip 1.0
- Hardware: 16× H100 (4 nodes × 4 GPUs) on MareNostrum 5 (BSC),
torchtitanwith FSDP (experts sharded through FSDP, no expert parallelism), ~17.5% MFU
Evaluation
All scores are from the paper. Each model is prompted in its native format; Baguettotron models get no system prompt and have <think>\n seeded after the assistant turn.
General benchmarks
| Model | Training tokens | MCQ (22 tasks) | Open-ended (8 tasks) |
|---|---|---|---|
| Baguettotron-MoE | 50B | 39.6 | 23.8 |
| Baguettotron-600M | 158B | 42.2 | 24.3 |
| Baguettotron-350M | 200B | 36.3 | 16.0 |
| Qwen3-0.6B | ~36T | 46.9 | 25.0 |
| LFM2.5-350M | 28T | 42.3 | 20.8 |
| SmolLM2-360M-Instruct | 4T | 22.8 | 21.6 |
| Gemma-3-270M-IT | 2T | 22.5 | 14.9 |
Factual precision (FActScore, 500 Wikipedia entities)
| Model | Training tokens | S/(S+C) | Macro |
|---|---|---|---|
| Baguettotron-MoE (13B / 1B active) | 50B | 81.7 | 46.3 |
| DeepSeek-MoE-16B-Chat (16B / 2.8B active) | ~2T | 82.9 | 38.3 |
| OLMoE-1B-7B-Instruct (7B / 1B active) | ~5.1T | 82.6 | 33.4 |
- S/(S+C): share of atomic facts supported by Wikipedia among supported + contradicted
- Macro: per-entity supported / all facts, with inconclusive facts counted in the denominator
- Held-out entities: on 600 Wikipedia Good Articles outside the seed corpus, the model abstains on 67% (vs. 20% in-seed), and precision on attempted answers drops from 82% to 62%. OLMoE abstains on 7% and keeps 81%, since web-scale data covers more of these entities.
Prompt format
The model was trained on ChatML with a <think> block and no system prompt. The bundled chat template opens the assistant turn with <think>\n:
<|im_start|>user
What do you know about the Treaty of Westphalia?<|im_end|>
<|im_start|>assistant
<think>
The model writes its reasoning, closes it with </think>, answers, and ends the turn with <|im_end|>.
RAG. Pass sources inside the user turn. The answer then cites them with <ref>[quote]</ref>:
<|im_start|>user
{question}
<source_1>[…]</source_1>
<source_2>[…]</source_2><|im_end|>
<|im_start|>assistant
<think>
Reasoning notation. Traces use SYNTH's compact stenographic style (→ derivation, ↺ backtracking, ∴ conclusion, …). The confidence markers ● (high), ◐ (partial) and ○ (low) are informative: on traces dominated by uncertain markers, FActScore drops from 0.50 to 0.41, and the model commits to ~15–20% fewer facts. See the Baguettotron-350M card for the full notation.
Inference
The bf16 weights are 26.6 GB and fit on a single 80 GB GPU.
vLLM
vllm serve PleIAs/baguettotron-moe \
--trust-remote-code \
--model-impl transformers \
--reasoning-parser deepseek_r1 \
--enforce-eager
--trust-remote-code --model-impl transformersloads the custom router through vLLM's Transformers backend, which still swaps in fused MoE kernels for the experts.--reasoning-parser deepseek_r1moves the<think>trace into a separatereasoningfield.- Tested with vLLM 0.24.0 and transformers 5.13.0.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="PleIAs/baguettotron-moe",
messages=[{"role": "user", "content": "What do you know about the Treaty of Westphalia?"}],
temperature=0.1,
top_p=0.95,
presence_penalty=0.1,
max_tokens=1536,
)
print(resp.choices[0].message.reasoning) # <think> trace
print(resp.choices[0].message.content) # answer
Recommended sampling: temperature=0.1, top_p=0.95, presence_penalty=0.1. Prompt and generation share the 2,048-token context, so keep max_tokens below 2,048 minus the prompt length.
Transformers
Requires transformers>=5 (the modeling code subclasses MixtralTopKRouter).
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "PleIAs/baguettotron-moe"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True, dtype="bfloat16", device_map="auto"
)
messages = [{"role": "user", "content": "What do you know about the Treaty of Westphalia?"}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=1536, do_sample=True, temperature=0.1, top_p=0.95)
print(tokenizer.decode(out[0, inputs.shape[1]:]))
Limitations
- 2,048-token context shared by prompt, reasoning and answer
- Knowledge is bounded by the seed corpus (~58.7k Wikipedia articles): recall is precise on seed topics, and the model mostly abstains outside them
- Weaker on broad web-knowledge benchmarks (e.g. ARC-Challenge, GeoBench) than models trained on trillions of web tokens
- Reasoning traces are English-only, even when the prompt and answer are in another language
- No preference tuning or safety alignment; not intended for high-stakes use without further evaluation
Citation
@inproceedings{langlais2026itsalltraining,
title = {It's All Training: A Fully Synthetic Single-Stage Recipe for {LLMs}},
author = {Langlais, Pierre-Carl and Delobelle, Pieter and Detrois, Yannick and Chizhov, Pavel and Rosas-Hinostroza, Carlos and Si Smail, Neil and Burtin, Benjamin and Shcharbakova, Hanna and Yamshchikov, Ivan and Stasenko, Anastasia},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
eprint = {2609.37891},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.37891}
}
Acknowledgements
Trained on MareNostrum 5 (BSC) through the EuroHPC Extreme Scale Access call (EHPC-EXT-2025E01-092, JULIP: Powerful LLMs made in Europe). Parts of this research received funding from SPRIN-D, the German Federal Agency for Breakthrough Innovation.
- Downloads last month
- 17