ZIPA Large CR-CTC β€” Luxembourgish + IPAPack++ 40% (finetune)

Fine-tuned from anyspeech/zipa-large-crctc-ns-800k (ZIPA Zipformer CR-CTC large with diacritics, NS 800k).

This checkpoint is a Luxembourgish MFA-phone IPA ASR model trained at the University of Luxembourg (ASR group, Department of Humanities). It is fine-tuned on Luxembourgish phone data mixed with ~40% IPAPack++ multilingual phone replay.

Item Value
Architecture Zipformer large CR-CTC (CTC + consistency-regularized CTC)
Base model anyspeech/zipa-large-crctc-ns-800k (with diacritics)
Diacritics Kept (remove_diacritics=False)
Output Space-separated MFA phone IPA tokens
Tokenizer lb_phones.model (SentencePiece, vocab size 71)
IPAPack++ replay ratio 0.40 (~23.9k replay cuts vs ~59.7k LB train cuts)
Replay languages English, German, Dutch, French (balanced quotas)
Epochs 10
Best valid loss ~0.145 (epoch 10)
Base LR 5e-4
License MIT (aligned with ZIPA)

Files

File Description
best-valid-loss.pt Lean PyTorch checkpoint (model / model_avg + architecture metadata; no optimizer)
config.json Hyperparameters and finetune metadata
lb_phones.model / lb_phones.vocab SentencePiece phone tokenizer
phones.json / phones.vocab Phone inventory helpers
tokens.txt Token id table

Intended use

Phone-level IPA recognition / forced-alignment pipelines for Luxembourgish, with extra robustness from IPAPack++ multilingual phone replay. Not a word-level orthographic ASR model.

Training data

  • In-domain: Luxembourgish speech with MFA phone IPA labels (lb_shar_phones)
  • Replay: IPAPack++ subset (anyspeech/datasets), remapped to the Luxembourgish MFA phone inventory at coverage β‰₯ 0.7

Inference notes

Load with the ZIPA / icefall Zipformer CR-CTC stack. The checkpoint prefers the averaged weights under the model_avg key (also mirrored as model).

Example (Hugging Face Hub):

from huggingface_hub import hf_hub_download
from zipa_ctc_inference import initialize_model

ckpt = hf_hub_download(
    "pgilles/zipa-large-crctc-luxembourgish-ipapack40",
    "best-valid-loss.pt",
)
bpe = hf_hub_download(
    "pgilles/zipa-large-crctc-luxembourgish-ipapack40",
    "lb_phones.model",
)
model = initialize_model(ckpt, bpe)
hyps = model.inference(audio_waveforms_16khz)  # list of phone-token lists

Architecture hyperparameters match the ZIPA large CR-CTC encoder (num_encoder_layers=4,3,4,5,4,4, etc.); see config.json.

Citation / attribution

License

MIT, matching the upstream ZIPA license. The Hugging Face base-weight repo currently declares no license field; if AnySpeech clarifies a more restrictive term for the pretrained weights, revisit this card.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for pgilles/zipa-large-crctc-luxembourgish-ipapack40

Finetuned
(2)
this model