plawanrath/static-student-modernbert-large-civil

answerdotai/ModernBERT-large (revision 45bb4654) fine-tuned for binary toxicity classification on Civil Comments, from the static-student project (Weights as Constants). It was trained as a realistic-size test model for compiling an encoder's weights into a native binary as a read-only constant and certifying a decision threshold on the kernel that ships. Code and results: https://github.com/plawanrath/static-student.

Training

data google/civil_comments@f2970eb3, split train
label toxic if toxicity >= 0.5, else non_toxic
rows 40,000, positive share 0.25 (balanced subsample)
epochs / steps / batch 1 / 1250 / 32
learning rate / max length 2e-05 / 128
seed 20260926
final training loss (last 100 steps) 0.241
wall time (Apple silicon, MPS) 29 min

Compiled-kernel results (from the repository's results/)

bits weight bytes as a constant kernel bit-exact to integer reference accuracy fp32 → kernel argmax agreement
8 439,499,648 1000/1000 93.15% → 93.33% 99.42%
4 267,402,112 1000/1000 93.15% → 93.20% 98.90%

Accuracy is on a held-out Civil Comments test pool; see results/w03_modernbert_toxicity_large_w{8,4}/summary.json for pool sizes, hashes and the guarantee records.

How to use

from transformers import AutoTokenizer, AutoModelForSequenceClassification
tok = AutoTokenizer.from_pretrained("plawanrath/static-student-modernbert-large-civil")
model = AutoModelForSequenceClassification.from_pretrained("plawanrath/static-student-modernbert-large-civil")
logits = model(**tok("you are wonderful", return_tensors="pt")).logits
print(model.config.id2label[int(logits.argmax())])

Intended use and limitations

  • A test model for the compile-as-constants pipeline, released so the repository's numbers can be reproduced. It is not a moderation system.
  • One epoch on 40k rows of English comments; toxicity labels are crowd annotations with known demographic biases (see the Civil Comments dataset card). Expect false positives on identity terms and dialect.
  • Max sequence length 128 tokens at training time.
  • The research project was concluded without a publication; the model is not maintained.

Citation

If you use this checkpoint, please cite the repository:

@software{rath2026staticstudent,
  author  = {Rath, Plawan Kumar},
  title   = {static-student: Weights as Constants --- compiling distilled task models into executables as read-only data},
  year    = {2026},
  doi     = {10.5281/zenodo.23202887},
  url     = {https://github.com/plawanrath/static-student},
  license = {Apache-2.0},
  version = {0.1.0}
}

CITATION.cff in the repository carries the same entry (GitHub's "Cite this repository" button).

Please also cite ModernBERT (Warner et al., 2024) and the Civil Comments dataset (Borkan et al., 2019).

Disclaimer

Plawan Kumar Rath (@plawanrath). This work was conducted in the author's personal capacity. The views expressed here are those of the author and do not necessarily reflect the views of Meta.

Downloads last month
30
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for plawanrath/static-student-modernbert-large-civil

Finetuned
(395)
this model

Dataset used to train plawanrath/static-student-modernbert-large-civil

Collection including plawanrath/static-student-modernbert-large-civil