Gemma 4 26B A4B GRPO subques adapter

Final LoRA adapter for the subques condition of the VLM benchmark extension. This repository contains adapter and processor files, not the full base model. Load the base model unsloth/gemma-4-26B-A4B-it at revision 60941ad6341d0b7af91277ff25c4175f08b56819, then attach this adapter with PEFT. Evaluation uses the same condition's 100-row test split; the same adapter is also evaluated on the separate 200-row OOD dataset.

Training dataset: yobro4619/final_common_train at revision 4e9742593fa8f7f9e17615a8d856e49458f4bd11 (440 train rows). Benchmark source commit: d7204c5d37e6c2cb33894911d0d0e28ef25b8627. Training: three epochs, 660 optimizer steps, eight generations per prompt, per-device batch eight, gradient accumulation two, 4-bit QLoRA, 8192 maximum sequence length, 4096 maximum completion length, frozen vision tower. Gemma's MoE expert LoRA requires dropout 0 instead of the original Qwen 0.1. The training prompt has no VQA system instruction. Evaluation retains the benchmark VQA system instruction. See the run manifest for all parameters, trainable modules, environment, and command.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yobro4619/gemma4_26b_a4b_grpo_subques

Adapter
(14)
this model