Instructions to use bytkim/Qwen3.8-27B-pi-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bytkim/Qwen3.8-27B-pi-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bytkim/Qwen3.8-27B-pi-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("bytkim/Qwen3.8-27B-pi-FP8") model = AutoModelForMultimodalLM.from_pretrained("bytkim/Qwen3.8-27B-pi-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bytkim/Qwen3.8-27B-pi-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bytkim/Qwen3.8-27B-pi-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bytkim/Qwen3.8-27B-pi-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bytkim/Qwen3.8-27B-pi-FP8
- SGLang
How to use bytkim/Qwen3.8-27B-pi-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bytkim/Qwen3.8-27B-pi-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bytkim/Qwen3.8-27B-pi-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bytkim/Qwen3.8-27B-pi-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bytkim/Qwen3.8-27B-pi-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use bytkim/Qwen3.8-27B-pi-FP8 with Docker Model Runner:
docker model run hf.co/bytkim/Qwen3.8-27B-pi-FP8
Quickstart · Benchmarks · Advanced · License
Built for pi
Qwen3.8-27B-pi builds on Qwen3.8-27B for coding work in the Pi agent harness. It is tailored to the loop of reading a repository, editing files, running tools, and responding to their feedback. The aim is more completed work with less generated text, while retaining the base model's familiar interface.
Fine-tuned for the Pi agent harness, the model is designed to work through coding tasks as an iterative process: inspect the code, make a change, check the result, and use tool feedback to guide the next step. Adjustable reasoning effort lets you balance responsiveness with deeper problem-solving, from focused edits to more involved debugging and implementation work.
Qwen3.8-27B-pi Highlights
- Built for Pi: Fine-tuned for the everyday coding loop—reading repositories, editing files, running tools, and working through feedback in the Pi harness—with an emphasis on turning plans into working, checked implementations.
- Curated coding experience: Supervised fine-tuning on filtered, successful Pi sessions helps the model learn complete coding workflows, rather than isolated answers or code snippets, including how to adapt to existing environments and check results against task requirements.
- Task quality matters: Development also included reviewing, repairing, and exploring more demanding coding tasks—with a focus on clear requirements and checks that distinguish working solutions from incorrect ones.
- Refined through reinforcement learning: A second training stage pairs verified task outcomes with a custom reasoning-efficiency reward built on GRPO, encouraging successful low- and medium-effort solutions to reason more economically while leaving xhigh focused on correctness.
- Available in practical formats: Choose a deployment format that fits your hardware, including GGUF quantizations calibrated using complete Pi coding sessions.
Pi’s development followed a staged process, from trajectory curation and supervised fine-tuning to reinforcement learning and checkpoint evaluation. Each stage was assessed against the same practical goal: helping the model complete useful work in Pi while managing the resources it spends. Checkpoints were compared using actual agent outcomes alongside generated tokens, tool calls, and completion time—not training loss alone.
Performance by reasoning level
Curated Pi sessions and reinforcement learning emphasized completed, checked coding work. Pi shows a steadier rise in completion from low to xhigh than Base, with fewer output tokens at every matched effort level. In these selected results, Pi’s medium setting matches Base’s xhigh completion rate with about 41% fewer output tokens.
Pi’s RL stage encouraged economical low/medium reasoning while keeping xhigh focused on correctness. The graph shows a smoother rise in Pi’s attempt success as effort increases, while Base peaks at medium. Pi achieves the highest xhigh score, but Base retains the medium-effort edge.
Pi’s development prioritized successful solutions alongside resource use—not shorter responses alone. Both models improve with higher effort, but Pi solves more subproblems at every level. At xhigh, Pi scores higher with about 23% fewer output tokens; at medium, its higher score requires more tokens.
Quickstart
vllm serve bytkim/Qwen3.8-27B-pi-FP8 \
--host 127.0.0.1 \
--max-model-len 262144 \
--max-num-seqs 1 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
By default, vLLM loads this repository’s
generation_config.json, which sets thinking-mode sampling totemperature=1.0,top_p=0.95, andtop_k=20. No additional sampling flags are needed; client-provided values override these defaults, and--generation-config vllmuses vLLM’s built-in defaults instead.
Non-thinking mode
To serve with non-thinking mode and its recommended sampling settings as defaults, use this command instead of the one above:
vllm serve bytkim/Qwen3.8-27B-pi-FP8 \
--host 127.0.0.1 \
--max-model-len 262144 \
--max-num-seqs 1 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--default-chat-template-kwargs '{"enable_thinking":false}' \
--override-generation-config '{"temperature":0.7,"top_p":0.8,"presence_penalty":1.5}'
Alternatively, keep the thinking-mode server above and override individual API requests:
{
"temperature": 0.7,
"top_p": 0.8,
"presence_penalty": 1.5,
"chat_template_kwargs": {
"enable_thinking": false
}
}
These overrides follow Qwen’s recommended non-thinking settings. The remaining sampling settings are inherited from the model and vLLM defaults; turning thinking off alone does not change sampling.
Advanced
SGLang serving is also supported, including DFlash2 speculative decoding and FP8 KV cache.
DFlash2 companion weights are mirrored unchanged from Inco AI, with attribution and license included in
dflash2/.
The examples below use DFlash2 instead of the Quickstart's MTP configuration, with FP8 KV cache and a 256K context limit. Use a runtime build with DFlash2 support and sufficient GPU memory; these are documentation-based examples.
Download the DFlash2 companion
Download the mirrored DFlash2 companion once before starting either server. The ~3.85 GB BF16 draft works as the companion for either target format.
hf download bytkim/Qwen3.8-27B-pi-FP8 \
--include "dflash2/*" \
--local-dir models/pi-fp8
SGLang: DFlash2 + FP8 KV cache
python -m sglang.launch_server \
--model-path bytkim/Qwen3.8-27B-pi-FP8 \
--host 127.0.0.1 \
--port 8000 \
--context-length 262144 \
--max-running-requests 1 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--kv-cache-dtype fp8_e4m3 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path models/pi-fp8/dflash2 \
--speculative-num-draft-tokens 8
vLLM: DFlash2 + FP8 KV cache
vllm serve bytkim/Qwen3.8-27B-pi-FP8 \
--host 127.0.0.1 \
--port 8000 \
--max-model-len 262144 \
--max-num-seqs 1 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--kv-cache-dtype fp8_e4m3 \
--speculative-config '{"method":"dflash","model":"models/pi-fp8/dflash2","num_speculative_tokens":7}'
Run one server at a time. SGLang's eight-token verification block corresponds to seven draft tokens in the DFlash2 example; vLLM specifies the seven draft tokens directly. FP8 KV cache is independent of weight precision. For default-precision KV cache, change --kv-cache-dtype fp8_e4m3 to --kv-cache-dtype auto.
References: DFlash2 model and launch examples, SGLang speculative decoding, SGLang server arguments, and vLLM FP8 KV cache.
License
Qwen3.8-27B-pi is fine-tuned from Qwen3.8-27B, whose base-model weights are licensed under the Apache License 2.0. See LICENSE for the full terms and NOTICE for upstream attribution and Pi modification details. Mirrored DFlash2 companions retain their license and provenance in dflash2/.
Acknowledgements
Built on Qwen3.8-27B from the Qwen team and adapted for the Pi agent harness. Thanks to the open-source training, inference, quantization, and evaluation projects—and dataset contributors—that supported its development.
- Downloads last month
- 104
Model tree for bytkim/Qwen3.8-27B-pi-FP8
Base model
Qwen/Qwen3.8-27B