Qwen3 Coder 30B A3B Instruct GGUF

Providerunsloth
Categorytext-generation
Licenseapache-2.0
Downloads108.1K
Stars28

Overview

Qwen3 Coder 30B A3B Instruct is a specialized Mixture-of-Experts (MoE) model optimized for high-performance programming tasks. By utilizing an active parameter count of 3B within a 30B total parameter architecture, it delivers the reasoning capabilities of a large model with the inference speed and memory efficiency of a much smaller one. For developers, this means a significant reduction in VRAM requirements without sacrificing complex logic handling or multi-language syntax accuracy. It is particularly effective for autonomous code generation, refactoring legacy systems, and acting as a local copilot. This GGUF quantization makes it highly accessible for local deployment via llama.cpp or Ollama, allowing seamless integration into IDEs without relying on cloud APIs.

Highlights

  • MoE architecture balances high reasoning with low latency.
  • Optimized for local deployment via GGUF quantization.
  • Strong multi-language support for complex software engineering.
  • Reduced VRAM overhead compared to dense 30B models.
  • Apache-2.0 license ensures flexible commercial integration.

Usage

Install
# Install Hugging Face transformers
pip install transformers torch
SDK Usage
# Load model with transformers
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF")
tokenizer = AutoTokenizer.from_pretrained("unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF")

Hugging Face Download

We recommend downloading the model via the Hugging Face CLI or Hub SDK.

Guidance:Before downloading, install huggingface_hub with:

Guidance
pip install -U huggingface_hub

CLI Download

Download the full repository

Download the full repository
huggingface-cli download unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF

Download a single file to a local folder (e.g. config.json into ./dir)

Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF config.json --local-dir ./dir

See the official docs for more CLI options

SDK Download

SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF

Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.

PyTorch / Transformers Usage

Install Transformers

Install Transformers
pip install -U transformers torch

Load the model and run inference

Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF')
tokenizer = AutoTokenizer.from_pretrained('unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF')

Model Download

We recommend downloading the model via the ModelScope CLI or SDK.

Guidance:Before downloading, install ModelScope with:

Guidance
pip install modelscope

CLI Download

Download the full repository

Download the full repository
modelscope download --model unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF

Download a single file to a local folder (e.g. README.md into ./dir)

Download a single file to a local folder (e.g. README.md into ./dir)
modelscope download --model unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF README.md --local_dir ./dir

See the docs for more CLI options

SDK Download

SDK Download
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://www.modelscope.cn/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF.git

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook Quickstart

Install the ModelScope library

Install the ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

Load the model and run inference

Load the model and run inference
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF')

Full Documentation

来源: HuggingFace

---
tags:

  • unsloth

  • qwen3

  • qwen

base_model:
  • Qwen/Qwen3-Coder-30B-A3B-Instruct

library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct/blob/main/LICENSE
pipeline_tag: text-generation
---
<div>
<p style="margin-bottom: 0; margin-top: 0;">
<strong>See <a href="https://huggingface.co/collections/unsloth/qwen3-680edabfb790c8c34a242f95">our collection</a> for all versions of Qwen3 including GGUF, 4-bit & 16-bit formats.</strong>
</p>
<p style="margin-bottom: 0;">
<em>Learn to run Qwen3-Coder correctly - <a href="https://docs.unsloth.ai/basics/qwen3-coder">Read our Guide</a>.</em>
</p>
<p style="margin-top: 0;margin-bottom: 0;">
<em>See <a href="https://docs.unsloth.ai/basics/unsloth-dynamic-v2.0-gguf">Unsloth Dynamic 2.0 GGUFs</a> for our quantization benchmarks.</em>
</p>
<div style="display: flex; gap: 5px; align-items: center; ">
<a href="https://github.com/unslothai/unsloth/">
<img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="133">
</a>
<a href="https://discord.gg/unsloth">
<img src="https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png" width="173">
</a>
<a href="https://docs.unsloth.ai/basics/qwen3-coder">
<img src="https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png" width="143">
</a>
</div>
<h1 style="margin-top: 0rem;">✨ Read our Qwen3-Coder Guide <a href="https://docs.unsloth.ai/basics/qwen3-coder">here</a>!</h1>
</div>

  • View the rest of our notebooks in our docs here.
| Unsloth supports | Free Notebooks | Performance | Memory use | |-----------------|--------------------------------------------------------------------------------------------------------------------------|-------------|----------| | Qwen3 (14B) | ▶️ Start on Colab | 3x faster | 70% less | | GRPO with Qwen3 (8B) | ▶️ Start on Colab | 3x faster | 80% less | | Llama-3.2 (3B) | ▶️ Start on Colab-Conversational.ipynb) | 2.4x faster | 58% less | | Llama-3.2 (11B vision) | ▶️ Start on Colab-Vision.ipynb) | 2x faster | 60% less | | Qwen2.5 (7B) | ▶️ Start on Colab-Alpaca.ipynb) | 2x faster | 60% less |

Qwen3-Coder-30B-A3B-Instruct

<a href="https://chat.qwen.ai/" target="_blank" style="margin: 2px;"> <img alt="Chat" src="https://img.shields.io/badge/%F0%9F%92%9C%EF%B8%8F%20Qwen%20Chat%20-536af5" style="display: inline-block; vertical-align: middle;"/> </a>

Highlights

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:

  • Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks.
  • Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding.
  • Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.

!image/jpeg

Model Overview

Qwen3-Coder-30B-A3B-Instruct has the following features:

  • Type: Causal Language Models

  • Training Stage: Pretraining & Post-training

  • Number of Parameters: 30.5B in total and 3.3B activated

  • Number of Layers: 48

  • Number of Attention Heads (GQA): 32 for Q and 4 for KV

  • Number of Experts: 128

  • Number of Activated Experts: 8

  • Context Length: 262,144 natively.

NOTE: This model supports only non-thinking mode and does not generate `<think></think> blocks in its output. Meanwhile, specifying enable_thinking=False is no longer required.

For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation.

Quickstart

We advise you to use the latest version of transformers.

With transformers<4.51.0, you will encounter the following error:

code
KeyError: 'qwen3_moe'

The following contains a code snippet illustrating how to use the model generate content based on given inputs.

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen3-Coder-30B-A3B-Instruct"

load the tokenizer and the model

tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype="auto", device_map="auto" )

prepare the model input

prompt = "Write a quick sort algorithm." messages = [ {"role": "user", "content": prompt} ] text = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, ) model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

conduct text completion

generated_ids = model.generate( model_inputs, max_new_tokens=65536 ) output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()

content = tokenizer.decode(output_ids, skip_special_tokens=True)

print("content:", content)

Note: If you encounter out-of-memory (OOM) issues, consider reducing the context length to a shorter value, such as 32,768.

For local use, applications such as Ollama, LMStudio, MLX-LM, llama.cpp, and KTransformers have also supported Qwen3.

Agentic Coding

Qwen3-Coder excels in tool calling capabilities.

You can simply define or use any tools as following example.

python
# Your tool implementation
def square_the_number(num: float) -> dict:
return num
2

Define Tools

tools=[ { "type":"function", "function":{ "name": "square_the_number", "description": "output the square of the number.", "parameters": { "type": "object", "required": ["input_num"], "properties": { 'input_num': { 'type': 'number', 'description': 'input_num is a number that will be squared' } }, } } } ]

import OpenAI

Define LLM


client = OpenAI(
# Use a custom endpoint compatible with OpenAI API
base_url='http://localhost:8000/v1', # api_base
api_key="EMPTY"
)

messages = [{'role': 'user', 'content': 'square the number 1024'}]

completion = client.chat.completions.create(
messages=messages,
model="Qwen3-Coder-30B-A3B-Instruct",
max_tokens=65536,
tools=tools,
)

print(completion.choice[0])

Best Practices

To achieve optimal performance, we recommend the following settings:

1. Sampling Parameters:
- We suggest using
temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05`.

2. Adequate Output Length: We recommend using an output length of 65,536 to

Join our Telegram