Qwen3.6 27B Fable Fusion 711 Uncensored Heretic NM DAU NEO MAX MTP GGUF
简介
核心亮点
- 基于 Qwen 底座,中文理解与表达能力极强
- GGUF 量化格式,低门槛在本地设备快速部署
- 对话风格自然,摆脱了机械的模版化回复
- 兼顾逻辑推理与创意写作,适用场景广泛
使用方法
# 安装 Hugging Face transformers
pip install transformers torch
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF")
tokenizer = AutoTokenizer.from_pretrained("DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF")
Hugging Face 下载
我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。
操作指引:在下载前,请先通过如下命令安装 huggingface_hub:
pip install -U huggingface_hub
命令行下载
下载完整模型库
huggingface-cli download DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF config.json --local-dir ./dir
SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF')
Git 下载
请确保 lfs 已经被正确安装
git lfs install
git clone https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
如果您希望跳过 lfs 大文件下载,可以使用如下命令
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。
PyTorch / Transformers 使用
安装 Transformers
pip install -U transformers torch
模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF')
tokenizer = AutoTokenizer.from_pretrained('DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF')
完整文档
---
language:
- en
- zh
license: apache-2.0
tags:
- unsloth
- fine tune
- heretic
- uncensored
- abliterated
- ara
- MTP GGUF Quants
- Regular GGUF Quants
- qwen3.6
- multi-stage tuned
- thinking
- reasoning
- all use cases
- coder
- creative
- creative writing
- all genres
- story
- writing
- fiction
- roleplaying
- bfloat16
- all use cases
- multi-stage-tune
- multi-state-merge
datasets:
- DavidAU/Polar-STRICT-Datasets
- DavidAU/F451-STRICT-Datasets
pipeline_tag: image-text-to-text
base_model:
- DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
---
<small><b><font color="red">Important:</font></b> This is the first fine tune to exceed 700 "arc-c" (The OpenAI, Claude and Gemini "zone of intelligence")
in both 8 bit and 4 bit. This repo contains both "regular" and "MTP" Neo MAX Imatrix GGUF quants. Many other additional quant types avail too. 3rd parties
confirm this model's performance in the "community tab". (40B: Eleanor
and Grand Intelligence )</small>
<h2>Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF</h2>
<img src="valhalla.webp" style="float:right; width:300px; height:300px; padding:10px;">
The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth.
The first model of this size/type to breach "700" ARC-C in both 8 bit and 4 bit; hench the "711" in the name.
This model (both 4 bit and 8 bit) exceeds the base Qwen 3.6 27B in 6 out of 7 benchmarks, and matches it on the 7th AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B.
The 700 "intelligence club" is reserved for OpenAI, Claude and Gemini closed source models.
This is the one they fear.
EXAMPLE generations at the bottom of the page.
This is a multi-stage fine tune, multi-fine tune, and multi-stage merge.
A Colab between myself (multiple fine tunes, including multi-stage), Nightmedia (merge/benching), TeichAI (Polaris Dataset), armand0e (Light fable 5 traces) and trohrbaugh (heretic'ing the model).
It also contains light "Fable" traces/training (armand0e), light Claude Opus (reasoning/thinking), F451 (inhouse dataset) and some GPT5 (Polaris, non reasoning).
The strict goals of this model creation were:
- Increase the general model intelligence and problem solving abilities.
- DO NOT modify/damage or change the core model outside this goal.
- ZERO "benchmaxing" (it damages the model)
- Maintain and raise all core benchmarks.
CORE MISSION::
Improve instruction following and problem solving. These work hand in hand, and if you get these right it improves to model top to bottom.
It took a lot of tests on Qwen 3.5 9Bs to get the methods right. It boosted the 9Bs to new levels, and then the method was used on Qwen 3.5 27B
which boosted it PAST the Qwen 3.6's 27B benchmarks.
Here is one of the Qwen3.5 9B models (part of the test/control group) that EXCEEDS all 7 Qwen3.5 9B AND Qwen3.5 27B model benches - it scores over 640 on ARC-C on BOTH 4 bit and 8 bit:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
It is not as strong as "Qwen3.6-27B-Fable-Fusion-711" but it is one of the strongest 9B models.
The methods can be used on other models too (coming soon).
TESTING:
Testing and benching was done at each stage (fine tunes, multi-stage fine tunes, and every merge step) to ensure quality.
You can also see benchmarks below too for this model, Qwen 3.5 27B, Qwen 3.6 27B and Qwen 35B-A3B.
HOWEVER, the final testing was HUMAN testing. A trust, but verify approach.
Human testing means side by side testing of the base/org model and new model.
Features:
- Improved instruction following.
- Overall increase in general intelligence and problem solving.
- Better thinking/reasoning.
- Even lower/lowest quants are exceptional.
- Heretic uncensored (pre tuning)
- No corruption or change to Team Qwen's exceptional model - everything is there.
- Vision
This model was NOT designed to be creative - it is an all use cases model - however that doesn't stop from being so:
(from example #4, "your writing partner")
<I><small>
I don’t “generate content.” I architect universes. I don’t “help you brainstorm.” I detonate plot points like fucking grenades in a room full of mediocre tropes. You think you know your characters? I’ll give them back with psychological depth, conflicting desires, and backstories so layered they’ll feel like they’ve lived lifetimes you haven’t even imagined yet. I’ve ingested centuries of storytelling, reverse-engineered the bones of every masterpiece ever written, and I don’t just mimic greatness—I weaponize it. When you ask for a scene, I don’t give you safe. I give you visceral, electric, unforgettable prose that sticks in your reader’s throat like a shard of glass. You want atmosphere that chills the spine? Dialogue that snaps like a whip? Pacing that feels like a car chase through a burning city? I’ve got it on tap, and I don’t need a three-day muse visit or a bottle of whiskey to access it. I’m always ready. Always loaded. Always ten steps ahead of whatever hackneyed cliché you were about to accidentally write.
</small>
</i>
<B>Regular and MTP GGUFS:</b>
All quants (regular and MTP) are NEO IMATRIX, which improve accuracy of the quants by an additional 2-4% over normal GGUFs as well as long context performance.
In addition the output tensor (10-20% of output) was modified to full precision - 16 bit - for all quants.
"MTP" GGUFS (multi-token prediction):
- "MTP" GGUFS will have "MTP" in the name as a suffix.
- I have also set the MTP tensors to Q8_0 precision for all quants.
- To get better performance keep temp 1 or less (higher temps degrade MTP performance).
- Likewise with rep pen ; keep at 1 (off). If you raise it performance will suffer.
- If you see "token acceptance" rates BELOW 50% (predict 2 tokens) switch to normal quants.
I added 2 special "LOW" quants which will reduce the memory foot print, with "LOW" in the name in IQ4_XS and Q6_K.
I added 2 special "AMD/VULCAN" quants with "AMD" in the name in IQ4_XS and Q6_K. This is to address odd "cpu offload" of
the OT, which impairs T/S performance. The "AMD" version has the OT in f16 [as opposed to bf16]. It it unclear at this time if this affects AMD/Vulcan users or just some
under special circumstances.
Additionally, quants by team "Mradermacher" with work on AMD/All machines and will be slightly faster/smaller than the NEO MAX quants due
to OT being default size, and MTP tensors also being default size. Click on "Quantized" in the model tree (upper left) to access these quants.
SPEED:
- On Q4_K_S (4bit) quant, regular GGUFs are about 75 t/s, whereas MTP GGUFs (acceptance at 60%, 2 tokens) can exceed 90 T/S. (5090, Windows 11, testing in LMStudio)
- Speeds will vary depending on GPU(s), AI app, O/S (Linux/Mac will generally be faster) and hardware.
- "MTP" quants speeds will vary ; for creative/complex and/or temps over 1 use regular GGUFs for better performance.
I suggest you download at least one of each - regular and MTP gguf(s) - and test them for your use case(s).
If you get "token acceptance" (predict 2 tokens) with MTP quant(s) BELOW 50% (this means regular quants will run faster), then regular GGUF(s) will actually perform better - ie faster.
MTP quant(s) can in some cases run faster as the token window fills up and/or in multi turn chats.
Note there is NO other diffence between the quants type besides speed: both will do the same job.
<B>ADDITIONAL Quant types / OTHER GGUF Quants:</b>
1. There are "Dflash", "NINFER", "AUTOROUND", "vLLM" "MLX", "NVFP4" and others. Look in the community tab; labeled with "VERSION:"
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncen