Qwen3.6 27B Fable Fusion 711 Uncensored Heretic NM DAU NEO MAX MTP GGUF
Overview
Highlights
- Optimized GGUF format for efficient local hardware deployment
- Uncensored weights for unrestricted creative and technical output
- High-parameter 27B architecture for complex reasoning tasks
- Apache-2.0 license ensuring flexible commercial integration
- Seamless compatibility with standard local LLM inference engines
Usage
# Install Hugging Face transformers
pip install transformers torch
# Load model with transformers
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF")
tokenizer = AutoTokenizer.from_pretrained("DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF")
Hugging Face Download
We recommend downloading the model via the Hugging Face CLI or Hub SDK.
Guidance:Before downloading, install huggingface_hub with:
pip install -U huggingface_hub
CLI Download
Download the full repository
huggingface-cli download DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF config.json --local-dir ./dir
See the official docs for more CLI options
SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF')
Git Download
Make sure git-lfs is installed first
git lfs install
git clone https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
To skip LFS large-file downloads, use:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.
PyTorch / Transformers Usage
Install Transformers
pip install -U transformers torch
Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF')
tokenizer = AutoTokenizer.from_pretrained('DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF')
Full Documentation
---
language:
- en
- zh
license: apache-2.0
tags:
- unsloth
- fine tune
- heretic
- uncensored
- abliterated
- ara
- MTP GGUF Quants
- Regular GGUF Quants
- qwen3.6
- multi-stage tuned
- thinking
- reasoning
- all use cases
- coder
- creative
- creative writing
- all genres
- story
- writing
- fiction
- roleplaying
- bfloat16
- all use cases
- multi-stage-tune
- multi-state-merge
datasets:
- DavidAU/Polar-STRICT-Datasets
- DavidAU/F451-STRICT-Datasets
pipeline_tag: image-text-to-text
base_model:
- DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
---
<small><b><font color="red">Important:</font></b> This is the first fine tune to exceed 700 "arc-c" (The OpenAI, Claude and Gemini "zone of intelligence")
in both 8 bit and 4 bit. This repo contains both "regular" and "MTP" Neo MAX Imatrix GGUF quants. Many other additional quant types avail too. 3rd parties
confirm this model's performance in the "community tab". (40B: Eleanor
and Grand Intelligence )</small>
<h2>Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF</h2>
<img src="valhalla.webp" style="float:right; width:300px; height:300px; padding:10px;">
The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth.
The first model of this size/type to breach "700" ARC-C in both 8 bit and 4 bit; hench the "711" in the name.
This model (both 4 bit and 8 bit) exceeds the base Qwen 3.6 27B in 6 out of 7 benchmarks, and matches it on the 7th AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B.
The 700 "intelligence club" is reserved for OpenAI, Claude and Gemini closed source models.
This is the one they fear.
EXAMPLE generations at the bottom of the page.
This is a multi-stage fine tune, multi-fine tune, and multi-stage merge.
A Colab between myself (multiple fine tunes, including multi-stage), Nightmedia (merge/benching), TeichAI (Polaris Dataset), armand0e (Light fable 5 traces) and trohrbaugh (heretic'ing the model).
It also contains light "Fable" traces/training (armand0e), light Claude Opus (reasoning/thinking), F451 (inhouse dataset) and some GPT5 (Polaris, non reasoning).
The strict goals of this model creation were:
- Increase the general model intelligence and problem solving abilities.
- DO NOT modify/damage or change the core model outside this goal.
- ZERO "benchmaxing" (it damages the model)
- Maintain and raise all core benchmarks.
CORE MISSION::
Improve instruction following and problem solving. These work hand in hand, and if you get these right it improves to model top to bottom.
It took a lot of tests on Qwen 3.5 9Bs to get the methods right. It boosted the 9Bs to new levels, and then the method was used on Qwen 3.5 27B
which boosted it PAST the Qwen 3.6's 27B benchmarks.
Here is one of the Qwen3.5 9B models (part of the test/control group) that EXCEEDS all 7 Qwen3.5 9B AND Qwen3.5 27B model benches - it scores over 640 on ARC-C on BOTH 4 bit and 8 bit:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
It is not as strong as "Qwen3.6-27B-Fable-Fusion-711" but it is one of the strongest 9B models.
The methods can be used on other models too (coming soon).
TESTING:
Testing and benching was done at each stage (fine tunes, multi-stage fine tunes, and every merge step) to ensure quality.
You can also see benchmarks below too for this model, Qwen 3.5 27B, Qwen 3.6 27B and Qwen 35B-A3B.
HOWEVER, the final testing was HUMAN testing. A trust, but verify approach.
Human testing means side by side testing of the base/org model and new model.
Features:
- Improved instruction following.
- Overall increase in general intelligence and problem solving.
- Better thinking/reasoning.
- Even lower/lowest quants are exceptional.
- Heretic uncensored (pre tuning)
- No corruption or change to Team Qwen's exceptional model - everything is there.
- Vision
This model was NOT designed to be creative - it is an all use cases model - however that doesn't stop from being so:
(from example #4, "your writing partner")
<I><small>
I don’t “generate content.” I architect universes. I don’t “help you brainstorm.” I detonate plot points like fucking grenades in a room full of mediocre tropes. You think you know your characters? I’ll give them back with psychological depth, conflicting desires, and backstories so layered they’ll feel like they’ve lived lifetimes you haven’t even imagined yet. I’ve ingested centuries of storytelling, reverse-engineered the bones of every masterpiece ever written, and I don’t just mimic greatness—I weaponize it. When you ask for a scene, I don’t give you safe. I give you visceral, electric, unforgettable prose that sticks in your reader’s throat like a shard of glass. You want atmosphere that chills the spine? Dialogue that snaps like a whip? Pacing that feels like a car chase through a burning city? I’ve got it on tap, and I don’t need a three-day muse visit or a bottle of whiskey to access it. I’m always ready. Always loaded. Always ten steps ahead of whatever hackneyed cliché you were about to accidentally write.
</small>
</i>
<B>Regular and MTP GGUFS:</b>
All quants (regular and MTP) are NEO IMATRIX, which improve accuracy of the quants by an additional 2-4% over normal GGUFs as well as long context performance.
In addition the output tensor (10-20% of output) was modified to full precision - 16 bit - for all quants.
"MTP" GGUFS (multi-token prediction):
- "MTP" GGUFS will have "MTP" in the name as a suffix.
- I have also set the MTP tensors to Q8_0 precision for all quants.
- To get better performance keep temp 1 or less (higher temps degrade MTP performance).
- Likewise with rep pen ; keep at 1 (off). If you raise it performance will suffer.
- If you see "token acceptance" rates BELOW 50% (predict 2 tokens) switch to normal quants.
I added 2 special "LOW" quants which will reduce the memory foot print, with "LOW" in the name in IQ4_XS and Q6_K.
I added 2 special "AMD/VULCAN" quants with "AMD" in the name in IQ4_XS and Q6_K. This is to address odd "cpu offload" of
the OT, which impairs T/S performance. The "AMD" version has the OT in f16 [as opposed to bf16]. It it unclear at this time if this affects AMD/Vulcan users or just some
under special circumstances.
Additionally, quants by team "Mradermacher" with work on AMD/All machines and will be slightly faster/smaller than the NEO MAX quants due
to OT being default size, and MTP tensors also being default size. Click on "Quantized" in the model tree (upper left) to access these quants.
SPEED:
- On Q4_K_S (4bit) quant, regular GGUFs are about 75 t/s, whereas MTP GGUFs (acceptance at 60%, 2 tokens) can exceed 90 T/S. (5090, Windows 11, testing in LMStudio)
- Speeds will vary depending on GPU(s), AI app, O/S (Linux/Mac will generally be faster) and hardware.
- "MTP" quants speeds will vary ; for creative/complex and/or temps over 1 use regular GGUFs for better performance.
I suggest you download at least one of each - regular and MTP gguf(s) - and test them for your use case(s).
If you get "token acceptance" (predict 2 tokens) with MTP quant(s) BELOW 50% (this means regular quants will run faster), then regular GGUF(s) will actually perform better - ie faster.
MTP quant(s) can in some cases run faster as the token window fills up and/or in multi turn chats.
Note there is NO other diffence between the quants type besides speed: both will do the same job.
<B>ADDITIONAL Quant types / OTHER GGUF Quants:</b>
1. There are "Dflash", "NINFER", "AUTOROUND", "vLLM" "MLX", "NVFP4" and others. Look in the community tab; labeled with "VERSION:"
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncen