Unsloth 发布免费 notebook,教用户把 Qwen3.5-4B 训练成生成决策而非文本的决策模型,本地仅需 8GB VRAM。内容涵盖数据准备(state、questions、gold answers)、训练与服务;配套指南见 unsloth.ai/docs。其引用帖提到用 Clef head 加 LoRA(r=64)微调一 epoch,把 Qwen3.5 0.8B 在 3 个决策基准上的总体准确率从 20.7% 提到 74.3%,仅需 4GB VRAM。Notebook:https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_5_(4B)-Decision.ipynb
You can now train your own Decision model with our free notebook! 💡
Qwen3.5-4B will generate decisions instead of text on just 8GB VRAM locally.
Learn to data prep (state, questions, gold answers), train, serve.
Notebook: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_5_(4B)-Decision.ipynb
Guide: https://unsloth.ai/docs/basics/train-your-own-decision-model-with-unsloth
You can now train your own Decision model like Jev locally! We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM. Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo. We fine-tuned with a Clef head using Unsloth and LoRA (r=64) for one epoch, increasing downstream accuracy from 30–37% to 78%. GitHub: https://github.com/unslothai/unsloth Guide and Notebooks: https://unsloth.ai/docs/basics/train-your-own-decision-model-with-unsloth在 X 查看被引用的帖子
来源:Unsloth AI · x.com