Unsloth Docker: Train and Run 500+ AI Models Locally With One Command

2026-09-20
The Unsloth Docker image packs PyTorch, a GUI, and notebooks into one command. Here's how Unsloth local AI lets you run and train AI models locally on your own GPU.
Setting up local AI used to be a weekend project. You matched your CUDA version to PyTorch, matched PyTorch to your driver, then fought bitsandbytes and xformers until something finally imported. Plenty of people buy a GPU full of enthusiasm and lose it to dependency errors before a single model runs.
Unsloth's Docker image is an attempt to delete that whole ritual. On September 17, 2026, the team posted that you can now train and run 500+ models locally with one command, no setup required, on NVIDIA and AMD hardware. Why does that matter to anyone who isn't a machine learning engineer? Because the setup was the wall, not the training. No launch event, no demo reel. Just a command and a link to the docs.
What Unsloth Docker actually is
Unsloth Docker is the official container image unsloth/unsloth, published on Docker Hub. Everything the project needs to fine-tune and run models ships inside it: PyTorch, Unsloth itself, unsloth-zoo, bitsandbytes, TRL, PEFT, and xformers. The image is roughly 9.2 GB.
You pull it and run it. That's the entire install story. The container was updated in September 2026 with the Unsloth Studio interface and AMD support, which is what turned a training environment into something you can also use without touching code.
There are two images to know about. The NVIDIA one is unsloth/unsloth. AMD gets a separate unsloth/unsloth-rocm, and it connects differently, reaching the GPU through the kernel driver's device nodes rather than a container toolkit, so the launch command looks a little different.
How Unsloth local AI works without the setup
The reason this saves time is simple. Your CUDA version, your PyTorch version, and your driver have to agree with each other. In a container, they already do, because the image was built with a known-good combination and tested before it shipped.
Unsloth shares the same model cache as its notebooks and scripts, so you don't re-download a model you already have. The three volume flags in the standard command keep your files, your models, and your Unsloth Studio data on your machine rather than inside the container. Delete the container and your work survives.
The mechanism under the hood matters less to most users than the result: you stop debugging your environment and start running models.
The one command that starts Unsloth Studio
Here's the run command on Linux or WSL with an NVIDIA GPU:
docker run -d --name unsloth --gpus all --ipc=host \
--ulimit memlock=-1 --ulimit stack=67108864 \
-p 8000:8000 -p 8888:8888 \
-e JUPYTER_PASSWORD="mypassword" \
-v "$PWD":/workspace/host \
-v "$HOME/.cache/huggingface":/workspace/.cache/huggingface \
-v unsloth-studio:/opt/unsloth-studio \
unsloth/unsloth
Run it from your project folder, because that folder becomes /workspace/host inside the container. Port 8000 is Unsloth Studio and port 8888 is JupyterLab. The only hard requirement is an NVIDIA driver of 570.26 or newer.
On Windows, the same command works in PowerShell with backticks instead of backslashes, and Docker Desktop with the WSL 2 engine is enough. You don't need a separate NVIDIA Container Toolkit install on Windows.
Unsloth GUI and the 500+ models claim
The Unsloth GUI is called Unsloth Studio, and it's a browser-based interface that lets you run and fine-tune models without writing code. In the Docker image it's pre-installed, and you reach it at http://localhost:8000, signing in as unsloth.
Studio wraps the training pipeline in a clean UI: model loading, dataset formatting, hyperparameter configuration, and live training monitoring. If you'd rather stay in code, JupyterLab is running on the other port with example notebooks ready to go.
The 500+ number is the headline Unsloth put on the release. The project supports a wide spread of model families, including Qwen, GLM, Kimi, DeepSeek, Gemma, Llama, gpt-oss, MiniMax, and Mistral, along with GGUF, MLX, diffusion, embedding, and audio models. The docs don't list a fixed maximum. What matters is GPU memory. A model stays loaded in VRAM until you unload it, and if you hit an out-of-memory error, you unload the model in Unsloth or stop the notebook kernel.
Training models locally with Unsloth
Running a model is the easy half. Training is where Unsloth built its reputation, and the Docker image carries the full fine-tuning stack.
The project claims fine-tuning that's 2× faster with 70% less VRAM and no accuracy loss. The container supports LoRA, QLoRA, full fine-tuning, pretraining, reinforcement learning, GRPO, DPO, and FP8. When you're done, you can export to GGUF, NVFP4, FP8, and other formats, or push the model straight into a deployment target.
For datasets, Unsloth Studio includes a feature called Data Recipes that builds training data from PDFs, CSVs, DOCX files, and other sources. That's the part that usually stops beginners, and putting it behind a form in a web UI lowers the wall considerably.
There's a catch worth naming. Unsloth Studio and JupyterLab share the same GPU, so you can't run a training job and a chat session on a large model at the same time without thinking about memory. On a single consumer card, that's a real constraint.
Practical examples of running AI models locally
The commands get more interesting once you connect the container to other tools.
If you use Claude Code or OpenAI Codex, Unsloth Start connects a local model to those agents with a line like unsloth start claude --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL. You keep the agent you like and point it at a model running on your own machine. The container also serves models through an OpenAI-compatible API, so anything that speaks that format can talk to it.
For remote access, unsloth studio --secure spins up a free Cloudflare HTTPS link, which means you can reach your local models from your phone. On a home network, LAN access is a settings toggle. Server-side tools are on by default, so the project's own docs warn you to keep the password safe or disable them when you expose the instance. That's a security call, not a formality: an exposed Studio can run code and read files on the host, so treat the password like a house key.
The image tags give you a choice too. latest includes Studio, JupyterLab, and notebooks. core is the training stack plus JupyterLab with no Studio, meant for scripts and CI where you want a slimmer pull.
Benefits and limits of Unsloth Docker
The benefits are concrete. You skip environment setup entirely. You get a GUI and notebooks in the same image. Your models and training data persist across container rebuilds. The image covers NVIDIA, AMD, Intel GPUs, CPUs, and the Vulkan backend, and it works on Windows, Linux, WSL, and macOS.
The limits deserve equal billing. The AMD image carries the training stack only, so there's no Studio or JupyterLab in it, and it needs native Linux because WSL exposes /dev/dxg rather than the /dev/kfd that ROCm wants. The core tag drops the GUI. And the 500+ model figure describes what the project supports, not what your card can hold at once. A 9.2 GB image is also a real download on a slow connection.
What Unsloth Docker means for you
Unsloth Docker turns local AI from an installation problem into a command. That's a smaller claim than it sounds and a bigger deal than it looks, because the installation problem is exactly what stops most people from ever training a model.
If you have a supported GPU and you've been putting off local fine-tuning, this is the cheapest on-ramp yet. Start with the latest tag, run a small model to confirm the GPU passthrough works, and only then point it at something big. The setup you don't have to do is the time you get to spend on the model.