Quantizing DeepSeek R1-0528-Qwen3-8B
In May 2025, DeepSeek released an impressive update to their R1 model: the DeepSeek-R1-0528-Qwen3-8B. The full baseline currently ranks in the Top-10 on the LLM leaderboard.
I set out to run this on a low-cost OrangePi 5 — a single board computer with just 8GB of RAM and an ARM SoC.
Background
Earlier this year, I used the original DeepSeek-R1 model for a Stanford Continuing Studies class and documented the setup in this GitHub repo.
Challenge: Model Requires Quantization
The 8B parameter model with bfloat16 precision weighs in at ~16GB, far beyond the 8GB RAM capacity of the OrangePi 5. Quantization is necessary to make it run.
1. Baseline Colab Test (unquantized)
I first verified the bfloat16 8B model on Google Colab using a 40GB A100 GPU with a classic reasoning riddle. The model responded correctly with “heroine” — ideal behavior.
Riddle: I’m a word in the English language. My first two letters signify a male. My first three letters signify a female. My first four letters signify a great man. And my entire meaning signifies a great woman. What am I?
2. Ollama’s Quantized Version
Ollama provides a quantized version of DeepSeek R1-8B. While it ran on the OrangePi, the model got stuck in a loop and failed to converge on “heroine”:
“I'm tired of these riddles... I’m overthinking…”
This prompted me to try my own quantized build.
3. My Quantized Version (via llama.cpp + Ollama)
I used llama.cpp to quantize the model to 4-bit (Q4_K_M), reducing the file
size to ~4.7GB. I published it on
HuggingFace.
ollama run hf.co/guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M --verbose
This run consumed about 76% of the 8GB RAM. It showed better convergence than the built-in Ollama model, repeatedly returning to the word “heroine”, though it still looped somewhat:
"So I think HEROINE is the word. Final Answer... Another thought..."
I manually stopped the run.
4. llama-run with My Quantized Model
To get better control over inference, I switched to using llama-run from
llama.cpp, built for ARM. The command --ngl enables layer offload which
prioritizes the fast cores on the ARM SoC.
Here’s the command I used:
./build/bin/llama-run \
./deepseek-r1-0528-qwen3-8b-q4_k_m.gguf \
-c 10000 \
--ngl 99 \
--prompt "I’m a word in the English language. My first two letters signify a male. My first three letters signify a female. My first four letters signify a great man. And my entire meaning signifies a great woman. What am I?"
This run completed cleanly with less looping and produced the expected result:
So the word is heroine.
</think>
The word that satisfies all the conditions is **heroine**.
- **First two letters: "HE"** – male pronoun
- **First three letters: "HER"** – female pronoun
- **First four letters: "HERO"** – great man
- **Entire word: "heroine"** – great woman
This word works because it leverages common pronouns and titles in an English language puzzle.
Success!
Discussion
This single riddle prompt illustrates how the 8B model behavior can vary significantly depending on:
- Quantization precision
- Inference framework
- Runtime configuration
| Framework (Precision) | Result | Notes |
|---|---|---|
| Colab (bfloat16) | Correct | Fast on A100 GPU and accurate |
| Ollama (Q4_K_M) | Loops | Some convergence, but repetitive |
| llama-run (Q4_K_M) | Correct | Clean result |
Framework and quantization compatibility matter — especially on resource-constrained devices like the OrangePi 5.
Update
Unsloth provide a quantized version of the same DeepSeek model. It performs well with the riddle example.
ollama run hf.co/unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF:Q4_K_M
Inspecting ollama downloads the native model has a different size.
orangepi@orangepi5-desktop:~$ ollama list
NAME ID SIZE MODIFIED
hf.co/unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF:Q4_K_M ecc092d5e10a 5.0 GB 47 hours ago
hf.co/guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M 0417db1664a3 5.0 GB 2 days ago
deepseek-r1:8b 6995872bfe4c 5.2 GB 2 days ago
Next Steps
- Check model temperature and test other Ollama hyperparameters
- Try a broader test set of riddles and reasoning prompts
Want to try it yourself? Grab the quantized model on HuggingFace: DeepSeek-R1-0528-Qwen3-8B_Q4_K_M