Topic: Local LLM Infrastructure & Inference Engines
Explore our technical articles and implementation guides tagged with Local LLM Infrastructure & Inference Engines.
01
Tutorials
Running Qwen 3.5 27B on Consumer GPUs: VRAM Setup Guide
Hardware setup tutorial to run Qwen 3.5 27B locally on consumer GPUs. Includes INT4 GGUF quantization, FlashAttention-2 compilation, and multi-GPU tensor parallelism.
02
Tutorials
How to Run DeepSeek R1 Locally with Ollama: Complete Guide
Complete developer setup guide to run DeepSeek R1 locally using Ollama. Includes VRAM memory formulas, GGUF quantization comparisons, CLI integration, and ChromaDB Python code.
03
Guides
Complete Guide to Ollama Model Quantization Formats (Q4_K_M vs Q8_0 vs EXL2)
Technical comparison guide evaluating Ollama GGUF quantization formats. Compares Q4_K_M, Q5_K_M, Q8_0, and EXL2 across perplexity, VRAM savings, and decode speed.
That's all, Love 🧡