Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090

381 · · July 24, 2026, 10:47 p.m.
Summary
This blog post provides a detailed benchmarking analysis of the LLM Qwen 3.6 35B MoE model using an RTX 3090 GPU, discussing performance implications of quantization, context lengths, and the efficiency of various configurations. The author shares personal insights and experiences in optimizing the model for both Vulkan and CUDA, highlighting the intricacies of using Mixture of Experts and memory management strategies.