Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090

· · July 24, 2026, 10:47 p.m.
Summary
This blog post provides a detailed benchmarking analysis of the LLM Qwen 3.6 35B MoE model using an RTX 3090 GPU, discussing performance implications of quantization, context lengths, and the efficiency of various configurations. The author shares personal insights and experiences in optimizing the model for both Vulkan and CUDA, highlighting the intricacies of using Mixture of Experts and memory management strategies.
AUTHOR
Sponsored
Zulip logo Zulip
Organized team chat for people who take work seriously. Topic-based threading keeps conversations focused.
Try Zulip
Become a sponsor →
BLOG POST FEATURED ON

Add this plugin to your blog