DIFF.BLOG
New Following Discover Jobs
More
Top Writers Suggest a blog Upvotes plugin
Report bug Contact About
Sign up
Topics
Follow your own topics →
Menu
New Following Discover Jobs Top Writers
More
Suggest a blog Upvotes plugin Report bug Contact About
Sign up
The home for great developer writing.
We surface the best developer writing from thousands of independent blogs, updated daily.
Join Diff.blog
TOPICS

Accelerating LLMs on Debian 13: Setting up CUDA for llama.cpp

127 · özkan pakdil · March 20, 2026, 2:04 a.m.
CUDA large language models Debian 13 GPGPU
Summary
This blog post provides a detailed guide on setting up NVIDIA CUDA on Debian 13 to enhance the performance of Large Language Models (LLMs) using llama.cpp, sharing personal experiences and insights from the author's journey.
Read full post on ozkanpakdil.github.io →
MORE POSTS LIKE THIS
Accelerating LLMs on Debian 13: Setting up Vulkan for llama.cpp
özkan pakdil · Mar 22, 2026
large language models Vulkan
Atom #hdxty3k
Brandur Leach · Aug 11, 2026
Performance Optimization large language models
Getting Fortran running on GPU's natively
Fortran Lang Discourse · Jul 28, 2026
Fortran GPU programming
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090
gpjt · Jul 24, 2026
benchmarking llm
llama.cpp vs. vLLM: Choosing the right local LLM inference engine
Red Hat · Jun 15, 2026
AI Inference Engines llama.cpp
High Performance Distributed Inference with Ray Serve LLM
Anyscale · Jun 18, 2026
distributed inference Ray Serve
Discover more posts →
AUTHOR
RECENT POSTS FROM THE AUTHOR
Choose how you want to continue.
Continue with GitHub Continue with Google