Topics
Follow your own topics →
DIFF.BLOG
New Following Discover Jobs
More
Top Writers Suggest a blog Upvotes plugin
Report bug Contact About
Sign up
Menu
New Following Discover Jobs Top Writers
More
Suggest a blog Upvotes plugin Report bug Contact About
Sign up
The home for great developer writing.
We surface the best developer writing from thousands of independent blogs, updated daily.
Join Diff.blog
TOPICS

Accelerating Apache Parquet Scans on Apache Spark with GPUs

15 · NVIDIA Corporation · April 3, 2025, 4:37 p.m.
Data Center / Cloud Data Science Development & Optimization accelerated data analytics Apache Parquet apache spark GPU acceleration data-processing
Summary
This blog post discusses the optimization of Apache Parquet scans on Apache Spark using GPU acceleration. It explores the advantages of utilizing GPUs for processing data stored in Parquet format, particularly in terms of performance improvement in handling large data sets typical in enterprise environments.
Read full post on developer.nvidia.com →
MORE POSTS LIKE THIS
Accelerating JSON Processing on Apache Spark with GPUs
NVIDIA Corporation · Jan 29, 2025
Data Science Development & Optimization
A Fast Path for Fixed-Length Lists in Parquet
gunnarmorling · Jul 22, 2026
Apache Parquet data storage
Polars for Machine Learning: Zero-Copy to PyTorch and XGBoost
Ahmed Nabil · Jul 15, 2026
Data Science 2026
Introducing Apache Spark 4.2
Mooncake · Jul 16, 2026
announcements engineering
Do concurrent selective accelerator? host, device execution
Fortran Lang Discourse · Jul 6, 2026
Fortran GPU acceleration
Processing Data Larger Than RAM: The Polars Streaming Engine (sink_parquet)
Ahmed Nabil · Jun 26, 2026
Data Science 2026
Discover more posts →
AUTHOR
RECENT POSTS FROM THE AUTHOR
Choose how you want to continue.
Continue with GitHub Continue with Google