Awesome-GPU
Awesome resources for GPUs
https://github.com/Jokeren/Awesome-GPU
Last synced: 18 days ago
JSON representation
-
Algorithms
-
BLAS
- DEVELOPING CUDA KERNELS TO PUSH TENSOR CORES TO THE ABSOLUTE LIMIT ON NVIDIA A100
- Demystifying Tensor Cores to Optimize Half-Precision Matrix Multiply
- A Coordinated Tiling and Batching Framework for Efficient GEMM on GPU
- CUTLASS: CUDA TEMPLATE LIBRARY FOR DENSE LINEAR ALGEBRA AT ALL LEVELS AND SCALES
-
Scans
-
Stencils
-
-
Applications
-
Deep Learning
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUs
- Sparse GPU Kernels for Deep Learning
- SuperNeurons: Dynamic GPU Memory Management for Training Deep Neural Networks
- Towards Pervasive and User Satisfactory CNN across GPU Microarchitectures
- Understanding and bridging the gaps in current GNN performance optimizations
- E.T.: re-thinking self-attention for transformer models on GPUs
- Towards Pervasive and User Satisfactory CNN across GPU Microarchitectures
-
-
Architecture
-
Cache
-
Memory
- Improving Inter-kernel Data Reuse With CTA-Page Coordination in GPGPU
- Umpire: Application-Focused Management and Coordination of Complex Hierarchical Memory
- Reducing GPU Offload Latency via Fine-Grained CPU-GPU Synchronization
- In-Depth Analyses of Unified Virtual Memory System for GPU Accelerated Computing
-
Parallelism
- Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls
- Controlled Kernel Launch for Dynamic Parallelism in GPUs
- LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs
- Virtual Thread Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit
- Understanding Latency Hiding on GPUs
- Controlled Kernel Launch for Dynamic Parallelism in GPUs
- LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs
- Virtual Thread Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit
- COOPERATIVE GROUPS
-
Resources Management
- Dynamic GPGPU Power Management Using Adaptive Model Predictive Control
- Transparent Offloading and Mapping (TOM): Enabling Programmer-Transparent Near-Data Processing in GPU Systems
- Reducing Energy in GPGPUs through Approximate Trivial Bypassing
- Locality-Aware CTA Clustering for Modern GPUs
- Dynamic Resource Management for Efficient Utilization of Multitasking GPUs
- Locality-Aware CTA Clustering for Modern GPUs
- Dynamic Resource Management for Efficient Utilization of Multitasking GPUs
- Dynamic GPGPU Power Management Using Adaptive Model Predictive Control
- Transparent Offloading and Mapping (TOM): Enabling Programmer-Transparent Near-Data Processing in GPU Systems
-
White Papers
- NVIDIA TURING GPU ARCHITECTURE
- NVIDIA TESLA V100
- NVIDIA TESLA P100
- NVIDIA’s Next Generation CUDA Compute Architecture: Kepler
- NVIDIA’s Next Generation CUDA Compute Architecture: Fermi
- INTRODUCING AMD CDNA 2 ARCHITECTURE
- INTRODUCING AMD CDNA ARCHITECTURE
- NVIDIA A100 Tensor Core GPU Architecture
- NVIDIA H100 Tensor Core GPU Architecture
- NVIDIA TURING GPU ARCHITECTURE
- NVIDIA TESLA V100
- NVIDIA TESLA P100
- NVIDIA’s Next Generation CUDA Compute Architecture: Kepler
-
-
Code Generation
-
Binaries
-
Compilers
- Generating GPU Compiler Heuristics using Reinforcement Learning
- Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation
- Implementing implicit OpenMP data sharing on GPUs
- gpucc: An Open-Source GPGPU Compiler
- Offloading Support for OpenMP in Clang and LLVM
- Performance Analysis of OpenMP on a GPU using a CORAL Proxy Application
- Integrating GPU Support for OpenMP Offloading Directives into Clang
- Coordinating GPU Threads for OpenMP 4.0 in LLVM
-
Profile Guided Optimization
-
Programming Models
-
-
Runtime
-
Tools
-
Benchmarking
-
Models
- Instruction Roofline An insightful visual performance model for GPUs
- Performance Tuning of Scientific Codes with the Roofline Model
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Fundamental_Optimizations
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- Performance Analysis and Tuning for General Purpose Graphics Processing Units (GPGPU)
- VOLTA Architecture and performance optimization
-
Profilers
- Exposing Hidden Performance Opportunities in High Performance GPU Applications
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Lynx: A dynamic instrumentation system for data-parallel applications on GPGPU architectures
- **Vampir|Score-P**
- **TAU**
- **Open|SpeedShop**
- **HPCToolkit**
- **NVIDIA Nsight Systems**
- **NVIDIA Nsight Compute**
- **SASSI**
- **NVBit**
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- GPU Code Optimization using Abstract Kernel Emulation and Sensitivity Analysis
- CUDAAdvisor: LLVM-based runtime profiling for modern GPUs
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Effective sampling-driven performance tools for GPU-accelerated supercomputers
- Parallel Performance Measurement of Heterogeneous Parallel Systems with GPUs
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Monitoring Heterogeneous Applications with the OpenMP Tools Interface
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- Identifying Optimization Opportunities Within Kernel Execution in GPU Codes
- **Ingero** - eBPF-based GPU causal observability agent. Traces CUDA Runtime/Driver APIs via uprobes and host kernel events to build causal chains explaining GPU latency. <2% overhead, production-safe.
- **PAPI**
- **Vampir|Score-P**
- **Allinea MAP**
- **HPCToolkit**
-
Simulators
-