NVCC
nvcc predefines the following macros:
- NVCC (1): Defined when compiling C/C++/CUDA source files.
- CUDACC (1): Defined when compiling CUDA source files.
compile a single cuda file
1 | nvcc -gencode arch=compute_70,code=sm_70 -g -G kernel.cu -o kernel |
memory
- mapped memory: https://leimao.github.io/blog/CUDA-Zero-Copy-Mapped-Memory/
- Increase Performance with Vectorized Memory Access
https://developer.nvidia.com/blog/cuda-pro-tip-increase-performance-with-vectorized-memory-access/
data transfer
gdr
https://developer.nvidia.com/gdrcopy
pinned pageable memory
https://developer.nvidia.com/blog/how-optimize-data-transfers-cuda-cc/
nvlink command
1 | nvidia-smi nvlink -h |
Showing NVLINK Status For Different GPUs
To show active NVLINK Connections, you must specify GPU index via -i
1 | nvidia-smi nvlink --status -i 0 |
example output
1 | GPU 0: NVIDIA H100 PCIe (UUID: GPU-a1bf2ff6-a98d-edac-a422-0bd80fbd9724) |
Display & Explore NVLINK Capabilities Per Link
Allows you to query to ensure each link associated with the GPU Index (specified by -i #) has specific capabilities related to P2P, System Memory, P2P Atomics, SLI.
1 | nvidia-smi nvlink --capabilities -i 1 |
example output
1 | Link 0, P2P is supported: true |
NVLink Usage Counters
nvidia-smi nvlink -g N -i N allows you to view the data being traversed on the different NVLink Link.
1 | nvidia-smi nvlink -g 0 -i 0 |
benchmarking p2p mem_copy
code for testing p2p memCopy
1 | // P2P Test by Greg Gutmann |
some results
- on H100
Unidirectional Bandwidth: 253.631489 (GB/s) - on 4090
Unidirectional Bandwidth: 25 (GB/s)
profiling
nsight compute: https://leimao.github.io/blog/Docker-Nsight-Compute/
cuda with docker: https://leimao.github.io/blog/NVIDIA-Docker-CUDA-Compatibility/