GPU Hardware: Difference between revisions
Jump to navigation
Jump to search
No edit summary |
No edit summary |
||
Line 40: | Line 40: | ||
</pre> | </pre> | ||
===GPU Hardware on | ===GPU Hardware on Sapelo2=== | ||
* | * 14 NVIDIA Tesla (Kepler) K20Xm GPU cards (32 x 2688 = 86016 GPU cores). These cards are installed on 2 hosts each of which has dual 6-core Intel Xeon CPUs and 96GB of RAM; there are 7 GPU cards per host. The output of NVIDIA SDK deviceQuery for one such GPU card is | ||
<pre class="gscript"> | <pre class="gscript"> | ||
Device 0: "Tesla K20Xm" | Device 0: "Tesla K20Xm" | ||
CUDA Driver Version / Runtime Version | CUDA Driver Version / Runtime Version 9.0 / 9.0 | ||
CUDA Capability Major/Minor version number: 3.5 | CUDA Capability Major/Minor version number: 3.5 | ||
Total amount of global memory: | Total amount of global memory: 5700 MBytes (5976424448 bytes) | ||
(14) Multiprocessors | (14) Multiprocessors, (192) CUDA Cores/MP: 2688 CUDA Cores | ||
GPU Clock rate: | GPU Max Clock rate: 732 MHz (0.73 GHz) | ||
Memory Clock rate: 2600 Mhz | Memory Clock rate: 2600 Mhz | ||
Memory Bus Width: 384-bit | Memory Bus Width: 384-bit | ||
L2 Cache Size: 1572864 bytes | L2 Cache Size: 1572864 bytes | ||
Maximum Texture Dimension Size (x,y,z) 1D=(65536), 2D=(65536, 65536), 3D=(4096, 4096, 4096) | |||
Maximum Layered 1D Texture Size, (num) layers 1D=(16384), 2048 layers | |||
Maximum Layered 2D Texture Size, (num) layers 2D=(16384, 16384), 2048 layers | |||
Total amount of constant memory: 65536 bytes | Total amount of constant memory: 65536 bytes | ||
Total amount of shared memory per block: 49152 bytes | Total amount of shared memory per block: 49152 bytes | ||
Line 62: | Line 63: | ||
Maximum number of threads per multiprocessor: 2048 | Maximum number of threads per multiprocessor: 2048 | ||
Maximum number of threads per block: 1024 | Maximum number of threads per block: 1024 | ||
Max dimension size of a thread block (x,y,z): (1024, 1024, 64) | |||
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535) | |||
Maximum memory pitch: 2147483647 bytes | Maximum memory pitch: 2147483647 bytes | ||
Texture alignment: 512 bytes | Texture alignment: 512 bytes | ||
Line 73: | Line 74: | ||
Device has ECC support: Enabled | Device has ECC support: Enabled | ||
Device supports Unified Addressing (UVA): Yes | Device supports Unified Addressing (UVA): Yes | ||
Device PCI Bus ID / | Supports Cooperative Kernel Launch: No | ||
Supports MultiDevice Co-op Kernel Launch: No | |||
Device PCI Domain ID / Bus ID / location ID: 0 / 12 / 0 | |||
Compute Mode: | |||
< Exclusive Process (many threads in one process is able to use ::cudaSetDevice() with this device) > | |||
</pre> | </pre> | ||
* | * | ||
Revision as of 08:09, 15 June 2018
GPU Hardware on Sapelo
- 16 NVIDIA K40m GPU cards. These cards are installed on 2 hosts, each of which has dual 8-core Intel Xeon CPUs and 128GB of RAM; there are 8 GPU cards per host. The output of NVIDIA SDK deviceQuery for one such GPU card is
Device 0: "Tesla K40m" CUDA Driver Version / Runtime Version 6.5 / 6.5 CUDA Capability Major/Minor version number: 3.5 Total amount of global memory: 11520 MBytes (12079136768 bytes) (15) Multiprocessors, (192) CUDA Cores/MP: 2880 CUDA Cores GPU Clock rate: 745 MHz (0.75 GHz) Memory Clock rate: 3004 Mhz Memory Bus Width: 384-bit L2 Cache Size: 1572864 bytes Maximum Texture Dimension Size (x,y,z) 1D=(65536), 2D=(65536, 65536), 3D=(4096, 4096, 4096) Maximum Layered 1D Texture Size, (num) layers 1D=(16384), 2048 layers Maximum Layered 2D Texture Size, (num) layers 2D=(16384, 16384), 2048 layers Total amount of constant memory: 65536 bytes Total amount of shared memory per block: 49152 bytes Total number of registers available per block: 65536 Warp size: 32 Maximum number of threads per multiprocessor: 2048 Maximum number of threads per block: 1024 Max dimension size of a thread block (x,y,z): (1024, 1024, 64) Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535) Maximum memory pitch: 2147483647 bytes Texture alignment: 512 bytes Concurrent copy and kernel execution: Yes with 2 copy engine(s) Run time limit on kernels: No Integrated GPU sharing Host Memory: No Support host page-locked memory mapping: Yes Alignment requirement for Surfaces: Yes Device has ECC support: Enabled Device supports Unified Addressing (UVA): Yes Device PCI Bus ID / PCI location ID: 4 / 0 Compute Mode: < Exclusive Process (many threads in one process is able to use ::cudaSetDevice() with this device) >
GPU Hardware on Sapelo2
- 14 NVIDIA Tesla (Kepler) K20Xm GPU cards (32 x 2688 = 86016 GPU cores). These cards are installed on 2 hosts each of which has dual 6-core Intel Xeon CPUs and 96GB of RAM; there are 7 GPU cards per host. The output of NVIDIA SDK deviceQuery for one such GPU card is
Device 0: "Tesla K20Xm" CUDA Driver Version / Runtime Version 9.0 / 9.0 CUDA Capability Major/Minor version number: 3.5 Total amount of global memory: 5700 MBytes (5976424448 bytes) (14) Multiprocessors, (192) CUDA Cores/MP: 2688 CUDA Cores GPU Max Clock rate: 732 MHz (0.73 GHz) Memory Clock rate: 2600 Mhz Memory Bus Width: 384-bit L2 Cache Size: 1572864 bytes Maximum Texture Dimension Size (x,y,z) 1D=(65536), 2D=(65536, 65536), 3D=(4096, 4096, 4096) Maximum Layered 1D Texture Size, (num) layers 1D=(16384), 2048 layers Maximum Layered 2D Texture Size, (num) layers 2D=(16384, 16384), 2048 layers Total amount of constant memory: 65536 bytes Total amount of shared memory per block: 49152 bytes Total number of registers available per block: 65536 Warp size: 32 Maximum number of threads per multiprocessor: 2048 Maximum number of threads per block: 1024 Max dimension size of a thread block (x,y,z): (1024, 1024, 64) Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535) Maximum memory pitch: 2147483647 bytes Texture alignment: 512 bytes Concurrent copy and kernel execution: Yes with 2 copy engine(s) Run time limit on kernels: No Integrated GPU sharing Host Memory: No Support host page-locked memory mapping: Yes Alignment requirement for Surfaces: Yes Device has ECC support: Enabled Device supports Unified Addressing (UVA): Yes Supports Cooperative Kernel Launch: No Supports MultiDevice Co-op Kernel Launch: No Device PCI Domain ID / Bus ID / location ID: 0 / 12 / 0 Compute Mode: < Exclusive Process (many threads in one process is able to use ::cudaSetDevice() with this device) >