GPU Hardware: Difference between revisions
Jump to navigation
Jump to search
No edit summary |
No edit summary |
||
| Line 40: | Line 40: | ||
</pre> | </pre> | ||
===GPU Hardware on | ===GPU Hardware on Sapelo2=== | ||
* | * 14 NVIDIA Tesla (Kepler) K20Xm GPU cards (32 x 2688 = 86016 GPU cores). These cards are installed on 2 hosts each of which has dual 6-core Intel Xeon CPUs and 96GB of RAM; there are 7 GPU cards per host. The output of NVIDIA SDK deviceQuery for one such GPU card is | ||
<pre class="gscript"> | <pre class="gscript"> | ||
Device 0: "Tesla K20Xm" | Device 0: "Tesla K20Xm" | ||
CUDA Driver Version / Runtime Version | CUDA Driver Version / Runtime Version 9.0 / 9.0 | ||
CUDA Capability Major/Minor version number: 3.5 | CUDA Capability Major/Minor version number: 3.5 | ||
Total amount of global memory: | Total amount of global memory: 5700 MBytes (5976424448 bytes) | ||
(14) Multiprocessors | (14) Multiprocessors, (192) CUDA Cores/MP: 2688 CUDA Cores | ||
GPU Clock rate: | GPU Max Clock rate: 732 MHz (0.73 GHz) | ||
Memory Clock rate: 2600 Mhz | Memory Clock rate: 2600 Mhz | ||
Memory Bus Width: 384-bit | Memory Bus Width: 384-bit | ||
L2 Cache Size: 1572864 bytes | L2 Cache Size: 1572864 bytes | ||
Maximum Texture Dimension Size (x,y,z) 1D=(65536), 2D=(65536, 65536), 3D=(4096, 4096, 4096) | |||
Maximum Layered 1D Texture Size, (num) layers 1D=(16384), 2048 layers | |||
Maximum Layered 2D Texture Size, (num) layers 2D=(16384, 16384), 2048 layers | |||
Total amount of constant memory: 65536 bytes | Total amount of constant memory: 65536 bytes | ||
Total amount of shared memory per block: 49152 bytes | Total amount of shared memory per block: 49152 bytes | ||
| Line 62: | Line 63: | ||
Maximum number of threads per multiprocessor: 2048 | Maximum number of threads per multiprocessor: 2048 | ||
Maximum number of threads per block: 1024 | Maximum number of threads per block: 1024 | ||
Max dimension size of a thread block (x,y,z): (1024, 1024, 64) | |||
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535) | |||
Maximum memory pitch: 2147483647 bytes | Maximum memory pitch: 2147483647 bytes | ||
Texture alignment: 512 bytes | Texture alignment: 512 bytes | ||
| Line 73: | Line 74: | ||
Device has ECC support: Enabled | Device has ECC support: Enabled | ||
Device supports Unified Addressing (UVA): Yes | Device supports Unified Addressing (UVA): Yes | ||
Device PCI Bus ID / | Supports Cooperative Kernel Launch: No | ||
Supports MultiDevice Co-op Kernel Launch: No | |||
Device PCI Domain ID / Bus ID / location ID: 0 / 12 / 0 | |||
Compute Mode: | |||
< Exclusive Process (many threads in one process is able to use ::cudaSetDevice() with this device) > | |||
</pre> | </pre> | ||
* | * | ||
Revision as of 08:09, 15 June 2018
GPU Hardware on Sapelo
- 16 NVIDIA K40m GPU cards. These cards are installed on 2 hosts, each of which has dual 8-core Intel Xeon CPUs and 128GB of RAM; there are 8 GPU cards per host. The output of NVIDIA SDK deviceQuery for one such GPU card is
Device 0: "Tesla K40m"
CUDA Driver Version / Runtime Version 6.5 / 6.5
CUDA Capability Major/Minor version number: 3.5
Total amount of global memory: 11520 MBytes (12079136768 bytes)
(15) Multiprocessors, (192) CUDA Cores/MP: 2880 CUDA Cores
GPU Clock rate: 745 MHz (0.75 GHz)
Memory Clock rate: 3004 Mhz
Memory Bus Width: 384-bit
L2 Cache Size: 1572864 bytes
Maximum Texture Dimension Size (x,y,z) 1D=(65536), 2D=(65536, 65536), 3D=(4096, 4096, 4096)
Maximum Layered 1D Texture Size, (num) layers 1D=(16384), 2048 layers
Maximum Layered 2D Texture Size, (num) layers 2D=(16384, 16384), 2048 layers
Total amount of constant memory: 65536 bytes
Total amount of shared memory per block: 49152 bytes
Total number of registers available per block: 65536
Warp size: 32
Maximum number of threads per multiprocessor: 2048
Maximum number of threads per block: 1024
Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535)
Maximum memory pitch: 2147483647 bytes
Texture alignment: 512 bytes
Concurrent copy and kernel execution: Yes with 2 copy engine(s)
Run time limit on kernels: No
Integrated GPU sharing Host Memory: No
Support host page-locked memory mapping: Yes
Alignment requirement for Surfaces: Yes
Device has ECC support: Enabled
Device supports Unified Addressing (UVA): Yes
Device PCI Bus ID / PCI location ID: 4 / 0
Compute Mode:
< Exclusive Process (many threads in one process is able to use ::cudaSetDevice() with this device) >
GPU Hardware on Sapelo2
- 14 NVIDIA Tesla (Kepler) K20Xm GPU cards (32 x 2688 = 86016 GPU cores). These cards are installed on 2 hosts each of which has dual 6-core Intel Xeon CPUs and 96GB of RAM; there are 7 GPU cards per host. The output of NVIDIA SDK deviceQuery for one such GPU card is
Device 0: "Tesla K20Xm"
CUDA Driver Version / Runtime Version 9.0 / 9.0
CUDA Capability Major/Minor version number: 3.5
Total amount of global memory: 5700 MBytes (5976424448 bytes)
(14) Multiprocessors, (192) CUDA Cores/MP: 2688 CUDA Cores
GPU Max Clock rate: 732 MHz (0.73 GHz)
Memory Clock rate: 2600 Mhz
Memory Bus Width: 384-bit
L2 Cache Size: 1572864 bytes
Maximum Texture Dimension Size (x,y,z) 1D=(65536), 2D=(65536, 65536), 3D=(4096, 4096, 4096)
Maximum Layered 1D Texture Size, (num) layers 1D=(16384), 2048 layers
Maximum Layered 2D Texture Size, (num) layers 2D=(16384, 16384), 2048 layers
Total amount of constant memory: 65536 bytes
Total amount of shared memory per block: 49152 bytes
Total number of registers available per block: 65536
Warp size: 32
Maximum number of threads per multiprocessor: 2048
Maximum number of threads per block: 1024
Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535)
Maximum memory pitch: 2147483647 bytes
Texture alignment: 512 bytes
Concurrent copy and kernel execution: Yes with 2 copy engine(s)
Run time limit on kernels: No
Integrated GPU sharing Host Memory: No
Support host page-locked memory mapping: Yes
Alignment requirement for Surfaces: Yes
Device has ECC support: Enabled
Device supports Unified Addressing (UVA): Yes
Supports Cooperative Kernel Launch: No
Supports MultiDevice Co-op Kernel Launch: No
Device PCI Domain ID / Bus ID / location ID: 0 / 12 / 0
Compute Mode:
< Exclusive Process (many threads in one process is able to use ::cudaSetDevice() with this device) >