M6i.8xlarge Benchmarks [2407019-NE-M6I8XLARG24]

131 Results Shown

Neural Magic DeepSparse:
NLP Token Classification, BERT base uncased conll2003 - Synchronous Single-Stream:
ms/batch
items/sec
NLP Token Classification, BERT base uncased conll2003 - Asynchronous Multi-Stream:
ms/batch
items/sec
BERT-Large, NLP Question Answering, Sparse INT8 - Synchronous Single-Stream:
ms/batch
items/sec
BERT-Large, NLP Question Answering, Sparse INT8 - Asynchronous Multi-Stream:
ms/batch
items/sec
CV Segmentation, 90% Pruned YOLACT Pruned - Synchronous Single-Stream:
ms/batch
items/sec
CV Segmentation, 90% Pruned YOLACT Pruned - Asynchronous Multi-Stream:
ms/batch
items/sec
NLP Text Classification, DistilBERT mnli - Synchronous Single-Stream:
ms/batch
items/sec
NLP Text Classification, DistilBERT mnli - Asynchronous Multi-Stream:
ms/batch
items/sec
CV Detection, YOLOv5s COCO, Sparse INT8 - Synchronous Single-Stream:
ms/batch
items/sec
CV Detection, YOLOv5s COCO, Sparse INT8 - Asynchronous Multi-Stream:
ms/batch
items/sec
CV Classification, ResNet-50 ImageNet - Synchronous Single-Stream:
ms/batch
items/sec
CV Classification, ResNet-50 ImageNet - Asynchronous Multi-Stream:
ms/batch
items/sec
Llama2 Chat 7b Quantized - Synchronous Single-Stream:
ms/batch
items/sec
Llama2 Chat 7b Quantized - Asynchronous Multi-Stream:
ms/batch
items/sec
ResNet-50, Sparse INT8 - Synchronous Single-Stream:
ms/batch
items/sec
ResNet-50, Sparse INT8 - Asynchronous Multi-Stream:
ms/batch
items/sec
ResNet-50, Baseline - Synchronous Single-Stream:
ms/batch
items/sec
ResNet-50, Baseline - Asynchronous Multi-Stream:
ms/batch
items/sec
NLP Text Classification, BERT base uncased SST2, Sparse INT8 - Synchronous Single-Stream:
ms/batch
items/sec
NLP Text Classification, BERT base uncased SST2, Sparse INT8 - Asynchronous Multi-Stream:
ms/batch
items/sec
NLP Document Classification, oBERT base uncased on IMDB - Synchronous Single-Stream:
ms/batch
items/sec
NLP Document Classification, oBERT base uncased on IMDB - Asynchronous Multi-Stream:
ms/batch
items/sec
Mlpack Benchmark:
scikit_linearridgeregression
scikit_svm
scikit_qda
scikit_ica
oneDNN:
Recurrent Neural Network Inference - CPU
Recurrent Neural Network Training - CPU
Deconvolution Batch shapes_3d - CPU
Convolution Batch Shapes Auto - CPU
IP Shapes 3D - CPU
IP Shapes 1D - CPU
OpenCV:
Stitching
Graph API
Video
Core
OpenVINO:
Age Gender Recognition Retail 0013 FP16-INT8 - CPU:
ms
FPS
Handwritten English Recognition FP16-INT8 - CPU:
ms
FPS
Age Gender Recognition Retail 0013 FP16 - CPU:
ms
FPS
Person Re-Identification Retail FP16 - CPU:
ms
FPS
Handwritten English Recognition FP16 - CPU:
ms
FPS
Noise Suppression Poconet-Like FP16 - CPU:
ms
FPS
Person Vehicle Bike Detection FP16 - CPU:
ms
FPS
Weld Porosity Detection FP16-INT8 - CPU:
ms
FPS
Machine Translation EN To DE FP16 - CPU:
ms
FPS
Road Segmentation ADAS FP16-INT8 - CPU:
ms
FPS
Face Detection Retail FP16-INT8 - CPU:
ms
FPS
Weld Porosity Detection FP16 - CPU:
ms
FPS
Vehicle Detection FP16-INT8 - CPU:
ms
FPS
Road Segmentation ADAS FP16 - CPU:
ms
FPS
Face Detection Retail FP16 - CPU:
ms
FPS
Face Detection FP16-INT8 - CPU:
ms
FPS
Vehicle Detection FP16 - CPU:
ms
FPS
Person Detection FP32 - CPU:
ms
FPS
Person Detection FP16 - CPU:
ms
FPS
Face Detection FP16 - CPU:
ms
FPS
ONNX Runtime:
Faster R-CNN R-50-FPN-int8 - CPU - Standard
Faster R-CNN R-50-FPN-int8 - CPU - Parallel
super-resolution-10 - CPU - Parallel
ResNet50 v1-12-int8 - CPU - Standard
ResNet50 v1-12-int8 - CPU - Parallel
ArcFace ResNet-100 - CPU - Parallel
fcn-resnet101-11 - CPU - Parallel
CaffeNet 12-int8 - CPU - Standard
CaffeNet 12-int8 - CPU - Parallel
bertsquad-12 - CPU - Standard
bertsquad-12 - CPU - Parallel
T5 Encoder - CPU - Standard
T5 Encoder - CPU - Parallel
yolov4 - CPU - Standard
yolov4 - CPU - Parallel
GPT-2 - CPU - Parallel
Whisper.cpp:
ggml-medium.en - 2016 State of the Union
ggml-small.en - 2016 State of the Union
ggml-base.en - 2016 State of the Union
Llama.cpp
oneDNN
OpenCV:
DNN - Deep Neural Network
Object Detection
Image Processing
Features 2D
ONNX Runtime:
super-resolution-10 - CPU - Standard:
Inference Time Cost (ms)
Inferences Per Second
ArcFace ResNet-100 - CPU - Standard:
Inference Time Cost (ms)
Inferences Per Second
fcn-resnet101-11 - CPU - Standard:
Inference Time Cost (ms)
Inferences Per Second
GPT-2 - CPU - Standard:
Inference Time Cost (ms)
Inferences Per Second

m6i.8xlarge

Processor: Intel Xeon Platinum 8375C (16 Cores / 32 Threads), Motherboard: Amazon EC2 m6i.8xlarge (1.0 BIOS), Chipset: Intel 440FX 82441FX PMC, Memory: 1 x 128 GB DDR4-3200MT/s, Disk: 537GB Amazon Elastic Block Store, Graphics: EFI VGA, Network: Amazon Elastic

OS: Ubuntu 22.04, Kernel: 6.5.0-1017-aws (x86_64), Vulkan: 1.3.255, Compiler: GCC 11.4.0, File-System: ext4, Screen Resolution: 800x600, System Layer: amazon

Kernel Notes: Transparent Huge Pages: madvise
Compiler Notes: --build=x86_64-linux-gnu --disable-vtable-verify --disable-werror --enable-bootstrap --enable-cet --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-gnu-unique-object --enable-languages=c,ada,c++,go,brig,d,fortran,objc,obj-c++,m2 --enable-libphobos-checking=release --enable-libstdcxx-debug --enable-libstdcxx-time=yes --enable-link-serialization=2 --enable-multiarch --enable-multilib --enable-nls --enable-objc-gc=auto --enable-offload-targets=nvptx-none=/build/gcc-11-XeT9lY/gcc-11-11.4.0/debian/tmp-nvptx/usr,amdgcn-amdhsa=/build/gcc-11-XeT9lY/gcc-11-11.4.0/debian/tmp-gcn/usr --enable-plugin --enable-shared --enable-threads=posix --host=x86_64-linux-gnu --program-prefix=x86_64-linux-gnu- --target=x86_64-linux-gnu --with-abi=m64 --with-arch-32=i686 --with-build-config=bootstrap-lto-lean --with-default-libstdcxx-abi=new --with-gcc-major-version-only --with-multilib-list=m32,m64,mx32 --with-target-system-zlib=auto --with-tune=generic --without-cuda-driver -v
Processor Notes: CPU Microcode: 0xd0003d1
Security Notes: gather_data_sampling: Unknown: Dependent on hypervisor status + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Mitigation of Clear buffers; SMT Host state unknown + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected

Testing initiated at 1 July 2024 09:30 by user root.

m6i.8xlarge

Statistics

Graph Settings

Multi-Way Comparison

Table

Run Management

m6i.8xlarge

Neural Magic DeepSparse

Mlpack Benchmark

oneDNN

OpenCV

OpenVINO

ONNX Runtime

Whisper.cpp

Llama.cpp

oneDNN

OpenCV

ONNX Runtime

131 Results Shown

m6i.8xlarge