Generalaws.amazon.com·9d ago

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

Submitted by @

Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU infrastructure by 75% while holding sub-second latency at 92.1 requests per second per GPU.

ORIGINAL REPORTING
AWS Machine Learning
Published 9d ago ago · Iman Abbasnejad
Read full story at AWS Machine Learning

Original reporting by AWS Machine Learning · Iman Abbasnejad. GridIndex is an aggregation and intelligence layer — full credit to the original publisher.

CONTINUE READING

This page summarizes and tracks coverage of this developing story.

Read full story at AWS Machine Learning
0 views 0 upvotes 0 comments 0 shares

Discussion · 0

Sign in to join the discussion.