и посмотреть медиа
DevOps&SRE Library - статьи
и посмотреть медиа
Библиотека статей по DevOps и SRE для специалистов.
Библиотека статей по DevOps и SRE для специалистов.
Собрание статей по темам DevOps и Site Reliability Engineering. Глубокий анализ практик и инструментов для профессионалов.
What Does 4.4% GPU Utilization Actually Mean? A few weeks after publishing the 1M token/s post, I spent a weekend helping my good friend Milko Ilari set up vLLM on his shiny new DGX Spark with Gemma 4. My first in-person reaction was “It’s Champagne” (from the old days of PC Perspective) The Spark is a wild little machine — 128 GB of unified memory in a box you can hold with one hand, running the same Blackwell architecture as the datacenter B200s. But its memory bandwidth is 273 GB/s. The B200s in our cluster do 8,000 GB/s. Almost 30x less. Watching the numbers on that tiny machine got me thinking. The benchmark I ran on GKE Autopilot with 96 B200 GPUs had reported 4.4% FLOPS utilization. 10.9% memory bandwidth. Tensor cores active 1.5% of the time. The GPUs looked almost idle while pushing a million tokens per second. Was something wrong? No. And honestly, figuring out why turned out to be more interesting than the benchmark itself. That first post covers the journey — every optimization, and many failure 🫠. This one covers the physics. Ссылка скрыта
Открыть канал и посмотреть медиаOnly registered users can share their opinion.
Be the first to share your impression of this resource!
Professional analysis of the stock market, futures, options and indices Nifty. Banknifty from Market Guru.
IELTS-channel with essay samples, preparation tips, humor and personal stories from an experienced teacher with high scores.
A channel about races and friendly meetings. For lovers of speed and communication.