и посмотреть медиа
DevOps&SRE Library - статьи
и посмотреть медиа
Библиотека статей по DevOps и SRE для специалистов.
Библиотека статей по DevOps и SRE для специалистов.
Собрание статей по темам DevOps и Site Reliability Engineering. Глубокий анализ практик и инструментов для профессионалов.
What Does 4.4% GPU Utilization Actually Mean? A few weeks after publishing the 1M token/s post, I spent a weekend helping my good friend Milko Ilari set up vLLM on his shiny new DGX Spark with Gemma 4. My first in-person reaction was “It’s Champagne” (from the old days of PC Perspective) The Spark is a wild little machine — 128 GB of unified memory in a box you can hold with one hand, running the same Blackwell architecture as the datacenter B200s. But its memory bandwidth is 273 GB/s. The B200s in our cluster do 8,000 GB/s. Almost 30x less. Watching the numbers on that tiny machine got me thinking. The benchmark I ran on GKE Autopilot with 96 B200 GPUs had reported 4.4% FLOPS utilization. 10.9% memory bandwidth. Tensor cores active 1.5% of the time. The GPUs looked almost idle while pushing a million tokens per second. Was something wrong? No. And honestly, figuring out why turned out to be more interesting than the benchmark itself. That first post covers the journey — every optimization, and many failure 🫠. This one covers the physics. Ссылка скрыта
Открыть канал и посмотреть медиаТолько зарегистрированные пользователи могут делиться своим мнением.
Станьте первым, кто поделится своим впечатлением об этом ресурсе!
Канал с бесплатными сигналами по Forex и золоту. 4-10 сигналов в день, только полезная информация без платных опций.
Канал о прибыльной крипто-ферме: staking, farming и ликвидные пулы.