и посмотреть медиа
DevOps&SRE Library - статьи
и посмотреть медиа
Библиотека статей по DevOps и SRE для специалистов.
Библиотека статей по DevOps и SRE для специалистов.
Собрание статей по темам DevOps и Site Reliability Engineering. Глубокий анализ практик и инструментов для профессионалов.
VibeOps: A Secure read-only setup for AI-Assisted Kubernetes Debugging There is a lot of noise right now about letting AI "fix" your infrastructure. When production is acting up, you need to maintain a complete mental model of the system. If you let the AI be the driving force, you lose the overview. Ссылка скрыта
Открыть канал и посмотреть медиаStop Manually Generating Kubeconfigs: Meet KubeUser KubeUser is a kubernetes-native operator that turns user management into a declarative code experience. No more manual certificate juggling — just apply a YAML file, and the operator handles the rest. Ссылка скрыта
Открыть канал и посмотреть медиаYou Don't Have a GIL Problem — You Have a CPU Problem This article documents a real production investigation into latency variance in Python-based microservices running on Kubernetes, revealing how CPU throttling amplifies GIL contention into unpredictable response time spikes. Ссылка скрыта
Открыть канал и посмотреть медиаРазработчики получают инфраструктуру самостоятельно. DevOps — перестают выполнять однотипные запросы. 18 августа на бесплатном онлайн-вебинаре Orion soft покажет, как работает новая IDP-функциональность HyperDrive: self-service, GitOps, политики безопасности и управление инфраструктурой через Model Context Protocol. Реклама. ООО "Орион", ИНН: ИНН 9704113582, erid: 2Vtzqvqgags
Открыть канал и посмотреть медиаThe feedback loops behind Kubernetes For the last decade, Kubernetes has been the backdrop to most of my work: operating clusters, helping build hosted Kubernetes, and writing Kubernetes operators. At PlanetScale, that now means running stateful systems like Postgres and MySQL in production. Kubernetes has many faces, but here I want to talk about one face only: why it is so good at running workloads at scale. People ask me what an operator actually does. The canonical answer is: "it reconciles desired state." This is correct, but it also tells you almost nothing. An operator is a feedback controller. It's the same closed loop that runs a thermostat or keeps your car at a fixed speed on cruise control. In our case, the thing being controlled is a database. I have been building these loops for years, and the best way I know to make them click is to ignore Kubernetes at the beginning. Kubernetes is full of control theory, even if we don't call it that in the day-to-day. Before we look at a single line of Kubernetes, we're going to run a production database by hand and slowly let the feedback loop appear on its own. Then we'll map that loop to Kubernetes, with the pieces production needs: a store, watches, queues, retries, and more. At the end, we'll look at what one of these loops looks like in a real operator. Ссылка скрыта
Открыть канал и посмотреть медиаClient’s GKE Cluster Ate Their Entire VPC GKE pod IP exhaustion is one of the few failure modes that gives you no warning before it goes terminal. I recently stepped into a war room where a client’s primary scaling group had flatlined — workloads cordoned, deployments stuck in Pending, and the estimated cost of the stall nearing $15k per hour in lost transaction volume. The culprit wasn’t traffic. It was a /20 subnet that had quietly run out of address space, and a set of GKE allocation defaults nobody had questioned at design time. The IP Math I Uncovered During Triage: Ссылка скрыта The Class E Rescue: Ссылка скрыта
Открыть канал и посмотреть медиа🔥 Приглашаем на бесплатный открытый вебинар курса «Observability: мониторинг, логирование, трассировка»: «Системы логирования: ELK, EFK или Graylog?» 🗓 Когда: 17 августа, 20:00 (мск) Логи — один из ключевых источников информации о состоянии системы. Но без правильно выбранного инструмента они превращаются в хаотичный поток данных, в котором сложно найти причину проблемы. На вебинаре сравним популярные системы централизованного логирования и поможем вам выбрать оптимальное решение под вашу инфраструктуру. Что будет на вебинаре: - Чем отличаются ELK, EFK и Graylog и в каких сценариях каждый стек наиболее эффективен - Как устроен процесс сбора, обработки, хранения и поиска логов - Как организовать централизованное логирование для мониторинга и диагностики распределённых систем - На что обратить внимание при выборе системы логирования для своей инфраструктуры В результате вы: - Получите понимание сильных и слабых сторон ELK, EFK и Graylog - Научитесь выбирать подходящее решение под задачи проекта и инфраструктуры - Узнаете лучшие практики построения централизованной системы логирования - Сможете использовать логи для ускорения диагностики и повышения наблюдаемости сервисов Кому будет полезно: DevOps- и SRE-инженерам, системным администраторам, Backend-разработчикам и архитекторам, которым важно быстро находить причины сбоев и анализировать поведение систем. 👉 Зарегистрируйтесь Ссылка скрыта Бесплатное занятие приурочено к старту курса «Observability: мониторинг, логирование, трассировка», на котором вы научитесь строить современные системы наблюдаемости с Prometheus, Grafana, ELK, Tempo и другими инструментами. Реклама. ООО «Отус онлайн-образование», ОГРН 1177746618576, erid: 2VtzqvHEFpX
Открыть канал и посмотреть медиаWhat Does 4.4% GPU Utilization Actually Mean? A few weeks after publishing the 1M token/s post, I spent a weekend helping my good friend Milko Ilari set up vLLM on his shiny new DGX Spark with Gemma 4. My first in-person reaction was “It’s Champagne” (from the old days of PC Perspective) The Spark is a wild little machine — 128 GB of unified memory in a box you can hold with one hand, running the same Blackwell architecture as the datacenter B200s. But its memory bandwidth is 273 GB/s. The B200s in our cluster do 8,000 GB/s. Almost 30x less. Watching the numbers on that tiny machine got me thinking. The benchmark I ran on GKE Autopilot with 96 B200 GPUs had reported 4.4% FLOPS utilization. 10.9% memory bandwidth. Tensor cores active 1.5% of the time. The GPUs looked almost idle while pushing a million tokens per second. Was something wrong? No. And honestly, figuring out why turned out to be more interesting than the benchmark itself. That first post covers the journey — every optimization, and many failure 🫠. This one covers the physics. Ссылка скрыта
Открыть канал и посмотреть медиа❗️Небольшое уточнение к предыдущему посту: в нём была указана некорректная ссылка на бота. Актуальная ссылка для получения доступа к эфиру: @shortcut_devops_bot
Открыть канал и посмотреть медиаSolo los usuarios registrados pueden compartir su opinión.
¡Sé el primero en compartir tu experiencia con este recurso!
Reembolso de CanalSn0wdenCon propinas de devolución de dinero en inglés y francés. Guías, devoluciones de gastos y protección de compras.
Un canal con operaciones bursátiles en directo sin riesgo. Transacciones reales, análisis y estrategias para el éxito de las operaciones intradía en Telegram.
Actualizaciones, instrucciones y consejos actualizados en Microsoft Office365.