
Hot-Pluggable GPUs Join the CNCF: What CoHDI's Sandbox Debut Means for Your Self-Hosted GPU Pool
CoHDI entered the CNCF Sandbox promising hot-pluggable GPUs via Kubernetes DRA. A grounded accounting of the before/after math for a small self-hosted GPU pool, the Cluster API implications, and why DRA on fixed GPUs is the move today.

GPU Autoscaling on Kubernetes With KEDA: Scaling on the One Metric HPA Can't See
HPA can't see GPU utilization — nvidia.com/gpu is an integer device count, so a vLLM pod at 8% CPU and a saturated GPU never scales. A worked KEDA pattern: dcgm-exporter to Prometheus to ScaledObject YAML, a minimal external scaler serving utilization plus queue depth, the scale-to-zero cost math, and where KEDA stops and DRA begins.

Self-Hosted vLLM on Kubernetes: What Owning Your Inference Layer Actually Costs
CNCF's July 2026 vLLM walkthrough builds self-hosted inference in three Kubernetes resources. A worked breakeven — 10 to 15M tokens a day against Sonnet 5 on a €1,199/month GPU box — plus the four production gaps the lab skips and the five cases where owning the layer wins.

Open-Weight AI's Kubernetes Moment: Stress-Testing the Analogy Phase by Phase
Tobi Knaup argues open-weight AI sits where Kubernetes sat in 2016. We grade the analogy across all four phases of the Kubernetes decade — substrate, distro fight, hyperscaler absorption, and the self-hosting price — with the utilization math that decides it.

Your 8B Model Isn't Dumb. Its Harness Is: Forge's 53%-to-99% Guardrail Lesson for Platform Ops
Forge's May 2026 result took an 8B model from 53% to 99.3% on agentic tasks with guardrails alone — beating unguardrailed Claude Sonnet outright. What the retry-nudge math, the 75-point serving-backend swing, and a worked token-vs-hardware breakeven mean for running platform-ops agents on a self-hosted GPU.

Run a Claude Fable Security Scan with Bex Security
Run a Claude Fable security scan with Claude Code and Bex Security. The exact command, the model row that actually runs, and the sandbox policy Bex enforces around Claude.

Run GLM and Kimi Security Scans with Bex Security
Run GLM and Kimi security scans on a real codebase with Bex Security. Use one evidence-driven workflow for discovery, validation, remediation, and review.

OpenCost 1.121 Finally Answers: What Does Each Token Cost on Your Own GPUs?
OpenCost 1.121 adds Kubernetes-native per-token inference metering via llm-d — allocation vs. usage cost, KV-cache-corrected — turning self-hosted LLM spend from a quarterly guess into a Prometheus metric you can alert on.

vLLM vs Ollama in Production: What PagedAttention's 19x Throughput Gap Really Buys on Owned GPUs
Red Hat's 2026 benchmark put vLLM at 793 tokens per second and Ollama at 41 on the same GPU and model — a 19x gap. The same comparison at one concurrent user shows them within 20%. A concrete breakdown of where PagedAttention earns that gap, where it doesn't, and what each engine costs to run on an owned Hetzner GPU fleet.