Skip to main content

26 posts tagged with "LLM"

Articles about large language models

View all tags

Hot-Pluggable GPUs Join the CNCF: What CoHDI's Sandbox Debut Means for Your Self-Hosted GPU Pool
·Dora Noda·11 min

Hot-Pluggable GPUs Join the CNCF: What CoHDI's Sandbox Debut Means for Your Self-Hosted GPU Pool

CoHDI entered the CNCF Sandbox promising hot-pluggable GPUs via Kubernetes DRA. A grounded accounting of the before/after math for a small self-hosted GPU pool, the Cluster API implications, and why DRA on fixed GPUs is the move today.

self-hosting
Kubernetes
infrastructure
LLM
GPU Autoscaling on Kubernetes With KEDA: Scaling on the One Metric HPA Can't See
·Dora Noda·9 min

GPU Autoscaling on Kubernetes With KEDA: Scaling on the One Metric HPA Can't See

HPA can't see GPU utilization — nvidia.com/gpu is an integer device count, so a vLLM pod at 8% CPU and a saturated GPU never scales. A worked KEDA pattern: dcgm-exporter to Prometheus to ScaledObject YAML, a minimal external scaler serving utilization plus queue depth, the scale-to-zero cost math, and where KEDA stops and DRA begins.

Kubernetes
LLM
self-hosting
cost-optimization
+1
Self-Hosted vLLM on Kubernetes: What Owning Your Inference Layer Actually Costs
·Dora Noda·12 min

Self-Hosted vLLM on Kubernetes: What Owning Your Inference Layer Actually Costs

CNCF's July 2026 vLLM walkthrough builds self-hosted inference in three Kubernetes resources. A worked breakeven — 10 to 15M tokens a day against Sonnet 5 on a €1,199/month GPU box — plus the four production gaps the lab skips and the five cases where owning the layer wins.

LLM
self-hosting
Kubernetes
cost-optimization
+1
Open-Weight AI's Kubernetes Moment: Stress-Testing the Analogy Phase by Phase
·Dora Noda·10 min

Open-Weight AI's Kubernetes Moment: Stress-Testing the Analogy Phase by Phase

Tobi Knaup argues open-weight AI sits where Kubernetes sat in 2016. We grade the analogy across all four phases of the Kubernetes decade — substrate, distro fight, hyperscaler absorption, and the self-hosting price — with the utilization math that decides it.

LLM
Kubernetes
self-hosting
cost-optimization
Your 8B Model Isn't Dumb. Its Harness Is: Forge's 53%-to-99% Guardrail Lesson for Platform Ops
·Dora Noda·12 min

Your 8B Model Isn't Dumb. Its Harness Is: Forge's 53%-to-99% Guardrail Lesson for Platform Ops

Forge's May 2026 result took an 8B model from 53% to 99.3% on agentic tasks with guardrails alone — beating unguardrailed Claude Sonnet outright. What the retry-nudge math, the 75-point serving-backend swing, and a worked token-vs-hardware breakeven mean for running platform-ops agents on a self-hosted GPU.

PaaS
self-hosting
AI agents
LLM
+1
Run a Claude Fable Security Scan with Bex Security
·Dora Noda·9 min

Run a Claude Fable Security Scan with Bex Security

Run a Claude Fable security scan with Claude Code and Bex Security. The exact command, the model row that actually runs, and the sandbox policy Bex enforces around Claude.

changelog
product
security
AI agents
+3
Run GLM and Kimi Security Scans with Bex Security
·Dora Noda·11 min

Run GLM and Kimi Security Scans with Bex Security

Run GLM and Kimi security scans on a real codebase with Bex Security. Use one evidence-driven workflow for discovery, validation, remediation, and review.

changelog
product
security
AI agents
+2
OpenCost 1.121 Finally Answers: What Does Each Token Cost on Your Own GPUs?
·Dora Noda·13 min

OpenCost 1.121 Finally Answers: What Does Each Token Cost on Your Own GPUs?

OpenCost 1.121 adds Kubernetes-native per-token inference metering via llm-d — allocation vs. usage cost, KV-cache-corrected — turning self-hosted LLM spend from a quarterly guess into a Prometheus metric you can alert on.

AI
LLM
infrastructure
Kubernetes
+2
vLLM vs Ollama in Production: What PagedAttention's 19x Throughput Gap Really Buys on Owned GPUs
·Dora Noda·16 min

vLLM vs Ollama in Production: What PagedAttention's 19x Throughput Gap Really Buys on Owned GPUs

Red Hat's 2026 benchmark put vLLM at 793 tokens per second and Ollama at 41 on the same GPU and model — a 19x gap. The same comparison at one concurrent user shows them within 20%. A concrete breakdown of where PagedAttention earns that gap, where it doesn't, and what each engine costs to run on an owned Hetzner GPU fleet.

LLM
AI
self-hosting
Kubernetes
+2
Showing 10–18 of 26 posts