
Only LLM Spans Should Cost Money: Span-Class-Aware Billing for Agent Observability
Span-class-aware metering prices LLM spans while tracing tool calls free — the worked cost math, which vendors already do it, and an OTel Collector recipe for self-hosted agent observability.

Open-Weight AI Is Having Its Kubernetes Moment — Here's What Your Fleet Needs to Serve It
Open-weight models are becoming the default infrastructure layer for AI. This post maps the standardized Kubernetes inference stack — OCI weights, DRA, InferencePool, KServe, llm-d — and shows what a self-hosted fleet needs to serve models next to tenant apps.

Pangolin Puts SSO and WireGuard in Front of LLM Access Instead of API Keys: What Tunnel-Based Identity Means for Agent Credential Hygiene
Pangolin's AI Gateway authenticates LLM access with SSO-backed WireGuard tunnels instead of static API keys, and joins self-hosted models to the same gateway as public ones. Here's how the mechanism works, how it compares to Tailscale Aperture, and what it changes in your threat model.

66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS
CNCF survey data shows 66% of organizations run AI inference on Kubernetes while only 7% deploy models daily. Four mechanisms — versioned model artifacts, DRA accelerator requests, eval-gated promotion, and correlated inference telemetry — close the gap between running a model and shipping one.

SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server
SGLang now serves 400,000+ production GPUs, making it a real second default next to vLLM. Benchmarks, workload maps, and a decision matrix for choosing the model server on your own GPU pool.

Envoy AI Gateway Hits 1.0: What an LLM-Aware Kubernetes Gateway Actually Load-Balances On
Envoy AI Gateway reached 1.0 as the standard routing layer for self-hosted LLM traffic on Kubernetes. How InferencePool endpoint selection on KV-cache state and queue depth beats round-robin — with benchmarks — and what it means for fleets you run yourself.

Hetzner's €214 GEX45: What 24 GB of Blackwell VRAM Buys Self-Hosted Inference (and When It Beats Serverless GPUs)
Hetzner's €214/month GEX45 pairs an RTX PRO 4000 Blackwell card with 24 GB of VRAM. The crossover math against serverless GPU rentals, and which models newly fit.

llm-d Is Now a CNCF Sandbox Project: What Kubernetes-Native Distributed Inference Means for Self-Hosting Your AI Agents' Brains
llm-d brings distributed LLM inference — prefix-cache-aware routing and prefill/decode disaggregation — to Kubernetes as a CNCF Sandbox project. What it means for self-hosting AI-agent inference, what it costs versus hosted APIs, and how early the project still is.

TurboFieldfare Fits a 26B MoE Model in 2GB of RAM: What SSD-Streamed Experts Mean for Your Cheapest Inference Node
TurboFieldfare runs Gemma 4 26B-A4B in ~2GB of RAM by streaming 4-bit experts from SSD — 5–6 tok/s on an M2 Air, 31–35 on an M5 Pro. The per-token byte math behind those numbers, and how to size a €49 NVMe box as your cheapest inference tier.