
Open-Weight AI Is Having Its Kubernetes Moment — Here's What Your Fleet Needs to Serve It
Open-weight models are becoming the default infrastructure layer for AI. This post maps the standardized Kubernetes inference stack — OCI weights, DRA, InferencePool, KServe, llm-d — and shows what a self-hosted fleet needs to serve models next to tenant apps.

Kubernetes' Agent Sandbox Goes Upstream: Inside the gVisor-Isolated Sandbox CRD Behind 16x GKE Growth
The SIG Apps Sandbox API hit v1beta1 after 16x GKE growth in five months, with Langchain and Lovable running millions of agents on it. Here is what a self-hosted PaaS gets for free — warm pools, pod snapshots, pluggable gVisor/Kata isolation — and what it no longer needs to build.

66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS
CNCF survey data shows 66% of organizations run AI inference on Kubernetes while only 7% deploy models daily. Four mechanisms — versioned model artifacts, DRA accelerator requests, eval-gated promotion, and correlated inference telemetry — close the gap between running a model and shipping one.

SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server
SGLang now serves 400,000+ production GPUs, making it a real second default next to vLLM. Benchmarks, workload maps, and a decision matrix for choosing the model server on your own GPU pool.

Stop Letting GPU Requests Pend: Kubernetes 1.36's DRA Prioritized Alternatives
Exact-match GPU requests leave pods pending while other card types sit idle. Kubernetes 1.36's stable DRA prioritized lists let one claim say H100, A100, or T4 — here is the claim-template design, the scoring behavior, and four gotchas for mixed-card fleets.

Stop Handing Agents Immortal Keys: Short-Lived Sandbox Credentials with Kubernetes 1.37 Pod Certificates
Kubernetes 1.37 graduates Pod Certificates and Cluster Trust Bundles to stable, replacing copyable bearer tokens with short-lived X.509 sandbox identity. Here is the pod-spec design, the migration off immortal service-account secrets, and the audit evidence to gather before agents deploy on their own.

gVisor vs Kata vs Firecracker: Picking Sandbox Isolation for an Agent Layer on Shared Nodes
An agent sandbox on a node shared with paying tenants must survive hostile model-generated code. This concrete comparison of gVisor, Kata Containers, and Firecracker covers cold starts, memory per sandbox, and blast-radius containment — plus a decision rule for self-hosted fleets.

Daytona and E2B Hit $0.05/vCPU-Hour Parity: What Sandbox Metering Really Costs a Deploy-From-Chat Platform
Daytona and E2B both charge $0.0504 per vCPU-hour, but the rate is the least important number on the invoice. A worked cost model — build plus preview dwell, plan fees, idle policies, and owned Hetzner capacity — shows where renting wins and where owning the base pays.

OpenAI Buys Ona (Formerly Gitpod) for Its Cloud Sandboxes, Not Its Dev Environments
OpenAI's acquisition of Ona (formerly Gitpod) confirms that agent sandboxes, not dev environments, are the durable product of the cloud-CDE era — and narrows the self-hostable path for agent compute. A concrete comparison of vendor-owned versus self-hosted execution.