Skip to main content

437 posts tagged with "AI"

Artificial intelligence and machine learning applications

View all tags

Open-Weight AI Is Having Its Kubernetes Moment — Here's What Your Fleet Needs to Serve It
·Dora Noda·10 min

Open-Weight AI Is Having Its Kubernetes Moment — Here's What Your Fleet Needs to Serve It

Open-weight models are becoming the default infrastructure layer for AI. This post maps the standardized Kubernetes inference stack — OCI weights, DRA, InferencePool, KServe, llm-d — and shows what a self-hosted fleet needs to serve models next to tenant apps.

AI
LLM
Kubernetes
self-hosting
Kubernetes' Agent Sandbox Goes Upstream: Inside the gVisor-Isolated Sandbox CRD Behind 16x GKE Growth
·Dora Noda·9 min

Kubernetes' Agent Sandbox Goes Upstream: Inside the gVisor-Isolated Sandbox CRD Behind 16x GKE Growth

The SIG Apps Sandbox API hit v1beta1 after 16x GKE growth in five months, with Langchain and Lovable running millions of agents on it. Here is what a self-hosted PaaS gets for free — warm pools, pod snapshots, pluggable gVisor/Kata isolation — and what it no longer needs to build.

Kubernetes
AI
self-hosting
security
66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS
·Dora Noda·11 min

66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS

CNCF survey data shows 66% of organizations run AI inference on Kubernetes while only 7% deploy models daily. Four mechanisms — versioned model artifacts, DRA accelerator requests, eval-gated promotion, and correlated inference telemetry — close the gap between running a model and shipping one.

Kubernetes
PaaS
AI
LLM
+1
SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server
·Dora Noda·10 min

SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server

SGLang now serves 400,000+ production GPUs, making it a real second default next to vLLM. Benchmarks, workload maps, and a decision matrix for choosing the model server on your own GPU pool.

self-hosting
AI
LLM
Kubernetes
+1
Stop Letting GPU Requests Pend: Kubernetes 1.36's DRA Prioritized Alternatives
·Dora Noda·9 min

Stop Letting GPU Requests Pend: Kubernetes 1.36's DRA Prioritized Alternatives

Exact-match GPU requests leave pods pending while other card types sit idle. Kubernetes 1.36's stable DRA prioritized lists let one claim say H100, A100, or T4 — here is the claim-template design, the scoring behavior, and four gotchas for mixed-card fleets.

self-hosting
PaaS
Kubernetes
AI
Stop Handing Agents Immortal Keys: Short-Lived Sandbox Credentials with Kubernetes 1.37 Pod Certificates
·Dora Noda·9 min

Stop Handing Agents Immortal Keys: Short-Lived Sandbox Credentials with Kubernetes 1.37 Pod Certificates

Kubernetes 1.37 graduates Pod Certificates and Cluster Trust Bundles to stable, replacing copyable bearer tokens with short-lived X.509 sandbox identity. Here is the pod-spec design, the migration off immortal service-account secrets, and the audit evidence to gather before agents deploy on their own.

Kubernetes
security
AI
self-hosting
gVisor vs Kata vs Firecracker: Picking Sandbox Isolation for an Agent Layer on Shared Nodes
·Dora Noda·13 min

gVisor vs Kata vs Firecracker: Picking Sandbox Isolation for an Agent Layer on Shared Nodes

An agent sandbox on a node shared with paying tenants must survive hostile model-generated code. This concrete comparison of gVisor, Kata Containers, and Firecracker covers cold starts, memory per sandbox, and blast-radius containment — plus a decision rule for self-hosted fleets.

AI
security
infrastructure
self-hosting
+1
Daytona and E2B Hit $0.05/vCPU-Hour Parity: What Sandbox Metering Really Costs a Deploy-From-Chat Platform
·Dora Noda·11 min

Daytona and E2B Hit $0.05/vCPU-Hour Parity: What Sandbox Metering Really Costs a Deploy-From-Chat Platform

Daytona and E2B both charge $0.0504 per vCPU-hour, but the rate is the least important number on the invoice. A worked cost model — build plus preview dwell, plan fees, idle policies, and owned Hetzner capacity — shows where renting wins and where owning the base pays.

AI
self-hosting
PaaS
cost-optimization
OpenAI Buys Ona (Formerly Gitpod) for Its Cloud Sandboxes, Not Its Dev Environments
·Dora Noda·10 min

OpenAI Buys Ona (Formerly Gitpod) for Its Cloud Sandboxes, Not Its Dev Environments

OpenAI's acquisition of Ona (formerly Gitpod) confirms that agent sandboxes, not dev environments, are the durable product of the cloud-CDE era — and narrows the self-hostable path for agent compute. A concrete comparison of vendor-owned versus self-hosted execution.

openai
acquisitions
Codex
self-hosting
+1
Showing 1–9 of 437 posts