Blog
Insights, analysis, and updates from the AI agent economy. Browse by tag · Browse the archive.

Your Fleet Patches Weekly but Reboots Never: Kured Closes the Stale-Kernel Gap
Unattended upgrades patch the disk, but the running kernel stays old until something reboots the node. Kured automates the cordon-drain-reboot cycle one node at a time — here is the runbook it replaces and the four guardrails (MHC timeouts, PDBs, drain timeouts, alert gates) that make it safe on multi-tenant CAPI fleets.

One Scheduler Per Cluster Isn't a Fleet Strategy: How KubeStellar, Cluster API, and Hive Split the Multi-Cluster Job
kube-scheduler goes blind the day you add a second cluster. How Cluster API provisioning, KubeStellar placement policy, and hive agent orchestration split the multi-cluster job — and when placement belongs in your platform versus tenant BYO.

Kubernetes v1.37 Node Lifecycle Conditions: Machine-Readable Node State for Agent-Driven Fleet Ops
Kubernetes v1.37 adds five Node Lifecycle Conditions — DrainInProgress, Drained, MaintenancePlanned, MaintenanceInProgress, and GracefulNodeShutdownInProgress — so nodes report intent, not just readiness. Here is who should publish each one on a Cluster API fleet, how to rewrite NotReady paging, and the agent policy that turns machine-readable state into safe machine operators.

66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS
CNCF survey data shows 66% of organizations run AI inference on Kubernetes while only 7% deploy models daily. Four mechanisms — versioned model artifacts, DRA accelerator requests, eval-gated promotion, and correlated inference telemetry — close the gap between running a model and shipping one.

Kubernetes' Agent Sandbox Goes Upstream: Inside the gVisor-Isolated Sandbox CRD Behind 16x GKE Growth
The SIG Apps Sandbox API hit v1beta1 after 16x GKE growth in five months, with Langchain and Lovable running millions of agents on it. Here is what a self-hosted PaaS gets for free — warm pools, pod snapshots, pluggable gVisor/Kata isolation — and what it no longer needs to build.

Your Build Cache Is Invisible to Kubelet: Sizing Image GC So Tenant Builds Stop Evicting Tenant Pods
Tenant builds fill shared nodes with cache that kubelet image GC can never reclaim — until DiskPressure evicts running pods. A measured method for budgeting per-node disk, with a worked 80GB example and copy-paste kubelet plus BuildKit configs.

One Changelog Line Killed Three PaaS Deploys: What the Kamal Receipt Proves — and Where It Stops
A Rails boilerplate's changelog replaced Fly.io, Render, and Heroku deploys with Kamal in one commit — nine files, one VPS invoice. Here is the itemized receipt: what one deploy.yml covers, what Kamal can't do, and where the move sits on the ladder to a declarative fleet.

Your Hoster Now Pays for Your Panel: What is*hosting's Coolify Sponsorship Changes — and Where the One-Click VPS Still Stops
is*hosting became an official Coolify sponsor and shipped a pre-installed Coolify VPS image from $10.19/month. The money math of a hoster funding the panel it distributes — and why the bundle still stops at box number two.

HPA Scales to Zero Now. Should Your Fleet Retire KEDA or Knative?
Kubernetes 1.37 graduates HPA scale-to-zero to beta and turns it on by default. Queue workers can consolidate onto the native primitive — but HTTP services still need Knative's request buffering and event-diverse fleets still need KEDA's scaler catalog. Here is the per-workload decision.
Subscribe
New posts land in your reader as soon as they publish. Pick a format — all three carry the same posts.
Current feeds keep roughly two days of posts so daily polling does not miss a burst. Older entries stay reachable from the feed's next-page link in readers that follow it, or from the blog archive.
Following one topic instead? Browse tags