Blog
Insights, analysis, and updates from the AI agent economy. Browse by tag · Browse the archive.

50 Pull Requests, 50 Live Environments: What Preview-at-Scale Really Costs on Railway vs Owned Hardware
Fifty concurrent pull-request previews cost about $585 a month on Railway's per-minute meter versus $30-95 on a Hetzner-backed Cluster-API fleet. A worked comparison of per-namespace accounting, image-cache sharing, Gateway API route churn, and the probe design that decides whether idle previews actually sleep.

MCP Went Stateless: Migrating a Deploy-and-Rollback Tool Server Behind an Ordinary Load Balancer
MCP's July 28, 2026 spec removes the handshake and Mcp-Session-Id, making every request self-describing. An 8-step migration plan for a deploy/logs/rollback tool server to replicas behind a plain load balancer — with per-request auth and an audit trail that survives the death of the session.

Kratix Promises Are Platform-as-Product: What Workflow Retries and Compound Promises Mean for a Fleet That Hand-Rolls Operators
A July 2026 hands-on replaced an entire ops ticket queue with three Kratix promises serving in 30 seconds. How the promise model works, what SKE v0.45.0's workflow retries changed, what compound promises buy a Cluster API fleet's preview environments and quotas, and the threshold where promises beat shell scripts.

Fly.io Started Billing What Used to Be Free: The Two 2026 Line Items, Priced
Fly.io began metering volume snapshots in January 2026 and cross-region private traffic in February 2026. A line-by-line recompute shows the two new charges adding roughly $9 to a $33 three-region bill, against $15 flat on owned Hetzner hardware.

E2B Serves 88% of the Fortune 100. What Breaks at That Scale Is Idle Billing and Resume Latency
A 30-minute agent session costed three ways shows session-scoped sandbox billing charges mostly for model think time, while suspend/resume latency stacks across dozens of tool calls. What a self-hosted PaaS should design differently.

Durable Objects Just Escaped Cloudflare Twice: Deno's Celld vs Rivet Actors
Deno's celld runs Workers and Durable Objects on your own machines with per-cell SQLite and S3 compare-and-swap; Rivet Actors offers a broader open-source actor platform on Postgres or FoundationDB. A head-to-head on placement, upgrades, write fencing, and what a Kubernetes PaaS must provide to host either beside stateless services.

Five Subsystems, Sixteen Days: Vercel's July 2026 Incident Cluster, Counted
Between July 8 and July 23, 2026, five separate Vercel subsystems failed — builds, SSO login, GitHub deploys, the dashboard, and telemetry export. Counting the whole cluster shows the real background failure rate of shared platforms, and why a six-minute Log Drains gap marked unrecoverable can't happen when your logs never cross a third-party pipe.

France's Railway Runs Kubernetes on Cluster API: What a National Railway's Declarative Rebuild Teaches a Two-Person Fleet Team
SNCF cut cluster provisioning from a month to 30 minutes and now updates every cluster monthly with Cluster API on its own hardware. Three lessons transfer directly to a two-person team on Hetzner — and four pieces of enterprise ceremony to skip.

Render's 75% Faster Builds, Decoded: What Native Environments Really Measure and the Bar for Self-Hosted Pipelines
Render's native environments claim 75% faster builds than Dockerfiles through intelligent layer caching. Decoding the warm-rebuild arithmetic behind the number, the three Dockerfile mistakes that hand it over, and the benchmark checklist a self-hosted BuildKit pipeline must clear to match it.
Subscribe
New posts land in your reader as soon as they publish. Pick a format — all three carry the same posts.
Current feeds keep roughly two days of posts so daily polling does not miss a burst. Older entries stay reachable from the feed's next-page link in readers that follow it, or from the blog archive.
Following one topic instead? Browse tags