
Render's $1,500 Agent Tier vs $100 of Owned Iron: Pricing 96 GB Per vCPU-GB
Render's August 26 overhaul added a 12-CPU/96 GB agent-platform tier at $1,500/month. A per-vCPU-GB recompute against matched Hetzner dedicated capacity lands at a 15–20x multiple — plus where memory-heavy agent loops earn the ratio and where they pay idle-RAM premium.

Karmada Just Graduated: What Multi-Cluster Kubernetes Maturity Means for Your Single-Region Fleet
Karmada graduated from the CNCF on September 8, 2026, with v1.19's AI-training scheduling in tow. For a Hetzner-only Cluster API fleet, the move is to borrow its propagation and failover patterns now and adopt the control plane when a second region arrives.

A Mini-PC With 128GB of Unified Memory Serves MoE Models at 63-97 Tok/s: What Strix Halo Does to Self-Hosted Inference Economics
A 128GB Strix Halo mini-PC serves MoE models at 63-97 tokens per second. The rent-vs-own math for agent-serving workloads — and why the box wins on privacy, not token arbitrage.

Istio's Agentgateway Gambit: The Service Mesh Is Coming for Your Agent Traffic
Istio's experimental agentgateway support puts MCP tool calls, agent-to-agent sessions, and LLM routing inside the service mesh. What the KubeCon EU 2026 announcement changes for platform teams, and how close it is to production-ready.

SEP-835 Scopes: Least-Privilege MCP Tokens for Agents That Hold Deploy Keys
MCP's SEP-835 adds native per-tool scopes and RFC 8707 audience binding, so a status-check agent holds read-only scopes while deploy and rollback stay behind explicit step-up — and a compromised tool call fails closed instead of shipping an unauthorized deploy.

60 Million AI Code Reviews Later: The CI Gate Your Self-Hosted PaaS Should Steal
GitHub processed 60M+ AI code reviews in 2026 while AI-written code shipped roughly 1.7x more defects per pull request. Here is the multi-agent review pattern behind the numbers, and how to wire it into a self-hosted PaaS pipeline as a proper CI gate.

Seven Months to Self-Host an AI Stack That Was Supposed to Take a Weekend: What OpenMake's Post-Mortem Bills Against a PaaS
OpenMake's seven-month self-hosting retrospective comes with receipts: four changelog bugs and a Node-plus-Postgres-plus-Docker stack. A worked Year 1 comparison shows VPS savings evaporating after two days of engineering time — and why Kubernetes wouldn't have shortened the seven months, but a PaaS deploy surface would have.

TurboFieldfare Fits a 26B MoE Model in 2GB of RAM: What SSD-Streamed Experts Mean for Your Cheapest Inference Node
TurboFieldfare runs Gemma 4 26B-A4B in ~2GB of RAM by streaming 4-bit experts from SSD — 5–6 tok/s on an M2 Air, 31–35 on an M5 Pro. The per-token byte math behind those numbers, and how to size a €49 NVMe box as your cheapest inference tier.

Borrow the Claim, Skip the Cloud Sprawl: What Crossplane's Modelplane Teaches a Self-Hosted PaaS About GPU APIs
Modelplane puts GPU inference fleets behind a Crossplane claim API so teams request an endpoint, not a machine. The before/after YAML, a six-row borrow-vs-skip table for a Hetzner CAPI fleet, and the €184/mo card math that decides it.