Skip to main content

694 posts tagged with "Self-Hosting"

Running your own PaaS and infrastructure on machines you own

View all tags

Read the Self-hosted PaaS guide

Kubernetes 1.36's Mixed Version Proxy Beta: What It Fixes for maxSurge:0 Fleets
·Dora Noda·8 min

Kubernetes 1.36's Mixed Version Proxy Beta: What It Fixes for maxSurge:0 Fleets

Kubernetes 1.36 graduates Mixed Version Proxy to Beta, closing the 404-on-your-own-CRD gap during rolling control-plane upgrades. Here's what changes for a maxSurge:0 fleet on bare metal, and the peer-auth flags you still have to configure yourself.

self-hosting
PaaS
infrastructure
engineering
Kubernetes 1.36's PSI Metrics Go GA (Memory QoS Tiering Stays Alpha): What It Takes to Bin-Pack a Fixed Fleet
·Dora Noda·11 min

Kubernetes 1.36's PSI Metrics Go GA (Memory QoS Tiering Stays Alpha): What It Takes to Bin-Pack a Fixed Fleet

Kubernetes 1.36 makes kubelet PSI pressure metrics GA but leaves Memory QoS tiering in Alpha — here's exactly what wiring real per-node stall data into scheduler placement takes when your fleet is four owned boxes, not an autoscaler.

infrastructure
self-hosting
PaaS
engineering
Your MCP Traffic and Your Tenants' LLM Calls Just Got a Real Kubernetes Primitive
·Dora Noda·9 min

Your MCP Traffic and Your Tenants' LLM Calls Just Got a Real Kubernetes Primitive

Kubernetes' new AI Gateway Working Group is formalizing token-based rate limiting, model-aware routing, and provider failover as Gateway API primitives. Here's what's already GA, what's still a proposal, and what it means for a platform routing both tenant and agent MCP traffic through one ingress layer.

infrastructure
self-hosting
PaaS
AI agents
+1
Kubernetes' Pod Checkpoint/Restore Is Headed for Alpha in v1.37 — What That Means for Preview Environments Right Now
·Dora Noda·9 min

Kubernetes' Pod Checkpoint/Restore Is Headed for Alpha in v1.37 — What That Means for Preview Environments Right Now

Kubernetes' new Checkpoint/Restore Working Group has pod-level restore alpha-targeted for v1.37 — but the CRIU mechanism already runs in production via a third-party shim. Here's what's actually shipped, what's still missing, and the real blockers for preview-environment resume.

self-hosting
PaaS
infrastructure
cost-optimization
Kubernetes Has Scheduled GPUs as Integers Since 2017 — What KAI Scheduler and Grove Actually Change for a Fleet Without Multi-GPU Nodes
·Dora Noda·10 min

Kubernetes Has Scheduled GPUs as Integers Since 2017 — What KAI Scheduler and Grove Actually Change for a Fleet Without Multi-GPU Nodes

Kubernetes has scheduled GPUs as opaque integers since 2017 — here's what KAI Scheduler and Grove's fragmentation-aware bin-packing actually fix, and why the classic multi-GPU fragmentation story doesn't apply to a fleet built on single-GPU Hetzner nodes.

self-hosting
PaaS
infrastructure
AI
+1
CVE-2026-33814: A Single Zero in One HTTP/2 Field Can Hang Every Go Client in Your Kubernetes Fleet
·Dora Noda·8 min

CVE-2026-33814: A Single Zero in One HTTP/2 Field Can Hang Every Go Client in Your Kubernetes Fleet

A malformed SETTINGS_MAX_FRAME_SIZE value can hang any unpatched Go HTTP/2 client forever — and in Kubernetes, that's the apiserver, the kubelet, and every Cluster API provider controller. Here's the mechanism, the exposed-component map, and the govulncheck commands to audit your own fleet.

security
self-hosting
PaaS
engineering
Kubernetes 1.35 Takes In-Place Pod Resize to GA: What the Restart-Based Workaround It Just Killed Was Actually Costing You
·Dora Noda·8 min

Kubernetes 1.35 Takes In-Place Pod Resize to GA: What the Restart-Based Workaround It Just Killed Was Actually Costing You

In-place pod resize reached GA in Kubernetes 1.35. Here's the before/after: what evicting and rescheduling a pod to change its CPU/memory actually cost versus a resize subresource PATCH that completes in seconds — and what a Cluster-API-managed PaaS's own autoscaling logic should do differently now.

self-hosting
PaaS
infrastructure
scalability
+1
Kubernetes' New Node Readiness Controller: Closing the Gap Between 'Node Joined' and 'Node Is Safe' on a Bare-Metal Fleet
·Dora Noda·10 min

Kubernetes' New Node Readiness Controller: Closing the Gap Between 'Node Joined' and 'Node Is Safe' on a Bare-Metal Fleet

A new kubernetes-sigs controller lets operators declare exactly which conditions a node must meet before it's schedulable — here's how it works, and exactly where the gap it closes shows up in a Cluster API Hetzner fleet's own bootstrap sequence.

PaaS
self-hosting
infrastructure
cloud computing
Let's Encrypt's 2.5-Hour Outage Broke Live Renewals: A Real Fallback-CA Design for Self-Hosted ACME
·Dora Noda·12 min

Let's Encrypt's 2.5-Hour Outage Broke Live Renewals: A Real Fallback-CA Design for Self-Hosted ACME

Let's Encrypt's May 2026 outage lasted 2.5 hours and still broke live renewals. The renewal-buffer math showing why short-lived certificates make it worse, and a concrete two-issuer failover design for self-hosted ACME automation.

self-hosting
PaaS
security
infrastructure
+1
Showing 550–558 of 694 posts
Prev62 / 78Next