Skip to main content

437 posts tagged with "AI"

Artificial intelligence and machine learning applications

View all tags

MCP Calls the Control Plane; A2A Coordinates the Investigation
·Dora Noda·9 min

MCP Calls the Control Plane; A2A Coordinates the Investigation

A concrete design for using MCP for audited PaaS actions and A2A for agent-to-agent deployment investigations without sharing a blanket infrastructure token.

AI agents
AI
cloud infrastructure
self-hosting
Platform Engineering 2.0 Has Five Good Ideas—and an Implementation Gap
·Dora Noda·10 min

Platform Engineering 2.0 Has Five Good Ideas—and an Implementation Gap

A practical readiness matrix for turning Platform Engineering 2.0’s AI, agent, FinOps, security, and composability pillars into operating designs a Kubernetes PaaS can verify.

AI
AI agents
Kubernetes
PaaS
+1
TRM’s Co-Case Agent Turns a Crypto Investigation Into a Reviewable Workflow—Not an Autonomous Verdict
·Dora Noda·9 min

TRM’s Co-Case Agent Turns a Crypto Investigation Into a Reviewable Workflow—Not an Autonomous Verdict

TRM’s Co-Case Agent speeds crypto case preparation with tracing, graph audits, and audit logs. See the human-review workflow that turns AI assistance into defensible investigation work.

AI
compliance
cybersecurity
blockchain
Power-Bound, Not GPU-Bound: Why the Grid — Not the Chip — Is the 2026 Bottleneck
·Dora Noda·11 min

Power-Bound, Not GPU-Bound: Why the Grid — Not the Chip — Is the 2026 Bottleneck

Gartner says 40% of AI data centers hit a power wall by 2027. With 24–36-month grid queues, 700W H100s, and 140kW racks, the grid — not the GPU — is the 2026 bottleneck. Here's what that means for a CPU-only Hetzner fleet.

self-hosting
PaaS
infrastructure
cost-optimization
+1
Fork in 3ms: What Kedge's Snapshot Tree Means for the Next Agent Sandbox
·Dora Noda·15 min

Fork in 3ms: What Kedge's Snapshot Tree Means for the Next Agent Sandbox

Kedge forks a hardware-isolated microVM in 3ms from a warm-pool snapshot tree — 10,000x faster than a Cluster API Machine on Hetzner. How the architecture works, what it costs, and what a self-hosted fleet must build to match it.

infrastructure
AI
engineering
scalability
SST Put Itself in Maintenance Mode to Build an AI Agent — 650K Devs in Five Months Shows Where Infra Builders Are Betting
·Dora Noda·12 min

SST Put Itself in Maintenance Mode to Build an AI Agent — 650K Devs in Five Months Shows Where Infra Builders Are Betting

SST has been in maintenance mode since 2025 while its team's terminal agent OpenCode hit 650K developers in five months. A timeline, market math, risk matrix, and migration checklist for teams still on SST.

AI
infrastructure
PaaS
self-hosting
+1
OpenCost 1.121 Finally Answers: What Does Each Token Cost on Your Own GPUs?
·Dora Noda·13 min

OpenCost 1.121 Finally Answers: What Does Each Token Cost on Your Own GPUs?

OpenCost 1.121 adds Kubernetes-native per-token inference metering via llm-d — allocation vs. usage cost, KV-cache-corrected — turning self-hosted LLM spend from a quarterly guess into a Prometheus metric you can alert on.

AI
LLM
infrastructure
Kubernetes
+2
The Idle-Compute Tax: What Vercel's Active CPU and Netlify's Durable Functions Reveal About Serverless AI Bills
·Dora Noda·12 min

The Idle-Compute Tax: What Vercel's Active CPU and Netlify's Durable Functions Reveal About Serverless AI Bills

Vercel's Active CPU pricing and Netlify's Durable Functions both fix serverless billing for AI workloads that idle 90% of wall-clock time — a worked cost recompute shows the idle tax, the fix, and what the same workload costs on a flat Hetzner box that never metered the wait.

PaaS
AI
cost-optimization
self-hosting
+1
vLLM vs Ollama in Production: What PagedAttention's 19x Throughput Gap Really Buys on Owned GPUs
·Dora Noda·16 min

vLLM vs Ollama in Production: What PagedAttention's 19x Throughput Gap Really Buys on Owned GPUs

Red Hat's 2026 benchmark put vLLM at 793 tokens per second and Ollama at 41 on the same GPU and model — a 19x gap. The same comparison at one concurrent user shows them within 20%. A concrete breakdown of where PagedAttention earns that gap, where it doesn't, and what each engine costs to run on an owned Hetzner GPU fleet.

LLM
AI
self-hosting
Kubernetes
+2
Showing 28–36 of 437 posts