Skip to main content

437 posts tagged with "AI"

Artificial intelligence and machine learning applications

View all tags

Deploying vLLM as Just Another Git-Push App: What GPU Scheduling Actually Adds
·Dora Noda·10 min

Deploying vLLM as Just Another Git-Push App: What GPU Scheduling Actually Adds

Placement mostly reduces to a manifest field and a node-pool label. Routing doesn't — a worked example against bex's actual Hetzner GPU inventory shows exactly where 'just another git-push app' breaks down for a model server.

self-hosting
PaaS
infrastructure
AI
+1
E2B vs Daytona vs Modal in 2026: Sandbox Pricing Hit Parity — Here's the Utilization Number That Decides If You Should Self-Host
·Dora Noda·10 min

E2B vs Daytona vs Modal in 2026: Sandbox Pricing Hit Parity — Here's the Utilization Number That Decides If You Should Self-Host

E2B and Daytona both landed on the exact same $0.0504/vCPU-hour price in 2026. Here's the actual utilization percentage where running your own Hetzner sandbox fleet gets cheaper than renting theirs.

AI
infrastructure
engineering
cost-optimization
+1
Fly.io Kills GPUs on August 1, 2026: Where Do Your Agent Sandboxes Go Now?
·Dora Noda·10 min

Fly.io Kills GPUs on August 1, 2026: Where Do Your Agent Sandboxes Go Now?

Fly.io fully deprecates GPU Machines on July 31, 2026. Here's the duty-cycle cost breakdown across RunPod, Modal, and a self-hosted Hetzner GPU box for the teams whose agent sandboxes and fine-tuning jobs have to move.

self-hosting
PaaS
AI
infrastructure
+1
What a Modal AI-Agent Sandbox Really Costs: The Multiplier Stack Comparison Sites Keep Getting Wrong
·Dora Noda·9 min

What a Modal AI-Agent Sandbox Really Costs: The Multiplier Stack Comparison Sites Keep Getting Wrong

Modal advertises $0.0000131 per CPU core-second, but AI-agent Sandboxes are forced non-preemptible and region-pinned, pushing the real bill to 3x-5.25x that rate — not the 3.75x figure recycled across pricing-comparison sites, which doesn't match Modal's current docs at all.

self-hosting
PaaS
infrastructure
cost-optimization
+1
The Trust Gradient Is Broken: Why AI Deploy Agents Need Revocable Capabilities, Not Permission Levels
·Dora Noda·11 min

The Trust Gradient Is Broken: Why AI Deploy Agents Need Revocable Capabilities, Not Permission Levels

A Cursor agent used a stray API token to delete a production Railway volume. New research shows why permission tiers like read-only or full-auto can't stop that — and what a capability that expires actually buys you.

AI
security
self-hosting
PaaS
Wasmer Built a Full Node.js Runtime in Two Weeks With Codex — What Edge.js Actually Buys a PaaS Over Docker
·Dora Noda·9 min

Wasmer Built a Full Node.js Runtime in Two Weeks With Codex — What Edge.js Actually Buys a PaaS Over Docker

Wasmer says Codex helped it build a full Node.js runtime in two weeks instead of a year. Here's what Edge.js's WASIX sandbox actually costs and buys a PaaS running MCP servers and agent-generated code, with real compatibility and cold-start numbers.

self-hosting
PaaS
security
AI
+1
Docker's MCP Gateway Caps Every Tool Call at 1 CPU / 2GB: The Container Security Model Your Deploy Bot Should Steal
·Dora Noda·8 min

Docker's MCP Gateway Caps Every Tool Call at 1 CPU / 2GB: The Container Security Model Your Deploy Bot Should Steal

A critical RCE in Anthropic's MCP SDK won't be patched at the protocol layer, so containment has to happen at the tool-server layer. Here's exactly what Docker's MCP Gateway locks down by default, and how to size the same model for a PaaS's own deploy/rollback/scale tools.

security
self-hosting
PaaS
AI
+1
Kubernetes 1.35's Job managedBy Field Ends the Reconciliation Fight Over Agent-Triggered Batch Work
·Dora Noda·8 min

Kubernetes 1.35's Job managedBy Field Ends the Reconciliation Fight Over Agent-Triggered Batch Work

Kubernetes 1.35 made the Job managedBy field GA, letting a platform's own controller claim a Job's status end-to-end instead of racing the built-in controller for it — here's what that means for AI agents triggering migrations and one-off tasks from chat.

self-hosting
PaaS
infrastructure
engineering
+1
The Self-Hosted GPU Breakeven Point: Why ~50M Tokens/Day Is Where Owning H100s Beats Renting Inference APIs
·Dora Noda·10 min

The Self-Hosted GPU Breakeven Point: Why ~50M Tokens/Day Is Where Owning H100s Beats Renting Inference APIs

A dedicated H100 breaks even against typical hosted LLM API pricing at roughly 50 million tokens a day of sustained use, but the real number depends entirely on which API tier you're comparing against and how well a platform can bin-pack tenant traffic to keep the GPU busy.

self-hosting
PaaS
cost-optimization
infrastructure
+1
Showing 91–99 of 437 posts