Skip to main content

26 posts tagged with "LLM"

Articles about large language models

View all tags

Only LLM Spans Should Cost Money: Span-Class-Aware Billing for Agent Observability
·Dora Noda·9 min

Only LLM Spans Should Cost Money: Span-Class-Aware Billing for Agent Observability

Span-class-aware metering prices LLM spans while tracing tool calls free — the worked cost math, which vendors already do it, and an OTel Collector recipe for self-hosted agent observability.

AI agents
LLM
Model Context Protocol
self-hosting
+1
Open-Weight AI Is Having Its Kubernetes Moment — Here's What Your Fleet Needs to Serve It
·Dora Noda·10 min

Open-Weight AI Is Having Its Kubernetes Moment — Here's What Your Fleet Needs to Serve It

Open-weight models are becoming the default infrastructure layer for AI. This post maps the standardized Kubernetes inference stack — OCI weights, DRA, InferencePool, KServe, llm-d — and shows what a self-hosted fleet needs to serve models next to tenant apps.

AI
LLM
Kubernetes
self-hosting
Pangolin Puts SSO and WireGuard in Front of LLM Access Instead of API Keys: What Tunnel-Based Identity Means for Agent Credential Hygiene
·Dora Noda·10 min

Pangolin Puts SSO and WireGuard in Front of LLM Access Instead of API Keys: What Tunnel-Based Identity Means for Agent Credential Hygiene

Pangolin's AI Gateway authenticates LLM access with SSO-backed WireGuard tunnels instead of static API keys, and joins self-hosted models to the same gateway as public ones. Here's how the mechanism works, how it compares to Tailscale Aperture, and what it changes in your threat model.

security
AI agents
LLM
self-hosting
+1
66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS
·Dora Noda·11 min

66% Run AI Inference on Kubernetes, 7% Deploy Daily: The Golden Path to an AI-Ready PaaS

CNCF survey data shows 66% of organizations run AI inference on Kubernetes while only 7% deploy models daily. Four mechanisms — versioned model artifacts, DRA accelerator requests, eval-gated promotion, and correlated inference telemetry — close the gap between running a model and shipping one.

Kubernetes
PaaS
AI
LLM
+1
SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server
·Dora Noda·10 min

SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server

SGLang now serves 400,000+ production GPUs, making it a real second default next to vLLM. Benchmarks, workload maps, and a decision matrix for choosing the model server on your own GPU pool.

self-hosting
AI
LLM
Kubernetes
+1
Envoy AI Gateway Hits 1.0: What an LLM-Aware Kubernetes Gateway Actually Load-Balances On
·Dora Noda·10 min

Envoy AI Gateway Hits 1.0: What an LLM-Aware Kubernetes Gateway Actually Load-Balances On

Envoy AI Gateway reached 1.0 as the standard routing layer for self-hosted LLM traffic on Kubernetes. How InferencePool endpoint selection on KV-cache state and queue depth beats round-robin — with benchmarks — and what it means for fleets you run yourself.

Kubernetes
self-hosting
LLM
infrastructure
Hetzner's €214 GEX45: What 24 GB of Blackwell VRAM Buys Self-Hosted Inference (and When It Beats Serverless GPUs)
·Dora Noda·10 min

Hetzner's €214 GEX45: What 24 GB of Blackwell VRAM Buys Self-Hosted Inference (and When It Beats Serverless GPUs)

Hetzner's €214/month GEX45 pairs an RTX PRO 4000 Blackwell card with 24 GB of VRAM. The crossover math against serverless GPU rentals, and which models newly fit.

self-hosting
cost-optimization
LLM
PaaS
llm-d Is Now a CNCF Sandbox Project: What Kubernetes-Native Distributed Inference Means for Self-Hosting Your AI Agents' Brains
·Dora Noda·9 min

llm-d Is Now a CNCF Sandbox Project: What Kubernetes-Native Distributed Inference Means for Self-Hosting Your AI Agents' Brains

llm-d brings distributed LLM inference — prefix-cache-aware routing and prefill/decode disaggregation — to Kubernetes as a CNCF Sandbox project. What it means for self-hosting AI-agent inference, what it costs versus hosted APIs, and how early the project still is.

Kubernetes
self-hosting
AI agents
LLM
+1
TurboFieldfare Fits a 26B MoE Model in 2GB of RAM: What SSD-Streamed Experts Mean for Your Cheapest Inference Node
·Dora Noda·10 min

TurboFieldfare Fits a 26B MoE Model in 2GB of RAM: What SSD-Streamed Experts Mean for Your Cheapest Inference Node

TurboFieldfare runs Gemma 4 26B-A4B in ~2GB of RAM by streaming 4-bit experts from SSD — 5–6 tok/s on an M2 Air, 31–35 on an M5 Pro. The per-token byte math behind those numbers, and how to size a €49 NVMe box as your cheapest inference tier.

self-hosting
AI
LLM
cost-optimization
+1
Showing 1–9 of 26 posts