A typical deploy-agent run makes 2 model calls and 30 tool calls. Under flat per-span tracing bills, those 30 tool spans cost 15x what the 2 model calls cost — even though the model calls are the only spans that explain what the agent spent. That meter taxes exactly the behavior you want: agents that check state, read logs, and verify before they act.
The fix is span-class-aware metering: meter the LLM spans, trace everything else free. It is not hypothetical. Datadog's Agent Observability already bills LLM spans only — every other span in the trace rides free — and Langfuse's cost engine prices only generation observations while tool and agent spans stay unpriced by design. This post works the math for a realistic agent workload, shows where class-aware metering wins (and where it loses), and gives you the OpenTelemetry Collector recipe to build the same distinction on infrastructure you already run.
The number first: a deploy agent's trace under two meters
Take a concrete workload: an agent that deploys apps through MCP-exposed tooling. Each run calls the model twice (plan, then verify) and calls tools ~30 times (read logs, check rollout status, describe resources, run health checks). At 50,000 runs a month, that is 100,000 LLM spans and 1.6M spans total — a 16:1 ratio of traced spans to model calls.
Now price it two ways, using published list rates as illustrative inputs: a flat meter at $8 per 100K spans (Langfuse Cloud Core's overage rate on its units model of traces + observations + scores) versus an LLM-span-only meter at $3.50 per 10K LLM spans (Datadog Agent Observability's list overage past the included 100K):
| Workload shape (per run) | Total spans/mo | LLM spans/mo | Flat meter | LLM-only meter |
|---|---|---|---|---|
| Bare chat: 1 LLM + 1 other | 100K | 50K | ~$8 | ~$18 |
| Balanced: 2 LLM + 8 other | 500K | 100K | ~$40 | ~$35 |
| Deploy agent: 2 LLM + 30 other | 1.6M | 100K | ~$128 | ~$35 |
Three things fall out of this table. First, the breakeven sits near 4 non-LLM spans per model call — above that ratio, class-aware metering wins, and a 16:1 agentic workload traces at roughly a quarter of the flat-meter price. Second, the counter-case is real: a bare chat workload with almost no tool calls is cheaper under flat billing, because the LLM-span rate is higher per unit. Anyone selling you span-class metering without showing that crossover is cherry-picking. Third, the flat meter creates a perverse incentive: every tool call your agent makes to be careful — re-checking rollout status instead of assuming — raises your observability bill. The meter should reward verification, not tax it.
That third point is the actual argument for "only LLM spans should cost money." It is not that tool spans are cheap to store — a span is a span on disk. It is that the meter should align with the scarce resource (model spend, which teams already budget per agent) instead of punishing the agentic pattern itself.
Note: the two rates come from different vendors' published lists (Langfuse's unit overages vs Datadog's LLM-span overages), so treat the table as directional arithmetic for the ratio effect, not a vendor shootout. The crossover math — breakeven near 4:1 — is what survives any particular price.
Who already meters this way
Datadog bills LLM spans only. Agent Observability (the current name for its LLM Observability product) meters one thing: LLM inference spans, where a span is one call to a model provider. List pricing runs about $160/month for the first 100K spans on the annual plan, then ~$3.50 per additional 10K, with a free tier around 40K LLM spans a month. Everything else in the trace — tool calls, retrieval, orchestration — is captured without touching the meter. That is the commercial proof that span-class billing exists at scale, not just as a blog-post idea.
Langfuse prices generations, not spans — for model cost. Here the distinction matters and most summaries blur it. Langfuse Cloud's platform billing counts every observation as a unit (monthly units = traces + observations + scores, from $29/month for 100K units). But its model-cost attribution — the dashboard that answers "which agent ran up the bill" — attaches cost only to generation and embedding observations. Plain spans carry no cost; tool and agent spans are unpriced by design. The project takes this classification seriously enough to ship fixes like #14808, which stopped Vercel AI SDK invoke_agent spans from double-counting model-call usage as generations. So the honest statement is: Langfuse meters all observations for platform billing, but its cost model already embodies the span-class distinction — only model-touching observations carry money.
Phoenix charges nothing and supplies the vocabulary. Arize Phoenix is free to self-host with no event caps (source-available under Elastic License 2.0, which permits internal self-hosting but not offering it as a managed service). There is no meter to be class-aware — but Phoenix's OpenInference semantic conventions define the span-kind taxonomy the whole ecosystem classifies with: openinference.span.kind takes values like LLM, EMBEDDING, TOOL, AGENT, CHAIN, RETRIEVER, RERANKER, GUARDRAIL, and EVALUATOR. That attribute is the field your own meter keys on.
| Platform | What the meter counts | Tool/agent spans |
|---|---|---|
| Datadog Agent Observability | LLM spans only | Free |
| Langfuse Cloud billing | All observations (units) | Billed as units |
| Langfuse cost attribution | Generations/embeddings only | Unpriced by design |
| Phoenix (self-hosted) | Nothing — no meter | Free (you pay infra) |
The vocabulary that makes it possible: span kinds
Every span-class scheme needs one reliable classifier, and in 2026 there are two that interoperate. OpenInference's openinference.span.kind is the agent-native one, emitted by LangChain/LlamaIndex/Vercel AI SDK instrumentations and consumed directly by Phoenix. The OTel-native one is the GenAI semantic convention's gen_ai.operation.name (chat, execute_tool, invoke_agent, …), which the OpenInference project maps onto span kinds: chat/text_completion/generate_content become LLM, execute_tool becomes TOOL, invoke_agent/create_agent become AGENT. Datadog added native support for the OTel GenAI conventions in late 2025, so OTel-instrumented agents can feed a class-aware meter without proprietary SDKs.
A meter, then, is a one-predicate function: does this span's kind (or operation name) indicate a model call? Everything downstream — cost attribution, sampling policy, retention tier — hangs off that predicate.
Build it on the Collector you already run
If you already operate an OTel Collector for tenant workloads, you do not need a second tracing product for agent-ops. You need a routing rule. The sketch below splits one incoming agent trace stream into two pipelines by span kind: LLM spans go to the metered/cost-attributed pipeline (your Langfuse, your cost rollup, your short-retention hot store), while tool/agent/chain spans go to cheap long-retention storage (a generic Tempo/ClickHouse backend you already pay for by the byte, not the span):
# Sketch: route by span class inside the Collector you already run.
connectors:
routing/agent_spans:
match_once: true
table:
- conditions:
- attributes["openinference.span.kind"] == "LLM"
pipelines: [traces/llm_metered]
default_pipelines: [traces/tool_cheap]Falling back across conventions costs one more condition: match gen_ai.operation.name in [chat, text_completion, generate_content, embeddings] as LLM-equivalent for instrumentations that emit GenAI conventions instead of OpenInference kinds. The metered pipeline can then apply strict sampling or per-project accounting, while the tool pipeline keeps everything at raw fidelity — because re-checking rollout status thirty times should be free, and now it is.
One gotcha to carry over from Langfuse's experience: agent frameworks love to duplicate usage. The Vercel AI SDK's invoke_agent span once carried aggregate token counts that duplicated the per-call usage on the grandchild model-call span, doubling every trace's cost until Langfuse reclassified it. Your routing rule should therefore prefer the leaf model-call span (gen_ai.operation.name == "chat" on the span that actually holds the provider response) as the single cost carrier, and treat aggregate parent spans as orchestration. Dedupe by construction, not by dashboard filter.
Why extend the Collector instead of buying agent-observability SaaS? Three reasons: the Collector already sees every span (no second agent to deploy), it already speaks both classification dialects (attribute routing, no code changes), and the marginal cost of the tool-span pipeline is object storage, not a per-span meter. Langfuse self-hosted, for reference, is a six-container stack (web, worker, Postgres, ClickHouse, Redis, S3-compatible store) — very runnable, but it is still a second system to operate. The Collector rule is a config change.
Applied: metering MCP deploy/rollback tooling
Now the concrete application. Suppose your PaaS exposes deploy, rollback, log-tail, and status probes as MCP tools, and tenants point their own agent frameworks at them. Each tenant agent session produces exactly the 16:1 trace shape from the opening table: a few model calls the tenant's framework makes, dozens of tool calls against your platform.
With span-class metering in your Collector, the policy writes itself: trace every tool call at full fidelity for free (your debugging lifeline when a deploy goes sideways), and attribute cost only to the model-call spans. If you ever charge for agent-ops, the meter reads "model calls per tenant," a number tenants already understand from their own provider bills — not "spans," a number that punishes them for your tools being chatty. And because the classifier is a standard attribute, a tenant bringing LangChain, the Vercel AI SDK, or raw OTel GenAI instrumentation all lands in the same two buckets without per-framework integration work.
The meter is the message
Observability pricing shapes agent behavior. A flat per-span meter quietly tells agent builders to make fewer tool calls — to guess instead of verify. An LLM-span meter tells them the opposite: check state as often as you like; the model calls are what count. Datadog proved the model commercially, Langfuse proved the classification in its cost engine, and Phoenix gave the ecosystem the vocabulary for free. The remaining step is operational, not theoretical: one routing rule in the Collector, keyed on span kind, turning the trace shape agentic workloads already produce into a bill that finally makes sense.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with APIs and MCP tooling designed for agents as first-class operators. Star the repo on GitHub or deploy your first app today.



