In March 2026, CNCF Ambassador Ana Margarita Medina published the sharpest four-word thesis of the agent-ops era: "AI needs APIs, not UIs." Her piece, "Crossplane and AI: The case for API-first infrastructure," argues that agents stall on today's platforms not for lack of capability but because the platform was built for humans who compensate for inconsistency — desired state in Git, actual state in the cloud console, policy in pipeline configs, and operational knowledge in the heads of engineers who eventually leave. An agent, she writes, needs a unified, structured, machine-readable interface — and Crossplane's declarative control plane is what that looks like: the agent declares intent as a Kubernetes object, and controllers reconcile reality to match.
Meanwhile, the rest of the ecosystem spent 2026 wiring agents to infrastructure the other way: through the Model Context Protocol. The containers/kubernetes-mcp-server project became the de facto standard for letting Claude, Cursor, and Copilot drive clusters through natural language, and at SUSECON 2026 in Prague, SUSE put MCP servers on both Rancher Prime and Multi-Linux Manager so agents can find faults, correlate logs, and restart services. Same goal — agents operating infrastructure — but a structurally different bet: the agent invokes discrete named tools and gets results back, one call at a time.
TL;DR for platform builders: these are two answers to one question — who owns infrastructure state, the API or the agent? Crossplane's answer is that state lives in the platform: the agent writes a desired-state object and reads convergence back from status, so a crashed or forgetful agent changes nothing about what the system does next. MCP's answer is that state lives in the conversation: the agent holds the plan in context and issues tools step by step.
The table below is the whole story; the rest of this post works the same task through both models, traces a mid-provision crash through each, and ends with what a platform's own MCP server should steal from the reconciliation loop.
| Same task: "give service X a Postgres database" | API-first (Crossplane) | Tool-call (MCP) |
|---|---|---|
| What the agent produces | One YAML object declaring the desired end state | A sequence of tool calls: create, poll, configure, verify |
| Where state lives between steps | In the object itself (spec vs status), on the API server | In the agent's context window, on the other end of each call |
| Who retries a failed step | The controller's reconcile loop, forever, unprompted | The agent — if it notices, and if its session survives |
| Who fixes drift a week later | The same controller, automatically | Nobody, until someone or something looks |
| What the agent must understand | The resource schema (one API shape for everything) | Each tool's arguments (simpler per call, more calls) |
| Governance | RBAC + admission + policy engines at the API, before anything persists | Whatever the server implements per tool (approvals, read-only modes) |
| Audit trail | The object and its event history: intent and outcome in one place | The tool-call transcript — complete only if someone kept it |
Model 1: the agent writes YAML and walks away
Here is the entire Crossplane interaction for the database task, taken from the shape Medina's piece uses:
apiVersion: example.crossplane.io/v1
kind: Database
metadata:
name: service-x-db
spec:
engine: postgres
storage: 100GiThe agent submits this object and its creative work is done. Behind the Database kind sits a Composition — the platform team's encoded knowledge of what "a Postgres database" means here: the RDS instance or Helm release, the security group, the credentials Secret, the monitoring resources. The agent never sees that expansion unless it asks; it reads progress the way every Kubernetes client does, from status: conditions flipping to Ready, the connection secret's name appearing when provisioning completes.
Three properties fall out of this that matter enormously for non-human operators. First, desired state and actual state are separated by construction — spec says what should be, status says what is, and the gap between them is inspectable and diffable by anything that can read the API. Second, convergence is somebody else's job: controllers observe the difference and reconcile continuously, so the agent doesn't orchestrate step-by-step logic, doesn't implement retry with backoff, and doesn't need to stay alive until the database exists.
Medina's line for this is the one worth quoting back at every agent-ops design review: "Without a control plane, agents become fragile orchestrators. With one, they become declarative participants." Third, governance is architectural rather than procedural. RBAC decides who can act, admission controllers validate before anything persists, OPA or Kyverno enforce constraints at runtime — the same enforcement path for every change, which means the agent operates inside boundaries the platform guarantees rather than boundaries its prompt begs it to respect.
Crossplane 2.0, released in August 2025 ahead of the project's CNCF graduation that November, widened this surface in exactly the agent-relevant direction. Compositions can now assemble any Kubernetes object, not just managed cloud resources, so a single composite API can provision infrastructure, deploy the app, configure networking, and set up observability in one object. And the new Operation types bring day-two work — scheduled upgrades, backups, maintenance — under the same declarative pattern: a CronOperation for weekly database maintenance is an API object the agent can inspect, trigger, watch, and propose edits to, instead of a runbook page it has to parse and execute by hand.
The price of all this is schema burden. The agent must understand the resource model — which kinds exist, what their fields mean, what status.conditions signal. Discovery is mechanical (kubectl api-resources style introspection works for agents too), but a model reasoning about a bespoke Database XRD is doing more conceptual work per action than one calling a tool named create_database. Crossplane's bet is that this cost is paid once per resource type and amortized across every operation forever, while the tool-call model's simplicity is paid per step, forever.
Model 2: the agent calls tools and watches results
Now the same task through a Kubernetes MCP server. The agent doesn't declare an end state; it performs a procedure, roughly:
create_database(or the lower-level resource-creation tool) with engine, size, and network arguments — and gets back an identifier or an error.get_database_status, polling until the backing resource reports ready — the agent decides the interval, the timeout, and what "stuck" means.create_secret/update_configmapto wire credentials into service X's namespace.get_events/get_logsto verify the app connected, then summarize the outcome to the user.
Nothing about this is wrong — it mirrors exactly what a human does at a terminal, which is why it works so well with today's models. Each tool is a small, legible action with typed arguments and an immediate result, and the agent can interleave investigation (read logs, describe the failing pod) with action in one fluid session. The containers/kubernetes-mcp-server project earned its de facto-standard status by sweating the details this model needs: a native Go Kubernetes client rather than a kubectl wrapper, RBAC enforcement on its tools, automatic secret redaction in outputs, a non-destructive mode for safe exploration, and eight switchable toolsets so operators expose only the surface their agents need. SUSE's 2026 move — MCP interfaces on Rancher Prime for cluster operations and Multi-Linux Manager for fleet operations — shows where the enterprise market thinks this goes: the tools your agent calls are the same operations your console performs.
But notice where state lives in the four steps above: in the agent's context window. The plan ("I created the database, now I'm waiting, next I wire credentials") exists only as conversation history. The platform sees four unrelated calls. That asymmetry is the whole critique, and it bites in three specific ways.
First, sessions are fragile: compact the context, hit a token limit, restart the harness, and the half-finished procedure evaporates — the database may exist with no record of what step comes next. Second, idempotency is the agent's problem: re-running "create the thing" after an ambiguous timeout requires the agent to check existence first, every time, for every tool, with no shared convention. Third, nothing converges afterward: six days later, when someone resizes the database by hand in a console, no loop notices or repairs it, because there is no declared desired state to drift from. The tool-call model gives the agent maximum flexibility per step and zero memory between sessions — fine for investigation, hazardous for ownership.
Same task, two bills: trace a crash through both
The comparison table at the top compresses the argument; the failure case is what makes it concrete. Take the database task and kill the operator halfway through — the agent process dies after the database exists but before credentials are wired.
In the API-first world, the crash is a non-event. The Database object still sits on the API server with spec declaring the full intent and status showing exactly how far convergence got. The controller keeps reconciling without anyone asking; when a fresh agent (or a human, or a GitOps pipeline) comes along, it reads one object and knows both what was wanted and what remains. Recovery is not a procedure to remember — it is the default behavior of the system.
In the tool-call world, the crash orphans the operation. The database exists; the transcript describing step 3 and step 4 is gone with the session. A new agent starts from "the user says service X needs a database," discovers a database-like resource already exists, and must now reconstruct intent from archaeology: is this the database from the interrupted run, or something a teammate made? Is it safe to adopt, or should it be replaced? Every answer is a guess informed by naming conventions and timestamps rather than a declared record. Multiply by dozens of agent-driven changes a week and you get the failure mode Medina's piece names precisely: autonomy stalls not from incapability but from missing structure.
This doesn't mean MCP is the wrong interface — it means purely imperative MCP is. The honest correction, and the one the ecosystem is already converging on, is to put something stateful behind the tools. That is exactly what the hybrid projects in the next section do, and it is the shape a platform's own MCP server should copy.
The synthesis is already shipping
The most interesting 2026 evidence isn't either camp's pitch — it's that builders keep bolting the two models together, each side reaching for what it lacks.
From the MCP side reaching toward declarative state: purpose-built Crossplane MCP servers now exist that are aware of the control plane rather than just proxying the Kubernetes API. They walk a claim's resourceRefs to show the agent the full resource tree its one YAML object expanded into, and diagnose routines find the deepest failing resource instead of surfacing the top-level symptom — read tools shaped around status, in other words, not around verbs. One widely referenced setup pairs this with requireApproval gates so the agent can propose reconciliation-affecting changes but cannot trigger them unprompted: declarative intent flowing through an imperative tool, with the human as the admission controller.
From the declarative side reaching toward agents: kagent, Solo.io's CNCF Sandbox framework, declares agents themselves as Kubernetes resources — Agent, ModelConfig, Memory custom resources, rolled out via GitOps like any workload — while its kmcp subproject runs MCP servers on-cluster as MCPServer objects. The agent's tools are declarative infrastructure; the agent's behavior is a reconciled object. Even the agent lifecycle gets the spec/status split. And Upbound's Modelplane, launched on top of Crossplane, extends the same control-plane pattern to AI inference itself — the thing serving the models is managed with the same declare-and-reconcile loop the models are being asked to use.
The pattern in all three: state lives in the API, and tools are how agents touch it. Nobody who has operated both keeps the plan only in the conversation.
What your platform's MCP server should steal
If you run a platform — self-hosted PaaS, internal developer platform, agent sandbox fleet — and you're building the MCP surface your agents will actually use, the verdict from this comparison is not "replace MCP with Crossplane." It's narrower and more actionable: keep the tool-call interface agents reason about so well, but back it with three borrowings from the reconciliation world.
First, write-tools should land intent and return a watchable handle, not just a result. create_database should create a record of the desired end state — a CR, a row, an intent object, anything addressable — and hand the agent an identifier it (or its successor session) can poll. The difference between "the tool returned success" and "here is the object tracking your request" is the difference between the crash scenario above being a non-event and being archaeology.
Second, shape read-tools like status reads. Agents investigate constantly; give them tools that answer "what did I ask for, how far has it converged, and what is blocking it" in one call, the way status.conditions does. The Crossplane diagnostic servers are the template: deepest-failure-first, tree-aware, no multi-hop guessing across generic getters. Every investigation your agent performs with five generic calls is context burned and a chance to misread the chain.
Third, run a reconcile loop behind the tool surface, even a dumb one. You don't need Crossplane itself to get 80 percent of the value: a controller that notices "intent object says three replicas, reality says two" and repairs it covers the drift case, the crashed-mid-provision case, and the hand-edited-in-console case alike. Tools stay imperative at the edge — agents still call verbs — but every verb lands in a system that converges rather than merely executes.
Keep as pure imperative surface exactly what reconciliation handles badly: one-shot investigation (read logs, describe state, run diagnostics), genuinely irreversible actions that should require approval rather than convergence, and escape hatches for states no schema anticipated. Those are tools because they are procedures, not because the platform forgot to model them.
The deeper point survives the tooling churn: agents are about to become the heaviest users of your platform's API, heavier than your dashboard's users and possibly heavier than your CLI's. Design the surface for a consumer that never sleeps, never remembers across sessions unless you make it, and reads exactly what you expose — desired state, actual state, and the policy between them, all machine-readable. Whether that surface is CRDs, MCP tools backed by intent objects, or both, the test is the same: kill the agent halfway through and check whether the platform still knows what it was doing.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with agents as first-class operators through a machine-readable API. Star the repo on GitHub or deploy your first app today.



