In March 2026, a Hacker News comment delivered the single most damaging sentence ever written about the two most popular single-box PaaS tools: "Currently Coolify and Dokploy are not designed for AI Agent auto deploy." Six months later, that verdict is still quoted — most recently in the kelson project's competitive analysis, which answers it by noting that Canine already ships 13 MCP tools for agent-driven deploys. But is the verdict still true? Both tools have APIs, both have shipped releases since March, and a small cottage industry of community MCP wrappers has sprung up around them. This post interrogates the verdict against September 2026 reality: a concrete scorecard, an honest March-to-September delta, and what "designed for agent auto deploy" actually requires.
The agent-operability scorecard
Before arguing about vibes, here is the concrete artifact: six dimensions that determine whether an AI agent can operate a PaaS, scored against Coolify v4, Dokploy, Canine, and an MCP-first surface.
| Dimension | Coolify v4 (Sept 2026) | Dokploy (Sept 2026) | Canine | MCP-first surface |
|---|---|---|---|---|
| First-party API docs | Partial — /api/v1 REST exists but ships disabled; no official API reference, community docs fill the gap | Partial — x-api-key + deploy hooks documented in auto-deploy guide; tRPC routes largely undocumented | Yes | Yes — OpenAPI + MCP schemas |
| First-party MCP server | No — community coolify-mcp-paas wraps the REST API | No — community skills (dokpilot, agent SKILL.md files) wrap tRPC | Yes — 13 MCP tools | Yes — 35–57 tools (AppCrane, nerdit) |
| Typed, idempotent operations | Mixed — REST verbs exist; idempotency and error shapes are whatever Laravel returns | Weak — tRPC mutations mirror dashboard clicks (application.deploy, compose.deploy); exact service keys required | Yes — tools designed as agent calls | Yes — idempotent writes, structured errors |
| Scoped agent tokens | Coarse — API tokens with broad scopes (read, write, read:sensitive); one token per human | Coarse — one API key per user via better-auth | Fine-grained per deploy target | Fine-grained — per-agent scoped, revocable, audited tokens |
| Machine-readable status + logs | Partial — log endpoints exist (/applications/{uuid}/logs) but the agent must first hunt UUIDs across list endpoints | Partial — status via tRPC queries; no single "is it healthy?" primitive | Yes — status as a first-class tool result | Yes — status, logs, and health as typed tool output |
| Audit trail | Dashboard-centric — who clicked what, not which agent called which tool with which args | Dashboard-centric | Yes | Yes — every tool call logged |
The pattern is unmistakable: both single-box tools score "partial" or worse on every row that matters to a non-human operator, while the tools designed for agents score "yes" down the column. The gap is not "has an API" — it is everything around the API that turns endpoints into something an LLM can reliably drive at 2am.
What the verdict still gets right, and the March-to-September delta
Fairness first: the March 2026 verdict was never literally "these tools have no API." Dokploy already had programmatic API-key setup via its better-auth apiKey plugin in early March 2026, and Coolify v4's /api/v1 surface predates the verdict too. The claim was narrower and sharper — no agent interface — and on that reading, here is what each tool looked like then versus now.
Coolify, March 2026 → September 2026. At verdict time, Coolify v4 was still in the 4.0.0-beta.4xx era: a REST API existed but lacked deployment cancellation, per-resource log endpoints, rollback, storage and scheduled-task management, and the instant_deploy/force parameters. By v4.3.17 (September 4, 2026), all of those exist, alongside a read:sensitive scope that gates secret reads. That is a genuine delta — an agent today can cancel a stuck deploy or fetch app logs over REST, which it could not do in March.
What has not changed: the API still ships disabled until an admin flips "Allow API Access" in Settings, there is still no official API reference (the best endpoint inventory lives in a community member's personal repo), and there is still no first-party MCP server. The community coolify-mcp-paas wrapper exists, but its own changelog confesses the underlying friction: no single-item GET endpoints (it lists everything and filters client-side) and inconsistent integer-ID-versus-UUID conventions across resources.
Dokploy, March 2026 → September 2026. Dokploy's story moved less. Then and now, programmatic deploys go through one of three doors: deploy-hook URLs (/api/deploy/<token>), an x-api-key header against tRPC procedures (application.deploy, compose.deploy, domain.create), or the documented auto-deploy API for CI pipelines. The tRPC surface mirrors dashboard mutations one-to-one — including the sharp edges. Community skill authors warn that serviceName must match the exact service key in docker-compose.yml, not the app's display name in Dokploy, a distinction no agent can infer without reading someone's reverse-engineered notes.
Six months later there is still no versioned REST API, no official API reference beyond the auto-deploy guide, and no first-party MCP server. The ecosystem response — dokpilot skill packs, agent SKILL.md files committed to individual repos — proves the demand while underscoring the point: every team hand-rolls its own agent glue because the platform does not ship any.
So the verdict, read precisely, still holds in September 2026: neither tool ships a first-party agent interface. What changed is the cost of the gap. In March it was a missing feature; today it is a tax every team pays in UUID-hunting, response-scraping, and maintaining forked community wrappers that break on each upstream release.
What "designed for agent auto deploy" concretely means
"Agent interface" is vague enough to mean anything, so here is the concrete checklist — five requirements, each with a tool that already ships it:
- A typed tool catalog the agent can discover. MCP's
tools/listgives the agent a machine-readable menu —list_services,deploy,get_logs,set_env— with typed parameters, instead of a wiki page of curl examples. Canine ships 13 such tools; AppCrane exposes 35appcrane_*tools behind oneclaude mcp add; nerdit ships 57. - Idempotent writes with structured errors. An agent retries; a dashboard human does not. Agent-designed surfaces make
deploysafe to call twice and return errors as structured data (which field, which constraint, what to do) rather than a Laravel stack trace or a toast notification. Nerdit's MCP server explicitly documents idempotent writes and structured errors. - Scoped, revocable agent tokens. A token minted for "Claude Code on the billing repo" should deploy exactly that repo, expire, and be revocable without rotating a human's credentials. Better-PaaS issues separate scoped tokens for Cursor, Claude Code, CI, and scripts; the single-box tools hand the agent a broad user-level key.
- Machine-readable status, logs, and health. "Did the deploy work?" must be one typed call returning a verdict, not three list endpoints plus log-grep plus vibes. Agent-first servers return deployment state, health checks, and log tails as tool results the agent can branch on.
- An audit trail of tool calls. When the agent ships a bad deploy at 2am, the postmortem needs "which tool, which args, which result" — not "someone clicked Deploy." Nerdit and the MCP-first cohort log every tool invocation; dashboard-centric tools log clicks.
Note the common thread: none of these is "more endpoints." All five are about the contract around the endpoints — discovery, safety, scoping, readability, accountability. That is why a REST API bolted onto a dashboard does not close the gap, and why the community wrappers keep hitting the same walls: they can translate transports, but they cannot retrofit idempotency, scoping, or audit into a server that never designed for them.
The 2am rollback test
Abstract checklists are cheap, so run the worked scenario both sides must survive: the on-call agent (LLM, not human) gets paged that production is serving 500s after an auto-deploy, and must roll back. Here is what each surface demands of it.
On a dashboard-first single-box tool, the agent's runbook looks like this: find the application's UUID by listing all applications and fuzzy-matching on name; find the previous deployment by listing deployments and guessing which SHA was last-known-good from timestamps; trigger a redeploy of that SHA through a tRPC mutation or REST POST whose exact parameters live in a community skill file; then verify by polling a status endpoint that returns raw container state, deciding for itself what "healthy" means; and if the redeploy hangs, discover — mid-incident — that cancellation only exists on versions newer than the one this box runs. Every step is a guess the agent must get right while the site is down, and the failure mode at each step is a screenshot-and-guess loop: open the dashboard, squint, retry.
On an MCP-first surface, the same incident is four typed calls: get_status returns "deployment #482 failing health checks, last healthy #471"; rollback takes a deployment ID and is idempotent — calling it twice rolls back once; get_logs streams the new revision's boot logs as structured output the agent can pattern-match; get_status confirms green. Scoped tokens mean the agent holds rollback-but-not-destroy permission; the audit log records every call for the morning postmortem. No UUID archaeology, no version-sniffing, no dashboard.
This is the concrete cost of the gap: not that agents cannot drive Coolify or Dokploy — clearly they can, with enough bespoke glue — but that every incident becomes custom integration work instead of a typed operation. The team that hands its 2am rollback to an agent on a dashboard-first platform is not automating operations; it is automating improvisation.
What to do Monday morning
If your deploy operator is already partly an LLM — and if your team uses Claude Code, Cursor, or any coding agent with production access, it is — evaluate your self-hosted PaaS with three questions. First, can a fresh agent go from zero to "deployed and verified" using only the platform's own documented interface, with no community wrapper and no dashboard? Second, is every mutating operation the agent can reach idempotent, scoped to least privilege, and recorded in an audit log you can read the next morning? Third, does "is it healthy" have a typed answer, or does the agent have to assemble one from container state and hope?
Six months after the HN verdict, Coolify and Dokploy have better APIs and a livelier wrapper ecosystem — genuine progress, honestly measured above. But neither has crossed the line from "scriptable dashboard" to "agent-operable platform," and the community glue filling that gap is proof of the gap, not its closure. The self-hosted PaaS that treats machine-readable state and agent-callable deploy, rollback, and status tools as first-class features — not a dashboard with an API bolted on later — is the one whose tenants can actually hand the 2am rollback to an agent.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with a Render-compatible API and machine-readable state designed for agent operators from day one. Star the repo on GitHub or deploy your first app today.



