Skip to main content

Kubernetes Dashboard Is Dead — Headlamp's Knative and Volcano Plugins Turn It Into a Platform Workbench

9 min readDora NodaDora Noda
Share
On this page

On January 21, 2026, the Kubernetes Dashboard project was archived and moved to kubernetes-retired/dashboard — no more security patches, bug fixes, or feature updates, ever. (hardening guide, 2026) The Kubernetes SIG UI group now points everyone at Headlamp, the CNCF Sandbox, RBAC-aware UI that runs as a desktop app or in-cluster. That alone would be a routine changing-of-the-guard story.

What makes it interesting is what happened five months later: on June 25, 2026, the Kubernetes blog published not one but three Headlamp plugin announcements on the same day — Cluster API, Knative, and Volcano — all built through CNCF LFX mentorship. The Dashboard's replacement is no longer just a prettier pod list. It is becoming the one browser tab where a small ops team debugs node lifecycle, serverless rollouts, and batch scheduling without touching three CLIs.

Here is the payoff up front, so the rest of the post has somewhere to land:

Debugging workflowBefore: the CLI shuffleAfter: the Headlamp tab
Knative canary: is the new revision actually serving traffic?kn for revisions, kubectl for route objects, Grafana for per-revision request ratesOne KService page: traffic split, readiness, and per-revision metrics inline
Knative autoscaling surprise: why did this service scale to zero (or not)?Cross-reference KService annotations against config-autoscaler / config-defaults ConfigMaps by handEffective autoscaling view showing explicit vs. cluster-default values in context
Volcano job stuck pending: is it queued or gang-blocked?Volcano CLI + kubectl across Job, PodGroup, Queue, and PodsMap view linking Job → PodGroup → Pods → Queue, with the blocker state highlighted
Volcano queue contention: who deserves what capacity?Raw Queue YAML: capacity, deserved, guaranteed, reservationsQueue detail page with allocation context and child queues
Cluster lifecycle: which machine is wedging the rollout?Raw kubectl plus deep familiarity with CAPI ownership hierarchiesCAPI dashboard with health cards, remediation guidance, and scale actions

The punchline: Headlamp's plugin system just absorbed the two workload types — serverless and batch — that a git-push PaaS reaches for the day its tenants outgrow plain Deployments. If you are a two-person ops team wondering whether adopting Knative or Volcano means also adopting a second console, the answer as of June 2026 is no.

Knative: see your serverless​

Start with the concrete inventory, because "a Knative plugin" could mean a revision list and nothing else. What Mudit Maheshwari and Kahiro Okina shipped (current release 0.3.0-beta) is a full operator surface over Knative Serving.

A KService is the top-level resource in Knative: it owns the lifecycle of Routes, Configurations, Revisions, and everything needed to run and expose the app. The plugin gives each KService a detail view with an Edit Mode toggle for live changes to traffic splits and autoscaling annotations, and surfaces common actions — view YAML, open logs, trigger a redeploy, restart backing pods — in the header, gated by your current RBAC permissions. That last clause matters more than it sounds: because Headlamp adapts the UI to what your role may actually do, the same KService page is safe to hand to a junior engineer or even a tenant with scoped access. The buttons they must not press simply are not there.

The traffic-splitting view is the one that earns the "workbench" label. Knative routes traffic across multiple Revisions of one service for canaries, gradual rollouts, tagged preview URLs, and A/B tests. The plugin shows the traffic assigned to each Revision, the latest ready Revision, readiness status, age, and configured tags — and in edit mode you adjust percentages and tags inline, with validation that traffic sums to 100% and tags are unique before anything is saved. Tagged routes with a reported URL render as clickable links, so checking the canary is one click, not a kn route describe plus a copy-paste.

Then there is the autoscaling view, which solves a genuinely annoying Knative sharp edge. The effective autoscaler behavior for any workload is a merge of KService-level annotations and cluster-wide ConfigMaps (config-autoscaler, config-defaults): concurrency targets, target utilization, min/max scale, stable window, scale-down delay, and more. The plugin reads both and shows the effective configuration per KService in context, so you see at a glance whether a setting is explicitly configured or silently falling back to the cluster default. Anyone who has debugged a "why did this scale to zero during the demo" incident by diffing annotations against ConfigMaps in two terminal panes knows exactly which afternoon this view gives back.

Pair it with the Headlamp Prometheus plugin and the KService and Revision pages render request rate, latency, and resource utilization graphs inline — the per-revision request-rate breakdown being the one you want open during a traffic split. Round it out with list and detail views for Revisions, DomainMappings, and ClusterDomainClaims plus a cluster-level Networking overview (effective ingress class and gateway settings read from config-network and config-gateway), and the old loop of "kn says this, kubectl says that, Grafana says a third thing" collapses into one page.


Volcano: see your batch​

Mahmoud Magdy's Volcano plugin (0.1.0-alpha) attacks the same fragmentation problem from the batch side. Kubernetes was designed around long-running services; Volcano extends it with queues, priorities, quotas, and gang scheduling for HPC, AI/ML, and batch workloads that arrive dynamically, compete for limited resources, and need multiple workers to start together before useful work begins. Operating that means constantly hopping between four related resources — Job, PodGroup, Queue, Pods — and, as the launch post puts it, "all of that is possible with CLI tools like kubectl and the Volcano CLI, but it can become fragmented very quickly."

The Job view is the center of the plugin. The list shows status, queue, running-versus-minimum-available values, task count, and age at a glance; the detail page keeps task details, Pod status, related Queue and PodGroup links, conditions, and events on one screen. Supported lifecycle actions (Suspend, Resume) fire directly from the UI in appropriate states. And Job logs open without leaving the page, with single-Pod and all-Pods views plus container selection, line count, previous logs, timestamps, and follow — the flags you would otherwise retype on every kubectl logs invocation.

Queues get the treatment batch operators actually need. The Queue page surfaces capacity, allocated resources, deserved and guaranteed resources, reservation details, and child queues — the full "who is entitled to what, and who is actually holding it" picture that raw Queue YAML buries. When two tenants' training jobs contend for one GPU pool, this is the page that settles the argument.

PodGroups are where gang scheduling lives or dies, and the plugin's PodGroup view highlights progress, conditions, and minimum resource requirements — a direct answer to "is this workload blocked because it hasn't met the scheduling conditions to run as a group." Walk the canonical incident: a training job sits pending. Before, you described the Job, listed its Pods, found half of them unscheduled, described the PodGroup, then described the Queue to learn the gang minimum can't be satisfied. After, the map view shows the Job, its PodGroup, the created Pods, and the Queue context in one graph, with warning and error states marking whatever needs attention. That map — Jobs, PodGroups, Queues, and Pods as one relationship graph — is the single view kubectl's flat pod list can never give you.

One honest caveat, stated plainly in the launch post: Prometheus integration is future work for the Volcano plugin ("richer scheduling insights" is on the roadmap), where the Knative and CAPI plugins already embed metrics. If your batch debugging starts from GPU-utilization graphs rather than scheduling state, you still need the second tab for now.

What this actually changes for a two-person ops team​

Three plugins landing the same day, all through LFX mentorship, is a velocity signal worth reading carefully. It says Headlamp's plugin API is now approachable enough that mentored contributors — not just the core team — can ship a full CRD operator surface in one mentorship cycle. And the trajectory kept going: on July 13 the Kubernetes blog published both a Headlamp plugin for Kubeflow for AI/ML workloads and a step-by-step Dashboard-to-Headlamp migration guide. The pattern is unmistakable: every workload family gets its own visual surface inside the same RBAC-aware shell.

But maturity labels are maturity labels, and a small team should read them before betting a production workflow on any of this:

PluginRelease at launchMetrics storyHonest status
Cluster API (Chayan Das)0.1.0-alphaPrometheus inline alreadyTry on your management cluster; scale actions and remediation guidance are genuinely useful
Knative (Maheshwari, Okina)0.3.0-betaPrometheus inline alreadyClosest to daily-driver; traffic-split editing is the standout
Volcano (Magdy)0.1.0-alphaFuture workTry for gang-scheduling visibility; keep Grafana for utilization

The launch posts are also explicit that none of this replaces kubectl or the Knative/Volcano CLIs for automation, scripting, and raw object inspection. What the plugins replace is the interactive part: discovering related resources, reading structured detail instead of raw YAML, and moving from scheduling state to runtime output without switching tools. That is precisely the part that eats a small team's on-call hours.

So what does this buy a self-hosted platform considering workload types beyond git-push web services? Two things. First, adopting Knative for idle-mostly services or Volcano for batch/GPU jobs no longer drags a second console into your ops story — the same Headlamp instance your team already uses for Cluster API machine lifecycle now speaks serverless revisions and gang-scheduled PodGroups too. Second, RBAC-gated actions mean the console is delegable: tenants and juniors get the views their roles permit, with dangerous actions absent rather than merely documented as forbidden.

Trying it is a ten-minute errand, not a migration project: install Headlamp (desktop app or in-cluster), open the Plugin Catalog, search for Knative or Volcano, install, and connect to a cluster where the workload system is already running. If you hit a gap, the maintainers are explicitly asking for operator feedback in headlamp-k8s/plugins — alpha software shaped by real Volcano and Knative users beats beta software shaped by nobody.

The Dashboard died in January. What replaced it isn't a dashboard at all — it's a workbench that now covers nodes, serverless, batch, and ML training behind one RBAC boundary. For teams running clusters they own, that consolidation is the whole game: every workload type you adopt without adopting another console is operational leverage you keep.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex