One tenant ships a deploy with an unbounded label — a request ID, a pod hash, something with a million distinct values — and by morning every other tenant's dashboards have gaps. The ingester didn't crash. It did exactly what an unpartitioned metrics backend does under a cardinality explosion: it spent the shared memory budget on whoever screamed loudest. If your tenants pay by usage, that incident is not an observability gap. It is a billing incident twice over: one tenant's runaway series consumed capacity everyone else paid for, and the usage numbers you would invoice from are the corrupted ones.
That is the lens for this comparison. The day tenants pay by usage, the metrics backend stops being observability decoration and becomes billing infrastructure, and billing infrastructure gets judged on three questions: can one noisy tenant's series explosion evict another tenant's data, can you attribute usage and query cost back to whoever caused it, and what does the whole thing cost to run on machines you own. Grafana Mimir, Cortex, and Thanos answer those three questions very differently, and the right pick depends on which answer you need most.
The verdict up front: for a greenfield multi-tenant platform that meters usage, run Mimir — native per-tenant isolation plus first-class cost-attribution machinery no other backend in this trio ships. Pick Thanos when you already operate a fleet of Prometheus servers and want its object-store-backed global view with downsampled long retention bolted on. Pick Cortex only when you specifically want its lineage — the freshly OSTIF-audited, still-maintained codebase that Amazon Managed Prometheus is built on — without tracking Mimir's faster-moving releases. The rest of this post is the evidence behind those three sentences, one dimension at a time.
What each system is, in one paragraph
Mimir is Grafana's horizontally scalable, multi-tenant long-term store for Prometheus metrics, forked from Cortex and developed at a pace Cortex no longer matches. Tenancy is native: every request carries an X-Scope-OrgID header, and isolation, limits, and retention are per-tenant concepts from the ingester down to the object store. The 3.x series reworked the core — Mimir 3.0 decoupled reads from writes through a Kafka-based ingest-storage layer and introduced the streaming Mimir Query Engine, which Grafana Labs says cuts query memory usage by up to 92% — and 3.1 made the V2 ingest record format the default, cutting Kafka throughput by roughly 70% and write-path CPU by about 8% in Grafana's own measurement.
Cortex is the project Mimir descends from: the original horizontally scalable, multi-tenant Prometheus store, still maintained under CNCF governance with the same X-Scope-OrgID tenancy model. Its headline 2026 event is a security story, not a feature story: Cortex completed an OSTIF security audit in August 2026, performed by Quarkslab over March 30 to April 16, 2026, with the fund publishing custom review and hardening documentation for the project. One recent head-to-head comparison summarizes the maintenance reality bluntly: Cortex is maintained while new work happens in Mimir, and Amazon's managed Prometheus offering stays on the Cortex lineage.
Thanos is the odd one out architecturally. Rather than replacing Prometheus with a push-based platform, it extends fleets of ordinary Prometheus servers: sidecars upload blocks to object storage, a global querier fans out across sidecars, receivers, and historical stores, and a compactor downsamples old data for cheap long-range queries. Multi-tenancy exists — receivers route on a THANOS-TENANT header with a per-tenant TSDB instance each — but it reads as bolted on next to Mimir and Cortex, where tenancy is the day-one design.
That same comparison table rates Thanos multi-tenancy "basic" against full marks for the other two, and the honest summary from practitioners is consistent: Thanos is for existing Prometheus deployments you want to federate, Mimir and Cortex are for greenfield multi-tenant platforms fed by remote write.
Dimension 1: per-tenant cardinality isolation
The noisy-tenant test is simple: tenant A starts emitting a million unexpected series, and tenant B must notice nothing — no eviction, no throttled ingestion, no missing samples. This is where native tenancy pays off, because the enforcement point has to sit in the write path with per-tenant counters, not in a dashboard filter applied after the damage.
Mimir enforces limits per tenant through its runtime configuration, reloadable without restarts: max_global_series_per_user caps a tenant's total active series cluster-wide, ingestion_rate and ingestion_burst_size bound their sample throughput, max_label_names_per_series rejects the wide-label accidents, and even retention (compactor_blocks_retention_period) can be overridden per tenant. Breaching tenants get their own writes rejected while everyone else's samples flow. The primitives compose into a genuine noisy-neighbor policy rather than a single global tripwire, which is exactly what a platform with untrusted tenant cardinality needs.
Cortex shares the model — same header, same per-tenant limit family inherited from the common lineage — so on isolation mechanics alone it stands roughly where Mimir does. The difference is velocity: new guardrails land in Mimir first (per-tracker cardinality configuration and similar refinements track Mimir releases), while Cortex holds the stable, now-independently-audited line. If your threat model values an audited write path over the newest knob, that tradeoff can favor Cortex; if you want the knob the day a novel cardinality incident invents the need for it, it favors Mimir.
Thanos answers the same test with receiver-side limits that are real but narrower. Receivers accept a per-tenant head_series_limit capping each tenant's active head series (settable to zero for unlimited), and the project's own mixins ship alerts for tenants approaching that limit. Tenancy routing itself runs through hashring configuration with soft (shared receivers) or hard (dedicated receiver sets) tenancy per tenant group.
On the read path, the querier gained tenancy enforcement via prom-label-proxy, injecting tenant labels into queries so one tenant cannot read another's data. What Thanos lacks is Mimir's breadth of per-tenant write-path policy: ingestion-rate shaping, per-label-name guards, and per-tenant retention are not first-class receiver concepts, so a platform on Thanos ends up reimplementing the missing policy in an admission layer in front of the receivers — or accepting coarser protection.
Dimension 2: usage and query-cost attribution for chargeback
Isolation keeps tenants from harming each other; attribution tells you who owes what. A billing pipeline needs per-tenant usage signals it can trust — ingested samples and active series at minimum, plus query and rule-evaluation cost if chargeback extends past ingestion — exported in a form a batch job can join against infrastructure cost.
Mimir is the only backend here with an explicit cost-attribution feature, and it keeps getting more expressive. The original design gave each tenant a single unnamed tracker (implicitly called cost-attribution); Mimir 3.2 added support for multiple named cost-attribution trackers per tenant, each independently configured with its own labels, cardinality limit, and cooldown. These trackers exist precisely so operators can slice a tenant's usage by team, product, or workload — the sub-tenant breakdown a chargeback model needs. Beneath that, the standard per-tenant usage metrics (ingested samples, active series, rule evaluations) give the raw inputs for per-tenant invoicing.
Cortex exposes the same cortex_* per-tenant usage-metric family both projects share, so samples-ingested and series-active signals per tenant are available for a billing join. What it does not have is the tracker concept: no first-class, cardinality-bounded, labeled usage breakdown within a tenant. A Cortex-based billing pipeline aggregates the raw per-tenant counters itself and stops at the tenant boundary, which suffices for per-tenant invoices but leaves sub-tenant chargeback (team-level showback inside one tenant) as custom recording-rule work rather than a configured tracker.
Thanos leaves the most for you to build. Per-tenant signals exist at the receiver — per-TSDB head stats, compactor metadata — but there is no attribution surface: no cost trackers, no per-tenant query-cost accounting, no usage API shaped for billing. Query cost in particular is hard to attribute because the querier fans out across stores on behalf of whoever asked, with nothing tagging the work by tenant for later accounting. Platforms billing on Thanos typically meter at the edge instead — counting remote-write bytes per tenant at an ingress gateway — which measures what tenants sent, not what the backend spent storing and querying it. That gap is acceptable for rough usage tiers and unacceptable for cost-plus chargeback.
Dimension 3: footprint on owned capacity
On rented object storage and autoscaled compute, backend efficiency is somebody else's margin. On a fixed set of owned Hetzner machines, every always-on component is a line item against the capacity tenants could have used. So count the moving parts and their sizing rules, not just the feature lists.
Mimir publishes unusually concrete capacity-planning guidance: plan roughly one compactor instance per 20 million active series (about 1 CPU core and 4 GB of memory each), and size ingesters around 1 core plus 2.5 GB of memory per 300,000 in-memory series. Store-gateways add memory proportional to the index of the blocks they serve, and the query-frontend/queriers scale with query load.
Two caveats matter for small fleets. First, the Kafka-based ingest-storage architecture that makes 3.x resilient is another distributed system to operate — worth it past real scale, heavy below it. Second, Mimir keeps raw resolution for the full retention window rather than downsampling old data, so long retention buys query fidelity at object-storage cost. For starting small, Mimir ships a monolithic mode that collapses the microservices into one binary, which is the honest on-ramp: one pod and object storage beats a dozen underutilized components.
Cortex sizes in the same order — same component family (distributor, ingester, store-gateway, compactor, query-frontend), same memory-dominated ingester math — with a simpler floor and a lower ceiling: no Kafka tier to operate, but also no streaming query engine delivering the 92%-class memory savings on the read path, and no V2 ingest-format efficiency work. Its single-process all-in-one target is the equivalent small-fleet starting point. For a platform that expects to stay at modest series counts indefinitely, Cortex's footprint story is arguably the cleaner of the two; for one that expects to grow into Mimir's efficiency work, starting on Cortex just to migrate later is a tax with no lasting benefit.
Thanos has the fewest mandatory pieces and the cheapest long tail. A minimal setup is receivers (or sidecars on existing Prometheus servers), one querier, one store-gateway, and the compactor — which must run as a singleton, a scaling ceiling worth knowing about before you need it. The compactor's downsampling is Thanos's genuine economic advantage: old blocks get 5-minute and 1-hour resolutions with independent retentions, so a year of history costs a fraction of Mimir's raw-resolution equivalent in both object storage and long-range query time. The counterweight is receiver memory — each tenant's TSDB head lives in receiver RAM — and store-gateway memory for block indexes, which grow with the historical surface you keep queryable. Thanos wins the footprint argument when retention is long and tenants are few; it loses it as tenant count grows and per-tenant receiver overhead starts to dominate.
The decision table
| Mimir | Cortex | Thanos | |
|---|---|---|---|
| Cardinality isolation | Per-tenant series, rate, label, and retention limits via runtime config; breaching tenant rejected alone | Same limit family, slower-moving; independently audited write path | Per-tenant head_series_limit + hashring tenancy; no rate/label/retention policy per tenant |
| Chargeback attribution | Named cost-attribution trackers (3.2) + per-tenant usage metrics | Per-tenant usage counters; sub-tenant breakdown is custom work | No attribution surface; meter at ingress or build it yourself |
| Owned-capacity footprint | Heaviest ceiling config (Kafka at scale), raw-resolution retention, monolithic on-ramp | Same-order sizing, simpler floor, no MQE/Kafka efficiency work | Lightest long-retention cost via downsampling; singleton compactor; receiver RAM per tenant |
| Development velocity | Active; 3.x released through 2026 | Maintained; OSTIF-audited Aug 2026 | Active; v0.42.x current |
| Pick it when | Greenfield multi-tenant platform with usage billing | You want the audited, stable lineage (or AMP compatibility thinking) | Existing Prometheus fleet, long retention, few tenants |
Two sensitivity notes before the table hardens into dogma. First, fleet size moves the answer: below roughly a million active series, Mimir's monolithic mode and Thanos's minimal topology are both one-pod-plus-storage problems, and the decision should rest on tenancy needs (many untrusted tenants favors Mimir) rather than efficiency numbers measured at 20-million-series scale. Second, existing investment dominates greenfield logic: if ten Prometheus servers already scrape everything you own, Thanos's sidecar-and-querier story delivers a global view in days, while migrating ingestion to Mimir's remote-write platform is a quarter-long project — no chargeback feature is worth that gap unless billing is the actual requirement this year.
The through-line is the one from the opening: the moment tenants pay by usage, the metrics backend is the cash register, and cash registers get audited. Run the backend whose isolation you can explain to an angry tenant, whose attribution your invoice can cite, and whose footprint your capacity plan already accounts for. For most self-hosted PaaS builds starting today, that is Mimir with per-tenant limits set before the first tenant deploys — because cardinality incidents, like billing disputes, are always cheaper to prevent than to reconcile.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



