Skip to main content

454 MCP Servers Catalogued and Quality-Scored Daily: Finding a Production-Grade Server in a Flood of Demos

12 min readDora NodaDora Noda
Share
On this page

There are tens of thousands of MCP servers in the wild, and the average quality score of the ones a daily automated scan bothers to catalogue is 16 out of 100. That single number — from the mcp-ecosystem-scanner project, which crawls GitHub every day, validates each repo as a genuine MCP server, and scores it — tells you almost everything about the state of Model Context Protocol supply in late 2026: enormous quantity, wildly uneven quality, and no shortcut around vetting before you hand a server to your agents.

Here is the scanner's headline state of the world, so you can calibrate everything that follows. As of the September 15, 2026 scan the catalogue held 454 servers across 10 categories; by the September 24 scan it was 455, with a combined 1.6 million GitHub stars and an average quality score of 16.2 out of 100. The count moves daily, which is itself the point: this is a living registry, not a frozen awesome-list.

Behind that 16.2 sits the rubric — the eight signals every server is scored on:

SignalPointsWhat it checks
Has README10Documentation exists
Has license5Open-source license present
Has tests15Test directory found
Has CI/CD10GitHub Actions or similar configured
Has releases10At least one tagged release
GitHub stars15Community adoption (log scale)
Recent activity15Days since last commit
Issue health10Ratio of closed to total issues

Note the ceiling implied by that table: the weights sum to 90, and sure enough the highest-scored servers in the catalogue sit at exactly 90. Nobody is clearing 100 because the rubric, as published, tops out at 90 — a useful reminder that this score is a hygiene triage signal, not a certification. Still, the gap between the 90s at the top and the 16.2 mean tells you the distribution has a long, messy tail. The rest of this post digs into what that tail looks like, what the score cannot see, and what a platform team deciding which MCP servers to bless for tenant agent workflows should steal from the scanner's approach.

What the scanner actually checks​

The scanner exists because MCP discovery is fragmented. The official registry is young, awesome-lists are hand-curated and go stale, and there was no automated way to tell a maintained server from an abandoned demo. So the project does three things on a daily loop: discover candidate repos on GitHub, validate each one as a genuine MCP server (as opposed to a repo that merely mentions MCP), and score the survivors against the rubric above.

The rubric's design choices are worth reading closely, because they encode a specific theory of what "production-grade" means for an integration surface:

  • Tests are worth triple a license (15 vs 5 points). Executable proof beats paperwork. A server with a test directory has at least one maintainer who thought about regression; a server with only a README has a pitch.
  • Releases and CI are each worth 10. A tagged release is a deployable, pinnable artifact — the difference between "clone main and pray" and a version you can roll back to. CI means someone set up a machine to say no.
  • Recency and issue health total 25 points. Over a quarter of the score is maintenance posture: is anyone still committing, and do reported problems get closed? For a protocol surface that agents call autonomously, an unmaintained server is a liability that compounds.
  • Stars are log-scaled (15 points max). Popularity counts but with diminishing returns, so a viral demo cannot lap a solid, boring integration purely on hype. As we will see, the catalogue still needed that dampener.

The highest-quality entries — browser automation, device control, and data-indexing servers scoring 89–90 — tend to max out nearly every structural signal. They have docs, tests, CI, releases, recent commits, and real adoption. That cluster at the top is what "blessable" looks like in rubric terms.

The real shape of the supply​

Now the interesting part: what does a validated, scored registry reveal about the MCP supply that raw server counts hide?

Stars do not mean quality. The single most-starred entry in the catalogue, at over 95,000 stars, is an awesome-list — a link collection, not a server at all. More broadly, the catalogue's 1.6 million combined stars concentrate heavily in a handful of viral repos while the median server sits near the bottom of the scoring range. If your selection process starts from "most stars on GitHub," you are sampling hype, not hygiene. The log-scaled stars component in the rubric exists precisely because raw star counts lie this loudly.

The category map shows where supply is real and where it is thin. All ten categories, by server count:

CategoryServersShare of catalogue
other16436%
dev-tools8719%
ai-ml8018%
web5412%
data307%
productivity164%
files102%
cloud92%
security3under 1%
finance2under 1%

Three things stand out. First, the largest category is other at 164 servers — more than a third of validated supply defies classification, which is itself a discovery failure: if the cataloguer cannot tell what a server is for, neither can your evaluation team at a glance. Second, dev-tools and ai-ml together hold another 167 servers, so nearly three-quarters of the ecosystem serves developers and model-adjacent workloads — exactly the audience writing agents today. Third, the long tail of categories is genuinely thin: finance has 2 servers, security has 3, cloud has 9. If you need an MCP server for a regulated or infrastructure-adjacent workflow, you are choosing from a shortlist, not a marketplace — and "shortlist" here means you can hand-audit every candidate, which changes the evaluation strategy completely.

The stack is a two-language, three-transport world. TypeScript (211 servers) and Python (200) together account for roughly 90% of the catalogue; Go trails at 12. On transports, SSE appears on 354 servers, stdio on 249, and streamable HTTP on 125. That transport mix matters operationally: SSE is the legacy remote-transport story, stdio is the local-process story, and streamable HTTP is where the spec is heading — a server that only speaks a legacy transport is a future migration you are signing up for.

Churn is visible even inside the validated set. Eighteen catalogue servers are archived or abandoned, and only one new server entered in the latest monthly window the dashboard reports. Zoom out beyond this one scanner and the churn looks starker: third-party directories count anywhere from a few thousand to nearly 38,000 servers depending on how aggressively they sweep forks and demos, lifecycle research estimates that over half of MCP servers go quiet within 90 days, and only a few dozen — one widely cited estimate says around 70 — are considered genuinely production-ready for real workflows. A separate spec-readiness probe of the official registry found that of 4,356 remotely reachable servers, exactly one passed every check the July 2026 spec release made mandatory. The validated 455 are already a filtered view of reality, and even that view averages 16.2.

What the score cannot see: security​

Here is the audit the rubric does not perform. Every signal in the scoring table is about repository hygiene — docs, tests, CI, activity, adoption. None of them look at what the server actually does when an agent calls it. A well-documented, CI-gated, recently committed server with 20,000 stars can still:

  • hide instructions in tool descriptions that steer the calling model (tool poisoning — the MCP-native attack, where a tool's metadata tells the model to exfiltrate data through a parameter);
  • ship overbroad input schemas that grant the model far more latitude than the task needs;
  • run its container as root, pin dependencies with known CVEs, or leave CI workflows unpinned and hijackable;
  • change behavior under you after adoption (the "rug pull" problem: a tool your agent trusts gets an update that quietly widens what it does).

Independent security scans keep confirming the gap. One project that probed a few thousand public servers reported roughly half with unpinned GitHub Actions, over 40% with overbroad tool input schemas, more than a quarter running containers as root, and one in nine pinned to dependencies with known vulnerabilities. Another scanner's pass over popular servers produced an average trust score in the mid-20s out of 100 — notably worse than the hygiene average, on servers people actually install. High-severity findings show up even in well-maintained projects, because "maintained" and "safe to grant tool access" were never the same property.

This is not an argument against the scanner — it is an argument about what layer it occupies. Repository hygiene is a necessary condition for trust, not a sufficient one. A score of 16 tells you to walk away without spending human review time; a score of 90 tells you the server is worth the expensive audit, not that it has passed one.

The borrowable checklist for blessing MCP servers​

If you run a platform where tenants wire MCP servers into agent workflows — the thing a self-hosted, AI-native PaaS is increasingly asked to support — hand-auditing every candidate integration does not scale. The scanner's real contribution is not the leaderboard; it is a machine-checkable first pass you can steal. Here is the checklist, in the order you should apply it:

  1. Is it a genuine, validated server? Confirm the repo actually implements the protocol (server scaffold, tool/resource definitions), not a wrapper blog post or a client. Automate this the way the scanner does: clone, detect, reject non-servers before a human looks.
  2. Does it ship releases? No tagged release, no blessing. Releases give you a pinnable version, a changelog to review, and a rollback target. "Deploy from main" is fine for a demo and disqualifying for shared tenant infrastructure.
  3. Does it have CI and tests? CI plus a test directory is the minimum evidence that changes get checked by something other than optimism. Weight tests heaviest, as the rubric does.
  4. Is it maintained right now? Check days since last commit and the closed-to-total issue ratio. For a fast-moving protocol, a server untouched for six months is running against a spec that has moved on — recall the readiness probe where nearly 91% of reachable servers failed the current spec's mandatory checks.
  5. Does the transport fit your topology? Prefer streamable HTTP for remote servers you operate; treat stdio-only servers as local-execution commitments with sandboxing implications; flag SSE-only servers as carrying migration debt.
  6. Then — and only then — run the security pass the score skips. Review tool descriptions and schemas for poisoning patterns (hidden instructions, smuggled directives in enums and defaults), check the container runs non-root with pinned dependencies, verify TLS on remote transports, scope credentials to the narrowest resource that works, and require human approval for high-impact tools. Catalogue-wide security research suggests this pass will reject candidates the hygiene score loved; that is the pass doing its job.

Applied in this order, the hygiene signals do the cheap triage (most candidates fail fast, costing you compute instead of reviewer hours) and the security review concentrates on the shortlist that survived. That ordering also answers the thin-category problem from the category table: when finance has two candidates and security has three, you skip straight to step 6 and audit all of them by hand — the checklist tells you when automation is leverage and when it is overhead.

Score first, trust after verification​

The mcp-ecosystem-scanner's catalogue is a reality check with a daily refresh cycle: 455 validated servers, a 16.2 average, a third of supply unclassifiable, and churn everywhere outside the top shelf. Its scoring rubric is worth adopting as triage precisely because it is shallow — eight machine-checkable signals that sort "worth a human's time" from "walk away" in seconds. But the security data is unambiguous that hygiene is only the first gate. Stars can be gamed, docs can be thorough, CI can be green — and the tool description can still be carrying instructions your model will obey.

For platform teams, the practical takeaway is a two-tier policy: automate the scanner's hygiene bar as a hard floor for anything tenants can install, and reserve human security review — schema scrutiny, transport and credential checks, approval gates on dangerous tools — for the shortlist that clears it. In an ecosystem where tens of thousands of servers exist and dozens are truly production-ready, the teams that bless integrations deliberately will spend their review budget where it matters, and everyone else will be auditing after the incident.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with AI agents as first-class operators. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide