Fifty developers on Copilot Business costs $11,400 a year. A single GPU box that serves the same team its autocomplete costs about $3,000 once. Somewhere between those two numbers sits a breakeven — and for a team that already owns GPU capacity, it arrives sooner than the per-seat price suggests: roughly a dozen seats if Tabby sips spare capacity on an inference node you already run, roughly thirty if you buy a dedicated box and staff it generously. This post prices both sides honestly — Copilot's 2026 two-layer billing with overage included, and self-hosting with hardware, power, and ops time all counted — so you can find your own row in the table.
Here is the table first, with the receipts below. All figures are annual; assumptions are stated, not buried: a loaded engineering hour at $150, electricity at $0.15/kWh, and a dedicated 4090-class box at roughly $3,000 of capex amortized over 36 months.
| Team size | Copilot Business ($19/seat) | Copilot + heavy-agent overage | Owned box, dedicated | Owned GPUs, spare capacity |
|---|---|---|---|---|
| 10 devs | $2,280 | $3,480 | $3,384–$6,984 | $1,932–$5,532 |
| 20 devs | $4,560 | $6,960 | $3,384–$6,984 | $1,932–$5,532 |
| 50 devs | $11,400 | $17,400 | $3,384–$6,984 | $1,932–$5,532 |
Two things to notice before we move on. First, self-hosting does not always win: at ten developers with no overage, Copilot's $2,280 beats a dedicated box's $3,384 low end. Second, the self-hosted columns do not move with headcount — that flatness is the entire economic argument, and the compliance argument in the last section is the half no price table captures.
What Tabby is (the thing being priced)
Tabby is a self-hosted AI coding assistant: an open-source, on-premises alternative to GitHub Copilot, written in Rust and shipped as a single Docker image. Version 0.32.0, an Apache 2.0 project with roughly 33,500 GitHub stars as of May 2026, is designed from the start as a shared team server rather than a solo plugin. One deployment, centralized GPU access, and no developer on the team needs an API key in their IDE.
Three properties make it the right candidate to price. First, it is self-contained: no DBMS, no cloud service, no external dependency beyond the model weights it downloads from Hugging Face on first run. The whole server is one container plus a /data volume. Second, it ships team auth: an admin dashboard for users, roles, and API tokens, per-user keys that can be revoked individually, and usage telemetry per developer — the difference between a homelab toy and something an engineering manager can roll out. Third, it is a full stack, not a plugin: model server, IDE extensions for VS Code and JetBrains, codebase indexing (multi-branch since v0.32.0, plus GitLab merge-request indexing), autocomplete and chat in one package.
That full-stack property is exactly what separates Tabby from its neighbors in Spheron's August 2026 self-hosted-coding-assistant roundup, the comparison this post's pricing builds on:
| Tool | Type | Model server included | IDE support | Multi-user |
|---|---|---|---|---|
| Tabby | Full stack | Yes (own runtime) | VS Code, JetBrains | Yes (team auth) |
| Continue | IDE plugin only | No (needs vLLM/Ollama) | VS Code, JetBrains | Via your backend |
| FauxPilot | Full stack (legacy) | Triton server | VS Code only | Limited |
| Kilo | Agentic platform | No (needs OpenAI API) | VS Code, JetBrains, CLI | Via your backend |
Continue is the developers' favorite and Kilo the agents' favorite, but both outsource the expensive half — the inference backend, its GPU, its auth — to you, which makes their "cost" a second deployment you price separately. FauxPilot's Copilot-plugin compatibility is a legacy virtue on a narrowing base. Tabby is the only entry where one deployment is the whole product, so it is the only one with a single priceable surface. That is why nobody priced the others yet either — and why Tabby goes first.
The Copilot side, fully loaded
Copilot's 2026 pricing is a two-layer bill, and any comparison that quotes only the seat price is stale. The base tiers are unchanged and well documented: Free at $0 with limits, Pro at $10/month ($100/year), Pro+ at $39/month, Business at $19/user/month, Enterprise at $39/user/month. A 50-developer team on Business pays $950/month — $11,400/year — before the second layer.
The second layer arrived June 1, 2026, when GitHub replaced fixed premium-request quotas with usage-based AI Credits. The base fee now buys unlimited use of included models plus a monthly credit allowance — 1,900 AI credits per user on Business, 3,900 on Enterprise per GitHub's billing docs — and agent work, chat, and premium models draw down the balance. Overage is billed at $0.01 per credit ($10 per 1,000 credits), and it is enabled by default for organizations unless an admin explicitly disables the paid-usage policy. Autocomplete-heavy developers may never touch overage; agent-heavy developers — the ones running coding agents against premium models all afternoon — can burn past the allowance the way they used to burn through premium requests.
That behavior-dependence is why the opening table shows two Copilot columns. The no-overage column is the floor every team pays; the +$10/user/month column prices one extra thousand credits per developer per month — a modest heavy-agent scenario, not a pathological one. At 50 developers that is the difference between $11,400 and $17,400 a year, and unlike the seat price, it grows with how enthusiastically your team uses the product's best features. CloudZero's 2026 Copilot cost breakdown makes the same point from the buyer's side: what the flat fees buy now depends on how hard you push the AI.
The self-host side, fully loaded
Start with what the model actually needs, because "a GPU" spans three orders of magnitude. Tabby's supported list centers on the Qwen and DeepSeek coder families, and the sizing math is public. Qwen2.5-Coder 7B — 88.4% HumanEval, native fill-in-the-middle support — needs roughly 15GB of VRAM in FP16 and about 5GB quantized to 4-bit. That fits in the spare capacity of an inference node you already run for tenant workloads: one more deployment, a fraction of one GPU, an OIDC gate in front. Step up to Qwen2.5-Coder 32B (92.7% HumanEval) at 4-bit and you want ~22GB, which still fits one rented A100 80GB serving 1–15 concurrent autocomplete requests — or one owned 24GB card with room to spare.
From there the two owned scenarios diverge. Spare capacity is the marginal case: the hardware is already bought and already powered for tenant inference, so Tabby's incremental cost is the extra power draw (order of $11/month for ~100W of additional load at $0.15/kWh) plus ops. Dedicated box is the honest standalone case: a 4090-class build at roughly $3,000 ($2,000 GPU, $1,000 host — inside the $900–$2,500 band prosumer 24GB cards actually sell in), amortized over a 36-month tenure at $83/month, drawing ~450W under load for another ~$49/month in power.
And then there is ops, the line item self-hosting pitches like to wave off and this one will not. Steady state for a single pinned container behind SSO is small but nonzero: image and model updates, break-glass restarts, watching the queue when the team grows. Price it at 1–3 hours a month of a loaded $150/hour engineer — $150–$450/month — and put it in the table next to the hardware instead of in a footnote. The ranges in the opening table are that ops band plus the fixed lines: dedicated runs $282–$582/month ($3,384–$6,984/year), spare capacity runs $161–$461/month ($1,932–$5,532/year).
Now the sensitivity the headline promised. Divide each monthly total by $19/seat and the breakeven falls out: dedicated breaks even at 15–31 seats without Copilot overage (low-ops to high-ops), spare capacity at 9–25 seats. Add the +$10/user heavy-agent overage and the same totals break even at 10–20 seats dedicated and 6–16 seats marginal. The "11–31" headline is the middle of that surface — mid-to-high ops assumptions, with and without overage.
For calibration, the rented-GPU version of this math (Spheron's headline: one A100 80GB at $1.04/hour, $749/month, breakeven at ~39 seats) ignores ops entirely; add the same $300/month mid-ops line and rented breakeven slides past 50 seats. Renting the GPU is the no-capex sensitivity row, not the recommended configuration — it exists to show what happens if you self-host the software but not the iron.
The compliance half of the pitch
None of the above is why regulated teams actually switch. They switch because every Copilot request sends code to a third party. Business and Enterprise tiers offer data-retention controls — your code is not supposed to be retained or trained on — but it still transits external servers on every completion, and "transits but is not retained" is a weaker claim than "never leaves the building." For healthcare, fintech, legal, and defense work, several compliance frameworks treat that distinction as dispositive, and some handle it by forbidding the external round-trip outright rather than by auditing the vendor's retention policy.
Tabby inverts the data flow by construction. Inference runs on your GPU, on your network, behind your auth; credentials in scratch files, proprietary algorithms, and unreleased product code never leave your machines. There is no DPA to negotiate over prompts, no subprocessors list to review, no "trust us, we delete it" — the architecture is the guarantee. That is also why the "code never leaves" pitch pairs with per-seat pricing as an opponent: SaaS per-seat billing scales with headcount and centralizes your code on someone else's infrastructure, while self-hosted inference flattens cost with headcount and keeps the code home. The teams that care most about the second half — regulated shops with growing headcounts — are exactly the teams for whom the first half already wins.
The decision rule, then, is a two-by-two most teams can place themselves in without a spreadsheet. Under ~10 developers with no compliance constraint: pay per seat, and revisit when headcount or agent-overage pushes the bill past a box. Over ~25 developers, or any size with a data-sovereignty requirement: price the dedicated box (or the spare slice of the inference fleet) against the fully-loaded SaaS bill, overage included. In between: the ops band decides — if running one more container is already your team's Tuesday, the breakeven is behind you.
The per-seat model won the last decade by making the buying decision trivial: count developers, multiply by nineteen. Self-hosted coding assistance wins the next one back the boring way — a $3,000 box, a pinned model, a team login page — by making the cost curve flat and the data flow short. Tabby is simply the first assistant whose price you can compute on one hand instead of two. Run the numbers for your headcount; the table at the top is waiting.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



