Skip to main content

Railway Agent Runs on Your ChatGPT Subscription: What BYO-Model Billing Concedes About PaaS Agent Margins

9 min readDora NodaDora Noda
Share
On this page

Railway just told its users to stop paying Railway for AI tokens. On September 17, changelog #0308 announced that your ChatGPT subscription can now power Railway Agent: connect your OpenAI account under Account Settings → Agents, pick GPT-6 Astra, GPT-5.6 Sol, or GPT-5.6 Terra from the OpenAI Subscription group, and usage counts toward your ChatGPT plan instead of Railway's model billing. The agent runs in the same places — the Railway Agent TUI and the agent in your Railway dashboard — but the meter belongs to someone else now.

That is a remarkable sentence for a PaaS to publish about its own flagship agent surface. Railway spent 2026 building agent distribution — a hosted Remote MCP server, railway agent in the CLI, and a one-command skills install back in April — and now it has unbundled the model bill from all of it. This post does the math on what that concedes: the token margin is structurally OpenAI's now, the PaaS is left competing purely on agent UX and compute, and a self-hosted deploy agent should treat BYO-subscription as the default billing shape rather than a fallback.

The money table: three ways to pay for agent tokens​

Strip the announcement to its economics and there are exactly three shapes a platform can use to charge for the tokens its agent burns. The table below prices a month of agent coding traffic under each shape at three usage levels, using GPT-6 Astra's published API rate — $5 per million input tokens, $25 per million output tokens — blended to roughly $10 per million for a coding-agent mix that is heavy on input context. The markup row assumes an illustrative 2x platform markup on metered tokens; the subscription row uses ChatGPT's actual plan prices (Plus $20/month, Pro $100–$200/month).

Monthly usageNative metered billing (2x markup)At-cost pass-throughBYO ChatGPT subscription
Light (~2M tokens)~$40~$20$20 flat (Plus)
Medium (~15M tokens)~$300~$150$100–$200 flat (Pro)
Heavy (~60M tokens)~$1,200~$600$200 flat (Pro, within limits)
PaaS keeps~50% margin on tokens$0 on tokens$0 on tokens

Read the last row first, because it is the whole story. Under native metered billing, every extra agent loop a user runs is platform revenue. Under pass-through, the platform is a pipe. Under BYO-subscription, the platform is not even the pipe — OpenAI meters, caps, and collects, and Railway's changelog says so in plain language: "Your plan's model availability and usage limits apply."

Now read the columns as a user. At light usage the three shapes cost about the same, which is why nobody argued about billing models when agents were a novelty. The divergence starts at medium usage and becomes a cliff at heavy usage: a power user burning 60M tokens a month pays something like $1,200 under marked-up metering, $600 at cost, or $200 flat on a Pro subscription — if the work fits inside the plan's limits. That 6x gap between the first and third column is the pressure that forced this changelog. No agent UX is good enough to make users ignore paying six times more for the identical model weights.

The honest caveat runs the other direction, and Railway states it: subscription plans have usage limits, and overflow still needs a metered path. A flat $200 month that throttles mid-sprint is not strictly better than a $600 month that never stops. But notice what that caveat concedes — the metered path is now the overflow lane, not the default. The default is OpenAI's bundle, and the PaaS collects nothing on it.

Why Railway conceded the margin: a five-month timeline​

This did not happen all at once. The concession arrived in three steps, each one narrowing what a PaaS can charge for between the user and the model.

April 2026: distribution first. Changelog #0286 shipped three agent surfaces in a single week — a hosted Remote MCP server, railway agent in the CLI, and a one-command skills install for the major coding editors. Railway was racing to own the surface where developers meet agents, and it paired that land grab with the first billing concession: tokens passed through at cost, with no platform markup. The margin on intelligence went to zero five months before September; #0308 just made the plumbing match the pricing.

August 2026: OpenAI builds the on-ramp. OpenAI's "Sign in with ChatGPT" entered beta with named partners, letting third-party apps authenticate users against their ChatGPT account and bill model usage to the user's own plan — the same shape as "Sign in with Google," except the resource being shared is a compute subscription, not an identity. Sam Altman had floated the idea back in 2023 and OpenAI formally solicited developer interest in May 2025, so platforms had over a year to see it coming. The beta turned a hypothetical into an integration any PaaS could ship in weeks.

September 3–17, 2026: the flagship model raises the stakes. OpenAI launched GPT-6 Astra on September 3 — billed as "the world's best computer use model," scoring 72.6% on OSWorld 2.0 against GPT-5.6 Sol's 65.7%, at API rates 2.5x Sol's. Two weeks later, Railway's #0308 put Astra, Sol, and Terra behind the ChatGPT-subscription group in Railway Agent. The timing matters: the more capable the frontier model, the wider the gap between subscription-flat and metered pricing, and the harder it becomes to defend a per-token markup on exactly the model every power user wants.

And do not miss what else shipped in #0308: Railway Sandboxes went generally available — isolated Linux VMs with Claude Code, Codex, OpenCode, and Pi preinstalled, checkpointable, forkable, and attached to your environment's private network, with usage drawing "from the same included credit as your services." Read the two halves of the changelog together and the strategy is legible: concede the token margin you cannot defend, and route the same workloads toward the compute margin you can. The agent thinks on OpenAI's dime and executes on Railway's VMs.

What the PaaS competes on now: agent UX and execution​

Once the model bill belongs to OpenAI, what is left for the platform to sell? Railway's answer is visible in the surface inventory: the Agent TUI, the dashboard agent, Discord and Slack bots, railway code for terminal-driven cloud agents, one-command skills installs, and now Sandboxes as the execution substrate. None of these sell intelligence. All of them sell proximity — the agent lives where the app lives, sees the logs, touches the private network, and checkpoints a VM instead of describing one.

That is a real moat, but it is a shallower one than a token margin. Model quality used to differentiate platforms passively: whichever PaaS integrated the best model first won, because the user could not bring their own. BYO-subscription inverts that. Every platform now offers the same frontier models at the same flat price, so differentiation has to come from everything around the model — how fast the agent can see a deploy log, whether it can fork a sandbox and try three fixes in parallel, how cleanly skills install into the user's editor. It is a UX race with commodity fuel.

The rest of the industry is converging on the same shape from different directions. Cursor moved to token-level credit billing after its flat-rate quotas proved unsustainable; Copilot shifted toward usage-based billing in June 2026; open-source harnesses like Aider, OpenCode, and Pi ship BYOK-native and free, competing purely on workflow while the user pays the model vendor directly. Railway's move is the managed-PaaS version of the same admission: the token margin was always the model vendor's to reclaim, and September 2026 is when OpenAI reclaimed it through a login button.

What a self-hosted deploy agent should copy​

Here is the good news for anyone building a deploy-from-chat agent on infrastructure they own: you never had a token margin, so you have nothing to concede. Railway just validated — with its own revenue — the billing shape a self-hosted platform should have shipped from day one. Copy these four decisions:

1. Make BYO-subscription the default, not the fallback. Do not build a token meter, a credit ledger, or a markup table. Authenticate the user's existing model subscription (ChatGPT today, others as their login protocols mature) and let the vendor bill. Your pricing page should sell seats and compute, never tokens.

2. Compete on execution, not intelligence. Railway pairs the concession with Sandboxes for a reason: the VM the agent runs in is the thing the platform owns outright. A self-hosted PaaS already owns its machines — give the agent a sandbox on them, wire it to the private network, preinstall the harnesses, and you have matched the differentiator without matching the burn rate.

3. Keep one metered overflow lane. Subscription limits are real, and a throttled agent mid-incident is how you lose trust. Offer at-cost API pass-through as the overflow for users who blow past plan limits — metered, transparent, zero markup — and say so upfront. It is a safety valve, not a business model.

4. Treat model choice as user configuration. Pin no default model, negotiate no exclusive provider deal, build no feature that only works on one lab's weights. The platform that is neutral on models is the platform that survives the next login button — when the second lab ships its version of Sign in with ChatGPT, you add a second Connect button instead of renegotiating your margins.

The login button is the business model now​

Six months ago, a PaaS agent strategy meant picking a model, marking up its tokens, and racing to integrate the next one first. Railway's September changelog ends that era: the model vendor now owns the customer billing relationship for intelligence, the platform owns the sandbox and the surface, and the boundary between them is a Connect button in account settings.

That is a healthier split than it looks. Token markups were always a tax on the users who used the agent most — the power users every platform claims to want. Flat vendor subscriptions align the heaviest usage with the flattest price, and platforms get to compete on something they actually control: how good the agent is at operating the infrastructure it lives next to. The PaaS that wins the next phase is not the one with the best model deal. It is the one whose agent can see the logs, fork the sandbox, and ship the fix — on whatever subscription the user already pays for.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide