Microsoft now sells a machine that rewrites your agent while you sleep. At Build 2026, the Foundry team demoed an Agent Optimizer that reads your agent's production traces, writes its own evaluation rubric from what it sees, generates improved prompts, skills, tool descriptions, and model picks, scores every candidate against that rubric, and deploys the winner as a new agent version — then feeds the new traces straight back into the next round. The key line from the demo: every run feeds the next eval, and every eval tells the optimizer where to improve next.
That is the observe → evaluate → optimize → deploy loop, closed and productized. If you run hosted agents on Foundry, it arrives as a feature. If you run a deploy-from-chat platform on machines you own, it arrives as a parts list you have to assemble yourself. This post inventories exactly what the managed loop contains, maps each stage to the self-hosted components that reproduce it, names the gaps honestly in both directions, and then argues that the deploy half of Foundry's loop — the automatic prompt rewrite — is the one place where a Git-backed, human-reviewed pull request is the better design, not the consolation prize.
What Foundry actually shipped
Start with the receipts, because "self-improving agents" has been demo-ware for years and this time the shipping details matter. The Agent Optimizer was announced June 3, 2026 in private preview with public preview roughly 30 days later, inside Foundry Agent Service. It sits on top of two foundations that already went GA: a unified OpenTelemetry pipeline carrying every model call, tool invocation, sub-agent hop, and handoff (with Application Insights auto-injected), and an Evaluations surface with out-of-the-box evaluators, custom evaluators, and continuous production monitoring piped into Azure Monitor.
The loop itself works in four stages. Observe: production traces accumulate in the Foundry control plane with evaluations linked back to the trace. Evaluate: the agent is scored against criteria — and in the Build demo, Foundry generated personalized evaluation criteria from the production traces themselves, covering governance and outcome quality, rather than asking the team to write a rubric from scratch.
Optimize: the optimizer reads the agent's current prompts and skills, searches for configurations that increase quality on the team's scenarios and constraints, and surfaces ranked candidates with full diffs, lineage, audit trail, and rollback. It tunes four things: the model, the instructions, the tool descriptions, and the skills. Every candidate is shown side by side — what improved, what regressed, and why. Deploy: promote the winner as a new agent version, and the new traces feed back into evaluation. Microsoft even published a hands-on agent-optimization workshop with one lab per node of the loop, which tells you how seriously they want this to become the default Agent DevOps motion: ship, observe, improve, re-ship, without leaving Foundry.
Note what makes this a loop rather than a tool: the output of every stage is the input of the next, and the flywheel spins without a human in the middle. That is the bar. Now here is the self-hosted version of the same machine.
The self-hosted parts list, stage by stage
There is no single open-source Agent Optimizer. There is, however, a mature OSS component for every stage — the work is the wiring between them. Here is the map:
| Loop stage | What Foundry gives you managed | Self-hosted substitute | What you wire yourself |
|---|---|---|---|
| Observe | Unified OTel trace pipeline, App Insights auto-inject, trace-linked evals | Self-hosted Langfuse or Arize Phoenix over OpenTelemetry GenAI conventions | Emit spans from your agent runtime; keep trace retention and PII redaction sane |
| Evaluate | Evaluations GA + auto-generated criteria from prod traces | Eval dataset seeded from prod traces + LLM-as-judge experiments runnable in CI | The rubric: which scenarios, which judges, what pass bar — this is the most human judgment in the loop |
| Optimize | Ranked candidate configs across model/instructions/tools/skills | GEPA reflective evolution (or DSPy) with an MCP tool-description adapter | The rollout harness: evaluate() runs your real agent, reflection proposes mutations, budget caps the search |
| Deploy | One-command winner promotion with diffs, lineage, audit, rollback | Git-backed PR promotion with human review + branch protection | The promotion gate: candidate lands as a PR, CI re-runs evals, a human merges |
Walk each row, because the table compresses real engineering decisions.
Observe is the most solved row. Langfuse is open-source and self-hostable (ClickHouse plus Postgres), ingests OpenTelemetry natively alongside its own SDKs, and ships datasets, evals, and prompt management as first-class features — plus a CLI and MCP server in the OSS product, so a coding agent can instrument your app and analyze production traces without leaving the editor. Arize Phoenix is the lighter alternative: a self-hostable trace viewer and eval workbench that runs in a notebook or as a server. Either way the standard is the OpenTelemetry GenAI semantic conventions, which is exactly what Foundry's own pipeline speaks — emit compliant spans and your traces are portable by construction. The wiring you own: actually emitting the spans (model calls, tool calls, sub-agent hops — the same three Foundry auto-injects), plus retention and redaction policy, because production traces contain user data and somebody has to decide what survives.
Evaluate is where the human judgment lives, and Foundry's demo half-admits it. Auto-generating criteria from traces is a starting point, not a conscience: the rubric encodes what "good" means for your agent on your scenarios under your constraints, and no vendor can write that for you. The self-hosted pattern is well-established: seed an eval dataset from production traffic (Langfuse's dataset abstraction exists precisely for this), define LLM-as-judge evaluators, and run the experiment suite in CI so every candidate gets a number, not a vibe. The wiring you own is the dataset curation loop — which traces become regression cases, how often the dataset refreshes as behavior drifts — and the pass bar. Spend your senior-engineer hours here. A great optimizer against a sloppy rubric just converges faster on the wrong thing.
Optimize is the row that recently got good. The technique to know is GEPA — reflective prompt evolution (arXiv:2507.19457): instead of mutating prompts blindly like DSPy's older MIPRO search, a reflector model reads the execution traces of failed cases and proposes targeted fixes for the specific semantic failures it observed. The standalone gepa-ai/gepa library ships a GEPAAdapter where you implement evaluate() (which runs your real agent as the rollout) and make_reflective_dataset(), and it orchestrates the propose-test-reflect cycle — including an MCP adapter that optimizes tool descriptions plus the system prompt, which maps directly onto the "tool descriptions" tuning Foundry demos. The wiring you own: the rollout harness that executes candidates against your eval dataset on your infra, and the budget that caps the search, because every candidate costs model calls and an unbounded optimizer is a billing incident with extra steps.
Deploy is the row where self-hosted should diverge from Foundry on purpose. Foundry promotes the winner with one command; the auditable move is landing the candidate as a Git pull request, re-running the eval suite in CI as a required check, and requiring human review before merge. This is the agentic GitOps pattern this blog has argued for: route agent-originated changes through a PR, not a live API call, so every promotion carries a diff, a discussion thread, branch-protection enforcement, and a revert button that works the same way for agent changes as human ones. A candidate prompt that passed evals but reads wrong to a human gets caught at review; a candidate that regresses in production gets git revert instead of a vendor-console rollback flow. Automatic rewrite optimizes for loop speed. The PR optimizes for the audit trail — and for infrastructure-adjacent agents, the audit trail is the product.
The honest gaps, in both directions
Assembling the parts list is not the same as matching the product. Four things Foundry does that DIY does not get for free:
- Criteria generation. Foundry writes the first draft of your rubric from your traces. Self-hosted, the empty eval dataset stares back at you until a human seeds it. Expect the first dataset curation pass to be the slowest week of the project.
- Always-on continuity. The managed loop spins continuously — every interaction feeds the next round. A DIY harness typically runs on schedule or on trigger (nightly optimization job, CI-gated promotion), which is calmer on cost but slower to catch drift.
- Governed rollback UX. Diffs, lineage, audit trail, and one-click rollback come integrated. Self-hosted, you get better primitives (Git) but you assemble the UX yourself: eval-linked PR comments, promotion dashboards, revert runbooks.
- Cost-to-quality ROI. Foundry's July follow-up connected traces, business-value evals, and operating cost in one view. Self-hosted, that ledger is yours to build — and as this blog argued when Anthropic's proposed Agent SDK meter made the same point, you should build the ledger before the meter changes, not after an automation run produces an invoice nobody can explain.
And three things the self-hosted loop does better, which are not consolation prizes:
- Your traces stay on your machines. Production traces are user conversations, tool payloads, and failure modes — exactly the data a regulated or paranoid team least wants piped into a vendor's optimization flywheel. The DIY loop's data boundary is your own VPC.
- No per-seat meter on improvement. Continuous optimization consumes model calls; on Foundry that spend compounds inside the platform bill. On owned infra, the same GPU-hours or API calls are a flat, predictable line you can budget and cap with your own harness.
- The promotion artifact is a PR, not a console action. Diffs, review threads, required checks, revert — the entire industry already knows how to audit a pull request. No new compliance story to write, no vendor audit-log export to wire into your SIEM.
Read both lists before picking a side. The managed loop wins on time-to-first-flywheel; the self-hosted loop wins on data boundary, cost shape, and auditability. Neither wins on everything.
Rent or build, by team shape
If you are a solo builder or a small team whose agents live entirely inside one vendor's ecosystem, rent the loop. The weeks you would spend seeding eval datasets and building a rollout harness are weeks your agent is not improving, and Foundry's private-to-public preview cadence means the managed optimizer is available to try now. Keep one discipline from the DIY side anyway: make the promotion a reviewed decision, even when the console offers one-click deploy. "The optimizer said so" is not a change record.
If your roadmap has deploy-from-chat on owned infrastructure — agents calling your MCP deploy tools, running under capped credits, touching machines you operate — build the loop, but build it in stage order: traces first, eval dataset second, optimizer third, promotion gate alongside all of it. The most common failure mode is buying the exciting row (optimization) before the boring rows (traces, rubric) exist. An optimizer with no eval dataset is a random prompt generator with confidence.
And if you already operate past one box, hold the optimizer to the same standard as any operator: full-surface accountability. The bar is the Better-PaaS principle turned on bigger platforms — anything the loop can change must be observable, reviewable, and revertible through the same Git history everything else flows through. A self-improving agent whose improvements bypass review is not autonomous operations. It is unreviewed deploys with a smarter author.
The loop is the right abstraction whoever runs it. Foundry proved the flywheel spins; the only question left is where yours spins — inside a vendor's control plane, or on machines you own, with a pull request standing between every candidate and production.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with agents as first-class operators of a real multi-node control plane. Star the repo on GitHub or deploy your first app today.



