"Why did CI take forty minutes yesterday?" On a GitHub Actions org of any real size, that question has no good answer. GitHub's own Actions insights are per-repo and shallow: no cross-org view of CI health, no way to slice slow or flaky workflows by team, no queue-time history.
The usage APIs are worse — billing-oriented minute counters, and a moving target at that: GitHub has spent the last year retiring per-workflow usage endpoints in favor of its enhanced-billing summary API. You can learn what CI cost. You cannot learn why it was slow.
The instinctive fix — instrument every workflow with tracing steps and SDK calls — works exactly until the org grows. Every team must opt in, every new repo starts blind, and the platform team ends up maintaining instrumentation scattered across hundreds of YAML files. It scales with everyone's diligence about something that isn't their job.
There is a better answer, and it was laid out in a September 8, 2026 CNCF post by George Sims: GitHub already emits almost everything you want as workflow_run and workflow_job webhook events. Listen once, at the org level, convert the events to OpenTelemetry spans, and every repo — including ones that don't exist yet — is traced. Here is the whole recipe, the span model it produces, and the operator's checklist for running it on your own infrastructure.
The setup in one picture
The architecture has three parts: one org-level GitHub webhook, one OpenTelemetry Collector running the contrib githubreceiver component, and any OTLP-compatible trace backend. Events flow in as webhooks and out as spans. Nothing is installed into any repository.
The GitHub side takes four clicks. As an org admin, go to Organization Settings → Webhooks → Add webhook. Set the payload URL to your collector's endpoint, pick application/json, set a secret, and — this is the step people fumble — under "Let me select individual events," check exactly workflow_run and workflow_job, nothing else. Those two event families carry run, job, and step timing for the entire org. A repo-level webhook would work too, but it repeats per repo and misses future repos; the org webhook is what makes the coverage total.
The collector side is one config file:
receivers:
github:
webhook:
endpoint: 0.0.0.0:19418
path: /events
secret: ${env:GITHUB_WEBHOOK_SECRET}
scrapers: # required even if you only want tracing, a dummy entry is enough
scraper:
github_org: ${env:GITHUB_ORG}
exporters:
otlp:
endpoint: ${env:TRACE_BACKEND_ENDPOINT}
headers:
authorization: ${env:TRACE_BACKEND_API_KEY}
service:
pipelines:
traces:
receivers: [github]
exporters: [otlp]Yes, that scrapers block looks like it wandered in from a different feature. It did — it belongs to the receiver's separate GraphQL/REST metrics scraping — but config validation fails without at least a dummy entry even when all you want is the webhook path. Budget twenty lost minutes for that the first time, or zero now that you've read this paragraph.
The span model is the part that makes it worth doing. Each workflow run becomes one outer span, each job a child span, each step a child of its job, plus queue spans capturing the wait before a job starts. What lands in your backend is a drillable trace tree, not a flat pile of events:
| Span | Source event | What it tells you |
|---|---|---|
| Workflow | workflow_run | Total run wall time, conclusion, triggering actor |
| Job | workflow_job | Per-job duration, runner label, conclusion |
| Step | workflow_job (steps array) | Which step ate the minutes inside a slow job |
| Queue | workflow_job queued → in_progress | Time waiting for a runner before any work began |
And the zero-instrumentation claim is literal, in both directions. No workflow file is created, modified, or even read by you — the telemetry source is GitHub's own event stream, which fires whether or not anyone opted in. And because the subscription point is the org, a repo created next Tuesday is observable from its first run with nobody remembering to switch anything on. Per-workflow instrumentation decays with every new repo; this approach appreciates.
What the spans buy a fleet running its own runners
Cross-org slow/flaky/queued-job visibility is the headline, but it lands differently depending on whose runners execute the jobs. If you rent GitHub-hosted runners, queue time is mostly GitHub's capacity problem to observe and complain about. If you run your own runners on hardware you own — static self-hosted runners or an autoscaled pool — queue time is your capacity signal, and before this setup you largely didn't have it.
Three queries become possible that GitHub's UI cannot express. Slowest workflows org-wide over the last 30 days, ranked by p95 run time, sliced by team or repo topic. Flakiest jobs, defined as jobs whose conclusion flips between success and failure across retries of the same commit — the shape of a test that fails one run in five and eats everyone's morning.
And queue-time distribution by runner label: the direct measurement of whether your runner pool is sized right, per pool, per hour of day. Queue depth tells you runners are busy; queue time tells you developers are waiting. Only the second one shows up in these traces, and it's the one that maps to money and morale.
One more detail compounds over time: span and trace IDs are deterministic, hashed from GitHub's own run and check-run IDs. Under the current default scheme, a job span ID is sha256("{check_run_id}-j") sliced to 16 hex characters, with step and queue spans following the same pattern. That means telemetry emitted from inside a step can compute the matching ID from ${{ job.check_run_id }} and attach to the same trace with no coordination with the collector. Today's zero-instrumentation traces become the skeleton that tomorrow's in-step spans hang off — profiling data, test timings, deployment markers — without re-architecting anything.
The backend is the least interesting decision
Because the collector exports plain OTLP, the backend is a one-line endpoint swap: Tempo, Jaeger, Datadog, Honeycomb, whatever the team already pays for. Sims makes the point bluntly and he's right — the moment you standardize on OTLP, the vendor choice stops being load-bearing.
For a team that self-hosts, the natural landing spot is Grafana Tempo, and the reason is architectural: Tempo needs only object storage to operate — S3, GCS, Azure Blob, or a self-hosted MinIO bucket — with no database to run. A single-binary Tempo pointed at an object bucket is a credible trace backend for a small platform team, and CI traces sitting in the same backend as application traces buys the correlation the CNCF post closes on: a slow deploy and a slow downstream service investigated in one tool instead of two dashboards owned by two people who don't talk until Thursday.
The operator's checklist: six things before you deploy
The config above will run. These six items are what make it production-worthy, drawn from the CNCF post and the receiver's own docs. Skipping them is easy precisely because the zero-instrumentation part feels so complete.
1. Pin the collector version. The githubreceiver is still at alpha stability in contrib, which means the config surface can shift between releases. Pin a specific collector version rather than tracking latest, and skim the changelog before bumping. This is cheap insurance against a field quietly changing shape under a webhook endpoint you can't afford to bounce.
2. Allowlist GitHub's webhook source ranges. The collector endpoint is publicly reachable — GitHub has to deliver to it — so don't lean on the shared secret as your only line of defense. GitHub publishes its webhook egress CIDRs at api.github.com/meta under the hooks key; restrict the webhook path to those ranges at your reverse proxy or WAF, refreshing the list on a schedule.
The receiver also supports required_headers, which rejects any request missing an agreed header — designed for exactly this topology, where a front-end proxy injects a header GitHub itself would never send. There is also a GitHub App authentication option if shared-secret rotation across many services is already a headache.
3. Confirm org-admin access early. Creating an org-level webhook requires org admin. That sounds trivial until rollout day, when the platform engineer doing the work discovers they were never granted it and the person who holds it is on leave. Confirm the permission before you schedule anything.
4. Validate on GitHub Enterprise Server. If your org lives on GHES rather than github.com, verify webhook delivery behaves the same in your version before assuming drop-in compatibility. Event payloads and delivery semantics track the cloud closely but not identically.
5. Respect step-name uniqueness. Under the default UseCheckRunID scheme (on by default since collector v0.151.0), step span IDs hash the check-run ID together with the step's raw name: field. Steps with duplicate or missing names inside one job can therefore produce colliding span IDs.
The legacy scheme had the analogous constraint one level up — unique job names per workflow — so this is a narrowed version of an old rule, not a new one. If your workflows reuse generic step names like "Run tests" twice in one job, disambiguate them or accept merged spans.
6. Name an owner. Someone owns the collector, the webhook secret rotation, and whatever alerting gets built on the trace data — going forward, not just on launch day. Sorting that out early saves the awkward moment three months later when delivery breaks and nobody is sure whose pager it is.
Size it before you build it
"Just turn on tracing for everything" is how a platform team ends up in an unwelcome conversation with whoever owns the tracing budget. The sizing method in the CNCF post is the part worth copying even if you adopt nothing else, because it generalizes to any org-wide telemetry rollout.
Start by scanning the org for total repo count — then throw that number away. Most orgs of real size carry a long tail of dormant, forked, archived, and abandoned repos that inflate the headline without generating CI traffic. What matters is the active slice: repos with genuine workflow activity in a representative week. From there, multiply outward:
active repos × average runs per repo × average steps per run × typical span payload ≈ daily trace volume
Then compare that volume against the application-trace volume your infrastructure already handles. In most orgs, CI trace volume turns out to be a rounding error next to it — and walking into the budget conversation with that comparison instead of "trust me, it's small" turns a drawn-out justification into a five-minute one. The method is the deliverable here, not any specific number: measure real activity, not headline repo count, and arrive with a ratio rather than an assertion.
What this doesn't replace: your runner fleet metrics
If you run self-hosted runners on the Actions Runner Controller, be precise about which layer this new telemetry covers, because the two look deceptively similar. ARC's metrics describe your runner fleet: how many runner pods exist, whether autoscaling is keeping up, how deep the pending-job queue is. The githubreceiver traces describe your workflows: why a specific run was slow, which jobs are flaky, where the minutes actually went.
Queue depth and queue time are the same question asked from two vantage points — the fleet's ("how backed up are we?") and the developer's ("how long did I wait?"). Worth running both, not picking one. The fleet metrics page whoever owns capacity; the traces answer whoever asks why yesterday's deploy took forty minutes.
Where it stops
Collection is only half an observability story, and Sims is explicit that his post covers just that half. Sitting on top of these traces should be alerting — queue-time thresholds per runner pool, flaky-test detection with enough statistical care to avoid crying wolf, conclusion-rate tracking per workflow — and that alerting has its own failure mode, which is noise. Thresholds set too tight train everyone to ignore the channel within a fortnight; set too loose and the 3 a.m. page arrives for something nobody can act on.
That is genuinely a separate problem — trace-derived alerting with calibrated sensitivity — and probably its own post. What the webhook-plus-collector setup buys you is the precondition: without queryable queue-time and conclusion history, there is nothing to alert on at all.
CI observability that covers every repo from a single webhook is the same shape as self-hosting done right: own the collection point once, and every workload inherits it. Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



