Software firm Atlassian has published an account of its shift from gostatsd to OpenTelemetry, overhauling a metrics platform that handles data from roughly 100,000 hosts spanning 14 regions. The existing service maintained a 99.95% service-level objective, forcing the team to swap out the platform without breaking the metrics tied to production alerts. In a post on the CNCF blog, engineers Iris Grace Endozo, Farzad Vazirnia and Albert Kerr explain that while gostatsd had performed reliably for years, its UDP-only architecture couldn't support traces or logs, and more services were generating OpenTelemetry data.

The platform now handles approximately 4.8 billion data points every minute and stores around 220 million, cutting volume by roughly 96%. The revised aggregation tier consumes about half the CPU for identical traffic levels. Combining metrics into the tracing sidecar saved an average of 3.9% CPU for each of Atlassian's most resource-intensive Micros services, and the team estimates it reduced sidecar cost by roughly 30% across the entire fleet. The old nomad proxy and gostatsd aggregation together accounted for about 38% of CPU requests in Atlassian's metrics clusters, with nomad alone consuming approximately 13% of total resources.

The engineers kept the StatsD-over-UDP interface intact and rebuilt the pipeline behind it, turning what could have been an organisation-wide application migration into a migration owned by the platform team. According to the authors, small tests and benchmarks weren't sufficient—continuous profiling in production actually revealed where to focus optimisation efforts. The report notes that adding another destination is now a configuration change rather than a separate integration project. Since upstream components didn't aggregate delta metrics the way Atlassian needed, the company developed and open-sourced its own aggregation processor.

Atlassian's approach avoided forcing thousands of services to switch immediately from StatsD clients to the OpenTelemetry SDK, which would have required reproducing features already being built by the OpenTelemetry Collector community. The new system uses the OpenTelemetry contrib load-balancing exporter to hash by stream ID, so individual time series stay together while one large service can spread across the pool, producing more even CPU distribution, fewer overloaded shards and better off-peak scaling. The ingest tier presented a challenge because metric aggregation is stateful—every point for a time series must reach the same aggregator, and the old system hashed each service and environment to one shard, concentrating the largest services on a few busy replicas.

The migration guidance centres on operational practice. Atlassian started with development and staging workloads, selected early adopters that stood to benefit, then increased rollouts through 1%, 10%, 50% and 100% stages. The team recommends preserving familiar operational workflows while old and new systems run side by side, and profiling under production load instead of relying on small benchmarks. Removing gostatsd aggregation and nomad will complete the end-to-end OpenTelemetry pipeline, and the next planned step is to migrate application instrumentation from StatsD, DogStatsD and vendor clients to OpenTelemetry SDKs. The compatibility layer addresses an environment where immediate re-instrumentation would have been a longer and riskier programme. Organisations with sprawling legacy instrumentation now face a pivotal choice between mandating SDK rewrites or building abstraction layers that defer the transition cost indefinitely. The success of this middle path may determine whether OpenTelemetry achieves ubiquity or remains confined to greenfield projects.