Engineers at Adobe have built an open-source system that lets teams access their own Prometheus metrics in shared Kubernetes environments without exposing data from other tenants, according to a detailed technical post published on InfoQ. The architecture is designed specifically for GPU-intensive settings, where teams need visibility into power draw and utilisation to determine whether costly accelerator resources are actually performing useful work. The approach uses existing Kubernetes and Cloud Native Computing Foundation technologies rather than introducing an entirely new metrics platform.
The system places a tenant-aware proxy between users and the central Prometheus server and optionally provides each tenant with its own smaller Prometheus instance. Requests flow through NGINX and kube-rbac-proxy for authentication and authorisation before reaching a multi-tenant Prometheus proxy, which identifies the tenant, discovers available Prometheus backends, and restricts queries to that tenant's namespace. A critical element is prom-label-proxy, which modifies incoming PromQL queries to enforce a namespace constraint before the query reaches Prometheus, preventing a tenant from deliberately constructing a query that accesses another namespace. The platform also introduces a Kubernetes custom resource called MetricAccess, allowing teams to declare which metrics they require using exact metric names, regular expressions, or PromQL selectors. Adobe reports that one example configuration reduced a tenant's stored series from more than 10,000 to roughly 300 when metricIsolation was enabled, meaning only the tenant's own series are collected.
The immediate driver is GPU visibility, the authors write. Accelerator capacity can represent a significant infrastructure cost, but allocation alone doesn't indicate whether those GPUs are doing useful work. The engineers describe finding a GPU that had remained at zero utilisation for 11 consecutive days despite being allocated and powered on. With access to metrics such as GPU utilisation, framebuffer memory, power consumption, and request rates, teams can build queries to identify idle GPUs, GPUs consuming power without corresponding application traffic, or workloads receiving traffic while their GPU capacity remains underutilised.
The problem is straightforward but difficult to solve safely, according to the report. A central Prometheus instance may contain metrics from thousands of namespaces, making it unsuitable for unrestricted tenant access. Giving every team query access could expose another team's data, while allowing large numbers of users to query the shared store can also create a performance and noisy-neighbour problem. Simply asking developers to include the correct namespace in their queries wouldn't provide a sufficient security boundary. The restriction must be applied by the proxy before the query reaches Prometheus. For teams requiring their own dashboards and alerts, the design can periodically remote-write a curated set of metrics into a tenant-specific Prometheus instance, creating another isolation boundary: data that is never collected into the tenant's store cannot subsequently be exposed through that store. This separation changes the operational model, with the central Prometheus continuing to provide infrastructure-wide collection while tenant-specific Prometheus instances handle the dashboards and queries that individual teams need, so the shared system doesn't have to serve every developer's routine observability queries.
The report notes that the underlying pattern—a tenant-aware observability layer built around authentication, query isolation, curated metric access, and optional per-tenant storage—is increasingly relevant as Kubernetes clusters become shared platforms for application teams, data workloads, and AI workloads. Platform teams need to provide sufficient observability for developers without turning a central telemetry system into either a security risk or a performance bottleneck. Several established approaches exist to the same underlying multi-tenancy problem, including Grafana Mimir and Cortex, which have multi-tenant isolation built into their architectures using tenant identifiers to scope metric queries. The Adobe approach differs in that it keeps a shared Prometheus environment while adding Kubernetes-aware access controls and namespace-based filtering around it. This pattern offers enterprises a blueprint for balancing developer autonomy with security when infrastructure becomes expensive enough that waste detection moves from operational hygiene to financial necessity.

