Red Hat this week released Red Hat AI 3.5, a platform update engineered to allow software teams to operate AI workloads with the same control and rigor as mission-critical enterprise applications. The release centers on expanded multi-tenancy capabilities that let AI service providers run multiple workloads on shared GPU infrastructure while maintaining complete hardware-to-software isolation for sensitive data, proprietary models, and regulated information. The platform also introduces priority-aware service routing, which ensures high-stakes workloads execute ahead of lower-priority tasks on the same compute resources.
The new release delivers several technical capabilities aimed at production AI deployment. Red Hat AI 3.5 now includes EvalHub for pre-deployment model verification through safety benchmarking and regulatory compliance certifications. Real-time observability dashboards provide platform teams with visibility into inference health, GPU utilization, and model performance metrics, while non-admin users can track per-user token consumption and distributed inference workloads. Fair-share GPU scheduling manages resource allocation across tenants, and priority-based request routing protects real-time inference while allowing background workloads to use available capacity. The platform also adds CPU offloading as a general release feature and storage offloading as a developer preview, letting models handle longer conversations and larger documents without requiring additional GPU hardware. Agent templates and starter kits offer pre-configured implementations for code review, document processing, and research workflows.
Tushar Katarki, Red Hat's Senior Director of Product for Red Hat AI, said operating enterprise AI without safety controls is "like driving a supercar blindfolded." He added that Red Hat AI 3.5 delivers "the operational guardrails, verifiable trust, and multi-tenant controls needed to run AI as a mission-critical service rather than an unpredictable experiment." The report notes that the platform unifies pre-deployment safety benchmarking, real-time observability, and GPU resource management to help platform teams transform isolated AI pilots into fully governed enterprise architecture. Applied mathematician Joshua Estrin told the publication that Red Hat's approach reflects current market needs because "every GPU request now becomes a priority decision," meaning one developer's internal experiment can't be treated with the same urgency as a financial close. Anindo Sengupta, VP of product management at Nutanix, stated that running multi-tenant AI at scale requires secure tenant isolation best achieved through virtualization, while real value emerges when agents can access both LLMs running on containers and enterprise systems on traditional infrastructure.
The report explains that GPUs remain expensive, and as organizations shift to live production use cases of agentic AI, they need to maximize GPU capacity across multiple workloads, teams, customers, and applications. This makes AI infrastructure efficiency the new bottleneck for agentic systems. Priority-aware services for native multi-tenancy on shared GPU infrastructure dynamically allocate GPU capacity based on workload priority, allowing lower-priority tasks to run on spare or cheaper capacity rather than requiring separately provisioned GPU resources. Isolation techniques enable AI services to run without accessing or interfering with another service's data, models, or compute environment, combining hardware consolidation with strong tenant separation. Yoram Novick, CEO of Zadara, previously noted that when teams scale AI, simply adding more GPUs without ensuring adequate interconnect bandwidth leads to diminishing returns. Red Hat aims for its platform to establish a state where GPU-based AI resources are managed as a policy-controlled infrastructure pool through priority-aware inference, tenant isolation, capacity sharing, and observability. The emphasis on verified safety before deployment, precise resource controls across shared GPU environments, and transparent usage metrics reflects the operational rigor required for mission-critical infrastructure. For teams navigating the shift from proof-of-concept to enterprise deployment, the success of this approach will hinge less on technical architecture and more on whether organizations can articulate clear service-level expectations before the first production incident forces the conversation. The real test isn't whether the platform can isolate tenants—it's whether leadership can decide which workloads deserve priority when capacity runs tight.

