The transition from experimental AI pilots to production-grade infrastructure requires a shift from simple model testing to rigorous operational control. Red Hat is attempting to address this gap with the release of Red Hat AI 3.5, a platform update designed to move AI workloads into the realm of mission-critical enterprise services. By integrating safety benchmarking, multi-tenancy, and granular observability, the company is positioning its portfolio to manage AI as a shared, governed resource across hybrid cloud environments. This release specifically targets the friction points encountered by platform engineering leaders, such as GPU resource contention, model risk, and the lack of transparent cost attribution in distributed AI architectures.
Red Hat AI 3.5 Safety and Observability Framework
Red Hat is introducing several mechanisms to provide what it describes as verifiable trust before models reach production. Through the release of EvalHub, the platform enables risk-focused safety benchmarking and the generation of regulatory compliance certifications for custom models, RAG, and agents. This is supported by an expanded model catalog that now includes over 20 validated models from providers such as Google, NVIDIA, and Alibaba Cloud. These models feature integrated Garak safety scores, PII exposure metrics, and toxicity risk assessments to assist enterprises in assessing model risk.
To manage the operational side of these deployments, Red Hat AI 3.5 introduces new observability dashboards. These tools provide platform teams with real-time metrics regarding inference health, GPU utilization, and model performance. Crucially, the update includes capabilities for non-admin users to access dashboards for per-user token consumption showback, addressing the growing need for cost attribution in shared AI environments. Additionally, the platform is introducing Inference-Time Scaling (ITS), which the company claims can optimize GPU expenditure by dynamically adjusting compute resources based on the complexity of incoming queries.
Multi-Tenancy and Agentic Development Capabilities
A core component of the Red Hat AI 3.5 update is the enhancement of hardware-to-software isolation for multi-tenant environments. The platform now officially supports running on Red Hat OpenShift hosted control planes deployed on Red Hat OpenShift Virtualization. This architecture allows each tenant to maintain a dedicated cluster control plane while consolidating the underlying hardware. By supporting AI workloads within Red Red Hat OpenShift Virtualization virtual machines, the company aims to provide robust VM-level isolation across shared, GPU-enabled infrastructure.
For developers building autonomous workflows, Red Hat is deploying AutoRAG to link enterprise data repositories directly to agentic applications. This includes multilingual document support, conversational testing, and contextual retrieval. To accelerate deployment, the AI Hub now provides agent templates and starter kits for common patterns such as research workflows, document processing, and code review. These templates are designed to integrate frameworks and deployment configurations within sandboxed environments to maintain security policies from initial deployment through to full production.
Key Takeaways
- Red Hat AI 3.5 introduces EvalHub to automate safety benchmarking and auditable compliance reporting for custom models and agents.
- The platform supports multi-tenancy through Red Hat OpenShift Virtualization, providing VM-level isolation for workloads sharing GPU-enabled infrastructure.
- New observability features include per-user token metering and real-time dashboards for GPU utilization and inference health.
TechInsyte's Take
In our view, Red Hat is making a calculated bet that the next phase of AI adoption will be won not by those with the most powerful models, but by those with the most disciplined infrastructure. By focusing heavily on "showback" capabilities, token metering, and VM-level isolation, Red Hat is signaling that the era of the "unmanaged AI experiment" is ending. The move to integrate safety scores directly into the model catalog suggests a recognition that enterprise legal and compliance departments are becoming primary gatekeepers for AI deployment. While the success of these features depends on how effectively they can reduce the manual overhead of managing fragmented AI tools, the emphasis on turning AI into a standardized, multi-tenant enterprise service is a logical evolution for hybrid cloud providers facing the complexities of shared GPU resource management.
Questions & Answers
How does Red Hat AI 3.5 address the challenge of GPU resource contention in shared environments?
The platform utilizes priority-aware serving and fair-share GPU scheduling. This allows for admission control and priority-based request routing, which protects real-time inference workloads while allowing background tasks to utilize available capacity.
What specific safety metrics are provided for the models in the Red Hat Catalog?
Validated models in the catalog now include integrated Garak safety scores, PII (Personally Identifiable Information) exposure metrics, and toxicity risk scores to provide transparency during the risk assessment process.
In what way does the update support cost management for AI services?
Red Hat AI 3.5 provides observability dashboards that allow for per-user token consumption showback, enabling organizations to track and attribute usage costs across different users or departments.
How can enterprises ensure isolation between different AI tenants?
Organizations can use Red Hat OpenShift hosted control planes on Red Hat OpenShift Virtualization to provide each tenant with a dedicated control plane and robust VM-level isolation on shared, GPU-enabled hardware.
Source: Businesswire