Red Hat has announced Red Hat AI 3.5, adding new safety, observability, multi-tenancy and agent development capabilities as enterprises move AI workloads from pilots into production environments.
The latest release is aimed at IT and platform engineering teams that need to manage AI with the same operational controls applied to other mission-critical infrastructure. Red Hat AI 3.5 brings together capabilities for evaluating models before deployment, monitoring inference and GPU usage, managing shared infrastructure and developing governed AI agents across hybrid cloud environments.
A key addition is EvalHub, which enables organisations to benchmark models for safety before deployment and generate compliance-related evidence. Evaluated models in the AI Catalogue now include integrated Garak safety scores, alongside information on personally identifiable information (PII) exposure and toxicity risks.
Red Hat has also added more than 20 validated models to its catalogue, including models from Google, NVIDIA and Alibaba Cloud. A selection of validated models has additionally been tagged for tool-calling, giving enterprises information when selecting models for agentic AI applications.
The release expands observability capabilities with dashboards covering inference health, AI model performance and GPU utilisation. Non-administrative users can also access per-user token consumption information to support usage monitoring and showback.
Multi-tenancy and GPU management
Red Hat AI 3.5 introduces expanded multi-tenancy capabilities for organisations sharing GPU infrastructure. Fair-share GPU scheduling manages resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing.
This allows real-time inference workloads to receive priority while background workloads can use available capacity.
Red Hat AI now also officially supports hosted control planes running on Red Hat OpenShift Virtualisation. Each tenant can have a dedicated cluster control plane while the underlying hardware is consolidated. AI workloads running in OpenShift Virtualisation virtual machines provide VM-level isolation across shared, GPU-enabled infrastructure.
The release also adds controlled model rollouts to manage traffic during model updates and minimise service disruption. CPU offloading is generally available, while storage offloading is available as a developer preview to help models handle longer conversations and larger documents without requiring additional GPU hardware.
“The conversation has moved from getting AI into production to running it at scale as trusted enterprise infrastructure, which requires safety evidence, governed agents, cost attribution and multi-tenancy,” said Joe Fernandes, Vice President and General Manager, AI Business Unit, Red Hat. “With Red Hat AI 3.5, we are delivering the operational controls, verifiable trust and agentic foundations IT leaders need to run AI as a safe, controlled and accountable enterprise AI architecture across the hybrid cloud.”
Building and governing AI agents
Red Hat is also expanding capabilities for enterprises developing agentic applications.
AutoRAG connects enterprise data repositories with agentic applications and adds support for multilingual documents, conversational testing and contextual retrieval. A visual pipeline allows teams to test RAG configurations before deployment.
The release also introduces agent templates and starter kits through AI Hub, with pre-configured implementations for use cases such as code review, document processing and research workflows. These are designed to integrate frameworks, tools and deployment configurations while maintaining security and operational controls.
Red Hat AI 3.5 includes general availability support for the Responses API and built-in RAG, providing an open-source interface for multi-turn agent conversations. Integrated NeMo Guardrails can intercept malicious tool calls.
For AI workloads requiring dynamic resource allocation, Inference-Time Scaling (ITS) adjusts compute based on query complexity. Red Hat said this can help organisations manage GPU spending by allocating additional resources to more complex queries while avoiding unnecessary capacity for simpler requests.
The platform also supports multimodal serving for text, audio and image generation through vLLM Omni in early access.
Extending AI infrastructure across clouds
Red Hat has expanded distributed inference capabilities beyond OpenShift to third-party Kubernetes environments. The LLM-distributed inference technology is now generally available on CoreWeave CKS and Microsoft Azure, with Amazon EKS available as a technology preview.
The release also adds the Kubeflow Spark Operator as a developer preview within the enterprise data workbench, bringing distributed data processing into the workbench environment alongside model-serving capabilities.
Red Hat AI 3.5 further expands GPU-as-a-Service capabilities with namespace isolation, hosted control plane support and dashboards providing visibility into hardware inventory, utilisation and available GPU capacity.
With these additions, Red Hat is positioning AI 3.5 as an infrastructure layer for enterprises seeking to move from individual AI experiments to shared, governed AI services across hybrid and multi-cloud environments.