CitiusCloud services rethinks infrastructure operations with Agentic AI

As enterprises scale across increasingly complex IT environments, keeping infrastructure running smoothly is becoming a business priority, not just an IT challenge. Teams are expected to identify and resolve issues quickly, minimise downtime and maintain consistent service levels, even as applications and workloads span multiple environments. Traditional automation can execute predefined tasks, but responding to unexpected issues often still requires significant human intervention and expertise.

This is where agentic AI is beginning to offer a new approach. Rather than simply executing predefined instructions, AI agents can analyse an issue, draw on operational knowledge, recommend the next course of action and, with appropriate approvals, execute the required remediation. For enterprises, this can help reduce troubleshooting effort while enabling engineering teams to focus on higher-value priorities. According to Gartner, organisations should start testing AI agents for mature IT-operations use cases, and predicts AI agents will be implemented in 60% of IT operations tools by 2028.

CitiusCloud Services, a technology transformation and managed services company, saw an opportunity to make this process more efficient. With experience supporting enterprises across BFSI, Pharma, ITeS and Manufacturing, the company wanted to reduce the manual effort involved in troubleshooting while retaining the checks and approvals needed for production environments.

Moving from task automation to operational reasoning

CitiusCloud worked with IBM to introduce an agentic automation capability into its InfraBOT platform using IBM watsonx Orchestrate, supported on-premise or as a subscription. The collaboration brings together CitiusCloud’s experience in infrastructure and managed services with IBM’s capabilities in agentic AI and orchestration. The solution is live and generally available on Red Hat OpenShift Container Platform and extensible to any Kubernetes platform.

The workflow brings health-check, compliance, troubleshooting, remediation and knowledge capture into a single operational loop. The system continuously scans the OpenShift environment for issues and creates a ticket when one is detected. AI agents then examine the issue against knowledge accumulated from previous incidents and propose a remediation for an engineer to review.

Once the proposed action is approved, the agents execute the remediation. The system then generates a root cause analysis (RCA) and adds it to the knowledge base, allowing future incidents to benefit from previous troubleshooting experience. This reduces the need to repeatedly build and maintain one-off troubleshooting scripts for recurring problems, while keeping remediation traceable through the ticket, diagnosis, approval and RCA.

Roshan Shetty, Co-Founder, CitiusCloud Services, said, “Our focus has always been on helping enterprises simplify infrastructure operations and enable their teams to respond to issues more efficiently. By bringing troubleshooting knowledge, automation and human oversight into a single workflow, we can help our clients reduce repetitive effort, improve consistency in incident resolution and allow their engineers to focus on more strategic priorities. Imagine a world with humans and AI agents working alongside each other to better managed services.”

Building knowledge into the operational workflow

A key part of the approach is capturing what the organisation learns from every incident. Once an issue is resolved, its RCA becomes part of the knowledge base. Over time, this creates a reusable source of operational knowledge that agents can draw on when diagnosing similar issues. This can help engineers move faster from detection to a potential fix without relying entirely on individual expertise or recreating troubleshooting steps for every incident.

Aditya Sharma, Head of AI Initiatives, CitiusCloud Services, said, “We see the value of agents in infrastructure operations as a combination of reasoning and action. Our collaboration with IBM has helped us validate and bring this approach into a production-ready environment, combining CitiusCloud’s infrastructure and AI expertise with IBM watsonx Orchestrate’s capabilities in orchestration and governance, with both teams working together on agent development.”.

As agents take on more complex IT tasks, governance becomes critical to ensuring greater autonomy does not come at the expense of control. In CitiusCloud’s implementation, proposed remediation is reviewed and approved by an engineer, bringing human in the loop, before a production change is executed. The workflow also provides traceability through diagnosis, approval and RCA, helping enterprises maintain oversight of automated actions.

A practical path toward autonomous IT operations

Siddhesh Naik, Executive Director, Growth Markets & Partner Ecosystem, Technology Sales, IBM India & South Asia said, “The CitiusCloud implementation demonstrates how agentic AI can bring reasoning, context and action together in infrastructure operations. By combining AI-driven recommendations and automated remediation with human approval, enterprises can resolve issues faster while maintaining the governance and oversight required in production environments.”

CitiusCloud’s implementation points to a practical path for enterprises looking to evolve their infrastructure operations—from predefined automation toward workflows that combine automation, contextual reasoning, accumulated knowledge and human oversight. By enabling agents to investigate issues, recommend remediation and execute approved actions, organisations can reduce operational effort while retaining control over higher-risk decisions.

Comments (0)
Add Comment