New Relic has introduced AI Evaluation, a capability within its AI Observability platform designed to help enterprises monitor AI response quality, identify security risks, and assess the business impact of AI workloads across the application lifecycle.
The capability extends beyond evaluating individual large language model (LLM) calls by linking AI performance and response quality to the underlying application transaction. It aims to give developers, site reliability engineers (SREs), and platform teams a unified view of AI behaviour, infrastructure performance, and resource consumption, from development through production.
As enterprises deploy generative AI applications in business-critical environments, traditional application performance metrics alone may not reveal issues such as hallucinations, prompt injection attacks, data leaks, or declining response quality following model updates.
New Relic AI Evaluation addresses these challenges by integrating AI quality assessments with application performance monitoring and distributed tracing. The platform uses an asynchronous LLM-as-a-judge service to evaluate sampled inputs and outputs, assigning quality scores to distributed traces to help engineering teams identify the source of failures.
By connecting AI responses with prompts, tool calls, retrieval systems, and backend infrastructure, teams can investigate whether an issue originates in the model, a retrieval-augmented generation (RAG) system, or another application component.
Brian Emerson, Chief Product Officer at New Relic, said evaluating response quality and model efficiency has become as important as monitoring application availability. He added that transaction-level visibility helps engineering teams understand AI performance, technical health, and business impact within their existing observability workflows.
AI quality, security and cost monitoring
New Relic AI Evaluation includes configurable guardrails to help detect prompt injection attempts, jailbreaks, personally identifiable information (PII) exposure, toxicity, and bias. Pre-built evaluators are intended to simplify setup and reduce the need for manual reviews.
For RAG-based applications, the capability assesses measures such as faithfulness and answer relevance, helping teams distinguish between problems with model-generated responses and those associated with information retrieval.
The platform also links qualitative response assessments with compute consumption and token costs. This allows teams to compare models on response quality and cost efficiency when selecting options for production workloads.
Pre-production testing and prompt management
The new capability also includes tools for testing and refining AI applications before deployment. A Prompt Playground enables engineers to compare prompts across models in a controlled environment, while reusable datasets support regression testing as prompts, models and configurations change.
Organisations can build versioned datasets from existing distributed traces or synthetic data and conduct controlled A/B tests to assess changes before releasing them into production.
Stephen Elliot, Group Vice President for I&O, Cloud Operations and DevOps at IDC, said enterprises moving generative AI from pilots to production need visibility into both application health and response quality, including accuracy, safety and model efficiency. He noted that combining live response evaluation, prompt lifecycle management and operational telemetry can help engineering and security teams manage risks and costs.
With AI Evaluation, New Relic aims to bring AI quality assessment and application observability into a common workflow, helping enterprises troubleshoot failures, strengthen safeguards, and evaluate the cost-effectiveness of production AI applications.