Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Sr. Evaluation Engineer

LogicMonitor
CompanyLogicMonitor
CategoryUncategorised
LocationBangalore
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted27 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Us:    We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers. This position is located in Bangalore.  You'll be working in a major tech center of Pune, India. Across the globe, our Centers of Energy serve as hubs where we accelerate productivity and collaboration, inspire creativity, and cultivate a culture of connection and celebration. Our teams coordinate their time in Centers of Energy to reflect how they work best. To learn more about life at LogicMonitor, check out our Careers Page . What You'll Do: LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise. Our customers love LogicMonitor's ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn's Best Places to Work for the seventh year in a row!  Edwin AI is LogicMonitor’s AI-powered observability and incident intelligence platform. It helps enterprise operations teams investigate incidents, identify root causes, recommend remediation, and automate operational workflows. As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents, retrieval systems, tool integrations, and complex investigation workflows Here's a closer look at this key role: Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness. Build offline and online evaluation pipelines in Python and integrate them with CI/CD, experimentation, model selection, prompt iteration, and release gating. Lead the creation and maintenance of golden datasets and regression suites using alerts, events, metrics, logs, traces, topology, configuration data, incident timelines, change records, ITSM workflows, and historical investigation outcomes. Build representative, customer-specific scenarios covering different technologies, failure modes, operational patterns, and environmental constraints. Use human-authored and AI-assisted methods to generate regression, edge, adversarial, rare, ambiguous, and incomplete-context test cases. Treat evaluation datasets and test suites as first-class components that evolve alongside Edwin AI. Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows. Assess both final outcomes and intermediate behavior, including planning, reasoning consistency, retrieval, evidence use, tool selection, tool parameters, state transitions, escalation decisions, and human-in-the-loop approvals. Identify whether failures originate from models, prompts, retrieval, data quality, tools, agent logic, orchestration, or infrastructure. Evaluate capabilities including incident investigation, on-
HOUSE AD991,236 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →