AI Agent Quality Engineer
Netskope
| Company | Netskope |
| Category | Engineering |
| Location | Taguig |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 22 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
About Netskope
Today, there's more data and users outside the enterprise than inside, causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed, one that is built in the cloud and follows and protects data wherever it goes, so we started Netskope to redefine Cloud, Network and Data Security.
Since 2012, we have built the market-leading cloud security company and an award-winning culture powered by hundreds of employees spread across offices in Santa Clara, St. Louis, Bangalore, London, Paris, Melbourne, Taipei, and Tokyo. Our core values are openness, honesty, and transparency, and we purposely developed our open desk layouts and large meeting spaces to support and promote partnerships, collaboration, and teamwork. From catered lunches and office celebrations to employee recognition events and social professional groups such as the Awesome Women of Netskope (AWON), we strive to keep work fun, supportive and interactive. Visit us at Netskope Careers. Please follow us on LinkedIn and Twitter @Netskope .
As a Senior AI Quality & Red Team Engineer at Netskope, you will lead the charge in testing, stress-testing, and breaking our AI agents before they ever reach production. From automating multi-turn prompt injections to tracking fleet-wide drift in CI/CD, you will own the automated harness that ensures our AI systems are secure, resilient, and compliant. If you love the idea of being the person who proves an agent isn't ready yet, welcome home.
Skills and competencies:
Build and grow the automated evaluation suite every agent runs against before it's approved for production, designed to run unattended and scale across a growing agent fleet, not something that needs a person babysitting each run.
Design adversarial test scenarios — prompt injection attempts, sycophancy checks where an agent has to correctly push back on a false premise, multi-attempt attacks rather than single-shot ones — and automate them so they run on every relevant change, not just before a big release.
Own the "break it on purpose" pass for every new agent: attempt to extract data it shouldn't expose, get it to act outside its registered tool boundaries, or get it to treat a synthetic test probe as real. As the fleet grows, build this into a repeatable, scriptable process rather than a manual exercise redone from scratch each time.
Partner with the Data Steward on data sensitivity classification for the systems agents touch, so your test scenarios reflect what's actually at stake, not a generic checklist.
Decide, for each agent capability, what "pass" actually means, and build that judgment into automated thresholds wherever possible so evaluation keeps up as the number of agents climbs into the hundreds.
Maintain the guardrail and negative-test catalog (fail-closed vs. fail-graceful behavior) across the platform, and add new cases as new failure modes get discovered in the wild.
Produce clear, audit-ready evidence for every agent's evaluation results, generated automatically as part of the pipeline rather than assembled by hand for each review.
Track drift over time across the whole fleet, not agent by agent, so a slow-moving problem in one corner doesn't go unnoticed just because no one's looking at that specific agent that week.
Must-Have:
At least 4 years in software quality, security testing, or a related discipline, with 1–2 years specifically evaluating or red-teaming LLM-based systems — not just running unit tests against traditional code.
Strong Python skills, since the evaluation harness, adversarial test scripts, and automated pipelines will mostly be built in it. Comfortable writing production-quality code, not just glue scripts.
Working knowledge of REST APIs and webhook/event-driven patterns, enough to build test harnesses that call an agent's tools directly and validate its inputs and out
991,236 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →