Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

AI Safety Expert for English and Indonesian Red Teaming

Saidgig
CompanySaidgig
CategoryData & Analytics
Location
Remote
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
First seen2 Aug 2026 (the employer did not state a posting date)
Last verified9 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Role Overview Probe conversational AI models with adversarial text-based inputs to surface vulnerabilities and produce reproducible red-team data that helps make AI systems safer. Work focuses on reviewing outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. Participation in higher-sensitivity reviews is optional and supported by clear guidelines and wellness resources, and topics will be communicated before you are exposed to any content. Key Responsibilities • Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation • Review AI outputs that involve sensitive topics, identify harmful or unsafe behavior, and flag areas of concern • Generate high-quality human data by annotating failures, classifying vulnerabilities, and calling out systemic risks • Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and structured • Document findings reproducibly, producing reports, datasets, and attack cases customers can act on Qualifications • Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing • Native fluency in English and Indonesian is required • Curious and adversarial mindset, with an instinct to push systems to breaking points • Structured approach, using frameworks or benchmarks rather than random experimentation • Strong communicator, able to explain risks to technical and non-technical stakeholders • Adaptable, able to move across projects and customer contexts Nice-to-Have Specialties • Adversarial machine learning, such as jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction • Cybersecurity skills, including penetration testing, exploit development, or reverse engineering • Socio-technical risk experience, including harassment, disinformation probing, abuse analysis, or conversational AI testing • Creative probing skills, such as psychology, acting, or creative writing for unconventional adversarial thinking What Success Looks Like • Uncovering vulnerabilities that automated tests miss • Delivering reproducible artifacts that customers use to strengthen their AI systems • Expanding evaluation coverage so more scenarios are tested and fewer surprises occur in production • Increasing customer trust in the safety of their AI through thorough adversarial probing Work Terms • Remote work • All tasks are text-based • Participation in higher-sensitivity projects is optional, with clear guidance and wellness resources provided • Topics to be reviewed will be disclosed before exposure Compensation 17 - 25 hourly Eligibility • Native fluency in both English and Indonesian is required • No other location, visa, or sponsorship details were provided; candidates must ensure they meet any applicable local work requirements