AI Safety Expert for English and Indonesian Red Teaming
Saidgig
| Company | Saidgig |
| Category | Data & Analytics |
| Location | — |
| Remote | — |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| First seen | 2 Aug 2026 (the employer did not state a posting date) |
| Last verified | 9 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
Role Overview Probe conversational AI models with adversarial text-based inputs to surface vulnerabilities and produce reproducible red-team data that helps make AI systems safer. Work focuses on reviewing outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. Participation in higher-sensitivity reviews is optional and supported by clear guidelines and wellness resources, and topics will be communicated before you are exposed to any content.
Key Responsibilities
• Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
• Review AI outputs that involve sensitive topics, identify harmful or unsafe behavior, and flag areas of concern
• Generate high-quality human data by annotating failures, classifying vulnerabilities, and calling out systemic risks
• Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and structured
• Document findings reproducibly, producing reports, datasets, and attack cases customers can act on
Qualifications
• Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing
• Native fluency in English and Indonesian is required
• Curious and adversarial mindset, with an instinct to push systems to breaking points
• Structured approach, using frameworks or benchmarks rather than random experimentation
• Strong communicator, able to explain risks to technical and non-technical stakeholders
• Adaptable, able to move across projects and customer contexts
Nice-to-Have Specialties
• Adversarial machine learning, such as jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction
• Cybersecurity skills, including penetration testing, exploit development, or reverse engineering
• Socio-technical risk experience, including harassment, disinformation probing, abuse analysis, or conversational AI testing
• Creative probing skills, such as psychology, acting, or creative writing for unconventional adversarial thinking
What Success Looks Like
• Uncovering vulnerabilities that automated tests miss
• Delivering reproducible artifacts that customers use to strengthen their AI systems
• Expanding evaluation coverage so more scenarios are tested and fewer surprises occur in production
• Increasing customer trust in the safety of their AI through thorough adversarial probing
Work Terms
• Remote work
• All tasks are text-based
• Participation in higher-sensitivity projects is optional, with clear guidance and wellness resources provided
• Topics to be reviewed will be disclosed before exposure
Compensation 17 - 25 hourly
Eligibility
• Native fluency in both English and Indonesian is required
• No other location, visa, or sponsorship details were provided; candidates must ensure they meet any applicable local work requirements