Bilingual AI Safety Evaluator
Unknown · Remote (Remote) · posted Sep 10, 2026
What this role actually asks for
Extracted by RemoteHuntMust have
- •Native/near-native fluency in a non-English language
- •Business-level written English
- •Strong written communication and reasoning skills
- •Excellent attention to detail and consistency
- •Sound judgment on sensitive information
Nice to have
- •Experience reviewing written content
- •Background in trust and safety
- •Experience with AI-generated content
Languages required
The full posting
We’ll ask you to help us improve safety for advanced AI models by evaluating sensitive content with care and consistency. Write expert-level prompts in your native or primary language across a range of sensitive subject areas. Apply structured guidelines to classify prompts and conversations. Identify adversarial phrasing, escalation patterns, and potential safety concerns. Document the reasoning behind your judgments clearly and accurately. Evaluate how AI models respond to sensitive and dual-use topics and help identify areas for improvement. You’ll work as an independent contractor and follow the workflow and instructions we provide during onboarding.
To succeed in this role, we’re looking for bilingual evaluators who can combine strong language ability with careful judgment and consistent written reasoning. Languages: Spanish, Russian, Thai, Vietnamese, Indonesian, German, Chinese, Arabic Native or near-native fluency in a non-English language, plus business-level written English. A bachelor’s degree (completed or in progress). Strong written communication and reasoning skills. Excellent attention to detail and consistency. Sound judgment when evaluating sensitive and dual-use information. Ability to follow detailed guidelines and apply them consistently. We also value applicants who take accuracy seriously, can distinguish nuance in wording, and can remain objective when assessing potentially risky content.
Based in a region where your target language is widely spoken (preference, not a requirement). Experience reviewing, grading, or red-teaming written or technical content. Background in trust and safety, content moderation, policy evaluation, quality assurance, or adversarial testing. Experience working with AI-generated content or evaluating model responses.
Is this one actually worth your time?
RemoteHunt scores every remote job 0–100 against your own resume, so you apply to the handful that fit instead of the hundred that don't. Free plan, no card required.