JobgetherRemote · India
JobgetherPosted 6d ago
AI Safety Policy Evaluator, Violence & Threats
AI Safety Policy Evaluator, Violence & Threats at Jobgether scores 50 out of 100 on AI centrality, which makes it an Level 2 role on this board.
AI in this role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Policy Evaluator, Violence & Threats based in United States.
This role sits at the intersection of AI safety, content policy, and expert human judgment, helping improve how advanced AI models handle violent and threatening content.
You’ll evaluate user requests, model responses, and conversation context to distinguish legitimate fictional, educational, historical, or defensive content from material that could enable real-world harm.
The work focuses heavily on nuanced edge cases where intent, context, and a single detail can materially change the appropriate policy decision.
You’ll contribute not only to evaluations, but also to adversarial testing, policy refinement, calibration, and the identification of emerging safety gaps.
This is a non-engineering role for someone with deep experience in areas such as violent fiction, military or emergency response, crisis intervention, threat assessment, trust and safety, or related fields.
You’ll work in a structured, feedback-rich environment where clear reasoning, consistency, and the ability to separate personal beliefs from policy standards are essential.
Because the role involves regular exposure to difficult material, candidates must be prepared to engage with sensitive content carefully, professionally, and sustainably.