All positions
AI Evaluator
AI Workforce Remote (Worldwide) Contract
We are hiring AI Evaluators to benchmark and stress-test LLM deployments. You will evaluate outputs against custom datasets, design validation criteria, and track model degradation.
What you'll do
- Execute automated and manual evaluation benchmarks across distinct model versions
- Write qualitative assessments of model performance in conversational tasks
- Design adversarial testing datasets (red teaming) to detect safety escapes
- Compile performance reports highlighting error distributions and regressions
What we're looking for
- Background in technical auditing, data analysis, or software quality assurance
- Familiarity with standard evaluation benchmarks (MMLU, GSM8K, HumanEval)
- Detail-oriented mindset with high standards for truthfulness and formatting
- Basic coding knowledge (Python/JSON) is a plus for automated benchmarks
Benefits & Perks
- Competitive hourly rates on global projects
- 100% remote workspace with flexible schedule
- Involvement in cutting-edge safety and trust research
- Direct collaboration with advanced technical architects
Ready to apply?
Takes about 5 minutes. We review every application personally.
Apply for this position