KrissDevHub
Technologies
0%
All positions

AI Evaluator

AI Workforce Remote (Worldwide) Contract

We are hiring AI Evaluators to benchmark and stress-test LLM deployments. You will evaluate outputs against custom datasets, design validation criteria, and track model degradation.

What you'll do

  • Execute automated and manual evaluation benchmarks across distinct model versions
  • Write qualitative assessments of model performance in conversational tasks
  • Design adversarial testing datasets (red teaming) to detect safety escapes
  • Compile performance reports highlighting error distributions and regressions

What we're looking for

  • Background in technical auditing, data analysis, or software quality assurance
  • Familiarity with standard evaluation benchmarks (MMLU, GSM8K, HumanEval)
  • Detail-oriented mindset with high standards for truthfulness and formatting
  • Basic coding knowledge (Python/JSON) is a plus for automated benchmarks

Benefits & Perks

  • Competitive hourly rates on global projects
  • 100% remote workspace with flexible schedule
  • Involvement in cutting-edge safety and trust research
  • Direct collaboration with advanced technical architects

Ready to apply?

Takes about 5 minutes. We review every application personally.

Apply for this position

Similar roles