AI Benchmark Quality Reviewer

Hace 7 días

Ecuador Braintrust Trabajo remoto Jornada completa $23 Indefinido
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment. What you’ll do
• Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
• Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
• Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
• Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.
• Write concise, evidence-backed findings and verify that revisions address the issues found. What we’re looking for
• At least five years of relevant technical or analytical experience.
• Ability to read Python, SQL, shell scripts, structured data and execution logs to understand task setup and grading behavior.
• Strong written English, analytical judgment and attention to detail.
• Ability to give specific, reproducible feedback and explain uncertainty clearly.
• Experience in AI evaluation, technical QA, data analysis or benchmark development is helpful, but not required. Familiarity with Harbor task setup is a plus. Engagement
• Remote contractor assignment for eight weeks.
• 40 hours per week, including eight hours of daily overlap with Pacific Time.
• Shortlisted applicants may be asked to complete an interest form before final review. Compensation
• $23 per hour.