AI Evaluation and Expert Projects
Human-Data QA Lead & AI Evaluation Calibrator
Build the quality system that turns expert human judgment into consistent, measurable and client-trusted AI evaluation.
About this path
This role designs and operates rubrics, gold answers, reviewer training, calibration, sampling, disagreement analysis, error taxonomies and customer reporting while protecting both quality and throughput.
You will build and operate the quality system that turns a team of expert reviewers into a consistent, trusted evaluation pipeline, from rubric design to disagreement analysis and client reporting. The role balances protecting quality with keeping throughput realistic, and it rewards people who can turn messy disagreement into a clear error taxonomy.
What you would own
- Design of evaluation rubrics and gold-standard answers
- Reviewer training and onboarding for evaluation cohorts
- Calibration exercises to keep reviewer judgment consistent
- Sampling strategy and disagreement analysis across reviewers
- Development of error taxonomies from real disagreement patterns
- Client-facing quality reporting on evaluation outcomes
You are likely a strong match if
- You have quality assurance, calibration or evaluation-program experience
- You can design a rubric that reduces reviewer disagreement
- You are comfortable analyzing disagreement data to find patterns
- You can train and onboard reviewers effectively
- You communicate quality findings clearly to clients or stakeholders
- You balance quality rigor with realistic throughput expectations
Helpful, not required
- Experience with human-data labeling or annotation programs
- Statistical literacy for sampling and inter-rater reliability analysis
- Experience managing distributed reviewer teams
- Familiarity with AI model evaluation practices
What success looks like
- Reviewer calibration stays consistent across large evaluation cohorts
- Disagreement analysis leads to concrete rubric or training improvements
- Clients trust the quality reports enough to act on them
- Throughput stays realistic without sacrificing quality standards
Practical proof that helps
- A description of a QA or calibration program you designed or ran
- Examples of rubrics or error taxonomies you have built
- Metrics showing improved inter-rater reliability from your work
What being in the network gives you
- Remote-first work with clients across the United States, Canada and Latin America.
- Human review of your profile, automation organizes information, people decide.
- One profile considered across current and future opportunities.
- Referral rewards when someone you refer directly is successfully placed.
- Full control over availability, matching and your data at any time.
One profile, many opportunities
Applying here creates a single reusable profile. If this path is not the right fit, you remain eligible for other suitable opportunities across the network.
Apply and join the network
Free to join • Start in about 60 seconds • One profile for multiple opportunities • No advanced AI experience required for many roles
Joining the network does not guarantee immediate work or placement. Opportunities depend on professional fit, location, availability and client demand.