All career paths

AI Evaluation and Expert Projects

Human-Data QA Lead & AI Evaluation Calibrator

Build the quality system that turns expert human judgment into consistent, measurable and client-trusted AI evaluation.

Remote, Time-Zone Requirements ApplyFull-time or project basedReferral reward $250
Apply for This Role and Join the Network

About this path

This role designs and operates rubrics, gold answers, reviewer training, calibration, sampling, disagreement analysis, error taxonomies and customer reporting while protecting both quality and throughput.

You will build and operate the quality system that turns a team of expert reviewers into a consistent, trusted evaluation pipeline, from rubric design to disagreement analysis and client reporting. The role balances protecting quality with keeping throughput realistic, and it rewards people who can turn messy disagreement into a clear error taxonomy.

What you would own

  • Design of evaluation rubrics and gold-standard answers
  • Reviewer training and onboarding for evaluation cohorts
  • Calibration exercises to keep reviewer judgment consistent
  • Sampling strategy and disagreement analysis across reviewers
  • Development of error taxonomies from real disagreement patterns
  • Client-facing quality reporting on evaluation outcomes

You are likely a strong match if

  • You have quality assurance, calibration or evaluation-program experience
  • You can design a rubric that reduces reviewer disagreement
  • You are comfortable analyzing disagreement data to find patterns
  • You can train and onboard reviewers effectively
  • You communicate quality findings clearly to clients or stakeholders
  • You balance quality rigor with realistic throughput expectations

Helpful, not required

  • Experience with human-data labeling or annotation programs
  • Statistical literacy for sampling and inter-rater reliability analysis
  • Experience managing distributed reviewer teams
  • Familiarity with AI model evaluation practices

What success looks like

  • Reviewer calibration stays consistent across large evaluation cohorts
  • Disagreement analysis leads to concrete rubric or training improvements
  • Clients trust the quality reports enough to act on them
  • Throughput stays realistic without sacrificing quality standards

Practical proof that helps

  • A description of a QA or calibration program you designed or ran
  • Examples of rubrics or error taxonomies you have built
  • Metrics showing improved inter-rater reliability from your work

What being in the network gives you

  • Remote-first work with clients across the United States, Canada and Latin America.
  • Human review of your profile, automation organizes information, people decide.
  • One profile considered across current and future opportunities.
  • Referral rewards when someone you refer directly is successfully placed.
  • Full control over availability, matching and your data at any time.

One profile, many opportunities

Applying here creates a single reusable profile. If this path is not the right fit, you remain eligible for other suitable opportunities across the network.

Apply and join the network

Free to join • Start in about 60 seconds • One profile for multiple opportunities • No advanced AI experience required for many roles

Applying to Human-Data QA Lead & AI Evaluation Calibrator. You stay eligible for other suitable opportunities.

Résumé or LinkedIn, either one is enough
Career paths, choose up to three

Applying creates one reusable Latino AI Talent Network profile. You can update your interests or pause matching at any time.

Joining the network does not guarantee immediate work or placement. Opportunities depend on professional fit, location, availability and client demand.