Stop applying to jobs that are already dead.
Every listing verified, aged honestly, expired when filled.

All listings

Anyone AI via Ashby

Human Data Evals Lead (Remote/US/LATAM)

lead Worldwide ⚠ ad says LATAM / United States
still open verified 22h ago posted 87d ago checked just now
Apply at jobs.ashbyhq.com

This is the employer's own posting, not a copy on a job board.

What we know

Is it still open?

Confirmed still open

Last checked 22h ago — checked against the employer's own applicant tracking system, which is the company answering directly.

We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.

Check this listing's status as JSON

How old is it?

Posted 87d ago

The date the source published, not the day we noticed it (2026-06-19). Last seen at its source just now.

Is it remote?

Argentina - Fully Remote, Ecuador - Fully Remote, Mexico - Fully Remote, Peru - Fully Remote

That is the location the employer filed this posting under. Quoted as written — we do not re-word the source's own location.

Who may apply?

Available worldwide

The listing is tagged as above, but its own description restricts the role to LATAM or United States. We show both rather than picking one. Read the ad before applying.

What the ad says
Location: Remote / Latam / US The role You will own Anyone AI’s data i…

Pay not stated

Similar roles pay $108.3k–165.6k/yr

Middle 50% of 10 listings that do state pay — Operations · all levels · Worldwide · USD/year. This employer has published no salary; this is what comparable listings we hold disclose, never converted between currencies or periods. How this is calculated.

Skills named in the ad

LLMProject Management

Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.

Carried by 1 source

The listing

Reports to: CEO

Owns: data proposals, sample development, quality, and pilot delivery

Location: Remote / Latam / US


The role

You will own Anyone AI’s data initiatives and proposals to AI labs, from the data proposal or responding to requests, through pilot delivery. You own how we build proposals and develop the sample packages and benchmarks: frontier-grade packages across reasoning, coding, agents, and tool use, multi-modal and others, produced in collaboration with subject-matter experts, with expert-verified ground truth, multi-model headroom results, and QC that survives buyer-side scrutiny. You are the person who designs the sample that demonstrates our quality, converts pilots into production engagements. On a small team, this is the operational center of the Human Data Division.

Responsibilities

  • Proposals & requests. Study public benchmarks and eval targets, and turn them into proposals and sample packages that demonstrate capability and win the work. Respond to lab data requests and pilots.

  • Sample & benchmark development. Design and build the sample packages, working with subject-matter experts. Every package meets the bar of our current sample set:

    • Expert-verified, exact-match-checkable ground truth and gold reasoning trajectories.

    • Multi-model evaluation showing real headroom, and proof the task discriminates the model, not just that it's hard.

    • Rigorous QC structure: calibration layers, severity-weighted rubrics, deterministic verifiers, evidence maps, etc.

  • Subject-matter experts. Recruit, brief, calibrate, and review a pool of experts across coding, agentic/tool-use, and STEM/reasoning. Raise their output to our standard and keep it there; be the arbiter of what "correct" and "frontier-difficulty" mean.

  • Lab relationships. Be a direct point of contact for lab partners on Slack and calls, with support from the CEO and the wider team. Keep senior lab contacts informed, surface what they actually need, and pull in the CEO and subject-matter experts when the conversation calls for it.

  • Pilot delivery. Own pilots end to end: scoping, SOW, staffing, production, QC, and delivery. Nothing ships before it's lab-ready, and nothing comes back rejected as "not frontier-level" without us already knowing why.

Experience

  • Originated data or benchmark proposals for AI labs, translated eval targets into sample tasks that demonstrate capability, and owned the engagement through delivery.

  • Deep evaluation and quality expertise: LLM benchmarking, with real strength in code-model evaluation.

  • Built QC processes and artifact standards that met enterprise or lab requirements, and set a quality bar a team of experts was held to.

  • Thrives in ambiguous, fast-moving environments where the rules are still being written, and delivers under pressure.

Qualifications

  • 5+ years in technical delivery, quality, or program management, with recent experience in AI/ML data, model evaluation, or benchmarking.

  • Hands-on experience delivering data or evaluation work to AI labs or enterprise ML teams, scoping through delivery.

  • Working fluency with how frontier models are evaluated: benchmarks, rubrics, pass rates, headroom, and what makes a task discriminate a model.

  • Proven people/vendor leadership, you've recruited, calibrated, and held a team or expert pool to a quality standard.

  • Fluent English. Spanish is a nice to have.

Apply at jobs.ashbyhq.com