Stop applying to jobs that are already dead.
Every listing verified, aged honestly, expired when filled.

All listings

Weekday AI via Workable

Technology AI Evaluation Expert

Level not stated $60–75/hr United States
still open verified 21h ago posted 68d ago checked just now
Apply at apply.workable.com

This is the employer's own posting, not a copy on a job board.

What we know

Is it still open?

Confirmed still open

Last checked 21h ago — checked against the employer's own applicant tracking system, which is the company answering directly.

We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.

Check this listing's status as JSON

How old is it?

Posted 68d ago

The date the source published, not the day we noticed it (2026-07-09). Last seen at its source just now.

Is it remote?

Marked remote on the employer's board

Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.

Who may apply?

United States

The description states no restriction of its own. This is the source's own tag.

Pay

$60–75/hr

Read out of the job description by us, not from a structured field. Shown in the posting's own currency and period; we never convert.

Skills named in the ad

Data AnalysisTechnical Writing

Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.

Carried by 1 source

The listing

This role is for one of our clients

Compensation: $60-$75 per hour

Join an advanced AI research initiative focused on improving how next-generation AI systems understand professional documents, execute complex instructions, and reason through real-world technical workflows. We are seeking experienced technology professionals to design high-quality benchmark tasks that evaluate AI performance across software engineering and data science domains.

In this role, you will create realistic, multi-step evaluation tasks based on technical documentation, code repositories, API references, architecture diagrams, and other workplace resources. Your work will help measure and improve the ability of AI models to interpret technical information, follow detailed instructions, and generate accurate, well-structured outputs.

This is a fully remote, independent contractor opportunity with flexible working hours.

Requirements

Key Responsibilities

Design AI Evaluation Tasks

  • Create realistic, multi-step benchmark tasks based on professional technology workflows.
  • Develop challenges using technical specifications, architecture documents, API documentation, codebases, web research, and code execution.
  • Ensure each task includes a clearly defined expected output and objective evaluation criteria.

Develop Evaluation Standards

  • Write comprehensive ground-truth solutions and structured scoring rubrics.
  • Design tasks that assess reasoning, technical understanding, instruction following, and output quality.
  • Maintain high standards of technical accuracy, clarity, and reproducibility.

Contribute Domain Expertise

  • Apply real-world knowledge from software engineering, data science, or analytics to create authentic evaluation scenarios.
  • Collaborate with research teams to improve benchmark quality and consistency.
  • Continuously refine tasks based on project feedback and evolving evaluation requirements.

Required Qualifications

  • Minimum 3 years of hands-on professional experience in one or more of the following areas:
    • Software Engineering
    • Data Science
    • Data Analytics
  • Strong understanding of technical documentation, software development workflows, and engineering best practices.
  • Experience working with codebases, APIs, technical specifications, or system architecture documentation.
  • Excellent analytical thinking and problem-solving skills.
  • Strong written communication with the ability to create clear technical instructions and evaluation criteria.
  • Ability to work independently while maintaining high standards of accuracy and consistency.

Engagement Details

  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of 15–20 hours per week.
  • Projects may be extended, shortened, or concluded based on business needs and performance.
  • Weekly payments processed through supported payment platforms.

Why Join

  • Help shape the next generation of AI systems for technical reasoning and document understanding.
  • Work on intellectually challenging projects involving real-world engineering and data science workflows.
  • Apply your technical expertise to improve advanced AI evaluation benchmarks.
  • Enjoy flexible remote work with meaningful impact on AI research.

Equal Opportunity Statement

We are committed to providing equal opportunities to all qualified applicants without regard to legally protected characteristics. Reasonable accommodations are available upon request.

Contract Information

  • Independent contractor engagement.
  • Fully remote work completed on your own schedule.
  • Weekly payments are processed based on approved work completed.
  • Work does not involve access to confidential or proprietary information from any employer, client, or institution.
  • Please note that visa sponsorship is not available for this opportunity.
Apply at apply.workable.com