Stop applying to jobs that are already dead.
Every listing verified, aged honestly, expired when filled.

All listings

Weekday AI via Workable

Applied Computer Science Benchmark Specialist

Level not stated $66–84/hr United States
still open verified 21h ago posted 29d ago checked just now
Apply at apply.workable.com

This is the employer's own posting, not a copy on a job board.

What we know

Is it still open?

Confirmed still open

Last checked 21h ago — checked against the employer's own applicant tracking system, which is the company answering directly.

We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.

Check this listing's status as JSON

How old is it?

Posted 29d ago

The date the source published, not the day we noticed it (2026-08-17). Last seen at its source just now.

Is it remote?

Marked remote on the employer's board

Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.

Who may apply?

United States

The description states no restriction of its own. This is the source's own tag.

Pay

$66–84/hr

Read out of the job description by us, not from a structured field. Shown in the posting's own currency and period; we never convert.

Skills named in the ad

API DesignDevOpsMachine LearningSRE

Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.

Carried by 1 source

The listing

This role is for one of our clients

Compensation: $66 - $84 per hour

We are seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. In this role, you will develop and validate rigorous multiple-choice questions across a broad range of computer science domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute through one of two primary task types:

Question Authoring — Develop original, challenging multiple-choice questions within your area of computer science expertise, assess their difficulty, and submit them for review.

Question Verification — Review existing questions for technical accuracy, clarity, completeness, and rigor. Make necessary edits, assess difficulty, and document the rationale behind your changes.

Requirements

Computer Science Domains

  • Accelerator / GPU Kernel Engineering
  • Formal Methods & Automated Reasoning
  • Computer Architecture & Accelerators
  • Distributed Systems
  • DevOps & Site Reliability Engineering
  • Data Engineering & Databases
  • Cloud Computing & Infrastructure
  • Operating Systems & Systems Kernel
  • Machine Learning Engineering
  • Web & API Development
  • Embedded Systems Engineering
  • Computer Graphics & Game Development
  • Mobile Engineering

Key Responsibilities

  • Create original computer science questions that evaluate deep conceptual understanding, technical reasoning, and problem-solving rather than surface-level recall.
  • Ensure every question is unambiguous, self-contained, technically accurate, and sufficiently specified for a qualified expert to solve.
  • Classify questions by difficulty:
    • Medium: Introductory undergraduate level
    • Hard: Advanced undergraduate level
    • Expert: Postgraduate level and above
  • Provide one correct answer alongside nine plausible but subtly incorrect alternatives designed to distinguish strong technical reasoning from superficial knowledge.
  • Develop clear, structured solution explanations that demonstrate the reasoning and technical principles required to reach the correct answer.
  • Provide 1–5 authoritative references per question, drawing from peer-reviewed research, academic publications, university resources, and other reputable technical sources.
  • For verification assignments, identify issues related to correctness, clarity, completeness, precision, or solvability and clearly explain the reasoning behind any recommended edits.
  • Apply consistent standards when evaluating questions and solutions to ensure benchmark quality and reproducibility.

Ideal Qualifications

  • PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or a closely related discipline.
  • A Master's degree may be considered for candidates with exceptional expertise in a specialized computer science domain.
  • Strong command of graduate-level computer science theory, algorithms, systems, software engineering, architecture, and/or machine learning.
  • Demonstrated depth in one or more of the listed technical domains.
  • Research publications, substantial industry experience at leading technology organizations, systems engineering experience, or competitive programming experience is a strong plus.
  • Excellent written English and the ability to communicate complex technical concepts clearly, accurately, and concisely.
  • Strong attention to detail and the ability to distinguish technically valid solutions from plausible but incorrect approaches.

More About the Opportunity

  • Expected commitment: 10+ hours per week
  • Fully remote and asynchronous
  • Flexible scheduling based on project requirements
  • Opportunity to contribute to the development of high-quality benchmarks for evaluating advanced AI systems
  • Strong contributors may be considered for additional review, evaluation, or subject-matter expert opportunities

Application Process

Submit your resume or a summary of your relevant academic and professional background.

Selected candidates may be asked to complete a short technical assessment or provide additional information about their area of expertise.

Equal Opportunity

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations throughout the application and engagement process.

Contract and Payment Terms

  • Engagement will be on an independent contractor basis.
  • This is a fully remote opportunity that can be completed on your own schedule.
  • Projects may be extended, shortened, or concluded early depending on project requirements and performance.
  • Work will not require access to confidential or proprietary information belonging to any current or former employer, client, or institution.
  • Payments are made weekly through Stripe or Wise, based on services rendered.
  • H-1B and STEM OPT candidates are not eligible for this opportunity at this time.
Apply at apply.workable.com