Binance
via Lever
Senior Evaluation Algorithm Engineer
This is the employer's own posting, not a copy on a job board.
What we know
Is it still open?
Confirmed still open
Last checked just now — checked against the employer's own applicant tracking system, which is the company answering directly.
We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.
How old is it?
Posted 1h ago
The date the source published, not the day we noticed it (2026-09-16). Last seen at its source just now.
Is it remote?
Marked remote on the employer's board
Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.
Who may apply?
Hong Kong
The description states no restriction of its own. This is the source's own tag.
Pay not stated
Similar roles pay $103.8k–173.1k/yr
Middle 50% of 23 listings that do state pay — Engineering · all levels · APAC · USD/year. This employer has published no salary; this is what comparable listings we hold disclose, never converted between currencies or periods. How this is calculated.
Skills named in the ad
Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.
Carried by 1 source
-
lever employer's own board first seen 1h ago · last seen just now
The listing
About the Role
Responsibilities
- Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.
- Lead the design and construction of evaluation datasets. Define evaluation dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.
- Analyze model capability boundaries and failure modes based on evaluation results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.
- Drive the automation and scaling of evaluation workflows. Build sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.
- Collaborate with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.
Requirements
- Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.
- Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.
- Familiar with mainstream evaluation methods (human evaluation, model-based automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.
- Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.
- Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.
- Strong business understanding and communication skills, able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration.
Bonus Qualifications
- Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs.
- Experience building high-quality AI training/evaluation data or data annotation systems.
- Familiarity with RLHF, reward models, or preference data-related work.