AI Research
This is the employer's own posting, not a copy on a job board.
What we know
Is it still open?
Confirmed still open
Last checked 1d ago — checked against the employer's own applicant tracking system, which is the company answering directly.
We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.
How old is it?
Posted 43d ago
The date the source published, not the day we noticed it (2026-08-28). Last seen at its source 1h ago.
We have tracked this listing since 9 Oct 2026 (1 days). The employer's own board has carried it every time we have read it, most recently 1 hour ago.
Is it remote?
Remote
That is the location the employer filed this posting under. Quoted as written — we do not re-word the source's own location.
Who may apply?
Not stated
The description states no restriction of its own. This is the source's own tag.
Skills named in the ad
Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.
Carried by 1 source
-
Greenhouse employer's own board first seen 1d ago · last seen 1h ago
The listing
About Vetto
Vetto builds the infrastructure for next-generation AI training data. We partner with the world’s top AI labs to push the frontier of what models can do — by designing, collecting, and delivering the highest-quality training data for the hardest problems in AI.
AI Researchers at Vetto figure out how to teach models new capabilities and how to tell whether it worked. You’ll design datasets and evaluations, run experiments, study model failures, and turn what you learn into better data and training strategies.
Most of our work is evaluation-driven, with some fine-tuning and post-training used to validate research hypotheses. You won’t be expected to build the infrastructure alone: you’ll work with an engineering team that builds the platforms and tools behind the research. You should still be comfortable writing code, analyzing data, and moving your own experiments forward.
Why Vetto?
- End-to-end ownership. You’ll take research questions from an initial idea to a dataset, benchmark, experiment, and useful conclusion.
- Direct collaboration with top AI labs. Your work will help frontier teams understand their models and decide what data or training approach to try next.
- Work that gets used. Your research can become training data, evaluations, technical reports, open benchmarks, and published papers.
- A wide range of hard problems. Depending on the project, you might work on agents, model evaluations, human or synthetic data, failure analysis, or post-training.
- Fast, flat team. You’ll ship experiments in days, not quarters.
What You’ll Do
- Design and run experiments to understand model behavior, test new ideas, and determine whether a dataset or training intervention actually works.
- Use fine-tuning and post-training experiments when useful to validate research hypotheses and measure improvements in the capabilities we care about.
- Create datasets, benchmarks, agent environments, rubrics, and evaluation methods for problems that are not well measured today.
- Work directly with AI labs and enterprise partners to turn open-ended model-development goals into concrete research and data projects.
- Dig into model outputs and trajectories to understand systematic failure modes and separate model limitations from problems in the data, grader, or evaluation setup.
- Write analysis code and build lightweight tools or prototypes for your own work, with support from engineers when a workflow needs to become reliable infrastructure.
- Share what you learn through clear technical reports, internal discussions, benchmarks, and research papers.
What We’re Looking For
Must-haves:
- Experience doing empirical research or similarly open-ended technical work with machine learning systems.
- Hands-on experience designing and running agentic systems for complex work, such as coding agents and automated research pipelines.
- Good experimental judgment: you can form a hypothesis, choose useful baselines and metrics, run a careful experiment, and make sense of noisy results.
- Strong Python and data analysis skills, plus enough software engineering ability to prototype your own research workflows.
- The independence to take an ambiguous question about a model and turn it into a dataset, benchmark, experiment, or other concrete contribution.
- Rigorous analytical thinking — you look at the underlying evidence, notice subtle failure modes, and question results that seem too simple or too good to be true.
- Excellent written and verbal communication with researchers, engineers, customers, and non-specialists.
- Comfort with ambiguity, fast iteration, changing priorities, and occasional forward-deployed work with partner teams.
Nice to have:
- Research or industry experience in LLM evaluation, post-training, agents, human or synthetic data, alignment, or AI safety.
- Experience designing datasets or benchmarks, validating annotation quality, or working with large collections of model outputs and trajectories.
- Published research, technical writing, open-source contributions, or strong independent projects.
- Experience building or managing human feedback pipelines.
- A background in cognitive science, linguistics, philosophy, statistics, or another field that helps you think clearly about evaluation.
- Experience at an AI lab, AI data company, or in a partner-facing research role.
We care more about evidence of excellent work than any single credential. A graduate degree, publications, lab experience, open-source work, independent research, and production engineering experience are all useful signals, but none is a requirement on its own.
What We Offer
- Competitive compensation and a stock option plan.
- Global, remote-first work with flexibility and periodic in-person on-sites every few months.
- Direct collaboration with cutting-edge AI labs on evaluation, data, and post-training challenges.
- High ownership and fast career growth in a founding-era team.
Location: Global / Remote, with periodic in-person on-sites every few months
Reports to: Co-founders