AI Engineer
This is the employer's own posting, not a copy on a job board.
What we know
Is it still open?
Confirmed still open
Last checked 11h ago — checked against the employer's own applicant tracking system, which is the company answering directly.
We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.
How old is it?
Posted 5d ago
The date the source published, not the day we noticed it (2026-09-09). Last seen at its source just now.
Is it remote?
Marked remote on the employer's board
Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.
Who may apply?
United States
The description states no restriction of its own. This is the source's own tag.
Pay not stated
Similar roles pay $160k–225k/yr
Middle 50% of 2161 listings that do state pay — Engineering · all levels · United States · USD/year. This employer has published no salary; this is what comparable listings we hold disclose, never converted between currencies or periods. How this is calculated.
Skills named in the ad
Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.
Carried by 1 source
-
ashby employer's own board first seen 5d ago · last seen just now
The listing
Location: Remote (Hybrid opportunity if live in Denver, CO)
Type: Full-Time
Clearance: Must pass FBI fingerprint and background check in multiple states (U.S. Citizenship is strongly preferred)
About GovWorx
GovWorx is helping public safety rise to today's greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-1-1 and emergency communications centers across the country by automating quality assurance, training, and real-time call evaluation—allowing agencies to strengthen their teams and better serve their communities. or the one you already have.
Position Overview
We're looking for an experienced AI Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science.
You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments.
Key Responsibilities
Design, build, and maintain automated AI evaluation pipelines for production LLM applications
Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods
Build offline evaluation datasets and regression testing frameworks to measure AI performance over time
Analyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvement
Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes
Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics
Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems
Investigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflows
Help establish best practices for Responsible AI, evaluation methodologies, and continuous model improvement
Qualifications
Must-Haves
Must pass FBI fingerprint and background check in multiple states (U.S. Citizenship is strongly preferred)
3+ years of experience in software engineering, machine learning, data science, or a related technical field
Experience designing evaluation metrics and interpreting AI model performance
Understanding of statistical methods including hypothesis testing and experiment design
Strong Python development experience
Strong SQL skills with experience analyzing large datasets
Experience building or supporting production LLM or Generative AI applications
Experience with prompt engineering and systematic prompt evaluation
Nice to Have
Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio
Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker
Experience building dashboards using Tableau, Sisense, Power BI, or similar tools
Knowledge of Responsible AI principles and evaluation methodologies
Why Join GovWorx?
Help build AI systems that directly support first responders and emergency communications professionals
Own AI quality, evaluation, and continuous improvement for production applications
Work on cutting-edge LLM technologies and help shape the future of Responsible AI
Collaborate with a high-performing team across AI, engineering, product, and data science
Solve technically challenging problems with real-world impact on public safety
Influence AI strategy and evaluation practices across a growing technology company