Mirrai Careers
Resume BuilderCareer Test
InsightsPricing
Get Started Free
Jobs/AI Evaluation Engineer

AI Evaluation Engineer

govworx

United States Remote Full-time$120k–$170k / year Posted 3d ago
Market rate. This role pays around the $180k median for similar USD roles (238 comparable postings in our corpus).
Apply on company site
AI EVALUATION ENGINEER Location: Remote (Hybrid opportunity if live in Denver, CO) Type: Full-Time Clearance: Must have US citizenship and pass FBI fingerprint and background check in multiple states ABOUT GOVWORX GovWorx is helping public safety rise to today's greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-1-1 and emergency communications centers across the country by automating quality assurance, training, and real-time call evaluation—allowing agencies to strengthen their teams and better serve their communities. or the one you already have. POSITION OVERVIEW We're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science. You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments. KEY RESPONSIBILITIES * Design, build, and maintain automated AI evaluation pipelines for production LLM applications * Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods * Build offline evaluation datasets and regression testing frameworks to measure AI performance over time * Analyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvement * Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes * Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics * Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems * Investigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflows * Help establish best practices for Responsible AI, evaluation methodologies, and continuous model improvement QUALIFICATIONS MUST-HAVES * Must have US citizenship and pass FBI fingerprint and background check in multiple states * 3+ years of experience in software engineering, machine learning, data science, or a related technical field * Experience designing evaluation metrics and interpreting AI model performance * Understanding of statistical methods including hypothesis testing and experiment design * Strong Python development experience * Strong SQL skills with experience analyzing large datasets * Experience building or supporting production LLM or Generative AI applications * Experience with prompt engineering and systematic prompt evaluation NICE TO HAVE * Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio * Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker * Experience building dashboards using Tableau, Sisense, Power BI, or similar tools * Knowledge of Responsible AI principles and evaluation methodologies WHY JOIN GOVWORX? * Help build AI systems that directly support first responders and emergency communications professionals * Own AI quality, evaluation, and continuous improvement for production applications * Work on cutting-edge LLM technologies and help shape the future of Responsible AI * Collaborate with a high-performing team across AI, engineering, product, and data science * Solve technically challenging problems with real-world impact on public safety * Influence AI strategy and evaluation practices across a growing technology company

See how well you match this job

Upload your resume and we’ll score your fit for this role and 6 similar roles — then tailor your CV to it with AI. Free, no credit card.

Check your match

Similar jobs

  • AI Engineer, Evaluation

    distyl

    Remote
  • System Operations Engineer

    govworx

    Remote$160k–$185k
  • AI Evaluation Engineer

    capitalrx

    Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United States
  • AI Engineer

    crogl

    Remote
  • Machine Learning Engineer

    10alabs

    Washington D.C.
  • Research Engineer

    firecrawl

    Remote$210k–$275k
Apply on company site

Want more roles like this? Browse fresh jobs or tailor your resume with AI.

Mirrai Careers

AI-powered career platform: build resumes, match jobs, and plan your career.

Product

  • All Tools
  • Resume Builder
  • Career Test
  • Pricing

Legal

  • Privacy Policy
  • Terms of Service
  • Fair Use Policy

Company

MIRRAI CHAT LTD (Company No. 16403306)

71-75 Shelton Street, Covent Garden

London, WC2H 9JQ, UNITED KINGDOM

[email protected]

© 2026 Mirrai Careers. All rights reserved.