Demo

Senior Machine Learning Engineer, AI Evaluation

Utah SHRM
Alexandria, VA Full Time
POSTED ON 8/21/2026
AVAILABLE BEFORE 9/18/2026
Company: SHRM

Industry: Human Resource Management

Level: Salary

Job Family: Data & Analytics

Location: Alexandria, VA

Compensation: $100,000 to $130,000 per year

About Us

SHRM is a member-driven catalyst for creating better workplaces where people and businesses thrive together.  As the trusted authority on all things work, SHRM is the foremost expert, researcher, advocate, and thought leader on issues and innovations impacting today’s evolving workplaces.  With nearly 340,000 members in 180 countries, SHRM touches the lives of more than 362 million workers and their families globally.

Are you interested in growing your career with us? If so, we encourage you to explore our available career opportunities and join our team at SHRM today!

Don’t see your dream job? Apply here to join our talent community!

To view our Statement of Accessibility, click here.

For More Information:

website: www.shrm.org

Facebook Created with Sketch. Twitter Created with Sketch. LinkedIn Created with Sketch. Instagram Youtube

Legal Disclaimer: SHRM is an equal opportunity employer (Minority/Female/Disabled/Veteran). view full text At SHRM, inclusion, equity, and diversity are fundamental to fulfilling our vision of building a better workplace and better world. From our hiring practices through the entire employee experience, we embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to the workplace. We encourage diverse points of view which allows us to develop innovative solutions to the ever-evolving world of work. SHRM strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront of helping us promote and sustain an inclusive workplace that works for all. view full text

  • photo_library 6 Photos
  • video_library 3 Videos

Share Created with bin/sketchtool.

Close Created with Sketch.

Facebook Created with Sketch.

Twitter Created with Sketch.

LinkedIn Created with Sketch.

Email Created with Sketch.

Overview

48_Advertising Created with Sketch.

Summary

The Senior Machine Learning Engineer, AI Evaluation builds and operates the measurement and engineering infrastructure supporting the organization's Applied AI Research (AAIR) function, a continuous experimental environment designed to evaluate how artificial intelligence models perform real-world HR and workplace-related tasks against established professional standards.

48_Advise Created with Sketch.

Job Description

This role is responsible for designing and maintaining the engineering infrastructure used to conduct rigorous, reproducible AI model evaluations and benchmarks. The Senior Machine Learning Engineer develops the systems that run multiple AI models against structured, domain-specific evaluations; builds scoring and evaluation frameworks; maintains reproducibility across model versions; and creates the data infrastructure necessary to analyze and track model performance over time.

48_Environment Created with Sketch.

Work Environment

Hybrid Schedule (3 Days In-Office/2 Days Remote). This position follows a hybrid work schedule, with Tuesday through Thursday in office and Monday and Friday remote. Employees must be available during standard business hours, with core hours beginning between 8:00–9:00 a.m. and concluding between 5:00–6:00 p.m. local time. Travel: Occasional 0 – 10%.

Responsibilities

48_Analyze Created with Sketch.

Evaluate

Identify and surface ambiguity, inconsistencies, or measurement limitations within proposed evaluation criteria and collaborate with subject matter experts to strengthen evaluation design. Evaluate emerging AI models, tools, technologies, and evaluation methodologies and recommend appropriate applications within the research environment.

48_Quality Assurance Created with Sketch.

Quality Assurance

Ensure evaluation methodologies align with research-defined validation standards and produce findings that are reproducible, transparent, and defensible. Ensure appropriate quality controls are incorporated throughout data collection, evaluation, scoring, storage, and reporting processes. Ensure AI evaluation systems and workflows comply with organizational requirements for data security, privacy, ownership, access, and responsible AI use.

48_Compliance Created with Sketch.

Maintenance

Maintain portable evaluation architecture across AI model providers to enable consistent and defensible cross-model comparisons as models and technologies evolve. Maintain complete technical documentation and metadata necessary to reproduce research findings and evaluation results. Maintain appropriate controls to protect proprietary, member, research, and other sensitive data from unauthorized access or use.

48_Agriculture Created with Sketch.

Develop

Develop and maintain a unified, provider-agnostic orchestration layer that enables consistent evaluation across multiple frontier model providers and architectures. Develop monitoring, reporting, and visualization capabilities using Looker, Looker Studio, or comparable tools to provide visibility into experiment status, model performance, and performance drift.

48_Artificial Intelligence Created with Sketch.

Artificial Intelligence

Apply knowledge of AI evaluation methodologies, benchmarking techniques, inter-rater reliability, and known limitations of automated and model-as-judge evaluation approaches.

48_Goals Created with Sketch.

Goals

Establish and maintain technical standards and engineering practices that support reliable, repeatable, and auditable AI evaluation. Establish processes for tracking changes in model behavior across model versions and over time.

48_Solutions Created with Sketch.

Build

Build systems and processes that support reproducible experimentation, including model-version pinning, comprehensive run logging, experiment tracking, and drift detection. Build and maintain data structures in BigQuery or comparable platforms that enable research results to be queried, analyzed, reproduced, and audited.

48_Teamwork Created with Sketch.

Teamwork

Collaborate with research leaders, HR subject matter experts, data professionals, engineers, and other internal stakeholders to translate research requirements into scalable technical solutions.

48_Design Created with Sketch.

Design

Design, build, and maintain scalable engineering infrastructure for conducting structured evaluations and experiments across multiple AI and large language model (LLM) families. Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring, model-as-judge methodologies with appropriate safeguards, partial-credit methodologies, and approaches for managing ambiguity.

48_Support Created with Sketch.

Support

Support the design of measurement methodologies when definitive ground truth is unavailable or requires expert interpretation. Support collaboration with university, affiliate, research, and other external partners when appropriate and within established organizational access controls and data-handling requirements.

48_Background Created with Sketch.

Technical

Contribute technical expertise to the design and continuous improvement of AI research experiments, benchmarks, and evaluation methodologies. Implement and maintain appropriate access controls and data-handling requirements when working with external research or partner organizations.

Requirements

48_Communication Created with Sketch.

Communication

Demonstrated commitment to reproducibility, including disciplined use of versioning, documentation, logging, experiment tracking, and drift detection. Ability to effectively communicate complex technical concepts, methodologies, limitations, and findings to technical and non-technical audiences.

48_ Decision Making Created with Sketch. ?

Knowledge

A Ph.D. is not required; demonstrated expertise in AI/ML evaluation engineering, research infrastructure, and production-grade systems is valued. Strong knowledge of machine learning, large language models, generative AI systems, and contemporary AI application architectures.

48_Experience Created with Sketch.

Experience

Seven (7) or more years of progressively responsible experience in ML/LLM engineering, applied AI, applied data science, or research infrastructure, including experience developing, implementing, and supporting production-grade systems. Demonstrated hands-on experience developing multi-model LLM applications and infrastructure, including provider-agnostic model access, APIs, prompt engineering, and evaluation frameworks.

48_Knowledge Created with Sketch.

Skills

Ability to translate complex, judgment-based requirements from subject matter experts into technically rigorous and measurable evaluation specifications without oversimplifying the underlying domain expertise. Strong analytical and problem-solving skills with the ability to identify technical, methodological, and data-quality issues and develop appropriate solutions.

48_Degree Created with Sketch.

Education

Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or a related quantitative or technical field, or relevant equivalent experience in lieu of degree. Master's degree in Computer Science, Data Science, Machine Learning, Artificial Intelligence, or a related field preferred.

48_Proficiency Created with Sketch.

Proficiency

Advanced proficiency in Python and strong software-engineering fundamentals, including the ability to develop reliable, maintainable, production-quality code. Ability to balance technical rigor, research requirements, scalability, and practical implementation considerations. Ability to effectively leverage AI tools and technologies to streamline workflows, enhance productivity, and improve overall work quality.

48_Understanding Created with Sketch. ?

Physical Requirements

Prolonged periods of sitting at a desk and working on a computer. Frequent use of hands and fingers for typing, handling documents, and using office equipment. Occasional standing, walking, bending, and reaching. Ability to lift and carry up to 30 pounds as needed. Clear verbal and written communication skills for effective interaction with colleagues and stakeholders.

For More Information:

website: www.shrm.org

Facebook Created with Sketch. Twitter Created with Sketch. LinkedIn Created with Sketch. Instagram Youtube

Legal Disclaimer: SHRM is an equal opportunity employer (Minority/Female/Disabled/Veteran). view full text At SHRM, inclusion, equity, and diversity are fundamental to fulfilling our vision of building a better workplace and better world. From our hiring practices through the entire employee experience, we embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to the workplace. We encourage diverse points of view which allows us to develop innovative solutions to the ever-evolving world of work. SHRM strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront of helping us promote and sustain an inclusive workplace that works for all. view full text

  • photo_library 6 Photos
  • video_library 3 Videos

Salary : $100,000 - $130,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Senior Machine Learning Engineer, AI Evaluation?

Sign up to receive alerts about other jobs on the Senior Machine Learning Engineer, AI Evaluation career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$149,493 - $192,976
Income Estimation: 
$149,493 - $192,976
Income Estimation: 
$184,796 - $233,226
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Utah SHRM

  • Utah SHRM Alexandria, VA
  • Company: SHRM Industry: Human Resource Management Level: Salary Job Family: Technology - General Location: Alexandria, VA Compensation: $100,000 to $120,00... more
  • 5 Days Ago

  • Utah SHRM Alexandria, VA
  • Company Industry: Level Job Family: Location: Compensation About Us Share Created with bin/sketchtool. Close Created with Sketch. Facebook Created with Ske... more
  • 11 Days Ago


Not the job you're looking for? Here are some other Senior Machine Learning Engineer, AI Evaluation jobs in the Alexandria, VA area that may be a better fit.

  • Scale AI Washington, DC
  • The goal of a Senior Machine Learning Engineer at Scale is to leverage techniques in the fields of generative AI, computer vision, reinforcement learning, ... more
  • 16 Days Ago

  • Evlo AI Washington, DC
  • About The Role The role owns the architecture, development, and scaling of machine learning systems, driving the transition of advanced AI models from rese... more
  • 15 Days Ago

AI Assistant is available now!

Feel free to start your new journey!