Demo

Head of Evaluations (Legal AI Benchmarking)

Newcode.ai
York, NY Full Time
POSTED ON 7/16/2026
AVAILABLE BEFORE 9/15/2026
Who are we?

At Newcode.ai, we're transforming how law firms and legal professionals harness AI for real-world impact. As part of our collaborative, high-growth team, you'll have the rare opportunity to work side-by-side with visionary founders at the bleeding edge of AI and legal innovation — shaping not just our product, but the future of legal work itself.

Note: We believe in being transparent about what it's like to work at Newcode. As a fast-growing startup, we're building and evolving every day. That means not every process, playbook, or framework is already in place, and priorities can shift quickly.

The people who thrive here are comfortable with ambiguity, take ownership, and don't wait for perfect direction. They are resourceful, proactive, and able to "figure it out"—solving problems, creating structure where needed, and helping build the company as they go. If you are good with this then, great! Keep reading to learn more.

Position Overview 

We are seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform. You will ensure our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate. 

Key Responsibilities 

  • Design Legal Benchmarks for: Contract Drafting, Information Extraction, Legal Research, and Contract Review 
  • Build, source and maintain relevant datasets 
  • Audit AI Output: Review and score complex AI-generated legal text, contract analyses, and statutory interpretations for accuracy and precision and lay out a strategy.  
  • Define Evaluation Metrics: Establish clear criteria for grading model performance, specifically focusing on logical reasoning, citation accuracy, and the model's ability to safely abstain from answering. 
  • Collaborate with Engineering: Partner directly with Engineering to translate legal errors into actionable technical feedback for model fine-tuning. 
  • PhD or Masters in statistics, mathematics, machine learning or equivalent 
  • Analytical Skills: Proven ability to break down complex statutory frameworks and case law into structured, logical data points. 
  • Tech-Savviness: python, panda, numpy, jupiter notebooks and similar statistical models 

Salary.com Estimation for Head of Evaluations (Legal AI Benchmarking) in York, NY
$396,680 to $552,782
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Head of Evaluations (Legal AI Benchmarking)?

Sign up to receive alerts about other jobs on the Head of Evaluations (Legal AI Benchmarking) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$162,804 - $221,695
Income Estimation: 
$183,562 - $247,678
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Not the job you're looking for? Here are some other Head of Evaluations (Legal AI Benchmarking) jobs in the York, NY area that may be a better fit.

  • Aaru Inc. York, NY
  • Head of Evaluations ResearchTechnical StaffResearchNYC$450K – $750K • Offers Equity## **About Aaru**Aaru builds simulations of human behavior. Each simulat... more
  • 12 Days Ago

  • Notion York, NY
  • Who We Are Notion is the collaborative AI workspace where teams and agents think together. We're building one place where your knowledge, projects, meeting... more
  • 11 Days Ago

AI Assistant is available now!

Feel free to start your new journey!