Demo

Data Scientist – Conversational AI & GenAI Evaluation

ClifyX
Johnston, RI Contractor
POSTED ON 9/19/2026
AVAILABLE BEFORE 10/18/2026

Job Title: Data Scientist – Conversational AI & GenAI Evaluation

Work Location : Johnston, RI or Westwood, MA (Hybrid Onsite)

Contract duration: 12 months Contract


Detailed job description - Skill Set:

We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.


Key Responsibilities

  • Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.
  • Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.
  • Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.
  • Perform prompt tuning and experimentation to improve model accuracy and reliability.
  • Conduct root-cause analysis of conversational failures and recommend remediation strategies.
  • Measure performance across different model configurations, prompts, and guardrail implementations.
  • Partner with AI Engineering and Product teams to validate solutions before production deployment.
  • Build dashboards and reports that communicate model effectiveness and operational impact.
  • Support fraud-related customer service use cases, including intent detection and multi-turn conversation flows.

Success Criteria

  • Develop reliable evaluation methodologies for conversational AI systems.
  • Quantify the effectiveness of fraud self-service voice agents.
  • Optimize prompts, retrieval strategies, and guardrails using empirical evidence.
  • Deliver actionable insights that improve customer experience and model performance.
  • Establish measurable KPIs for intent detection and multi-turn conversation success.


Must have:

  • Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field.
  • Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions.
  • Deep understanding of model evaluation techniques and metrics, including:
  • Precision@K
  • Recall@K
  • Mean Reciprocal Rank (MRR)
  • F1 Score
  • Retrieval and generation quality assessment
  • Experience performing experimentation, statistical analysis, and performance benchmarking.
  • Strong Python programming skills.
  • Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy, and related tools.
  • Ability to communicate technical findings succinctly to highly technical stakeholders.


Desired Skills:

  • Experience with:
  • Generative AI and LLM ecosystems
  • Multi-agent systems
  • RAG/Agentic RAG architectures
  • Amazon Bedrock
  • AWS SageMaker
  • Databricks
  • MLflow
  • LangSmith
  • Weights & Biases
  • Knowledge of conversational AI, IVR systems, digital assistants, and voice agents.
  • Experience in financial services, fraud detection, or customer service automation.


Important Note

This role is primarily a Data Science and AI Evaluation position, not an AI Engineering or deployment-focused role. The emphasis is on measuring, analyzing, validating, and improving AI system performance rather than building production deployment pipelines.


Hourly Wage Estimation for Data Scientist – Conversational AI & GenAI Evaluation in Johnston, RI
$45.00 to $55.00
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Data Scientist – Conversational AI & GenAI Evaluation?

Sign up to receive alerts about other jobs on the Data Scientist – Conversational AI & GenAI Evaluation career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$90,112 - $113,166
Income Estimation: 
$116,765 - $144,626
Income Estimation: 
$90,112 - $113,166
Income Estimation: 
$116,765 - $144,626
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at ClifyX

  • ClifyX Rapids, IA
  • Job Title: Senior Java Platform Engineer – Vulnerability Remediation Work Location: Cedar Rapids, Lowa, USA-Hybrid Contract duration: Long term contract Vi... more
  • 11 Days Ago

  • ClifyX Bentonville, AR
  • Job Title: Google Cloud Platform Data Solution Architect Work Location : Bentonville, AR (3 days onsite is a must) Contract duration: 12 months Contract Vi... more
  • 1 Day Ago

  • ClifyX Chandler, AZ
  • ClifyX group is an award winning IT Consultancy formed in 1998. Our Mission is to provide our clients with Optimal Technology solutions that are effective ... more
  • 1 Day Ago

  • ClifyX San Diego, CA
  • Job Title: IoT Integration Engineer – Smart Buildings Work Location: San Diego, CA (Hybrid) Contract duration: Long term contract Visa: /Only Job Details: ... more
  • 1 Day Ago


Not the job you're looking for? Here are some other Data Scientist – Conversational AI & GenAI Evaluation jobs in the Johnston, RI area that may be a better fit.

  • BURGEON IT SERVICES LLC Johnston, RI
  • Position : Data Scientist - Conversational AI & Gen AI Location : Johnston, RI or Westwood, MA Duration: Full Time Domain (Industry): Cards, Banking, FSI M... more
  • 2 Days Ago

  • BURGEON IT SERVICES LLC Johnston, RI
  • Role: Senior Data Scientist Generative AI / Conversational AI Evaluation Location: Johnston, RI or Westwood, MA Hybrid/Onsite preferred, though location fl... more
  • 2 Days Ago

AI Assistant is available now!

Feel free to start your new journey!