What are the responsibilities and job description for the Data Scientist - Conversational AI & GenAI position at Wise Skulls?
Title: Data Scientist - Conversational AI & GenAI
Location: Boston, MA (Hybrid)
Duration: 6 months (possibility of an extension)
Implementation Partner: Infosys
End Client: To be disclosed
JD:
Role Summary
We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.
Key Responsibilities
Required Qualifications
Preferred Qualifications
Success Criteria
Important Note
This role is primarily a Data Science and AI Evaluation position, not an AI Engineering or deployment-focused role. The emphasis is on measuring, analyzing, validating, and improving AI system performance rather than building production deployment pipelines.
Location: Boston, MA (Hybrid)
Duration: 6 months (possibility of an extension)
Implementation Partner: Infosys
End Client: To be disclosed
JD:
Role Summary
We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.
Key Responsibilities
- Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.
- Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.
- Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.
- Perform prompt tuning and experimentation to improve model accuracy and reliability.
- Conduct root-cause analysis of conversational failures and recommend remediation strategies.
- Measure performance across different model configurations, prompts, and guardrail implementations.
- Partner with AI Engineering and Product teams to validate solutions before production deployment.
- Build dashboards and reports that communicate model effectiveness and operational impact.
- Support fraud-related customer service use cases, including intent detection and multi-turn conversation flows.
Required Qualifications
- Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field.
- Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions.
- Deep understanding of model evaluation techniques and metrics, including:
- Precision@K
- Recall@K
- Mean Reciprocal Rank (MRR)
- F1 Score
- Retrieval and generation quality assessment
- Experience performing experimentation, statistical analysis, and performance benchmarking.
- Strong Python programming skills.
- Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy, and related tools.
- Ability to communicate technical findings succinctly to highly technical stakeholders.
Preferred Qualifications
- Experience with:
- Generative AI and LLM ecosystems
- Multi-agent systems
- RAG/Agentic RAG architectures
- Amazon Bedrock
- AWS SageMaker
- Databricks
- MLflow
- LangSmith
- Weights & Biases
- Knowledge of conversational AI, IVR systems, digital assistants, and voice agents.
- Experience in financial services, fraud detection, or customer service automation.
Success Criteria
- Develop reliable evaluation methodologies for conversational AI systems.
- Quantify the effectiveness of fraud self-service voice agents.
- Optimize prompts, retrieval strategies, and guardrails using empirical evidence.
- Deliver actionable insights that improve customer experience and model performance.
- Establish measurable KPIs for intent detection and multi-turn conversation success.
Important Note
This role is primarily a Data Science and AI Evaluation position, not an AI Engineering or deployment-focused role. The emphasis is on measuring, analyzing, validating, and improving AI system performance rather than building production deployment pipelines.