What are the responsibilities and job description for the Data Scientist position at Programmers.io?
We are seeking a Data Scientist to help build and maintain a robust evaluation framework for conversational AI systems. This role will focus on developing automated AI quality measurement solutions, including LLM-as-a-Judge systems, to assess hallucination rates, intent accuracy, transcript quality, and responsible AI metrics at scale.
As a key member of our AI Quality team, you will work closely with annotation specialists, product teams, and governance stakeholders to ensure our AI experiences are accurate, safe, compliant, and customer-centric.
- Design and develop LLM-as-a-Judge evaluation frameworks for conversational AI systems.
- Build automated monitoring solutions to evaluate millions of voice and chat-based customer interactions.
- Collaborate with user experience teams to identify core metrics and evaluation criteria.
- Validate model performance against human-labeled datasets using statistical best practices.
- Measure and report precision, recall, false-positive rates, and false-negative rates.
- Develop statistical sampling methodologies and quality measurement frameworks.
- Continuously calibrate and improve evaluation models as production AI systems evolve.
- Analyze hallucination rates, intent classification accuracy, transcript quality (WER), guardrail effectiveness, vulnerability to jailbreak attempts, and fairness metrics.
- Create dashboards and reporting to support governance, legal, compliance, and Responsible AI reviews.
- Collaborate with annotation teams to improve label quality and gold-standard datasets.
- Support future multilingual AI evaluation initiatives.
- Master's degree or PhD in Data Science, Statistics, Computer Science, Machine Learning, or related field, or equivalent experience.
- 3 years of experience in data science, machine learning, analytics, or AI model evaluation.
- Experience analyzing voice-based language data.
- Strong experience with Python and data science tooling.
- Experience designing experiments and evaluating machine learning models.
- Knowledge of statistical analysis, sampling methodologies, and model validation techniques.
- Experience communicating technical findings to both technical and non-technical stakeholders.
- Experience with Generative AI, LLMs, or conversational AI systems.
- Familiarity with AI safety, Responsible AI, fairness, bias, and governance frameworks.
- Experience developing automated evaluation systems.
- Experience with cloud-based AI platforms and experimentation environments.