Demo

Sr. Data Engineer

KPIT
Columbus, IN Full Time
POSTED ON 7/20/2026
AVAILABLE BEFORE 8/9/2026
Job Description

Responsibilities:

  • Designs and automates deployment of our distributed system for ingesting and transforming data from various types of sources (relational, event-based, unstructured).
  • Own end-to-end delivery of AI and ML solutions from problem definition to production deployment
  • Build and maintain data pipelines using PySpark and Spark in Azure Databricks
  • Designs and implements framework to continuously monitor and troubleshoot data quality and data integrity issues.
  • Implements data governance processes and methods for managing metadata, access, retention to data for internal and external users.
  • Designs and provide guidance on building reliable, efficient, scalable and quality data pipelines with monitoring and alert mechanisms that combine a variety of sources using ETL/ELT tools or scripting languages.
  • Perform feature engineering and data preparation for machine learning models
  • Deploy ML models and AI agents for real business use cases
  • Design and implement agentic AI workflows including multi-step reasoning and tool usage
  • Track experiments, manage models, and support deployment using MLflow
  • Define and execute model evaluation frameworks including both ML and AI agent performance
  • Designs and implements physical data models to define the database structure. Optimizing database performance through efficient indexing and table relationships.
  • Participates in optimizing, testing, and troubleshooting of data pipelines.
  • Designs, develops and operates large scale data storage and processing solutions using different distributed and cloud-based platforms for storing data (e.g. Azure , AWS, Spark , Pyspark, Scala , advanced SQL , Hadoop others).
  • Uses innovative and modern tools, techniques and architectures to partially or completely automate the most-common, repeatable and tedious data preparation and integration tasks to minimize manual and error-prone processes and improve productivity. Assists with renovating the data management infrastructure to drive automation in data integration and management.
  • Ensures the timeliness and success of critical analytics initiatives by using agile development technologies such as DevOps, Scrum, Kanban
  • Coaches and develops less experienced team members.
  • Define and execute model evaluation frameworks including both ML and AI agent performance
  • Work independently on ambiguous business problems and convert them into scalable solutions
  • Collaborate with other data engineers, analysts, solution architects and business teams to deliver solutions
  • Guide and support team members on data engineering, ML, and AI best practices
  • Write clean, production-ready, and well-documented code

Requirements:

  • System Requirements Engineering - Uses appropriate methods and tools to translate stakeholder needs into verifiable requirements to which designs are developed; establishes acceptance criteria for the system of interest through analysis, allocation and negotiation; tracks the status of requirements throughout the system lifecycle; assesses the impact of changes to system requirements on project scope, schedule, and resources; creates and maintains information linkages to related artifacts.
  • Collaborates - Building partnerships and working collaboratively with others to meet shared objectives.
  • Communicates effectively - Developing and delivering multi-mode communications that convey a clear understanding of the unique needs of different audiences.
  • Customer focus - Building strong customer relationships and delivering customer-centric solutions.
  • Decision quality - Making good and timely decisions that keep the organization moving forward.
  • Data Extraction - Performs data extract-transform-load (ETL) activities from variety of sources and transforms them for consumption by various downstream applications and users using appropriate tools and technologies.
  • Programming - Creates, writes and tests computer code, test scripts, and build scripts using algorithmic analysis and design, industry standards and tools, version control, and build and test automation to meet business, technical, security, governance and compliance requirements.
  • Quality Assurance Metrics - Applies the science of measurement to assess whether a solution meets its intended outcomes using the IT Operating Model (ITOM), including the SDLC standards, tools, metrics and key performance indicators, to deliver a quality product.
  • Solution Documentation - Documents information and solution based on knowledge gained as part of product development activities; communicates to stakeholders with the goal of enabling improved productivity and effective knowledge transfer to others who were not originally part of the initial learning.
  • Solution Validation Testing - Validates a configuration item change or solution using the Function's defined best practices, including the Systems Development Life Cycle (SDLC) standards, tools and metrics, to ensure that it works as designed and meets customer requirements.
  • Data Quality - Identifies, understands and corrects flaws in data that supports effective information governance across operational business processes and decision making.
  • Problem Solving - Solves problems and may mentor others on effective problem solving by using a systematic analysis process by leveraging industry standard methodologies to create problem traceability and protect the customer; determines the assignable cause; implements robust, data-based solutions; identifies the systemic root causes and ensures actions to prevent problem reoccurrence are implemented.
  • Values differences - Recognizing the value that different perspectives and cultures bring to an organization.
  • College, university, or equivalent degree in relevant technical discipline, or relevant equivalent experience required.
  • At least 5 years of experience in data engineering with a strong background on Azure Databricks and Scala/Python.
  • Experience in handling unstructured data processing and transformation with programming knowledge.
  • Hands on experience in building data pipelines using Scala/Python
  • Big data technologies such as Apache Spark, Structured Streaming, Advanced SQL, Databricks, Delta Lake, Azure/AWS
  • Strong analytical and problem-solving skills with the ability to troubleshoot spark applications and resolve data pipeline issues.
  • Familiarity with version control systems like Git, CICD pipelines.
  • Experience with Azure Databricks and MLflow
  • Good understanding of ML workflows, model development, and evaluation
  • Knowledge of MLOps fundamentals such as CI/CD, versioning, and monitoring
  • Ability to build end-to-end data and ML solutions
  • Exposure to production ML or AI systems
  • Understanding of data engineering and data modeling basics
  • Ability to work independently on loosely defined problems
  • Strong problem-solving and communication skills
  • Mentoring experience is a plus

Preferred

  • Experience with AI agents, LLMs, or agentic AI systems

Compensation and Benefits:

Along with competitive pay, as a full-time KPIT employee, you are eligible for the following benefits:

  • Geo Blue PPO and HSA plan.
  • MetLife – Dental and Vision plan.
  • Healthcare and Dependent care flexible spending account(FSA).
  • 401k with employer match.
  • Company-paid Basic Life and Long-term disability insurance.
  • Voluntary benefits include Critical Illness, Hospital indemnity, accident insurance, theft, and legal service.
  • Employee Assistance Program.
  • Paid Holidays.
  • Employee discounts and perks.
  • Gym benefit.

Required SkillsSystem Requirements Engineering, Requirements Management, Requirements Traceability,Requirements Analysis, Acceptance Criteria Definition, Change Impact Analysis,Stakeholder Management, Collaboration, Communication Skills, Customer Focus,ETL Development, Data Extraction Transformation and Loading (ETL), Azure Databricks, Scala, PythonSupported SkillsAI agents, LLMs,Agentic AI systems

Salary.com Estimation for Sr. Data Engineer in Columbus, IN
$120,723 to $151,683
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at KPIT

  • KPIT Peoria, IL
  • About KPIT KPIT is reimagining the future of mobility, forging ahead with group companies and partners to shape a world that is cleaner, smarter, and safer... more
  • 15 Days Ago

  • KPIT Auburn, MI
  • About KPIT KPIT is reimagining the future of mobility, forging ahead with group companies and partners to shape a world that is cleaner, smarter, and safer... more
  • 1 Day Ago

  • KPIT Novi, MI
  • Job Description Job Summary Industrial Engineer is responsible for improving manufacturing efficiency, productivity, and capacity through time studies, lin... more
  • 2 Days Ago

  • KPIT Novi, MI
  • Job Description Responsibilities: Diagnostics Data Engineering & Architecture Diagnostics Strategy & Transformation Define and execute the Data science, ma... more
  • 7 Days Ago


Not the job you're looking for? Here are some other Sr. Data Engineer jobs in the Columbus, IN area that may be a better fit.

  • Dorleco Columbus, IN
  • Location : Columbus, IN. About the Role We are hiring a Data Engineer with strong hands-on experience in building high‑performance data pipelines for a hea... more
  • 15 Days Ago

  • mgpru Edinburgh, IN
  • At M&G, our purpose is to give everyone real confidence to put their money to work. With a heritage dating back more than 175 years, we have a long history... more
  • Just Posted

AI Assistant is available now!

Feel free to start your new journey!