What are the responsibilities and job description for the Machine Learning Engineer – Databricks position at Aptino?
Role: Machine Learning Engineer – Databricks
Location: Malvern, PA (Hybrid – 3 days/week onsite)
Duration: 12 Months
Position Overview:
We are seeking an experienced Machine Learning Engineer with strong expertise in the Databricks Lakehouse Platform to develop, deploy, and optimize scalable AI/ML solutions. The ideal candidate will have hands-on experience building production-ready machine learning models, implementing MLOps best practices, and designing end-to-end ML pipelines using Databricks, Spark, and AWS cloud services.
Key Responsibilities:
- Design, develop, and deploy scalable machine learning models using the Databricks Machine Learning platform.
- Build and optimize end-to-end ML pipelines, including data ingestion, feature engineering, model training, validation, deployment, and monitoring.
- Utilize MLflow for experiment tracking, model versioning, lifecycle management, and production deployments.
- Develop high-performance data processing pipelines using PySpark, Apache Spark, and SQL to support large-scale analytics and machine learning workloads.
- Build and maintain production-grade applications and ML workflows on AWS, leveraging services such as Lambda, S3, Glue, ECS/EKS, Step Functions, SageMaker, and Bedrock.
- Implement feature engineering, model evaluation, hyperparameter optimization, and performance tuning to improve model accuracy and scalability.
- Collaborate with data engineers, data scientists, and business stakeholders to translate business requirements into production-ready ML solutions.
- Establish MLOps best practices, CI/CD processes, monitoring strategies, and governance standards for machine learning deployments.
- Optimize data architecture and machine learning workflows using Databricks Lakehouse, Delta Lake, and Unity Catalog.
- Contribute to AI innovation initiatives by evaluating emerging technologies, Generative AI use cases, and modern machine learning frameworks.
Required Qualifications:
- 5 years of experience in Machine Learning, Artificial Intelligence, or Data Science engineering.
- Strong hands-on expertise with Databricks Machine Learning and the Databricks ecosystem.
- Proven experience using MLflow for experiment management, model registry, and deployment.
- Strong programming skills in Python, PySpark, Apache Spark, and SQL.
- Experience developing supervised and unsupervised machine learning models for enterprise applications.
- Hands-on experience building, deploying, and supporting production-grade ML pipelines.
- Solid understanding of feature engineering, model validation, hyperparameter tuning, and model performance optimization.
- Experience developing cloud-native AI/ML solutions on AWS.
- Experience building scalable data pipelines using Spark, Glue, Airflow, dbt, or similar orchestration tools.
- Strong analytical, troubleshooting, and problem-solving skills with experience handling large datasets.
Preferred Qualifications:
- Experience designing solutions on the Databricks Lakehouse Architecture.
- Knowledge of Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), or AI-powered applications.
- Experience with orchestration frameworks such as Apache Airflow or Azure Data Factory.
- Familiarity with Docker, Kubernetes, CI/CD, and modern DevOps practices.
- Hands-on experience with Delta Lake, Unity Catalog, and enterprise data governance.
- Databricks certification is an added advantage.