What are the responsibilities and job description for the Databricks Data Engineer position at Holistic Partners, Inc?
Role: Databricks Data Engineer/Azure/AI
Location: Manassas, VA or Chesterfield, MO (Hybrid 4 DAYS A WEEK)
Role Type: Full time
Interview Required: Onsite
Requirement Notes:
This role will design, build, and maintain scalable data pipelines in a modern cloud environment to support analytics, AI, and ML initiatives. The engineer will work with Databricks, Delta Lake, PySpark, and AI/BI tools to turn raw data into usable insights, and will help convert legacy SSIS ETL processes into cloud native Databricks pipelines. The position also supports MLOps work, partnering with data scientists and product managers to train, evaluate, and deploy models using MLflow and Databricks Workflows, with a focus on governance through Unity Catalog and CI/CD automation
Job Description:
This role will design, build, and maintain scalable data pipelines in a modern cloud environment to support analytics, AI, and ML initiatives. The engineer will work with Databricks, Delta Lake, PySpark, and AI/BI tools to turn raw data into usable insights, and will help convert legacy SSIS ETL processes into cloud native Databricks pipelines. The position also supports MLOps work, partnering with data scientists and product managers to train, evaluate, and deploy models using MLflow and Databricks Workflows, with a focus on governance through Unity Catalog and CI/CD automation.
Key Responsibilities
Build and maintain ETL and ELT pipelines using Azure Databricks, Delta Lake, and PySpark. Modernize SSIS logic into PySpark notebooks and Delta Live Tables. Support batch and streaming pipelines with data quality checks. Improve performance using Photon, partitioning, and caching.
Develop Bronze, Silver, and Gold data layers within Delta Lake. Manage schema evolution and versioning. Apply Unity Catalog governance for access control and lineage tracking.
Partner with data science teams on feature pipelines, model training, and production scoring. Deploy models using MLflow and Databricks Model Registry. Use built in AI functions such as ai_query and ai_forecast to generate insights. Monitor pipeline health, data drift, and model performance. Build CI/CD workflows using GitHub Actions or Azure DevOps.
Maintain data security through Databricks secrets and key vault tools. Document architecture, workflows, and operational procedures. Serve as a subject matter expert on Databricks best practices.
Qualifications
At least 3 years of experience in Databricks, PySpark, Python, and data engineering. Databricks certification is a plus.
Strong SQL skills including stored procedures. Experience with SSIS, SSRS, or Power BI is a plus.
Experience with cloud platforms such as Azure, Databricks, or Spark, and with batch and streaming pipelines.
Familiarity with Databricks AI functions like AI Query, AI Classify, and AI Forecast.
Experience with Python ML frameworks such as PyTorch or TensorFlow.
Bachelor's degree in Computer Science, Engineering, Math, or a related field required. Master's degree preferred. Experience in commercial insurance is a plus.