What are the responsibilities and job description for the Azure DataBricks Lead position at Covetus, LLC?
Title: Azure Databricks Lead - Fulltime Role
Only local candidates to NYC, NY or Charlotte, NC
Only one In-Person Interview
Interview Process: Only one round of In-Person interview at any location Charlotte, NC or New York, NY as suitable for applicant.
Job Description:
Data Engineers are responsible for building reliable and scalable data infrastructure that enables organizations to derive meaningful insights, make data-driven decisions, and unlock the value of their data assets.
Job Description - Grade Specific:
- We are looking for an experienced Azure Databricks Engineer with strong expertise in cloud-based data engineering, ETL development, and distributed data processing.
- The ideal candidate should have solid hands-on experience with PySpark, Delta Lake, Azure Data Factory (ADF) and building scalable data pipelines on Azure.
- The engineer will work closely with business stakeholders, Data Architects, and cross-functional teams to design, develop, and optimize data pipelines for enterprise-grade analytics and reporting.
Key Responsibilities:
- Design, develop, and optimize ETL/ELT pipelines using Azure Databricks and PySpark.
- Build scalable data ingestion workflows from various structured and unstructured data sources.
- Implement transformation logic, data cleansing, enrichment, and data validation frameworks.
- Work with Delta Lake to build and maintain a Medallion Architecture (Bronze, Silver and Gold layers).
- Develop reusable Databricks notebooks and jobs for production-grade data workflows.
- Build and orchestrate data pipelines using Azure Data Factory (ADF).
- Integrate Databricks with other Azure services, including ADLS, Azure SQL, Event Hubs, Key Vault and Synapse Analytics.
- Optimize compute environments, including clusters, pools and auto-scaling configurations.
- Implement DevOps processes using Git, CI/CD and Azure DevOps.
- Optimize PySpark jobs for performance, scalability and cost efficiency.
- Implement best practices for data governance, security and access control.
- Troubleshoot production issues and perform root cause analysis.
- Conduct code reviews, ensuring adherence to coding standards and data quality requirements.
- Collaborate with Data Architects to define architecture standards and design patterns.
- Prepare technical documentation, solution diagrams, and operational runbooks.
- Collaborate with business stakeholders to understand requirements and translate them into technical solutions.
Mandatory Skills:
- Azure Databricks: Notebooks, Jobs, Workflows, Delta Lake, Spark SQL, performance optimization and debugging.
- PySpark: DataFrames, transformations, distributed data processing, and optimization.
- Azure Data Factory (ADF): Pipelines, triggers, integrations, runtime management, and monitoring.
- Azure Data Lake Storage Gen2 (ADLS Gen2): Storage architecture, folder structures, partitioning and security.
- Strong understanding of data partitioning strategies and performance tuning.
- Experience with security, access control and data governance best practices.
- CI/CD implementation using Git, branching strategies and Azure DevOps Pipelines.
- Strong SQL skills with proficiency in writing and optimizing complex queries.