What are the responsibilities and job description for the Azure Databricks Engineer position at NLB Services?
Job Role: Azure Databricks Lead
Location: Jersey City, NJ (Hybrid - 3 Days per Week)
Job Type: Full-Time Permanent role
Job Description:
We are looking for an experienced Azure Databricks Engineer with strong expertise in cloud-based data engineering, ETL development, and distributed data processing. The ideal candidate should have solid hands-on experience with PySpark, Delta Lake, Azure Data Factory, and building scalable data pipelines on Azure.
The engineer will work closely with business, Data Architects, and cross-functional teams to design, develop, and optimize data pipelines for enterprise-grade analytics and reporting.
Key Responsibilities
Data Engineering and Pipeline Development
- Design, develop, and optimize ETL/ELT pipelines using Azure Databricks and PySpark.
- Build scalable data ingestion workflows from various structured and unstructured sources.
- Implement transformation logic, data cleansing, enrichment, and validation frameworks.
- Work with Delta Lake to build medallion architecture: Bronze, Silver, and Gold layers.
- Develop reusable Databricks notebooks and jobs for production data workflows.
Azure Cloud Integration
- Build and orchestrate pipelines using Azure Data Factory (ADF).
- Integrate Databricks with other Azure services such as ADLS, Azure SQL, Event Hub, Key Vault, and Synapse.
- Optimize compute environments, clusters, pools, and autoscaling.
- Implement DevOps processes using Git, CI/CD, and Azure DevOps.
Performance, Quality, and Governance
- Optimize PySpark jobs for performance and cost efficiency.
- Implement best practices for data governance, security, and access control.
- Troubleshoot production issues and perform root-cause analysis.
- Conduct code reviews, ensuring coding standards and data quality.
Collaboration and Documentation
- Work with Data Architects to define architecture and design patterns.
- Prepare technical documents, solution diagrams, and runbooks.
- Collaborate with business stakeholders to understand requirements and translate them into technical solutions.
Mandatory Skills
- Azure Databricks notebooks, jobs, workflows, and Delta Lake.
- PySpark dataframes, Spark SQL optimization, and debugging.
- Azure Data Factory (ADF), triggers, pipelines, and integration runtime.
- Data Lake Storage (ADLS Gen2), folder structures, partitioning, and security.
- CI/CD, Git branching strategies, and Azure DevOps pipelines.
- SQL, with strong proficiency in writing optimized queries.
Good-to-Have Skills
- Azure Synapse Analytics.
- Azure Event Hub and Kafka.
- Azure Functions.
- Databricks REST APIs.
- Streaming pipelines and Structured Streaming.
- Experience with data modeling.
- Knowledge of Lakehouse architecture.
Behavioral and Soft Skills
- Strong analytical and problem-solving skills.
- Ability to work independently and in cross-functional teams.
- Good communication skills for stakeholder interaction.
- Comfortable working in Agile/Scrum models.