What are the responsibilities and job description for the Data Engineer position at Mind Ware Inc?
Hiring for Data Engineer @ Denver, CO for a long-term contract position on w2
Primary Skills: Azure Data Bricks, Data Factory, Pyspark and Master Data Management
Job Description :
A seasoned Data Engineer specialising in enterprise-scale data platform design across Databricks and Microsoft Fabric (Azure), with a technology-agnostic philosophy that delivers portable, future-proof solutions. Recognised for deep expertise in Medallion architecture, metadata-driven pipeline orchestration, and distributed processing with Apache Spark, paired with a strong command of data governance, Master Data Management, and enterprise catalog tooling. Rounds out a comprehensive engineering profile with disciplined CI/CD practices, Infrastructure-as-Code, and data quality observability frameworks that ensure reliable, production-grade data systems at scale.
• Proficient across Databricks and Microsoft Fabric (Azure) with a technology-agnostic approach, designing portable solutions that leverage the strengths of each platform interchangeably.
• Architects and implements Medallion (Bronze/Silver/Gold) data lake frameworks, enforcing clear separation of raw ingestion, conformance, and curated analytical layers.
• Designs and orchestrates fault-tolerant, scalable data pipelines using Azure Data Factory, Databricks Workflows, and Microsoft Fabric Pipelines — augmented by metadata-driven, configuration-as-code automation frameworks that enable dynamic pipeline generation, parameterization, and self-service onboarding of new data sources with minimal manual effort.
• Applies Apache Spark (PySpark / Spark SQL) for large-scale distributed data processing, transformation, and performance-tuned query optimization across batch and streaming workloads.
• Implements Master Data Management (MDM) solutions including golden-record creation, entity resolution, probabilistic/deterministic matching, and deduplication to ensure a single source of truth.
• Enforces data governance, stewardship, and data-contract standards — defining ownership, access policies, SLA commitments, and end-to-end lineage — while configuring enterprise data catalogs via Unity Catalog (Databricks) and Microsoft Purview (Fabric/Azure) for asset discovery, classification, sensitivity labeling, and access control.
• Establishes data quality frameworks and observability pipelines with automated profiling, anomaly detection, and SLA monitoring to proactively detect and remediate data issues in production.
• Applies CI/CD practices, Git-based version control, Infrastructure-as-Code (IaC), and rigorous unit/integration testing for Python and SQL codebases to ensure reliable, repeatable deployments.