What are the responsibilities and job description for the Sr. Data Engineer position at Tredence Inc.?
Role description
About Tredence¬: -
Tredence focuses on last mile delivery of insights into actions by uniting its strengths in business analytics, data science, and software engineering. The largest companies across industries are engaging with Tredence and deploying its prediction and optimization solutions at scale –empowering end users to improve decision making. Headquartered in the San Francisco Bay Area, the company serves clients in the US, Canada, Europe, and SE Asia. Learn more at www.tredence.com
Job Summary
We are seeking an experienced Databricks Architect to design, build, and optimize our modern Lakehouse platform. You will leverage the full Databricks ecosystem — including Delta Lake, Lakehouse architecture, and unified data governance — to deliver scalable, high-performance data solutions. You will be responsible for data modeling, quality enforcement, and ETL/ELT pipelines using PySpark, Python, and SQL.
Key Responsibilities
Databricks Platform & Architecture
Architect and implement end-to-end Lakehouse solutions on Databricks.
Design and optimize Delta Lake storage, including ACID transactions, time travel, and schema evolution.
Configure clustering, partitioning, Z-order, and vacuuming for performance tuning.
Data Modeling & Governance
Build data models (Kimball, Inmon, Data Vault, or medallion architecture: bronze/silver/gold).
Implement Data Catalog (Unity Catalog) for metadata management, lineage, and access control.
Define and enforce Data Quality rules using Great Expectations, DBT, or Databricks DLT (Delta Live Tables).
Development & Pipelines
Develop scalable ETL/ELT pipelines using PySpark, Python, and SQL.
Optimize Spark jobs for performance, cost, and reliability.
Automate workflows with Databricks Jobs, workflows, and orchestration tools (Airflow, Azure Data Factory, etc.).
Collaboration & Best Practices
Partner with data scientists, analysts, and business stakeholders to understand data requirements.
Establish CI/CD for Databricks notebooks and repositories (DBFS, Repos, Git integration).
Monitor and troubleshoot pipeline failures, data drift, and performance bottlenecks.
Required Qualifications
Area Skills
Databricks Core Lakehouse, Delta Lake, Unity Catalog, DLT, Workflows
Languages PySpark, Python, SQL
Data Modeling Star schema, slowly changing dimensions (SCD), medallion architecture
Data Quality Validation, anomaly detection, DQ rules implementation
Catalog & Governance Unity Catalog, Hive Metastore, lineage, access controls
Experience 5 years in data engineering; 2 years hands-on with Databricks
Preferred Qualifications
Databricks Certification (e.g., Data Engineer Professional or Associate).
Experience with cloud platforms (AWS S3/Glue, Azure Data Lake/ADF, GCP).
Knowledge of streaming (Kafka, Kinesis, Structured Streaming).
Familiarity with DBT, Terraform, or MLflow.
Soft Skills
Strong analytical and problem-solving abilities.
Excellent communication for technical and non-technical audiences.
Self-starter capable of leading architectural decisions.