What are the responsibilities and job description for the Data Engineer position at Evlo AI?
About The Role
The role owns the design, implementation, and scaling of core data infrastructure, high-throughput ETL pipelines, and analytical data models that drive company-wide decision-making.
The team works closely with data scientists, product managers, and analytics engineers to ensure data reliability, performance, and accessibility across all business units.
Key Responsibilities
The role owns the design, implementation, and scaling of core data infrastructure, high-throughput ETL pipelines, and analytical data models that drive company-wide decision-making.
The team works closely with data scientists, product managers, and analytics engineers to ensure data reliability, performance, and accessibility across all business units.
Key Responsibilities
- Architect, build, and maintain scalable batch and streaming data pipelines using Python, SQL, and Apache Spark
- Manage cloud data warehousing infrastructure (Snowflake, BigQuery, or Redshift) with a focus on cost optimization and query performance
- Implement robust data quality checks, schema evolution handling, and monitoring tools to guarantee data integrity across all sources
- Collaborate with analytics and engineering teams to model data for reporting, machine learning features, and operational workflows
- Write clean, modular, and well-tested infrastructure-as-code using Terraform and CI/CD deployment pipelines
- Document data lineage, schemas, and pipeline architectures to foster a culture of data transparency and self-service
- 3–6 years of professional experience in data engineering, backend development, or analytics engineering roles
- Advanced SQL and Python programming skills, with proven expertise in building production-grade data pipelines
- Hands-on experience with modern cloud data warehouses (Snowflake, BigQuery) and orchestrators (Airflow, Dagster, Prefect)
- Solid understanding of distributed data processing frameworks such as Apache Spark, Flink, or Ray
- Bachelor's degree in Computer Science, Statistics, Engineering, or equivalent practical experience
- Bonus: Experience with real-time streaming technologies (Kafka, Kinesis) and data lakehouse formats (Iceberg, Delta Lake)