What are the responsibilities and job description for the Java Spark Engineer position at Compunnel Inc.?
Key Responsibilities
Architect and build scalable, fault-tolerant data pipelines using Apache Spark and Java
Lead the design and implementation of batch and streaming ETL/ELT systems handling large data volumes
Perform deep performance tuning, including partitioning strategies, memory management, shuffle and skew optimization, and job cost reduction
Establish coding standards and lead code and design reviews across the team
Drive technical decisions related to data architecture, storage formats, and pipeline orchestration
Mentor mid-level and junior engineers and serve as a technical escalation point
Partner with product, analytics, and platform teams to translate requirements into scalable data systems
Own production reliability, including on-call responsibilities, incident response, and root-cause analysis for pipeline failures
Evaluate and introduce new tools and frameworks where they improve system performance, reliability, or maintainability
Contribute to capacity planning and cost optimization for cluster infrastructure
Communicate technical tradeoffs, risks, and developmental challenges effectively with technical and non-technical stakeholders
Work effectively both independently and as part of a collaborative team in a changing environment
Required Qualifications
7 years of experience
Bachelor's degree in Computer Science, Engineering, or a related field
Master's degree in Computer Science, Engineering, or a related field
Skills
Java Development
Apache Spark
Distributed Systems
SQL
Columnar and Modern Data Storage Formats
Cluster Managers
Kafka
CI/CD
Containerization
Infrastructure-as-Code
ETL/ELT Pipeline Development
Flink
Data Governance Frameworks
Workflow Orchestration
Multi-tenant and Multi-region Data Platform Design
Technical Leadership
Mentoring and Coaching
Code Review Facilitation
Data Architecture
Capacity Planning
Incident Response Coordination
Stakeholder Communication
Cross-functional Collaboration
Requirements Gathering
Analytical Problem Solving
Schedule
Start date: 2026-09-23
Architect and build scalable, fault-tolerant data pipelines using Apache Spark and Java
Lead the design and implementation of batch and streaming ETL/ELT systems handling large data volumes
Perform deep performance tuning, including partitioning strategies, memory management, shuffle and skew optimization, and job cost reduction
Establish coding standards and lead code and design reviews across the team
Drive technical decisions related to data architecture, storage formats, and pipeline orchestration
Mentor mid-level and junior engineers and serve as a technical escalation point
Partner with product, analytics, and platform teams to translate requirements into scalable data systems
Own production reliability, including on-call responsibilities, incident response, and root-cause analysis for pipeline failures
Evaluate and introduce new tools and frameworks where they improve system performance, reliability, or maintainability
Contribute to capacity planning and cost optimization for cluster infrastructure
Communicate technical tradeoffs, risks, and developmental challenges effectively with technical and non-technical stakeholders
Work effectively both independently and as part of a collaborative team in a changing environment
Required Qualifications
7 years of experience
Bachelor's degree in Computer Science, Engineering, or a related field
Master's degree in Computer Science, Engineering, or a related field
Skills
Java Development
Apache Spark
Distributed Systems
SQL
Columnar and Modern Data Storage Formats
Cluster Managers
Kafka
CI/CD
Containerization
Infrastructure-as-Code
ETL/ELT Pipeline Development
Flink
Data Governance Frameworks
Workflow Orchestration
Multi-tenant and Multi-region Data Platform Design
Technical Leadership
Mentoring and Coaching
Code Review Facilitation
Data Architecture
Capacity Planning
Incident Response Coordination
Stakeholder Communication
Cross-functional Collaboration
Requirements Gathering
Analytical Problem Solving
Schedule
Start date: 2026-09-23