What are the responsibilities and job description for the GCP Data Engineer position at Fixity Technologies?
Key Responsibilities:
- Design, develop, and optimize scalable data pipelines for large-scale data processing and transformation.
- Develop cloud-based data engineering solutions using GCP services, including BigQuery, DataProc, Airflow, and Pub/Sub.
- Design and optimize data models, databases, and SQL queries for performance and scalability.
- Build batch and real-time/event-driven data processing pipelines using Apache Kafka and/or GCP Pub/Sub.
- Develop data processing solutions using Apache Spark, Hadoop, and Hive.
- Write efficient and maintainable code using Python and Shell Scripting.
- Apply strong Object-Oriented Programming concepts and software design patterns.
- Design and develop solutions for distributed and multi-tiered systems.
- Build and maintain CI/CD pipelines using Jenkins, Git, and GitHub.
- Implement automated testing frameworks and quality practices for data engineering solutions.
- Follow software development lifecycle processes and Agile/Scrum methodologies.
- Collaborate effectively with product teams, architects, developers, and other cross-functional teams.
- Analyze complex data engineering problems and develop effective, scalable solutions.
- Identify opportunities for automation, process improvement, and continuous improvement.
- Learn and adopt enterprise frameworks and emerging technologies to build modern Big Data solutions.
Required Skills:
- 8 years of experience in software development and/or Big Data Engineering.
- Strong hands-on experience with Google Cloud Platform (GCP).
- Strong experience with BigQuery.
- Experience with Apache Airflow / Cloud Composer.
- Experience with Google Cloud DataProc.
- Experience with GCP Pub/Sub and/or Apache Kafka.
- Strong knowledge of Apache Spark, Hadoop, and Hive.
- Strong programming experience in Python.
- Experience with Shell Scripting.
- Strong SQL and relational database experience.
- Experience designing and optimizing data models and data pipelines.
- Experience with Jenkins, Git, and GitHub.
- Knowledge of CI/CD pipelines and automated testing frameworks.
- Strong understanding of distributed systems, algorithms, and software design patterns.
- Good understanding of Object-Oriented Programming concepts.
- Experience working in Agile/Scrum environments and understanding of SDLC processes.
- Strong analytical, communication, and problem-solving skills.
Preferred Qualifications:
- Google Cloud Professional Data Engineer Certification is a plus.
- Experience with enterprise Big Data frameworks and architecture.
- Knowledge of infrastructure automation and configuration management.
- Experience with continuous integration and deployment practices.
- Strong curiosity and willingness to learn new technologies.