What are the responsibilities and job description for the New contract opening of Python Pyspark Databricks -- Boston, MA position at Signature IT World Inc?
Greetings,
I hope you're having a great day!
I came across your profile and I have a position which you might be interested in. If you think you are a good fit for this role, then inbox me your updated resume at kamil.qureshi@sitwinc.com
Role: Python Pyspark Databricks
Location: Boston, MA - Onsite
Employment: Contract
Job Summary
We are seeking an experienced Python PySpark Databricks Developer to design, develop, and optimize scalable data engineering solutions on modern cloud-based analytics platforms. The ideal candidate will have strong expertise in Python, PySpark, Databricks, and big data technologies, with experience building robust ETL/ELT pipelines, processing large datasets, and supporting enterprise data initiatives.
Key Responsibilities
Design, develop, and maintain scalable data pipelines using Python, PySpark, and Databricks.
Build and optimize ETL/ELT workflows to ingest, transform, and process structured and unstructured data.
Develop high-performance Spark applications for large-scale data processing.
Implement data quality checks, validation frameworks, and error-handling mechanisms.
Work with Delta Lake to build reliable and efficient data lakes.
Optimize Spark jobs for performance, scalability, and cost efficiency.
Collaborate with data architects, data scientists, analysts, and business stakeholders to understand data requirements.
Integrate data from multiple sources including databases, APIs, cloud storage, and streaming platforms.
Participate in code reviews, testing, deployment, and production support.
Monitor data pipelines and troubleshoot performance issues.
Follow data governance, security, and best practices throughout the development lifecycle.
Create technical documentation and maintain data pipeline documentation.
Required Skills
Strong experience with Python programming.
Hands-on experience with PySpark for distributed data processing.
Extensive experience working with Databricks.
Strong understanding of Apache Spark architecture and optimization techniques.
Experience with Delta Lake.
Knowledge of SQL and relational databases.
Experience designing and maintaining ETL/ELT pipelines.
Familiarity with data modeling and data warehousing concepts.
Experience with Git and version control.
Understanding of CI/CD pipelines and DevOps practices.
I hope you're having a great day!
I came across your profile and I have a position which you might be interested in. If you think you are a good fit for this role, then inbox me your updated resume at kamil.qureshi@sitwinc.com
Role: Python Pyspark Databricks
Location: Boston, MA - Onsite
Employment: Contract
Job Summary
We are seeking an experienced Python PySpark Databricks Developer to design, develop, and optimize scalable data engineering solutions on modern cloud-based analytics platforms. The ideal candidate will have strong expertise in Python, PySpark, Databricks, and big data technologies, with experience building robust ETL/ELT pipelines, processing large datasets, and supporting enterprise data initiatives.
Key Responsibilities
Design, develop, and maintain scalable data pipelines using Python, PySpark, and Databricks.
Build and optimize ETL/ELT workflows to ingest, transform, and process structured and unstructured data.
Develop high-performance Spark applications for large-scale data processing.
Implement data quality checks, validation frameworks, and error-handling mechanisms.
Work with Delta Lake to build reliable and efficient data lakes.
Optimize Spark jobs for performance, scalability, and cost efficiency.
Collaborate with data architects, data scientists, analysts, and business stakeholders to understand data requirements.
Integrate data from multiple sources including databases, APIs, cloud storage, and streaming platforms.
Participate in code reviews, testing, deployment, and production support.
Monitor data pipelines and troubleshoot performance issues.
Follow data governance, security, and best practices throughout the development lifecycle.
Create technical documentation and maintain data pipeline documentation.
Required Skills
Strong experience with Python programming.
Hands-on experience with PySpark for distributed data processing.
Extensive experience working with Databricks.
Strong understanding of Apache Spark architecture and optimization techniques.
Experience with Delta Lake.
Knowledge of SQL and relational databases.
Experience designing and maintaining ETL/ELT pipelines.
Familiarity with data modeling and data warehousing concepts.
Experience with Git and version control.
Understanding of CI/CD pipelines and DevOps practices.