What are the responsibilities and job description for the AI Data Engineer position at QUANTUM TECHNOLOGIES LLC?
AI Data Engineer
Location: Tallahassee, FL, USA
Duration: 12 Months Extension
Bill Rate: $90/hr on C2C
Job Type: C2C/1099 Contract
Client: To Be Discussed Later
Work Authorization: US-Citizen, H-1B, OPT-EAD, GC-EAD
Job Description:
- Design, develop, and optimize scalable batch and real-time data pipelines using Apache Spark (PySpark/Spark SQL).
- Write complex, high-performance SQL queries for data extraction, transformation, analytics, and reporting.
- Build and maintain ETL/ELT pipelines to ingest, cleanse, transform, and integrate structured and unstructured data.
- Prepare, curate, and validate datasets for Machine Learning and Generative AI applications.
- Develop and optimize RAG (Retrieval-Augmented Generation) data pipelines using vector databases and document processing frameworks.
- Integrate enterprise data with Large Language Models (LLMs) such as OpenAI GPT, Azure OpenAI, Claude, or Gemini.
- Implement AI-powered data quality validation, anomaly detection, and automated monitoring solutions.
- Perform data validation, testing, and quality assurance to ensure data accuracy, completeness, and consistency.
- Optimize Spark jobs, SQL queries, and distributed processing for maximum performance and scalability.
- Collaborate with Data Scientists, AI Engineers, Business Analysts, and Application Developers to support AI initiatives.
- Monitor ETL workflows and AI data pipelines to ensure reliable and timely data delivery.
- Maintain technical documentation for data architecture, AI pipelines, metadata, and data models.
- Follow best practices for data governance, security, compliance, and AI model lifecycle management.
Preferred Qualifications: - 5 years of experience as a Data Engineer.
- Strong hands-on experience with SQL and query optimization.
- Extensive experience with Apache Spark (PySpark/Spark SQL).
- Strong Python programming experience.
- Experience building and maintaining scalable ETL/ELT pipelines.
- Strong understanding of relational databases and data modeling.
- Experience with Generative AI, LLMs, and AI-powered data engineering workflows.
- Knowledge of RAG architecture, embeddings, vector databases (Pinecone, ChromaDB, FAISS, or Weaviate), and semantic search.
- Experience with AI frameworks such as LangChain or LlamaIndex.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Experience with data testing, validation, and quality assurance.
- Strong analytical, troubleshooting, and communication skills.
- Ability to work effectively in an onsite, collaborative environment.
- Experience with Databricks.
- Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
- Knowledge of Kafka, event-driven architectures, or streaming data pipelines.
- Experience with Airflow or other workflow orchestration tools.
- Experience with Docker and Kubernetes.
- Familiarity with MLOps tools such as MLflow.
- Experience implementing enterprise AI governance and responsible AI practices.
Salary : $80 - $100