What are the responsibilities and job description for the Mid/Senior Data Engineer – AI Data Pipelines position at Pegasus Knowledge Solutions?
<>Mid/Senior Data Engineer – AI Data Pipelines <>NYC, NY<>Contract 6 Months with possible extension
We’re hiring a strong Data Engineer to turn messy customer support data into clean, AI‑ready intelligence. You’ll build high‑scale pipelines, vector DBs, and reliable data systems that power our next‑gen AI agents.
Role Highlights
- Clean, process & structure huge volumes of unstructured support data (Salesforce, Sprinklr, Bliss, JIRA).
- Build AI memory systems using vector databases (Pinecone, Milvus, Weaviate, pgvector).
- Create a central metrics hub for analytics & AI insights.
- Own high‑reliability pipelines with monitoring & zero‑downtime standards.
- Work closely with product/ops and push for better data logging at the source.
Tech Stack
- Python, PySpark
- Spark, Kafka, Flink, Hadoop, Hudi, Presto, Pinot
- Vector DBs (Pinecone/MilvWeaviate/pgvector)
- Cloud data warehouses
- AI tools (Claude, Codex, etc.) for automation & parsing
Ideal Candidate
- 5 years in data engineering with big data systems
- Strong Python PySpark
- Experience building AI‑ready data layers
- Proactive, independent, business‑focused
- Comfortable working in ambiguity and driving clarity
- Honest communicator who flags risks early