What are the responsibilities and job description for the LLM / GenAI Engineer position at Evlo AI?
About The Role
The role focuses on building production-grade generative AI systems, moving beyond simple prompting into robust architectures like advanced RAG, agentic workflows, and fine-tuning pipelines.
The team works at the intersection of applied research and scalable backend engineering to deliver reliable, low-latency, and context-aware AI applications for enterprise users.
Key Responsibilities
The role focuses on building production-grade generative AI systems, moving beyond simple prompting into robust architectures like advanced RAG, agentic workflows, and fine-tuning pipelines.
The team works at the intersection of applied research and scalable backend engineering to deliver reliable, low-latency, and context-aware AI applications for enterprise users.
Key Responsibilities
- Design, build, and optimize production-grade RAG pipelines utilizing advanced chunking, hybrid search, and re-ranking strategies
- Integrate and manage vector databases such as Pinecone, Weaviate, or pgvector for high-performance semantic retrieval
- Execute parameter-efficient fine-tuning (LoRA, QLoRA) on open-source foundation models using PyTorch and Hugging Face
- Develop automated LLM evaluation frameworks using LLM-as-a-judge patterns, deterministic test suites, and hallucination metrics
- Collaborate with backend engineers to deploy scalable inference endpoints with vLLM, Triton, or managed cloud APIs
- Monitor deployed LLM applications for latency, token costs, drift, and quality degradation using observability tools like LangSmith or Arize
- 3-6 years of software engineering or machine learning experience, with a minimum of 2 years dedicated to building and deploying LLM applications in production
- Expert proficiency in Python and deep familiarity with LLM orchestration frameworks such as LangChain or LlamaIndex
- Hands-on experience with vector embeddings, semantic search infrastructure, and prompt engineering techniques
- Solid understanding of cloud infrastructure (AWS, GCP, or Azure) and containerization tools like Docker and Kubernetes
- Bachelor's degree in Computer Science, Artificial Intelligence, or equivalent practical experience
- Bonus: Contributions to open-source AI projects, experience with multi-agent frameworks, or a published paper in NLP/ML