What are the responsibilities and job description for the LLM / GenAI Engineer position at Evlo AI?
About The Role
The LLM / GenAI Engineer will design, build, and operate production AI systems spanning retrieval-augmented generation, agentic workflows, model adaptation, and evaluation. The work goes beyond prompt engineering: it requires reliable data pipelines, measurable model behavior, scalable inference, and clear operational controls.
Based in Denver, CO with a remote working model, the role will partner with applied scientists, product engineers, and platform teams to turn emerging language-model capabilities into dependable user-facing products. Success means improving answer quality, latency, cost, and safety while maintaining systems that can be monitored and iterated in production.
Key Responsibilities
The LLM / GenAI Engineer will design, build, and operate production AI systems spanning retrieval-augmented generation, agentic workflows, model adaptation, and evaluation. The work goes beyond prompt engineering: it requires reliable data pipelines, measurable model behavior, scalable inference, and clear operational controls.
Based in Denver, CO with a remote working model, the role will partner with applied scientists, product engineers, and platform teams to turn emerging language-model capabilities into dependable user-facing products. Success means improving answer quality, latency, cost, and safety while maintaining systems that can be monitored and iterated in production.
Key Responsibilities
- Design and implement RAG and agentic application architectures using Python, LangChain, LlamaIndex, or custom orchestration services
- Build ingestion, chunking, embedding, retrieval, reranking, and citation pipelines using vector technologies such as Pinecone, Weaviate, Elasticsearch, or pgvector
- Develop evaluation frameworks with curated benchmark sets, LLM-as-judge workflows, human review processes, regression testing, and quality dashboards
- Fine-tune and adapt foundation models using supervised fine-tuning, LoRA, QLoRA, prompt optimization, and domain-specific training datasets
- Deploy and optimize model-serving systems on AWS, GCP, or Azure, balancing throughput, latency, context-window limits, and inference cost
- Instrument production AI services with tracing, logging, feedback capture, and monitoring for hallucinations, drift, safety issues, and performance regressions
- Collaborate on technical design reviews, write tested and maintainable software, and document model behavior, data lineage, and operational runbooks
- 3–8 years of software engineering, machine learning engineering, or applied AI experience, including at least 1 year delivering LLM or GenAI systems to production
- Strong Python skills with experience building REST or gRPC services, asynchronous workflows, automated tests, and production data-processing pipelines
- Hands-on experience with LLM application patterns including RAG, tool calling, structured generation, prompt versioning, and multi-step agent workflows
- Working knowledge of transformer architectures, tokenization, embeddings, attention mechanisms, fine-tuning strategies, and model evaluation methods
- Experience with cloud infrastructure and deployment tooling such as Docker, Kubernetes, CI/CD, managed model APIs, and at least one major cloud platform
- Bachelor’s or master’s degree in computer science, machine learning, data science, engineering, or a related technical field, or equivalent practical experience
- Bonus: Experience with open-source models such as Llama, Mistral, or Qwen; distributed training; GPU optimization; guardrails and red-teaming; vector search tuning; or observability tools such as LangSmith, OpenTelemetry, or Weights & Biases