What are the responsibilities and job description for the LLM / GenAI Engineer position at Scale.jobs?
About The Role
This role focuses on building and scaling production-grade Generative AI systems, moving beyond basic prototyping to deploy robust RAG pipelines, agentic workflows, and customized LLM architectures. The engineer will collaborate closely with product and data teams to transition advanced AI concepts into reliable, low-latency APIs that power user-facing applications.
The position demands a deep understanding of LLM orchestration, systematic evaluation, and performance optimization. It is an opportunity to shape the core AI infrastructure of a rapidly growing platform, ensuring high-throughput, cost-effective, and safe model deployment at scale.
Key Responsibilities
This role focuses on building and scaling production-grade Generative AI systems, moving beyond basic prototyping to deploy robust RAG pipelines, agentic workflows, and customized LLM architectures. The engineer will collaborate closely with product and data teams to transition advanced AI concepts into reliable, low-latency APIs that power user-facing applications.
The position demands a deep understanding of LLM orchestration, systematic evaluation, and performance optimization. It is an opportunity to shape the core AI infrastructure of a rapidly growing platform, ensuring high-throughput, cost-effective, and safe model deployment at scale.
Key Responsibilities
- Design and deploy production-grade Retrieval-Augmented Generation (RAG) pipelines using LangChain, LlamaIndex, or custom orchestration layers.
- Optimize vector database performance and hybrid search capabilities across millions of documents using tools like Pinecone, Weaviate, or pgvector.
- Develop and execute fine-tuning strategies (including LoRA, QLoRA) on open-source foundation models like Llama and Mistral to adapt them to domain-specific tasks.
- Establish rigorous LLM evaluation frameworks, implementing LLM-as-a-judge patterns and automated regression testing to continuously monitor generation quality.
- Implement guardrails, moderation APIs, and semantic caching layers to manage prompt injection risks and reduce inference costs.
- Collaborate with MLOps and backend engineers to package models into containerized microservices deployed via Kubernetes or AWS ECS.
- Write clean, async-optimized Python code, establishing robust CI/CD, unit testing, and monitoring pipelines for Generative AI workloads.
- 3-6 years of professional software engineering experience, with at least 1.5 years dedicated to building and deploying LLM applications in a production environment.
- Strong proficiency in Python, including async programming, FastAPI, and hands-on experience with PyTorch or Hugging Face Transformers.
- Direct experience building, evaluating, and fine-tuning applications with commercial APIs (OpenAI, Anthropic) and open-source models.
- Solid understanding of vector embeddings, indexing strategies, semantic search mechanics, and database design.
- Familiarity with cloud native infrastructure (AWS, GCP, or Azure), Docker, and MLOps tools for model tracking and deployment.
- Bachelor's or Master's degree in Computer Science, Data Science, or a related technical discipline.
- Bonus: Experience with agentic frameworks (CrewAI, AutoGen), DSPy for systematic prompt programming, or deep understanding of CUDA memory optimization.