C
Senior Generative AI Engineer
Job Description
Cogniify is seeking a hands-on Senior Generative AI Engineer to design, build, and deploy production-grade AI applications. The role focuses on LLM-powered products, RAG pipelines, and intelligent agent workflows, using Python alongside cloud infrastructure.
Role Focus
- Design and develop scalable LLM-powered applications in Python
- Build RAG pipelines using document processing, embeddings, vector databases, semantic search, and reranking
- Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration, and human approval steps
- Integrate LLMs with internal systems, external APIs, databases, and enterprise applications
Key Responsibilities
- Evaluate and select foundation models based on accuracy, latency, cost, security, and business needs
- Improve prompt quality, retrieval accuracy, response time, and token usage
- Implement safety guardrails, output validation, access controls, and fallback mechanisms
- Create automated evaluation frameworks to measure response quality, hallucination, relevance, and reliability
- Containerize applications with Docker and deploy on AWS, Azure, or GCP
- Apply monitoring and LLMOps practices for model performance, cost, latency, errors, and production usage
- Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions
- Document technical architecture, design decisions, APIs, and operational processes
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical software-development experience
- 6-9 years of professional software-development experience, including strong hands-on experience with Python
- Experience developing backend services and integrating REST APIs
- Hands-on experience building LLM or Generative AI applications
- Practical experience implementing RAG using embeddings, semantic search, and vector databases
- Experience with LangChain, LangGraph, LlamaIndex, or a comparable LLM application framework
- Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini, or Azure OpenAI
- Working knowledge of vector databases including Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, or pgvector
- Experience with Docker and deployment on at least one cloud platform: AWS, Azure, or GCP
- Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation
- Strong understanding of software engineering practices including Git, testing, debugging, and clean code
- Ability to communicate technical solutions clearly to technical and non-technical stakeholders
Technologies and Tools
- Python, Retrieval-Augmented Generation (RAG), embeddings, semantic search, reranking
- Frameworks: LangChain, LangGraph, LlamaIndex
- Model providers/APIs: OpenAI, Anthropic Claude, Gemini, Azure OpenAI
- Vector databases: Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, pgvector
- Cloud and DevOps: Docker, AWS, Azure, GCP, Kubernetes, CI/CD pipelines
- ML/Serving and ecosystems: Hugging Face Transformers, PyTorch, LoRA, QLoRA, vLLM, TGI, Ollama
- Observability/LLMOps: LangSmith, Langfuse, Arize Phoenix, MLflow, Weights & Biases
- Agent frameworks: AutoGen, CrewAI, Semantic Kernel
Benefits
- Unlimited PTO
- Very generous parental leave, above industry standards
- Entrepreneurial culture with frequent experimentation and risk-taking
- Open communication with management and company leadership
- Small, dynamic teams with high impact
- Medical, Dental and Vision coverage for employees
- Access to Disability & Life insurance
- Mental health and wellbeing support
- Annual bonus program
- Employer Stock Purchase Program (ESPP)
- Yearly team building experiences
- Mentorship and sponsorship opportunities
- Manager resources and support
What Success Looks Like
- Production-ready AI applications that are accurate, secure, and maintainable
- RAG systems that retrieve relevant information and reduce hallucinations
- AI-agent workflows that reliably complete business tasks and integrate with existing systems
- Measurable improvements in response quality, latency, and inference cost
- Clear monitoring of application performance, usage, errors, and model behaviour
Location and Compensation
- Location: Hybrid remote in the San Francisco Bay Area, CA
- Salary: USD 150,000 - 170,000 per year
- Salary range note: US East/West Coast: $150000 - $170000
Education
Bachelor’s or Master’s degree in Computer Science, Engineering or a related discipline.