Gen AI Engineer with Java & Microservices
Job Description
The Gen AI Engineer will design and implement production-ready AI services that use Java and a microservices architecture. The role focuses on building and deploying API layers for large language models, including RAG, monitoring, testing, and evaluation within an agile delivery environment.
Responsibilities
- Design, develop, and deploy production-grade Java microservices (such as with Spring Boot) that integrate with large language models and AI/ML platforms.
- Architect and implement RESTful APIs and event-driven communication patterns to deliver generative AI capabilities to internal and external clients.
- Develop prompt engineering strategies, including prompt templating, chaining, and dynamic context management to improve output quality for business use cases.
- Implement Retrieval-Augmented Generation (RAG), including integration with vector databases to generate grounded, context-aware responses from proprietary knowledge bases.
- Build resilient model integration layers with retry logic, timeouts, rate limiting, and fallback mechanisms to support high availability.
- Create services for asynchronous processing and streaming responses to support long-running inference tasks efficiently.
- Implement monitoring, logging, and observability capabilities to track AI service health, including token usage, latency, and model performance.
- Collaborate with data scientists and ML engineers to move experimental models into production, emphasizing architecture and performance optimization.
- Develop and execute unit, integration, and load tests to ensure reliability and quality under production conditions.
- Evaluate new generative AI models, libraries, and tools by providing technical assessments and recommendations.
- Write technical documentation and support maintenance of internal AI service frameworks and shared libraries.
- Stay current with developments in the Gen AI landscape and propose improvements to system architecture and development practices.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a related technical field (or equivalent practical experience).
- 5+ years of professional software development experience, with strong object-oriented programming and enterprise application development.
- Expert proficiency in Java, including concurrency, functional programming patterns, and robust error handling.
- Hands-on experience building and deploying microservices, including service discovery, API gateways, and distributed tracing.
- Strong working knowledge of Spring Boot or a similar Java microservice framework, including testing and configuration.
- Proven experience integrating with third-party AI/LLM APIs such as OpenAI, Anthropic, or open-source models in a server-side environment.
- Solid understanding of cloud platforms (AWS, GCP, or Azure) and experience with Docker and Kubernetes (highly desirable).
- Proficiency with SQL and NoSQL databases and practical vector database experience (Pinecone, Weaviate, pgvector) for semantic search.
- Working knowledge of observability tools and practices, including logging, metrics, and monitoring dashboards.
- Excellent problem-solving, debugging, and analytical skills, with a proactive and self-directed work ethic.
- Strong communication skills and the ability to collaborate effectively while translating complex technical concepts for non-technical stakeholders.
- Passion for learning new technologies, particularly in generative AI and machine learning engineering.
Technologies
- Java
- Spring Boot
- RESTful APIs
- OpenAI
- Anthropic
- Microservices architecture
- Service discovery
- API gateways
- Distributed tracing
- AWS, GCP, Azure
- Docker
- Kubernetes
- SQL, NoSQL
- Vector databases
- Pinecone, Weaviate, pgvector
- Retrieval-Augmented Generation (RAG)
Location and Work Mode
Atlanta, GA (onsite).