EngineerJobs.io
← Back to all jobs

Job Description

The Gen AI Engineer will design and implement production-ready AI services that use Java and a microservices architecture. The role focuses on building and deploying API layers for large language models, including RAG, monitoring, testing, and evaluation within an agile delivery environment.

Responsibilities

  • Design, develop, and deploy production-grade Java microservices (such as with Spring Boot) that integrate with large language models and AI/ML platforms.
  • Architect and implement RESTful APIs and event-driven communication patterns to deliver generative AI capabilities to internal and external clients.
  • Develop prompt engineering strategies, including prompt templating, chaining, and dynamic context management to improve output quality for business use cases.
  • Implement Retrieval-Augmented Generation (RAG), including integration with vector databases to generate grounded, context-aware responses from proprietary knowledge bases.
  • Build resilient model integration layers with retry logic, timeouts, rate limiting, and fallback mechanisms to support high availability.
  • Create services for asynchronous processing and streaming responses to support long-running inference tasks efficiently.
  • Implement monitoring, logging, and observability capabilities to track AI service health, including token usage, latency, and model performance.
  • Collaborate with data scientists and ML engineers to move experimental models into production, emphasizing architecture and performance optimization.
  • Develop and execute unit, integration, and load tests to ensure reliability and quality under production conditions.
  • Evaluate new generative AI models, libraries, and tools by providing technical assessments and recommendations.
  • Write technical documentation and support maintenance of internal AI service frameworks and shared libraries.
  • Stay current with developments in the Gen AI landscape and propose improvements to system architecture and development practices.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related technical field (or equivalent practical experience).
  • 5+ years of professional software development experience, with strong object-oriented programming and enterprise application development.
  • Expert proficiency in Java, including concurrency, functional programming patterns, and robust error handling.
  • Hands-on experience building and deploying microservices, including service discovery, API gateways, and distributed tracing.
  • Strong working knowledge of Spring Boot or a similar Java microservice framework, including testing and configuration.
  • Proven experience integrating with third-party AI/LLM APIs such as OpenAI, Anthropic, or open-source models in a server-side environment.
  • Solid understanding of cloud platforms (AWS, GCP, or Azure) and experience with Docker and Kubernetes (highly desirable).
  • Proficiency with SQL and NoSQL databases and practical vector database experience (Pinecone, Weaviate, pgvector) for semantic search.
  • Working knowledge of observability tools and practices, including logging, metrics, and monitoring dashboards.
  • Excellent problem-solving, debugging, and analytical skills, with a proactive and self-directed work ethic.
  • Strong communication skills and the ability to collaborate effectively while translating complex technical concepts for non-technical stakeholders.
  • Passion for learning new technologies, particularly in generative AI and machine learning engineering.

Technologies

  • Java
  • Spring Boot
  • RESTful APIs
  • OpenAI
  • Anthropic
  • Microservices architecture
  • Service discovery
  • API gateways
  • Distributed tracing
  • AWS, GCP, Azure
  • Docker
  • Kubernetes
  • SQL, NoSQL
  • Vector databases
  • Pinecone, Weaviate, pgvector
  • Retrieval-Augmented Generation (RAG)

Location and Work Mode

Atlanta, GA (onsite).

Similar Jobs