Senior AI Engineer II
Job Description
Floor & Decor is hiring a Senior AI Engineer II to design, build, and operate an agentic AI framework and platform on Microsoft Azure, and to deliver production-grade agentic solutions end to end. The role combines architecture and engineering execution with evaluation, observability, security, and cost control.
Core Responsibilities
- Design, build, and evolve a shared agentic framework that enables teams to build, deploy, and operate agentic solutions, including agent scaffolding, tool and integration interfaces, context and memory management, evaluation, guardrails, and observability.
- Convert one-off agent implementations into reusable components such as service templates, harness components, MCP server patterns, evaluation harnesses, and deployment pipelines.
- Design and maintain MCP integrations with enterprise systems (including inventory and merchandising and systems that follow) so agents can use live operational data without reworking integrations for each solution.
- Define and document standards, contracts, and reference architectures that agentic workloads follow.
- Make and document architectural decisions supporting reliability, scalability, security, and observability for AI services deployed in Azure.
- Build and operate production agentic solutions end to end, from problem framing and deployment through ongoing operation, using these deliveries to improve the underlying platform.
- Select appropriate approaches for each problem, including tool calling, multi-step planning, retrieval, memory, human-in-the-loop review, or deterministic service design when an agent is not appropriate.
- Where retrieval is warranted, design and tune retrieval strategies on Azure AI Search, including chunking, vector indexing, retrieval ranking, and context engineering.
- Contribute across the stack, including an Angular front end and a Python-based service with an LLMOps layer, taking ownership to finish features rather than deferring work.
- Support the transition of mature agentic products to partner teams through documentation, runbooks, and knowledge transfer so solutions remain maintainable after initial build-out.
- Establish evaluation and regression testing as a core platform capability, including eval sets, LLM-as-judge scoring, task-level success metrics, and regression gates in CI for prompt, model, tool, and retrieval changes.
- Own evaluation pipelines for retrieval-based components using Ragas, tracking faithfulness, answer relevance, and context precision across releases.
- Instrument agentic systems for observability using Langfuse alongside Azure Monitor, including tracing, latency, token and cost attribution, tool-call success rates, quality signals, and failure modes.
- Manage AI cost and performance through token budgeting, caching, model routing and right-sizing, and latency optimization.
- Own production services including on-call participation, incident response, and post-incident follow-up.
- Troubleshoot complex issues across agent behavior, retrieval quality, hallucinations, latency, and integration reliability.
- Use spec-driven development by translating ambiguous business requests into clear specifications and acceptance criteria before implementation, keeping specs aligned with the delivered solution.
- Define and model standards for AI-augmented development, including responsible use of agentic coding tools and the review discipline that accompanies them.
- Conduct code reviews and mentor peers through technical influence and continuous learning.
- Partner with product managers and business stakeholders to identify workflows suitable for automation, translate them into technical requirements, and help prioritize backlog work.
- Act as a technical consultant to other teams adopting the platform, helping them build within the platform rather than around it.
- Participate in agile ceremonies including sprint planning, retrospectives, and daily stand-ups.
- Communicate technical trade-offs and architecture decisions clearly to technical and non-technical audiences.
- Partner with security, data, and platform teams on governance, PII handling, prompt injection defense, and responsible use practices.
- Evaluate emerging agent capabilities, tooling, protocols, and Azure OpenAI / AI Foundry updates, and recommend and prototype improvements to keep the platform current.
- Establish and maintain engineering best practices, including CI/CD pipelines, infrastructure as code, code quality standards, and security practices for AI workloads.
- Continuously reduce the cost and time required to bring the next agentic solution into production.
Required Qualifications
- 7 to 10 years of professional software engineering experience with increasing scope and ownership.
- Deep backend engineering proficiency in at least one of: C#/.NET, Java, Node.js/TypeScript, or Python.
- Experience with service-based and microservice architectures, RESTful API design, and asynchronous communication, including API versioning, contract design, error semantics, and backward compatibility.
- Production experience with cloud-based serverless microservices (Azure Functions, Container Apps, or equivalent).
- Event-driven architecture experience including queues, pub/sub, idempotency, retries, and dead-letter handling.
- Strong data fundamentals including relational and NoSQL data modeling, query performance, and transactional correctness.
- Demonstrated ownership of code quality with automated testing, code review, and CI/CD as standard practice.
- Designed and deployed agentic solutions that shipped and were operated in production.
- Hands-on harness engineering, including context construction and management, tool and function-call interfaces, multi-step planning and control flow, memory and state, structured output, guardrails and validation, human-in-the-loop checkpoints, and graceful failure and fallback behavior.
- Azure OpenAI services experience including LLM APIs and Azure AI Search, plus associated Azure infrastructure (Azure AI Foundry and AWS Bedrock experience also counts).
- Experience with agent and LLM orchestration frameworks such as LangChain, LangGraph, or similar.
- Experience with MCP (Model Context Protocol) or other tool-calling/function-calling patterns for LLM-to-system integration.
- Knowledge of AI-specific failure and risk modes (hallucination, prompt injection, data leakage, non-determinism, runaway tool loops) and mitigation techniques.
- Judgment on when LLMs or agents are appropriate, and how to bound and validate outputs.
- Daily working fluency with agentic coding tools, with Claude Code as the standard.
- A credible point of view on where agentic coding tools improve delivery and where they do not, including responsible review and testing of model-generated code.
- Ability to discuss measurable impact on throughput and quality.
- Cloud fundamentals in Azure (compute, storage, networking, and IAM).
- Cloud security fundamentals including IAM, secrets management, and network boundaries.
- Experience with Agile/Scrum, including sprint ceremonies, story estimation, and backlog grooming.
- Working knowledge of Angular or comparable modern front-end frameworks.
- Strong written and verbal communication, comfort with ambiguity, and ability to move from a vague business problem to a scoped, specified, shippable increment.
Technologies
- C#/.NET, Java, Node.js/TypeScript, Python
- REST APIs
- Azure Functions, Container Apps
- Azure OpenAI, Azure AI Search, Azure AI Foundry, AWS Bedrock
- LangChain, LangGraph, MCP (Model Context Protocol)
- Claude Code, Cursor, GitHub Copilot, Devin
- Angular
- Ragas, Langfuse, Azure Monitor, Application Insights
- CI/CD pipelines, Infrastructure as code
- Bicep, Terraform, ARM
- RAG where warranted
Education
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience. Demonstrated capability is weighted above credentials.
Benefits
- Bonus opportunities
- Career advancement opportunities at every level
- 401k with company match
- Employee Stock Purchase Plan
- Referral Bonus Program
- Medical, Dental, Vision, Life, and other Insurance Plans (subject to eligibility criteria)
- Paid vacation and sick time for eligible associates
- Paid holidays plus a personal holiday
- Paid Volunteer Time Off that starts on Day 1
Work Environment and Schedule
- Hybrid position based at Atlanta, GA headquarters, with a standard office schedule Monday through Friday during core business hours.
- Collaborative, open-plan office environment within the IT department, with dedicated space for focused engineering work.
- Regular in-person collaboration with product managers, business stakeholders, and the engineering team.
- Occasional visits to retail store locations may be required to gather associate feedback and observe how the product is used in context.
- Some extended hours may be needed around major releases or on-call rotations for production incidents.
- Standard physical requirements of a professional office environment apply (prolonged sitting, use of a computer workstation, and participation in in-person and video meetings).
Travel, Environment, and Physical Requirements
- Travel may be required, including air and car travel.
- Noise level is typically quiet to moderate.
- Sedentary Work: ability to exert 10 to 20 pounds of force occasionally and/or negligible amount frequently, including lifting, carrying, pushing, pulling, or moving objects.
- Sedentary work involves sitting most of the time, with brief periods of walking or standing.
Nice to Have
- Spec-driven development and using specs to drive AI-assisted implementation
- Building internal developer platforms, frameworks, or SDKs for other engineering teams
- Building or publishing MCP servers
- LLM and agent evaluation frameworks such as Ragas, G-Eval, LLM-as-judge, or agent trajectory evaluation
- AI observability tooling such as Langfuse (or equivalent tracing and evaluation systems)
- Retrieval infrastructure beyond Azure AI Search (pgvector, Pinecone, Elastic, or hybrid search design)
- Multi-agent orchestration, agent-to-agent protocols, or durable/long-running workflow engines
- Infrastructure as code experience (Bicep, Terraform, or ARM)
- Fine-tuning and model adaptation (LoRA/PEFT, distillation, or evaluating fine-tuning against prompting and retrieval alternatives)
- LLMOps tooling such as MLflow, Weights & Biases, or Azure ML
- Retail systems familiarity (POS, OMS, inventory/merchandising platforms)
- Experience with a Center of Excellence or innovation team within a larger enterprise
- Open source contributions, technical writing, or speaking in the AI engineering space
Location
Atlanta, GA (onsite)