Staff AI Engineer/AI Solution Architect, Conversation Intelligence Systems
Job Description
Xenoss is hiring a Staff AI Engineer / AI Solution Architect to lead the applied AI architecture for a long-term In-Call Assistant initiative for a world-leading financial services organization. The work focuses on designing real-time conversational intelligence that detects customer needs, objections, and buying signals, then recommends the required process steps with compliance guardrails.
In this PoC-first engagement, you will set the technical direction for signal and intent understanding, model training and evaluation, and the move from offline experimentation to production-grade low-latency inference, all executed inside the client perimeter.
Role context and project scope
- Own the applied AI architecture and evaluation strategy for a complex enterprise AI program at Staff/Architect level.
- Lead the AI architecture for a long-term In-Call Assistant initiative that supports front-office employees during live customer conversations.
- Design a system that identifies customer needs, objections, buying signals, and required process steps, delivering concise, context-aware recommendations.
- Cover end-to-end system design including low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, confidence management, and compliance guardrails.
- Help define how the AI architecture, models, evaluation framework, and feedback loops evolve from an initial offline version to live production use.
Responsibilities
- Lead the applied AI architecture across the In-Call Assistant lifecycle, from data and taxonomy design to model training, evaluation, and production readiness.
- Design the end-to-end AI architecture for the In-Call Assistant.
- Define signal and trigger taxonomies for live conversations.
- Design training strategies for signal detection and specialist recommendation models.
- Shape data preparation, annotation, and SME validation workflows.
- Evaluate fine-tuning, post-training, RAG, and hybrid approaches.
- Design low-latency signal detection, routing, context preparation, and confidence management.
- Design evaluation frameworks, golden datasets, and model improvement cycles.
- Define grounding, guardrails, abstention, and policy-compliance behavior.
- Make trade-offs between model quality, latency, cost, explainability, and governance.
- Partner with AI engineers, data engineering, MLOps, and client SMEs.
- Define the AI approach for the conversation intelligence PoC, including event/intent/insight taxonomy and golden dataset strategy and annotation workflow.
- Establish evaluation frameworks and acceptance criteria, driving trade-offs between accuracy, explainability, latency, cost, and governance.
- Decide which modeling approaches are appropriate for each use case.
- Act as an escalation point for AI architecture, evaluation, and data strategy decisions.
- Work within a cross-functional team spanning AI engineering, data engineering, MLOps, solution architecture, and client stakeholders.
- Partner with domain SMEs on taxonomy, labeling, and validation.
- Mentor engineers working on extraction, evaluation, and data pipelines.
- Translate ambiguous business use cases into testable AI hypotheses and validation plans.
Requirements
- Strong hands-on experience with applied AI/ML systems in production-oriented environments.
- Experience with NLP, conversational AI, or transcript-based intelligence systems.
- Ability to design evaluation frameworks, not just run experiments.
- Experience building or validating structured datasets from unstructured text.
- Strong understanding of LLM-based extraction, classification, RAG, and fine-tuning trade-offs.
- Practical knowledge of classical ML or predictive modeling.
- Understanding of probability-based prediction, calibration, and outcome evaluation.
- Comfort working with messy enterprise data and incomplete labels.
- Ability to communicate with both technical teams and business stakeholders.
- Strong ownership of ambiguity, scope control, and PoC validation strategy.
- Financial services domain exposure.
- Experience with sales, call center, or customer conversation analytics.
- Speech/ASR pipeline familiarity.
- Model governance and auditability experience.
- Experience with real-time AI systems or low-latency inference.
- Experience combining unstructured conversation signals with structured CRM, transaction, or customer profile data.
- Experience designing golden datasets and SME review workflows.
Technologies
- LLM, SFT, DPO, preference optimization, LoRA, QLoRA, PEFT
- PyTorch, Hugging Face
- RAG, embeddings, retrieval, MLOps
Technology landscape
- LLM and smaller-model training for conversational AI
- SFT, DPO / preference optimization, LoRA / QLoRA, and PEFT
- PyTorch and Hugging Face ecosystem
- Signal extraction and multi-label classification
- RAG and knowledge-grounded recommendation generation
- Embeddings, retrieval, and context preparation
- Low-latency model serving and inference optimization
- Golden dataset creation, annotation, and SME validation
- Model evaluation, confidence calibration, and error analysis
- MLOps, monitoring, feedback loops, and model governance
Delivery and work environment
- Engagement structure: FTE-equivalent via long-term B2B contract.
- Work location: On-site or closely aligned with the client team in New York.
- Infrastructure: Client environment only, with no external training or data processing environments.
- Data residency: All work executed within the client perimeter.
- Delivery mode: PoC-first, with a path toward production-grade conversation intelligence and prediction systems.
Location: New York, NY (onsite)