EngineerJobs.io
← Back to all jobs

Job Description

Xenoss is hiring a Staff AI Engineer / AI Solution Architect to lead the applied AI architecture for a long-term In-Call Assistant initiative for a world-leading financial services organization. The work focuses on designing real-time conversational intelligence that detects customer needs, objections, and buying signals, then recommends the required process steps with compliance guardrails.

In this PoC-first engagement, you will set the technical direction for signal and intent understanding, model training and evaluation, and the move from offline experimentation to production-grade low-latency inference, all executed inside the client perimeter.

Role context and project scope

  • Own the applied AI architecture and evaluation strategy for a complex enterprise AI program at Staff/Architect level.
  • Lead the AI architecture for a long-term In-Call Assistant initiative that supports front-office employees during live customer conversations.
  • Design a system that identifies customer needs, objections, buying signals, and required process steps, delivering concise, context-aware recommendations.
  • Cover end-to-end system design including low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, confidence management, and compliance guardrails.
  • Help define how the AI architecture, models, evaluation framework, and feedback loops evolve from an initial offline version to live production use.

Responsibilities

  • Lead the applied AI architecture across the In-Call Assistant lifecycle, from data and taxonomy design to model training, evaluation, and production readiness.
  • Design the end-to-end AI architecture for the In-Call Assistant.
  • Define signal and trigger taxonomies for live conversations.
  • Design training strategies for signal detection and specialist recommendation models.
  • Shape data preparation, annotation, and SME validation workflows.
  • Evaluate fine-tuning, post-training, RAG, and hybrid approaches.
  • Design low-latency signal detection, routing, context preparation, and confidence management.
  • Design evaluation frameworks, golden datasets, and model improvement cycles.
  • Define grounding, guardrails, abstention, and policy-compliance behavior.
  • Make trade-offs between model quality, latency, cost, explainability, and governance.
  • Partner with AI engineers, data engineering, MLOps, and client SMEs.
  • Define the AI approach for the conversation intelligence PoC, including event/intent/insight taxonomy and golden dataset strategy and annotation workflow.
  • Establish evaluation frameworks and acceptance criteria, driving trade-offs between accuracy, explainability, latency, cost, and governance.
  • Decide which modeling approaches are appropriate for each use case.
  • Act as an escalation point for AI architecture, evaluation, and data strategy decisions.
  • Work within a cross-functional team spanning AI engineering, data engineering, MLOps, solution architecture, and client stakeholders.
  • Partner with domain SMEs on taxonomy, labeling, and validation.
  • Mentor engineers working on extraction, evaluation, and data pipelines.
  • Translate ambiguous business use cases into testable AI hypotheses and validation plans.

Requirements

  • Strong hands-on experience with applied AI/ML systems in production-oriented environments.
  • Experience with NLP, conversational AI, or transcript-based intelligence systems.
  • Ability to design evaluation frameworks, not just run experiments.
  • Experience building or validating structured datasets from unstructured text.
  • Strong understanding of LLM-based extraction, classification, RAG, and fine-tuning trade-offs.
  • Practical knowledge of classical ML or predictive modeling.
  • Understanding of probability-based prediction, calibration, and outcome evaluation.
  • Comfort working with messy enterprise data and incomplete labels.
  • Ability to communicate with both technical teams and business stakeholders.
  • Strong ownership of ambiguity, scope control, and PoC validation strategy.
  • Financial services domain exposure.
  • Experience with sales, call center, or customer conversation analytics.
  • Speech/ASR pipeline familiarity.
  • Model governance and auditability experience.
  • Experience with real-time AI systems or low-latency inference.
  • Experience combining unstructured conversation signals with structured CRM, transaction, or customer profile data.
  • Experience designing golden datasets and SME review workflows.

Technologies

  • LLM, SFT, DPO, preference optimization, LoRA, QLoRA, PEFT
  • PyTorch, Hugging Face
  • RAG, embeddings, retrieval, MLOps

Technology landscape

  • LLM and smaller-model training for conversational AI
  • SFT, DPO / preference optimization, LoRA / QLoRA, and PEFT
  • PyTorch and Hugging Face ecosystem
  • Signal extraction and multi-label classification
  • RAG and knowledge-grounded recommendation generation
  • Embeddings, retrieval, and context preparation
  • Low-latency model serving and inference optimization
  • Golden dataset creation, annotation, and SME validation
  • Model evaluation, confidence calibration, and error analysis
  • MLOps, monitoring, feedback loops, and model governance

Delivery and work environment

  • Engagement structure: FTE-equivalent via long-term B2B contract.
  • Work location: On-site or closely aligned with the client team in New York.
  • Infrastructure: Client environment only, with no external training or data processing environments.
  • Data residency: All work executed within the client perimeter.
  • Delivery mode: PoC-first, with a path toward production-grade conversation intelligence and prediction systems.

Location: New York, NY (onsite)

Similar Jobs