EngineerJobs.io
← Back to all jobs

Job Description

Join Realtor.com’s AI Integrations Team in Austin, TX (onsite) and help build the AI safety layer behind RealAssist, a consumer-facing LLM real-estate assistant. This role focuses on guardrails, evaluation pipelines, and classifiers that support compliance and content moderation while keeping the experience safe, accurate, and reliable for users.

What you’ll own

  • AI safety layer end-to-end, from design through deployment
  • Fair-Housing compliance and content moderation classifiers, tuned to target high recall on disallowed content without overblocking legitimate use
  • LLM-as-a-judge evaluation for safety and compliance decisions, using category-level performance metrics
  • Guardrail evaluation pipelines that include labeled dataset curation, train/test/validate splits, and offline prod-replay to identify overblocking
  • Runtime prompt-injection and jailbreak screening using cloud guardrail services with fail-open behavior and alerting
  • Safety regression checks in CI and red-teaming to reduce the risk of guardrails degrading as the product evolves
  • Guardrail infrastructure as code with Terraform (templates, IAM, and project shape), plus participation in technical design reviews
  • Risk tiering and launch support through collaboration with ML, backend, product, and legal/compliance stakeholders

How you’ll make decisions

  • Use confusion matrices and precision/recall/F1 per category to make safety decisions defensible
  • Ground tuning and safety tradeoffs in data-driven eval metrics
  • Keep evaluation datasets calibrated to real traffic and emerging attack patterns

What you’ll bring

  • AI Safety & Evaluation Expertise (Required)
  • 4+ years of professional software/ML engineering experience, including hands-on LLM application work
  • Bachelor’s degree or equivalent experience
  • Strong Python proficiency
  • Classification-metrics literacy: precision/recall/F1 (macro vs weighted), confusion matrix debugging, dataset curation, and calibration to production distribution
  • Experience evaluating LLM prompts/classifiers with an eval framework such as DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, or a homegrown harness
  • Ability to turn small seed sets into robust labeled eval datasets (including dataset-synthesis tooling and prompt-optimization loops)
  • Familiarity with prompt-injection/jailbreak defense: OWASP LLM Top 10, input/output filtering, least-privilege tool access, and adversarial testing

Bonus (helpful background)

  • Hands-on with cloud guardrail services such as Google Cloud Model Armor, AWS Bedrock Guardrails, or Azure AI Content Safety
  • Terraform / IaC for cloud infrastructure
  • Experience with Google Cloud Vertex AI / Gemini
  • Monitoring and observability experience (for example, New Relic or similar)
  • Exposure to regulated or compliance-sensitive domains (fair housing, fair lending, healthcare, finance, trust & safety)
  • Red-teaming or AI-security background

Technologies you may work with

Python, DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, Google Cloud Model Armor, Terraform, Google Cloud Vertex AI / Gemini, New Relic, OWASP LLM Top 10, CI, Google Cloud, AWS Bedrock Guardrails, Azure AI Content Safety.

Benefits

  • Inclusive and competitive medical, Rx, dental, and vision coverage
  • Family forming benefits
  • 13 Paid Holidays
  • Flexible Time Off
  • 8 hours of paid Volunteer Time off
  • Immediate eligibility into Company 401(k) with 3.5% company match
  • Tuition Reimbursement program for degreed and non-degreed programs
  • 1:1 personalized Financial Planning Sessions
  • Student Debt Retirement Savings Match program
  • Free snacks and refreshments in each office location

Expectations: Independently design, implement, and tune guardrail and evaluation systems; balance safety, latency, and user experience; keep eval datasets calibrated to real traffic and emerging attack patterns; stay current with AI safety research and jailbreak techniques; write clean, maintainable, well-tested code; deliver high-quality work consistently on schedule.

Similar Jobs