AI Engineer
Job Description
Join Realtor.com’s AI Integrations Team in Austin, TX (onsite) and help build the AI safety layer behind RealAssist, a consumer-facing LLM real-estate assistant. This role focuses on guardrails, evaluation pipelines, and classifiers that support compliance and content moderation while keeping the experience safe, accurate, and reliable for users.
What you’ll own
- AI safety layer end-to-end, from design through deployment
- Fair-Housing compliance and content moderation classifiers, tuned to target high recall on disallowed content without overblocking legitimate use
- LLM-as-a-judge evaluation for safety and compliance decisions, using category-level performance metrics
- Guardrail evaluation pipelines that include labeled dataset curation, train/test/validate splits, and offline prod-replay to identify overblocking
- Runtime prompt-injection and jailbreak screening using cloud guardrail services with fail-open behavior and alerting
- Safety regression checks in CI and red-teaming to reduce the risk of guardrails degrading as the product evolves
- Guardrail infrastructure as code with Terraform (templates, IAM, and project shape), plus participation in technical design reviews
- Risk tiering and launch support through collaboration with ML, backend, product, and legal/compliance stakeholders
How you’ll make decisions
- Use confusion matrices and precision/recall/F1 per category to make safety decisions defensible
- Ground tuning and safety tradeoffs in data-driven eval metrics
- Keep evaluation datasets calibrated to real traffic and emerging attack patterns
What you’ll bring
- AI Safety & Evaluation Expertise (Required)
- 4+ years of professional software/ML engineering experience, including hands-on LLM application work
- Bachelor’s degree or equivalent experience
- Strong Python proficiency
- Classification-metrics literacy: precision/recall/F1 (macro vs weighted), confusion matrix debugging, dataset curation, and calibration to production distribution
- Experience evaluating LLM prompts/classifiers with an eval framework such as DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, or a homegrown harness
- Ability to turn small seed sets into robust labeled eval datasets (including dataset-synthesis tooling and prompt-optimization loops)
- Familiarity with prompt-injection/jailbreak defense: OWASP LLM Top 10, input/output filtering, least-privilege tool access, and adversarial testing
Bonus (helpful background)
- Hands-on with cloud guardrail services such as Google Cloud Model Armor, AWS Bedrock Guardrails, or Azure AI Content Safety
- Terraform / IaC for cloud infrastructure
- Experience with Google Cloud Vertex AI / Gemini
- Monitoring and observability experience (for example, New Relic or similar)
- Exposure to regulated or compliance-sensitive domains (fair housing, fair lending, healthcare, finance, trust & safety)
- Red-teaming or AI-security background
Technologies you may work with
Python, DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, Google Cloud Model Armor, Terraform, Google Cloud Vertex AI / Gemini, New Relic, OWASP LLM Top 10, CI, Google Cloud, AWS Bedrock Guardrails, Azure AI Content Safety.
Benefits
- Inclusive and competitive medical, Rx, dental, and vision coverage
- Family forming benefits
- 13 Paid Holidays
- Flexible Time Off
- 8 hours of paid Volunteer Time off
- Immediate eligibility into Company 401(k) with 3.5% company match
- Tuition Reimbursement program for degreed and non-degreed programs
- 1:1 personalized Financial Planning Sessions
- Student Debt Retirement Savings Match program
- Free snacks and refreshments in each office location
Expectations: Independently design, implement, and tune guardrail and evaluation systems; balance safety, latency, and user experience; keep eval datasets calibrated to real traffic and emerging attack patterns; stay current with AI safety research and jailbreak techniques; write clean, maintainable, well-tested code; deliver high-quality work consistently on schedule.