EngineerJobs.io
← Back to all jobs

Job Description

Arena Intelligence, Inc. is building the infrastructure that powers online, real-world model evaluation, and the Platform team is central to that mission. In this hands-on role, you will help deliver low-latency, reliable services that route evaluation traffic, stream results, and provide the enterprise-grade controls and observability customers will expect as Arena scales.

Software Engineer, Platform will contribute across the core layers beneath Arena’s online evaluation systems, including the AI gateway, automated arena runtimes, and serving components that enable reliable and scalable evaluation at production load.

What you’ll do

  • Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
  • Own SSE and streaming response handling across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
  • Build infrastructure features enterprise customers will need as usage grows, including rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
  • Instrument the platform with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards to make system behavior visible to both customers and the team.
  • Integrate with the core evaluation platform, Arena data, and customer-specific benchmarks, partnering with the research team to ship novel ideas as full-featured products.
  • Contribute to the backend of the Leaderboards and Evals platforms as needed, helping unify public and private data architectures.

What you bring

  • ~5+ years of backend engineering experience, including meaningful work on distributed systems, infrastructure, or developer-facing platforms.
  • Strong Go proficiency (primary backend language for the role).
  • Experience working with LLM provider APIs (OpenAI, Anthropic, Google, etc.), with a practical understanding of streaming, token management, rate limits, and model-specific quirks.
  • A product-oriented mindset, focused on developer experience and asking “why” before “how.”
  • Comfort with ambiguity in a startup environment, where scope can shift and priorities change.

Technologies

  • Go, OpenAI, Anthropic, Google, SSE
  • Distributed tracing, Postgres, Redis, AWS, GCP, Azure
  • Kubernetes, Terraform, SOC 2, SSO, RBAC
  • vLLM, LiteLLM, LangChain, Stripe, Metronome, Orb

Nice to have

  • Cloud and infrastructure experience (AWS, GCP, or Azure), Kubernetes, Terraform, and databases such as Postgres and Redis.
  • Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
  • Background in AI/ML infrastructure, model serving, inference, or evaluation frameworks.
  • Experience delivering enterprise-ready features such as SSO, RBAC, audit logs, and multi-tenancy.
  • Experience building billing infrastructure around Stripe, Metronome, and Orb.
  • Familiarity with modern AI infrastructure tooling such as vLLM, LiteLLM, and LangChain.

Benefits

  • Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
  • Competitive compensation and equity aligned to the markets where team members are based.
  • Work on cutting-edge AI with a small, mission-driven team.
  • A culture centered on transparency, trust, and community impact.

Location and working model

  • Based in the San Francisco Bay Area, CA with hybrid expectations.
  • Minimum 3 days/week onsite is required.
  • Fully remote candidates will only be considered with a very strong endorsement.

About the role

  • Build the core infrastructure beneath Arena’s online evaluation systems, including AI gateways, automated arena runtimes, and serving layers for scaled real-world evaluation.
  • Help support arenas that route traffic across frontier models from many providers, manage bursty and unpredictable load, fail gracefully when upstream models degrade, and keep evaluation behavior fair and consistent.
  • Work with an AI gateway currently in private, gated launch, with individual developers as primary users today, and enterprise use cases planned for later.
  • Support near-term shipping priorities, including direct model access and new infrastructure features, plus usage trace collection and leaderboards shared with lab partners.
  • This is an individual contributor position. Arena is not hiring specifically for a tech lead or SRE function right now, and the team is focused on hands-on building.

Similar Jobs