OpenAI is looking to design and build an evals infrastructure that measures the quality of OpenAI’s support automation. This role will help to design and build robust systems and backend services that serve as the foundation for how knowledge is created, accessed, and applied across OpenAI.
Requirements
- Proficiency in backend technologies. Our tech stack includes Python, FastAPI, and Postgres
- Experience designing and scaling distributed systems, APIs, or data processing pipelines
- Have experience building AI agents or applications, including designing evals and improving performance through prompting or scaffolding
- Are familiar with evaluation methods for LLMs and have worked with patterns like multi-agent workflows, tool use, or long context.
- Experience creating production evals and/or measuring performance of ML/LLM models at scale
Responsibilities
- Design eval pipelines that are reliable, reproducible, and extendable
- Build the infrastructure for continuous eval monitoring frameworks (regression/drift monitoring, building robust golden datasets) along with feedback loops that ultimately strengthen support automation
- Design, build, and maintain backend services and APIs to support intelligent automation and knowledge systems
- Integrate and structure data across internal platforms, transforming it into formats optimized for use by downstream systems and AI workflows.
- Collaborate closely with data, research, and engineering teams to integrate OpenAI models into high-leverage workflows
- Own the full development lifecycle of new backend systems and internal platform capabilities
- Build with scale and maintainability in mind, while rapidly iterating on new ideas
Other
- 4+ years of backend engineering experience at product-driven companies (excluding internships)
- A pragmatic mindset. You’re comfortable shipping iteratively while building toward a long-term vision
- We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
- Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act.