Researcher Evaluation positions focus on delivering results in their domain. This page aggregates open Researcher Evaluation roles and what employers typically expect.
Epoch AI is looking for a researcher to evaluate frontier AI models on hard-to-grade tasks drawn from real-world scenarios. ## About the role We’re seeking a Researcher to lead a new effort evaluating how well frontier models perform on the kinds of open-ended tasks that make up real office work. You will curate a suite of realistic tasks to serve as a benchmark, design the grading rubrics for AI performance, and run newly-released models through the suite, assessing their performance both quantitatively and qualitatively. The focus is on how models handle messy, real-world work rather than on scientific knowledge or programming ability. The role makes heavy use of AI tools, but strong software engineering experience is not required. Comfort setting up AI-assisted automated workflows is a plus. If this role sounds interesting, we are also looking for [researchers on multiple other teams](https://docs.google.com/document/d/1cNXuKtMnT5olDofyuUOAKOERME4tdpQ2-RGQhas29UI/edit?tab=t.0). **Applications are rolling**. ### Key Responsibilities - **Create and curate an evaluation suite.** Find real-world tasks that serve as challenging tests for practical AI capabilities, and update the tasks over time as AI capabilities evolve. Devise rubrics for evaluating AI performance. - **Evaluate AI systems.** Regularly evaluate new, notable AI models and products on the task suite. Update tasks and rubrics to reflect the changing landscape of AI capabilities. - **Communicate your research**. Create public-facing reports, blog posts, and data visualizations with your observations. Ensure the evaluations feed into our other research topics and help keep our team informed. - **Conduct data analysis**. Analyze evaluation results and compare models across tasks. - **Improve the process**. You might automate parts of the workflow, and build out parts of the evaluation into standalone benchmarks. ### What we are looking for - **Analytical thinking.** You conduct experiments with rigor and care, making sure that findings are well-supported by evidence. - **Grounded, skeptical mentality**. You form your own well-reasoned view of what an AI system can do, distinguishing practical capabilities from hype. - **Comfort with AI agents and tools.** You have experience working with AI agents in the course of your own work, and are comfortable delegating tasks. - **Familiarity with AI benchmarks and evaluations**. You follow AI capabilities at least casually and have opinions on what benchmarks do and don’t tell us. - **Research and data-analysis** **experience,** including enough comfort with light coding to analyze your own results. - **Strong written communication skills:** You can convey nuanced observations clearly and precisely. ### Nice to have - Experience testing frontier models and writing assessments of their capabilities - Coding skills, including python proficiency If you don’t tick all these boxes but think you would be a great fit, please consider applying anyway! ### Compensation & Benefits - Annual salary between **$115,000 – $200,000 USD**, depending on location and experience. - Salaries are not restricted to USD, and contracts and payments are usually in local currencies. Conversions are based on one-year average exchange rates. - Fully remote environment, including flexible work hours. - Competitive global benefits program, including a comprehensive health insurance program—including supplemental benefits specific to a local country, as available and mandated by local law—and life insurance and a pension plan, if applicable in your country. - Generous paid time off (PTO), including no specific annual limit, with 30 days PTO per year protected, unlimited personal and sick leave, and 4 months paid parental leave for permanent staff with at least 12 months of tenure (prorated parental leave if less than 12 months). - A flexible and generous expense policy for you to spend on equipment and a large range of productivity tools or learning/developm…