Jobedly Post a Job

Senior AI Engineer (Evaluations - Canvas Agent)

instructure · Remote
RemoteFull-timeResearch & DevelopmentEducation$136,000–$184,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior AI Engineer role

Senior AI Engineer positions focus on delivering results in their domain. This page aggregates open Senior AI Engineer roles and what employers typically expect.

At Instructure, we believe in the power of people to grow and succeed throughout their lives. Our goal is to amplify that power by creating intuitive products that simplify learning and personal development, facilitate meaningful relationships, and inspire people to go further in their education and careers. We do this by giving smart, creative, passionate people opportunities to create awesome. And that's where you come in: We are looking for an AI Engineer to join the team behind the Canvas Agent. As we build out agentic workflows that can "act" upon Canvas and assist educators and students, ensuring the quality, safety, and accuracy of these interactions is our top priority. In this role, you will design and implement robust LLM-as-a-judge evaluation pipelines that automatically assess the Canvas Agent’s multi-step reasoning, tool usage, and conversational helpfulness. What You’ll Do - Design the Evaluation Framework: Build and maintain scalable LLM-as-a-judge pipelines to automatically score the Canvas Agent’s actions, responses, and tool usage across a variety of complex educational workflows. - Develop Rubrics & Datasets: Create comprehensive grading rubrics and curate high-quality "golden" datasets (both real and synthetically generated) to baseline and test the agent's performance. - Optimize Judge Prompts: Engineer and iterate on prompts for the judges, ensuring automated scoring aligns with high quality evaluations. - Full-Stack Contribution: Step beyond evaluation pipelines to participate in full-stack product development as needed, collaborating with the team to build and refine the core AI features, UI components, and application architecture. - Accelerate Iteration: Integrate your automated evaluations directly into our CI/CD pipelines, creating a "paved path" that allows our AI product teams to ship updates with high velocity and total confidence. - Analyze & Report: Monitor evaluation metrics to identify failure modes, hallucination rates, and regressions. Translate these subjective quality signals into objective, actionable engineering tasks. What You’ll Need - LLM Evaluation Experience: A background that clearly demonstrates hands-on experience designing and writing automated tests for LLMs. You must have a proven track record of writing and implementing LLM-as-a-judge evaluations in real-world scenarios. - Technical Stack: A strong background in Python is highly preferred, or a demonstrated willingness and ability to learn it quickly. You should also have professional experience in full-stack or backend engineering to support both the evaluation infrastructure and general product development needs (experience with TypeScript/Node.js is a plus). - Prompt Engineering Expertise: Deep understanding of how to reliably prompt models for classification, extraction, and grading tasks without falling prey to common biases (e.g., position bias, verbosity bias). - Evaluation Tooling: Familiarity with modern LLM observability and evaluation frameworks (e.g., LangSmith, Braintrust, Ragas, promptfoo, or similar). - Agentic Architectures: Strong conceptual understanding of how AI agents plan and execute tool calls (e.g., ReAct, tool-use APIs) so you can effectively evaluate multi-step workflows. - Analytical Mindset: An ability to translate highly subjective concepts (like "helpfulness" or "tone") into rigorous, trackable metrics. - Collaboration Skills: Ability to work across product and engineering teams to understand the Canvas Agent's use cases and align evaluation criteria with user needs. Onsite Collaboration Requirement: This role requires working onsite on Tuesday and Wednesday, with Thursday strongly encouraged as part of our company’s in-person collaboration model. Why Join Us? In this role, you aren't just testing a feature; you are the guardian of quality for how AI interacts with the world of education. Your work will be the catalyst that allows Instructure to move faster than ever, safely turning static tool…

Salary estimate

$136,000 – $184,000/yr
Provided by the employer.

Skills for this role

TypescriptReactNODEPythonGOCi/CdLLM

Resume tips for Senior AI Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior AI Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior AI Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About instructure

Instructure is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles