Jobedly Post a Job

QA Engineer - Gen AI

Sustainment · Austin, TX
Full-timeTechnologyTechnology$136,000–$184,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the QA Engineer Gen AI role

QA Engineer Gen AI positions focus on delivering results in their domain. This page aggregates open QA Engineer Gen AI roles and what employers typically expect.

Company Overview: Sustainment is an AI-native software platform that helps US-based manufacturers easily find and work with the critical suppliers they need to build and manage their supply chains. Our vision is to reimagine American manufacturing as a hyperconnected, secure, and resilient ecosystem of local and regional suppliers who can more easily connect, interact, and do business with the industry and government customers that rely on them. We are a dual-use technology platform that supports both DoD and commercial customers in pursuit of our vision. This is a contract opportunity Job Overview: We are seeking a QA Engineer to help ensure the reliability, accuracy, and robustness of our AI Agents. This role will focus on data quality, model evaluation, and regression testing frameworks to identify and mitigate common LLM failure modes. You will be responsible for designing automated and scalable quality assurance systems while working in an AWS-based infrastructure. If you have a strong background in LLM testing, data validation, and automated QA frameworks, this role is an excellent opportunity to contribute to cutting-edge AI systems. Responsibilities: Design and run regression test suites for LLM evaluation. Identify and track LLM failure modes, including hallucinations, biases, factual inconsistencies, and logical errors. Design data-quality checks to assess training and test datasets. Automate LLM performance monitoring using advanced metrics and validation strategies. Apply best practices for prompt-engineering testing, fine-tuning validation, and output-consistency analysis. Collaborate with ML engineers, data scientists, and product teams to align on quality benchmarks. Work within an AWS ecosystem, leveraging services such as EKS, S3, SageMaker, or Databricks for model testing and evaluation. Build tools and dashboards to track LLM quality over time. Curate and version the ground-truth datasets that serve as the accuracy baseline for document parsing, and translate business and domain requirements into written, testable field definitions (partnering with the labeling team on annotation guidelines). Evaluate structured extraction from real business documents (multi-page PDFs, scans, spreadsheets) by scoring model output field-by-field against ground truth, with tolerance-aware comparison for numbers, dates, free text, and repeated structures. Maintain the ground-truth corpus as a versioned, evolving test asset: keep existing annotations valid as extraction schemas change, preserve dataset provenance, and grow the corpus from real production failures so every customer-reported miss becomes a permanent regression case. Calibrate and validate automated scoring itself; confirm that semantic/LLM-judge scoring agrees with human judgment. Qualifications: 3+ years in software testing and quality assurance 2+ years with a focus on ML evaluation, NLP, LLMs, VLMs, etc. Deep understanding of LLM data quality challenges and common failure modes. Experience designing automated tests for AI/ML models. Familiarity with Python and testing frameworks such as PyTest, Hypothesis, or similar. Knowledge of evaluation metrics for LLMs (DeepEval, MLflow, LangSmith, or similar). Hands-on experience with automated data validation techniques. Strong debugging and analytical skills. Experience creating or working with labeled evaluation datasets (“golden” sets) for model evaluation. Working knowledge of evaluation metrics for structured information extraction: field-level precision, recall, and F1; exact vs. fuzzy matching; numeric tolerance; and alignment of repeated or nested records. Experience translating ambiguous business requirements into precise, documented field definitions in collaboration with non-technical subject-matter experts. Preferred Qualifications SQL proficiency, including seeding test data across Postgres environments (local/dev/staging/prod). Comfort with observability and incident-response tooling (e.g., Datadog monito…

Salary estimate

$136,000 – $184,000/yr
Provided by the employer.

Skills for this role

PythonSQLAWSNLPLLMQA

Resume tips for QA Engineer Gen AI applicants

Interview preparation

Prepare concrete STAR-format stories that show QA Engineer Gen AI outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical QA Engineer Gen AI problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Sustainment

Sustainment is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles