Staff Software Development Test Engineer positions focus on delivering results in their domain. This page aggregates open Staff Software Development Test Engineer roles and what employers typically expect.
About Tekion: Positively disrupting an industry that has not seen any innovation in over 50 years, Tekion has challenged the paradigm with the first and fastest cloud-native automotive platform that includes the revolutionary Automotive Retail Cloud (ARC) for retailers, Automotive Enterprise Cloud (AEC) for manufacturers and other large automotive enterprises and Automotive Partner Cloud (APC) for technology and industry partners. Tekion connects the entire spectrum of the automotive retail ecosystem through one seamless platform. The transformative platform uses cutting-edge technology, big data, machine learning, and AI to seamlessly bring together OEMs, retailers/dealers and consumers. With its highly configurable integration and greater customer engagement capabilities, Tekion is enabling the best automotive retail experiences ever. Tekion employs close to 3,000 people across North America, Asia and Europe. Roles & Responsibilities Generative AI & LLM Evaluation Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities across LLM-powered products Implement automated evaluations for RAG systems measuring context relevance, groundedness, and answer faithfulness using frameworks like RAGAS or DeepEval Design test beds to validate multi-agent workflows — tool-calling accuracy, multi-step reasoning, memory, and autonomous decision loops Build and run automated conversation simulations — scripted and synthetic user journeys — to stress-test agent behaviour across intents, edge cases, and multi-turn dialog flows Create prompt regression frameworks to assess how changes in system prompts, temperature, and sampling parameters impact output consistency Data Quality Assurance Statistically validate AI data outputs — distributions, precision/recall, error pattern analysis — to catch silent data quality failures before production Programmatically audit data ingestion, transformation, and feature store pipelines for schema drift and data corruption Validate vector DB indexing, embedding semantic similarity accuracy, and retrieval latency Verify quality, diversity, and privacy compliance of synthetic datasets used for model training and evaluation Classical ML & Deep Learning Validation Maintain automated suites tracking ML metrics — Precision, Recall, F1, ROC-AUC — and deep learning loss curves across model versions Implement continuous monitoring scripts to detect data and concept drift on live inference endpoints Automation Engineering & CI/CD Build and maintain scalable test automation frameworks for APIs, backend services, and model endpoints Embed AI evaluation and data QA suites into MLOps and CI/CD pipelines so quality failures block releases automatically Define and track AI quality KPIs and communicate release readiness to engineering and product teams Experience of 8+ years SDET role Technical Skills & Frameworks Core Programming Python — expert level; test automation, eval pipelines, data analysis (Pandas, NumPy, Pytest) SQL — data output validation, ground truth querying, pipeline data quality checks GenAI & Evaluation RAGAS / TruLens / DeepEval / Promptflow etc — LLM evaluation frameworks for measuring faithfulness, hallucination rate, and task success LangChain / LangSmith / LlamaIndex — agent workflow testing, prompt tracing, and LLM response debugging OpenAI / Anthropic / Hugging Face APIs — direct LLM endpoint testing and output consistency validation Vector DBs — retrieval quality testing, embedding validation, and latency benchmarking Pandas / NumPy etc. — statistical analysis for output validation and error pattern investigation, data profiling, schema validation, and pipeline integrity checks API & Automation Pytest — modular, reusable test framework for AI eval and automation suites Postman / REST Assured / Requests — API contract validation and service-level integration testing MLOps & CI/CD MLflow — tracking model versions and eval runs to detect regressions ac…