Senior Software Engineer Autotagging positions focus on delivering results in their domain. This page aggregates open Senior Software Engineer Autotagging roles and what employers typically expect.
About the Company At Torc, we have always believed that autonomous vehicle technology will transform how we travel, move freight, and do business. A leader in autonomous driving since 2007, Torc has spent over a decade commercializing our solutions with experienced partners. Now a part of the Daimler family , we are focused solely on developing software for automated trucks to transform how the world moves freight. Join us and catapult your career with the company that helped pioneer autonomous technology, and the first AV software company with the vision to partner directly with a truck manufacturer. Meet the Team The Auto Tagger team is the engine behind our data flywheel, responsible for translating petabytes of raw, multi-modal vehicle data into a highly curated library of critical driving scenarios. By mining driving logs for long-tail events, we provide the foundational data required for safe autonomous trucking. Leveraging Pegasus logical layers, this team structures and catalogs findings into an observations database that directly accelerates development across autonomous perception, sensor fusion, and generative simulation testing. What You'll Do Integrate and deploy automated event-tagger into production pipelines, running and monitoring tagging tasks at scale across petabytes of vehicle log data. Build and maintain the data engineering pipelines that organize, structure, and catalog tagged scenario data into the observations database. Own CI/CD for the Auto Tagger pipeline using GitHub Actions, keeping deployments reliable, tested, and repeatable. Write production grade code in Python across the pipeline, from data ingestion and transformation through model integration and deployment. Build and operate on Databricks for large scale data processing, interactive querying, and pipeline orchestration. Design, deploy, and scale AWS infrastructure (as code) to support high-volume, distributed processing of vehicle log pipelines — working with structured/tagged outputs and metadata. Instrument pipelines with logging, metrics, and alerting; own on-call response for tagging job failures and data quality regressions. Partner with ML engineers on the team to take tagging and classification models from development into a scalable, monitored production pipeline. Ensure data quality and metadata integrity as tagged events move from raw logs into the observations database used by perception, simulation, and systems teams. Troubleshoot and improve pipeline performance, reliability, and cost as data volume and model complexity grow. What You'll Need to Succeed BS or MS in Computer Science, Engineering, or a related field, with 5+ years of software engineering experience, including production data pipeline or ML infrastructure work. Strong Python skills, with experience building and maintaining production data or ML pipelines. Hands-on CI/CD experience, GitHub Actions required. Required experience with Databricks for large scale data processing and orchestration. Required experience with AWS, including infrastructure-as-code (Terraform or CloudFormation) for provisioning distributed processing infrastructure. Experience processing large scale time series or unstructured datasets. Experience with observability tooling (e.g., Datadog, Grafana, CloudWatch) for production pipeline monitoring and alerting. Experience integrating and deploying ML models into production systems — serving, monitoring, and rollback, not just training. Strong communication skills to work across ML, perception, and simulation teams. Bonus Points! Familiarity with auto-labeling pipelines, VLMs, or zero-shot classification for scenario extraction. Experience with distributed compute frameworks such as Ray, Spark, or Daft. Familiarity with robotics data formats (ROS bags, MCAP) and columnar storage formats (Parquet, Arrow). Experience with model serving frameworks such as vLLM or SGLang. Familiarity with scenario description standards like Pegasus layers. Perks o…