AI Infrastructure Engineer Agent ML System positions focus on delivering results in their domain. This page aggregates open AI Infrastructure Engineer Agent ML System roles and what employers typically expect.
ABOUT US: Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk. Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at Havoc: All-Domain Collaborative Autonomy http://havocai.com/ . ABOUT THE ROLE As an AI Infrastructure Engineer – Agents & ML Systems, you will help build the internal AI infrastructure that allows HavocAI teams to use modern AI systems safely, reliably, and effectively. You will develop tools, services, pipelines, and integrations that connect large language models, agentic workflows, internal data sources, engineering systems, and ML workflows. This role is ideal for a strong software or infrastructure engineer who is excited about the practical application of AI. You need not have worked on every part of the AI stack, but you should be curious, hands-on, and comfortable building production systems that connect models, tools, data, and users. You will work on systems that help internal teams search and reason over company data, automate engineering workflows, support simulation and autonomy development, curate data for future model training, and evaluate AI systems before they are trusted in critical workflows. This is a high-impact role at the intersection of software engineering, AI infrastructure, developer tooling, data systems, and applied ML. JOB RESPONSIBILITIES - Build internal AI infrastructure that connects LLMs and AI agents with internal tools, APIs, data sources, data lakes, telemetry stores, simulation tools, code repositories, documentation systems, logs, and engineering workflows. - Develop and maintain agentic AI systems for task automation, data analysis, engineering support, simulation workflows, and internal productivity. - Build tool integration and connector infrastructure for AI agents, including MCP and other emerging tool-use standards, spanning servers, tools, resources, prompts, connectors, and secure tool-use patterns. - Create pipelines for retrieval, RAG, context management, document processing, embeddings, and internal knowledge search. - Support ML infrastructure workflows such as data preparation, dataset curation, experiment tracking, model evaluation, fine-tuning support, and model deployment. - Build evaluation frameworks for agent performance, tool-use reliability, task success, model quality, regression testing, and failure analysis. - Develop observability, logging, tracing, auditability, monitoring, and debugging tools for AI agents, model calls, MCP tools, and ML pipelines. - Partner with Autonomy, Software, Data, Simulation, Product, and Operations teams to identify high-value AI use cases and turn them into reliable internal tools. - Secure agentic AI systems end-to-end with least-privilege tool access, sandboxed tool execution, prompt-injection and misuse mitigation, secrets management, human-in-the-loop approvals, and safe handling of sensitive and defense data. - Maintain documentation, reusable examples, templates, and best practices that help internal teams adopt AI tools safely and effectively. QUALIFICATIONS - Bachelor's degree in Computer Science, Engineering, Machine Learning, Data Science, Applied Mathematics, or a related technical field. - 3+ years of experience in software engineering, infrastructure engineering, ML infrastructure, backend systems, data engineering, developer tools, or related technical roles. - Strong programming experience in Python, TypeScript, Go, C++, or similar languages. - Experience building production software systems, APIs, services, data pipelines, or internal platforms. - Experience with, o…