Staff Infrastructure Engineer positions focus on delivering results in their domain. This page aggregates open Staff Infrastructure Engineer roles and what employers typically expect.
Kikoff: The Fintech Powering Financial Security at Scale Kikoff is a profitable, pre-IPO fintech company on a mission to empower everyone to achieve financial security. With record revenue growth in 2025 and a unicorn valuation, we've built a suite of products that help millions of people build credit, access liquidity, and save money. We're scaling fast. Join us if you want to build something meaningful and help millions of people move forward financially. Why Kikoff: This is a consumer fintech startup, and you will be working with serial entrepreneurs who have built strong consumer brands and innovative products. We value extreme ownership, clear communication, a strong sense of craftsmanship, and the desire to create lasting work and work relationships. Yes, you can build an exciting business AND have real-life real-customer impact. About the Role Kikoff's infrastructure team builds the systems that enable engineering teams to move quickly without sacrificing reliability, security, or cost discipline. The team owns five connected areas: Observability, Developer Productivity, Compute Infrastructure, Networking and Storage, and Data Infrastructure. As a Staff Infrastructure Engineer, you will own high-leverage infrastructure problems from design through production operation and help run infrastructure as a product for Kikoff engineers. You will build internal products and paved paths that turn ambiguous problems into durable systems. Your work should improve delivery speed and reliability, reduce cost, and reduce the amount of operational work product teams carry. This is a hands-on Staff IC role. You will foster relationships, write code, review designs, make architecture decisions, lead through incidents, and help engineers solve problems outside the runbook. You will have a primary area of depth and enough range to follow production problems across infrastructure boundaries. AI is increasing the amount of code and change moving through our systems. You will help build a platform that can absorb that growth without creating more incidents, unsustainable cost, or a team that depends on heroics. In This Role, You Will Build and Own Critical Platform Systems Design and implement self-service infrastructure on AWS using reusable code and infrastructure-as-code patterns. We use Pulumi, HCL, and TypeScript heavily. The company runs on Ruby. Own the systems you build in production, including reliability, security, capacity, cost, upgrades, incidents, and recovery. Give critical services an SLO, actionable alerts, a useful dashboard, a runbook, and a tested recovery path. Automate recurring operational work and eliminate failure-prone manual steps rather than allowing them to become permanent processes. Run Infrastructure as a Product Work directly with engineers to turn recurring friction into paved paths, self-service tools, and automated workflows that are faster and safer than one-off solutions. Measure outcomes for internal customers through adoption, developer feedback, delivery speed, reliability, cost, toil, and on-call load. Use the fastest responsible path when a team is blocked, then turn recurring friction into automation or a durable platform capability. Set Technical Direction Set technical direction for ambiguous infrastructure work, then carry it from problem framing and design through implementation, rollout, and production ownership. Make clear trade-offs among delivery speed, reliability, security, cost, and long-term operational complexity. Partner across Infrastructure, Security, Data, and Product Engineering to build security and compliance controls into normal engineering workflows. Raise the Engineering Bar Review code and designs, challenge weak assumptions, help other engineers make better technical decisions, and uphold high standards for reliability, testing, safe deployments, security, and maintainability. Lead through incidents and unfamiliar failure modes, and ensure the system is better after servic…