Software Engineer Production Management positions focus on delivering results in their domain. This page aggregates open Software Engineer Production Management roles and what employers typically expect.
Who We Are: Galaxy Digital Inc. (Nasdaq: GLXY) is a global leader in digital assets and data center infrastructure, growing the economy that runs on code. Galaxy delivers the onchain infrastructure that connects institutions to digital assets, including trading, advisory, asset management, staking, self-custody, and tokenization. Galaxy also develops and operates data center infrastructure to power AI and HPC workloads. Anchored by its Helios campus in Texas, Galaxy is building a multi-gigawatt pipeline of more than 5.7 GW of potential capacity, positioning it among the largest and fastest-growing data center developers in North America. The Company is headquartered in New York City, with offices across North America, Europe, the Middle East, and Asia. Additional information about Galaxy's businesses and products is available on www.galaxy.com . What We Value: We are a diverse team of free thinkers, and fast movers united to help investors and creators energize the global economy. We are looking for individuals who thrive in a culture of builders and overachievers and embrace high performance, transparent feedback, and a mission-first approach. Our culture shapes our way of working and gets us where we want to be. Seek Excellence. Be Selective To Be Effective. Be Highly Aligned, Loosely Coupled. Disagree Transparently. Encourage Independent Decision-Making. Build Dream Teams. Who You Are: You're a software engineer who gets energized by building tools that solve real operational problems. You're comfortable in a codebase, take pride in what you ship, and don't mind being close to production. You'll spend the majority of your time building and improving the internal tools, workflows, and automation that make our production support operation run more effectively, with the rest of your time working alongside teammates on monitoring, incident response, and cross-team coordination. What You’ll Do: Work on automation projects focused on reliability improvements, toil reduction, and incident response, building tools and scripts that make the team more effective Build and maintain internal tooling in Python, including scripts, bots, and lightweight applications that support monitoring, alerting, and related workflows Improve observability across systems; build dashboards, refine alerting thresholds, and surface signals that help the team stay ahead of issues Automate operational tasks across our Kubernetes/EKS environment, including deployments, health checks, alerting, and runbook automation Integrate with internal platforms (FireHydrant, DataDog, Jira, Confluence) to streamline incident workflows and reduce context-switching for the team Debug and troubleshoot complex applications written in Java and Python Collaborate with development teams to implement fixes and enhancements to improve system stability and performance Monitor system health and performance; identify and resolve issues promptly to ensure uninterrupted operations Triage and respond to production incidents across our cryptocurrency business lines Investigate issues using SQL and log analysis; escalate and coordinate with development teams as needed Coordinate with global counterparts to ensure seamless follow-the-sun support; some on-call is expected Document and maintain support procedures, runbooks, and knowledge base articles Embrace and champion the thoughtful adoption of AI to improve team performance and business outcomes Leverage AI tools (e.g., generative AI, automation platforms, data copilots) to improve productivity, decision-making, and output quality in your day-to-day work. What We’re Looking For: Proven software development experience; you have written production-quality code that others depend on Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field 5+ years of experience in software engineering or production support roles, preferably in finance, with direct exposure to trading system internals, post-trade proc…