Jobedly Post a Job

Sys/Cloud Admin/Incident Response Engineer

i4DM · Remote
RemoteFull-timeTechnologyMid Level$102,000–$138,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Sys Cloud Administrator Incident Response Engineer role

Sys Cloud Administrator Incident Response Engineer positions focus on delivering results in their domain. This page aggregates open Sys Cloud Administrator Incident Response Engineer roles and what employers typically expect.

**About Our Team** Our employees thrive in a culture that is fast-paced, collaborative, and ego-free, where innovation and teamwork are encouraged at every level. We provide Federal agencies with immediate access to highly skilled professionals who understand complex mission challenges and deliver efficient, scalable solutions. By continuously investing in talent, technology, and specialized capabilities, we maintain expert teams prepared to support evolving Federal missions through tailored technical solutions and modern service delivery approaches. We value diverse perspectives and strive to attract talent from all backgrounds. We are seeking professionals who are passionate about technology, mission success, and solving complex operational challenges with creativity and purpose. If you enjoy expanding your technical expertise while supporting impactful Federal initiatives, you will thrive within our organization. Veterans and military spouses are strongly encouraged to apply and bring their valuable experience to our team. **About the Role** We are seeking an experienced and highly motivated Sys/Cloud Admin/Incident Response Engineer to support enterprise monitoring operations, incident detection, response activities, and operational situational awareness for a mission-critical platform within the Department of Veterans Affairs (VA) environment. In this role, you will provide hands-on administration and operational support to help ensure monitoring and incident management processes effectively sustain system reliability, operational continuity, and rapid restoration of services across a large-scale, 24x7 enterprise healthcare platform. You will work closely with the Monitoring & Incident Management Manager, Program Manager, Technical Directors, DevSecOps & SRE teams, and VA stakeholders to identify, escalate, communicate, and help resolve incidents in alignment with strict service-level expectations and operational standards. **RESPONSIBILITIES** **Monitoring, Administration & Operational Support** - Administer, monitor, and support cloud and platform services, virtual infrastructure, and hosted applications to maintain system health, availability, and performance. - Configure, tune, and maintain monitoring, logging, and alerting solutions to improve visibility across infrastructure, applications, and service dependencies. - Validate alert accuracy, reduce noise, and help ensure operational issues are detected proactively through effective observability practices. - Perform routine system administration tasks such as environment checks, service restarts, access support, patch coordination, and operational maintenance activities. **Incident Response & Service Restoration** - Monitor incident queues and system alerts, perform initial triage, document impact, and execute defined escalation procedures for incidents affecting mission-critical services. - Participate in major incident response activities, including troubleshooting, log review, coordination with engineering teams, and support for service restoration efforts. - Follow incident response playbooks, severity models, and communication protocols to support timely resolution and accurate status reporting. - Document incident timelines, actions taken, recovery steps, and supporting evidence to enable post-incident review and continuous improvement. **Operational Coordination & Stakeholder Support** - Support coordination during operational events by working across infrastructure, application, DevSecOps, SRE, and service management teams. - Provide clear, timely updates on incident status, service impact, troubleshooting progress, and recovery actions to internal stakeholders. - Escalate issues appropriately based on impact, urgency, and established operational procedures. - Maintain accurate operational records in ticketing, incident, and knowledge management systems. **Observability, Automation & Continuous Improvement** - Partner with engineers and platform teams to im…

Salary estimate

$102,000 – $138,000/yr
Provided by the employer.

Skills for this role

CommunicationAutomation

Resume tips for Sys Cloud Administrator Incident Response Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Sys Cloud Administrator Incident Response Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Sys Cloud Administrator Incident Response Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About i4DM

i4DM is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles