Jobedly Post a Job

Production Reliability Engineer

Friedman Vartolo · Garden City, NY
Full-timeManufacturingMid Level$50,000–$70,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Production Reliability Engineer role

Production Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Production Reliability Engineer roles and what employers typically expect.

**The Company** Friedman Vartolo LLP is a fast-growing, New York-based law firm specializing in real estate and default services, with over 300 employees providing top-tier legal services to our clients in seven states. While our legal expertise sets us apart, it's our mindset that drives us forward. We bring a fresh, fast-paced energy that drives our momentum and shapes how we approach every challenge. We are a company that chooses to dig deeper, solving problems at the root instead of settling for surface fixes. Here, there are no passengers because every individual adds value, owns outcomes, and moves the firm forward. With an underdog mentality, we embrace constant elevation, always sharpening, always climbing, and never coasting. When challenges come, we row together and lean in as one team to get the job done, no matter what. **The Position** We are seeking a Production Reliability Engineer to own the operational health of our deployed internal application portfolio. The portfolio is growing rapidly - many new applications shipping every week - and we need a dedicated owner ensuring that what we deploy stays reliable, secure, and performant over its lifetime, and that the operational footprint stays manageable as the portfolio grows. This is a hands-on role for an engineer who thinks systematically about reliability, communicates clearly during incidents, and knows when to retire a deployed application that has outlived its usefulness. You will be the source of truth for "is our portfolio healthy" and the decision-maker on what gets maintained, what gets fixed, and what gets retired. You will work closely with the Senior Technology Manager, a DevOps Engineer (counterpart role), and a group of Production Engineers who own initial deployment and 30-day iteration of each application. **Key Responsibilities:** - Own the operational health of the firm's deployed internal application portfolio - Lead incident response for production issues; serve as on-call rotation lead - Manage security patching, dependency updates, base image refreshes, and certificate rotation across the deployed portfolio - Conduct monthly triage of every deployed application - usage levels, error rates, security posture, retire/keep recommendation - Maintain the application retirement queue and drive retire/deprecate decisions for low-usage or obsolete applications - Define and track Service Level Objectives (SLOs) for application uptime, performance, and reliability - Partner with the DevOps Engineer to ensure deployment pipelines incorporate reliability requirements from day one - Produce a weekly portfolio health report - uptime trends, open incidents, security posture, retire/deprecate decisions - Conduct incident postmortems and drive remediation actions to completion - Contribute to security and compliance posture work (SOC2 readiness, Vanta evidence collection) as it relates to operational reliability **Required Skills & Qualifications:** - 5+ years professional experience in Site Reliability Engineering, DevOps, production engineering, or equivalent - Strong incident-response experience - leading incidents, not just participating; comfortable owning the on-call rotation - Experience defining and tracking SLOs and error budgets for production systems - Strong Azure cloud experience (or directly comparable AWS / GCP experience) - Hands-on experience with monitoring, alerting, and observability tooling (Azure Monitor, Application Insights, Datadog, New Relic, or equivalent) - Comfortable reading code in multiple languages (TypeScript, Python, C# preferred) for incident triage and root-cause analysis - Understanding of cloud security fundamentals - patching cadence, dependency management, secret rotation, base image hygiene - Excellent written communication for incident comms, postmortems, and portfolio reporting **Preferred Skills & Qualifications** - Experience managing a portfolio of small-to-medium applications, rather than a single monolith - E…

Salary estimate

$50,000 – $70,000/yr
Provided by the employer.

Skills for this role

TypescriptPythonC#AWSAzureGCPCommunicationDevopsSecurity

Resume tips for Production Reliability Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Production Reliability Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Production Reliability Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Friedman Vartolo

Friedman Vartolo is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles