Jobedly Post a Job

Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

Tech Holding · Remote
RemoteFull-timeEngineeringTechnology$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Lead Site Reliability Engineer role

Lead Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Lead Site Reliability Engineer roles and what employers typically expect.

About us: Working at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-quality solutions to our clients. Our founders and team members have industry experience and have held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability. The Role: We are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives. This is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation. You will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling. Key Responsibilities: Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams Build capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost Lead load, stress, soak, spike, failure, and recovery testing in representative environments Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements Partner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates Own technical readiness assessments for major pilots, partnerships, and production launches Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures Lead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership Recommend capacity and reliability investments before they become production constraints Required Skills: Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements Deep understanding of observability, performance analysis, capacity planning, and reliability engineering Strong hands-on experience with cloud infrastructure and production distributed systems Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics Hands-on experience performing load, stress, soak, scalability, and resilience testing Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with en…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

LeadershipDevopsAutomation

Resume tips for Lead Site Reliability Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Lead Site Reliability Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Lead Site Reliability Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Tech Holding

Tech Holding is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles