Lead Site Reliability Engineer

Luxoft
Wilmington, NC

Responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others.

Responsibilities

Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.

Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.

Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.

Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.

Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.

Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.

Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.

Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.

Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.

Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).

Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.

Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.

Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.

Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.

Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.

Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.

Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.

Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.

Lead capacity planning, performance tuning, and workload optimization efforts across production environments.

Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.

Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.

Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.

Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.

Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.

Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.

Identify reliability, operational, and technology risks requiring escalation to management.

Promote an environment that supports a culture of belonging and reflects the Client brand.

Maintain Client internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.

Complete other related duties as assigned.

Skills

Must have

Strong experience in observability and monitoring, including hands-on expertise with:

Dynatrace

OpenTelemetry (OTel)

Distributed tracing

Metrics collection and analysis

Centralized logging and log aggregation

Alerting and dashboard development

Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.

Strong proficiency in Infrastructure as Code (IaC) using Terraform.

Experience with CI/CD pipelines, deployment automation, and operational tooling.

Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.

Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.

Cloud & Platform Expertise

Strong experience with Microsoft Azure, including:

Azure App Services

Resource Groups

Azure networking concepts

Scaling and performance optimization

Deployment and release management

Application lifecycle management

Experience leveraging Azure-native operational tooling such as:

Azure Monitor

Application Insights

Log Analytics

Azure dashboards and alerting

Experience supporting cloud-native and hybrid infrastructure environments.

Reliability & Engineering Practices

Demonstrated experience implementing and operating SRE practices, including:

Service Level Objectives (SLOs)

Service Level Indicators (SLIs)

Error budgets

Incident management

Problem management

Root Cause Analysis (RCA)

Reliability automation

Ability to improve system reliability through:

Performance tuning

Capacity planning

Observability-driven insights

Proactive issue detection

Reliability engineering initiatives

Experience developing automated recovery mechanisms and self-healing solutions.

Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.

Nice to have

Experience supporting large-scale enterprise applications in regulated environments.

Strong analytical and troubleshooting skills related to production systems and distributed architectures.

Experience working in Agile and DevOps operating models.

Ability to work autonomously and lead complex reliability initiatives.

Strong organizational and time management skills.

Advanced verbal and written communication skills.

Experience driving project milestones and delivery commitments.

Proven experience leading major incident response and post-incident improvement efforts.

Experience partnering with architecture, infrastructure, cybersecurity, and application development teams.

Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.

Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering preferred.

Other

Languages

English: C1 Advanced

Seniority

Lead

Posted 2026-07-31

Recommended Jobs

Office Manager

CER-MET INC
Charlotte, NC

Job Description Job Description Benefits/Perks Flexible Scheduling Competitive Compensation Careers Advancement Job Summary Supervises and/or performs secretarial, clerical and othe…

View Details
Posted 2026-06-20

KITCHEN MANAGER

Metro Services, LLC
Raleigh, NC

Job Description Job Description Position Description: We are looking for friendly folks like you to join our team! Metro Diner is known for warm, welcoming service, familiar faces, and award-winn…

View Details
Posted 2026-06-23

CNA/PCA

ADP - RNOOID0028057745
Raleigh, NC

Job Description Job Description Now Hiring: CNA / PCS Caregivers We’re a growing home health agency looking for compassionate, reliable CNAs and Personal Care Specialists (PCS) to join our tea…

View Details
Posted 2026-05-29

Nurse Practitioner - Pediatric Critical Care

Duke University Health System
Durham, NC

At Duke Health, we're driven by a commitment to compassionate care that changes the lives of patients, their loved ones, and the greater community. No matter where your talents lie, join us and discov…

View Details
Posted 2026-06-21

Regional Technical Manager - East Central

JM Huber Corporation
Charlotte, NC

Portfolio Business :   Huber Engineered Woods   J.M. Huber Corporation is one of the largest privately held, family-owned companies in the United States. Established in 1883, we are a diversifie…

View Details
Posted 2026-03-24

Full-time Food & Beverage Supervisor

Aileron Management LLC
Boone, NC

Job Description Job Description Description: The Horton Hotel is looking for a full-time, Bar Supervisor to help manage the bar and create memorable experiences for guests during their stay. 3…

View Details
Posted 2026-04-14

Production Team Partner - Mat Roller & Order Builder - UniFirst

UniFirst
Wilmington, NC

Our Production Team is Kind of a Big Deal! UniFirst is seeking a reliable and hardworking Production Team Partner to join our UniFirst Family. As a Team Partner in the Production Department, you w…

View Details
Posted 2026-07-31

Occupational Therapy Assistant / COTA / OTA

Broad River Rehabilitation
Gatesville, NC

Occupational Therapy Assistant / COTA / OTA - Assisted Living Facility - Gatesville, NC / North Carolina Broad River Rehab is seeking an Occupational Therapy Assistant to join our ALF facility in …

View Details
Posted 2026-05-11

Proposal Specialist - Federal or Commercial Energy

TEEMA
Charlotte, NC

Job Description Job Description Are you a meticulous coordinator with a knack for persuasive writing? Do you thrive in a fast-paced environment where your organization directly drives business gr…

View Details
Posted 2026-06-25

General Laborer

SGS Consulting
North Carolina

Job Responsibilities: Perform general manual labor tasks including loading, unloading, lifting, and moving materials. Be able to use simple power and hand tools Know how to read a tape measu…

View Details
Posted 2026-04-22