Site Reliability Engineer
Software Engineering
Remote
Site Reliability Engineer
Remote
Contract:1+ years
Overview
We are building a high-impact Site Reliability Engineering team to support 12+ mission-critical enterprise applications across a mix of legacy and modern environments.
This role is part of a strategic initiative focused on:
- Application instrumentation
- Observability adoption (OpenTelemetry, Dynatrace)
- Reliability engineering practices
- Platform standardization and automation
The ideal candidate combines software engineering expertise with SRE principles, capable of modifying application code, driving reliability improvements, and influencing cross-functional teams.
Key Responsibilities
- Partner with application teams to instrument services using OpenTelemetry and Dynatrace
- Improve reliability, availability, and performance of critical production applications
- Analyze system behavior and implement proactive monitoring and alerting
- Contribute to release engineering processes and CI/CD pipeline improvements
- Design and implement automation solutions to reduce manual operations
- Support hybrid environments (legacy + modern/cloud-based systems)
- Drive adoption of SRE best practices (SLIs/SLOs, error budgets, observability)
- Collaborate with engineering, platform, and operations teams to enhance delivery tooling
- Act as a technical influencer to improve engineering practices across teams
Required Qualifications (Must Have)
Strong software engineering background
1. Ability to read, debug, and modify application code (Java, .NET, Python, or similar)
- Experience with observability and instrumentation tools
- OpenTelemetry, Dynatrace (or equivalent)
- Solid understanding of Site Reliability Engineering principles
- Hands-on experience with:
- Monitoring & alerting
- Automation (scripting, tooling)
- Incident response & root cause analysis
- Familiarity with DevOps practices and CI/CD pipelines
- Experience working in enterprise-scale or complex environments
- Strong communication and collaboration skills
- Ability to work with diverse application teams and influence change
Preferred Qualifications (Nice to Have)
- Experience in financial services or regulated environments
- Exposure to large-scale platform rollouts or “factory model” transformations
- Experience with release engineering practices
- Familiarity with hybrid cloud environments (Azure preferred)
- Knowledge of modern DevOps tools and delivery platforms
- Experience modernizing legacy applications
Role Levels & Expectations
Senior SRE (2 positions)
- Lead instrumentation and observability strategy across applications
- Work independently on complex systems and reliability challenges
- Mentor junior engineers and guide best practices adoption
- Drive design decisions for automation, monitoring, and reliability frameworks
Mid / Junior SRE
- Support instrumentation, monitoring, and automation efforts
- Assist with incident management and reliability improvements
- Work within defined frameworks to gain SRE expertise
- Collaborate closely with senior engineers and application teams
Environment & Tools
- Observability: OpenTelemetry, Dynatrace
- Cloud: Azure (not primary but relevant)
- Systems: Hybrid (legacy + modern distributed systems)
- Focus Areas: Critical application support, reliability engineering, automation