Overview In this role you drive reliability by shaping observability across critical customer journeys. You collaborate with product, SRE, and engineering teams to translate business outcomes into measurable SLIs/SLOs and robust observability assets. You lead governance, dashboards, and incident analysis to shorten detection and reduce customer impact, while mentoring teams on best practices. This position offers impact at scale within a mission-driven financial services environment.
Compensation / Benefits- Healthcare (medical, dental, vision)
- Life insurance
- Disability coverage
- Parental leave
- 401(k) retirement plan
- Paid vacation and holidays
Responsibilities- Define and govern SLIs/SLOs, error budgets, and reliability metrics for enterprise applications
- Lead observability strategy across critical customer journeys and ensure alignment with business outcomes
- Develop dashboards, alerts, synthetic monitoring, and telemetry governance frameworks
- Instrument applications to validate reliability and production-readiness during design and release
- Provide senior guidance on tracing, logging, metrics, synthetic testing, and alert governance
- Analyze telemetry and incident trends to identify gaps and drive reliability improvements
- Maintain authoritative inventory of observability assets and evidence of compliance
- Collaborate with product, SRE, operations, and engineering teams to ensure instrumented, measurable reliability
Key requirements- Bachelor's degree or equivalent experience
- Six to eight years in business/risk analysis, IT service management, production support, product/project management, or application development
- Expertise in Observability/SRE or reliability engineering
- Strong knowledge of SLIs, SLOs, and customer journey monitoring
- Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry
- Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry
- Experience building and tuning dashboards and actionable alerts for service health and customer impact
- Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes
- Ability to use incident analysis, RCA, and performance data to drive reliability improvements
- Excellent stakeholder management, communication, and technical leadership
- Stakeholder management
- Communication
- Technical leadership
- Datadog
- Dynatrace
- Splunk