Site Reliability Engineer Director
Software Engineering · Full-time
Heredia Province, Heredia, Costa Rica
Who We Are
Boston Consulting Group partners with leaders in business and society to tackle their most important challenges and capture their greatest opportunities. BCG was the pioneer in business strategy when it was founded in 1963. Today, we help clients with total transformation-inspiring complex change, enabling organizations to grow, building competitive advantage, and driving bottom-line impact.
To succeed, organizations must blend digital and human capabilities. Our diverse, global teams bring deep industry and functional expertise and a range of perspectives to spark change. BCG delivers solutions through leading-edge management consulting along with technology and design, corporate and digital ventures—and business purpose. We work in a uniquely collaborative model across the firm and throughout all levels of the client organization, generating results that allow our clients to thrive.
What You'll Do
The Principal Site Reliability Engineer (SRE) is a senior technical leader responsible for shaping how reliability, automation, and operational excellence are engineered across the organisation. Operating across domains including traditional infrastructure, cloud engineering, network operations, identity, observability, security, AI-driven operations, and automated data workflows, the role focuses on designing scalable systems, reusable engineering patterns, and standardised controls that reduce operational toil, improve resilience, and embed reliability, governance, and compliance directly into delivery pipelines and operational platforms.
This role will drive organisational change towards automation-first, measurable, and repeatable practices. A key part of the role is building and evolving reusable CI/CD and Terraform modules, engineering guardrails, observability patterns, and automation frameworks that can be adopted across multiple teams and domains without requiring each team to solve the same problems independently.
The Principal SRE also plays an important enablement role beyond deeply technical teams, helping less technical areas of the business adopt structured, governed, and scalable ways of working. This includes translating complex engineering practices into practical standards, improving how governance is implemented through engineering controls rather than manual oversight, and driving operational maturity across a broad and diverse technology landscape.
The ideal candidate is a systems thinker who understands how services, networks, identity, data flows, and operational processes fail in real-world conditions, and can apply that understanding to build automation-first, reliability-focused operating models that scale across both technical and non-technical functions.
Core responsibilities
- Identify systemic risks and failure modes across platforms and services, and define engineering solutions to mitigate them.
- Ensure operational activities are embedded into delivery models through automation, CI/CD integration, and event-driven workflows.
- Design and evolve reliability patterns across cloud, network, identity, and security domains.
- Drive maturity of observability and ops automation across multiple business units, including AIOps integration, event correlation, noise reduction, and automated remediation.
- Drive enterprise-scale transformation toward secretless and least-privilege engineering.
- Define and own the secure engineering strategy and architecture across the organisation.
- Partner with security and engineering leadership to embed security engineering into delivery and operations.
- Drive enterprise-scale transformation toward automation-first, segmented, observable networks.
- Lead the design of automation frameworks that eliminate manual operational tasks across multiple domains.
- Translate incident learnings and operational inefficiencies into scalable automation and preventative controls.
- Drive adoption of automation-first principles, reducing dependency on human-driven processes.
- Contribute to AI-driven operational use cases, including event correlation, anomaly detection, noise reduction, operational insights, and automated remediation.
- Provide technical leadership across teams, influencing standards, architecture, and engineering practices.
- Mentor engineers on reliability engineering, automation, and systems thinking.
- Drive consistency through reusable patterns, frameworks, and documentation.
- Build and evolve reusable modules, patterns, and frameworks for CI/CD, Terraform, and operational automation.
- Embed governance, validation, and reliability controls into shared engineering assets by default.
- Translate governance requirements into practical engineering controls, automated checks, and repeatable standards.
What You'll Bring
- 10-12+ years of experience in Site Reliability Engineering, Platform Engineering, or related fields.
- Strong hands-on experience across multiple domains, including: cloud platforms (AWS, Azure); CI/CD and Infrastructure-as-Code (e.g. Terraform); observability tools (e.g. Datadog, Splunk); automation and scripting (e.g. Python).
- Experience designing and implementing scalable automation and reliability solutions.
- Deep understanding of distributed systems, failure modes, and resilience patterns.
- Experience integrating operational and security controls into engineering workflows.
- Strong stakeholder engagement and technical communication skills.
- Expertise in enterprise-scale observability architecture, signal engineering, and ops automation.
- Proven track record of driving SLO-driven engineering and signal-driven automation culture across multiple teams or business units.
- Deep understanding of telemetry pipeline architecture, AIOps integration, observability cost engineering, and event-driven ops automation.
- Expertise in multi-cloud platform engineering architecture and strategy across AWS, Azure, GCP, and/or Alibaba Cloud.
- Proven track record of driving organisation-wide cloud platform adoption, consistency, and governance.
- Deep understanding of multi-cloud landing zone design, governance, networking, and IaC at enterprise scale.
- Experience setting cloud platform strategy and standards across large, federated engineering organisations.
- Expertise in enterprise identity, access, and secrets engineering.
- Proven track record of driving Zero Trust and secret-less transformation at scale.
- Deep understanding of OIDC, SAML, workload identity, and modern identity architecture.
- Expertise in enterprise secure engineering architecture and strategy.
- Proven track record of driving secure-by-default and shift-left transformation at scale.
- Deep understanding of policy-as-code, threat modelling, and secure SDLC practices.
- Expertise in enterprise network engineering architecture and strategy.
- Proven track record of driving automated and Zero Trust network transformation at scale.
- Deep understanding of hybrid and cloud network architecture, segmentation, and observability.
Preferred qualifications
- Experience with identity and access management systems (e.g. Entra ID, HashiCorp Vault).
- Experience working in large, federated organisations with diverse technology stacks.
- Exposure to compliance and regulatory requirements (e.g. PCI, HIPAA, SOX).
- Experience creating strategies for observability tooling, signal frameworks, ops automation governance, and adoption at scale.
- Hands-on architecture experience across all four major cloud providers, including Alibaba Cloud.
- Experience with cloud FinOps strategy, unit economics, and cost engineering at scale.
- Experience setting identity engineering strategy and standards.
- Experience setting secure engineering strategy and standards.
- Familiarity with Zero Trust principles and security engineering practices.
- Experience setting network engineering strategy and standards.
- Experience with network architecture and security controls (e.g. firewalls, segmentation).
Who You'll Work With
Hybrid or on-site work model.
Operates as a senior individual contributor with broad cross-organisational influence.
Expected to balance hands-on technical leadership with strategic direction.
-
Occasional travel may be required for team or stakeholder engagement.
Boston Consulting Group is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, age, religion, sex, sexual orientation, gender identity / expression, national origin, disability, protected veteran status, or any other characteristic protected under national, provincial, or local law, where applicable, and those with criminal histories will be considered in a manner consistent with applicable state and local laws.
BCG is an E - Verify Employer. Click here for more information on E-Verify.