Global IT Software Engineer Manager

Boston Consulting Group
Boston Consulting Group

Software Engineering, IT

Gurugram, Haryana, India

Posted on Aug 11, 2026

Who We Are

Boston Consulting Group partners with leaders in business and society to tackle their most important challenges and capture their greatest opportunities. BCG was the pioneer in business strategy when it was founded in 1963. Today, we help clients with total transformation-inspiring complex change, enabling organizations to grow, building competitive advantage, and driving bottom-line impact.

To succeed, organizations must blend digital and human capabilities. Our diverse, global teams bring deep industry and functional expertise and a range of perspectives to spark change. BCG delivers solutions through leading-edge management consulting along with technology and design, corporate and digital ventures—and business purpose. We work in a uniquely collaborative model across the firm and throughout all levels of the client organization, generating results that allow our clients to thrive.



What You'll Do

We are seeking an automation-first Platform Automation Engineer Manager to join our IT Operations team. The role is accountable for engineering and continuously expanding the automation that runs our day-to-day operational processes — systematically reducing ticket volumes, eliminating manual toil, and moving the function from reactive support toward proactive, self-healing, autonomous operations.

The successful candidate will automate the core workflows of a 24x7 operations centre, from event detection and triage through runbook-driven auto-remediation and proactive anomaly detection, and will embed AI/LLM-powered intelligence to make operations increasingly autonomous.

A working understanding of core reliability principles — SLIs, SLOs, error budgets, and toil reduction — is valued in this role, as these concepts increasingly inform how automation priorities are set and measured across the operation.

Success is measured in operational outcomes: ticket deflection, reduced manual effort, faster resolution, higher SLA attainment, and a demonstrable shift from manual to autonomous operations.

You will bridge the gap between traditional IT Operations and next-generation AI/LLM-powered automation — driving measurable improvements in SLA attainment, MTTR reduction, and operational efficiency across the enterprise.

Operational Process Automation & Ticket Reduction

  • Own operational automation: lead the automation of core 24x7 operational processes — incident, request, monitoring-response, and routine operational tasks — with a primary mandate to eliminate tickets from the source with automation.
  • Build the runbook-automation library: convert manual operating procedures into reusable, event-driven, auto-remediation workflows, and continuously govern and expand the catalogue.
  • Deliver automation at scale: target measurable year-on-year reduction in manual tickets and operational toil across all managed services.

Proactive & Autonomous Operations

  • Detect before it tickets: implement proactive anomaly detection and predictive alerting to identify and resolve issues before they become incidents or reach the operations queue.
  • Engineer toward lights-out operations: design closed-loop, self-healing automation that detects, diagnoses, and remediates without human intervention, advancing the operation toward autonomous running.
  • Embed AI/LLM intelligence: integrate Claude pipelines, OpenAI APIs, or equivalents into operational workflows for anomaly detection, alert summarization, predictive triage, and auto-resolution.

Operational Toolchain & Integration

  • Unify the operational data fabric: develop and govern API-based integrations across ServiceNow, PagerDuty, Datadog, Splunk and equivalents for seamless end-to-end process orchestration.
  • Orchestrate cross-system workflows: design event-driven and scheduled automation — including human-in-the-loop approvals — using enterprise RPA/workflow platforms such as Automation Anywhere, Microsoft Power Automate, or n8n, or any additional BCG-approved toolsets.

Reliability & Metrics-Driven Automation

  • Apply reliability metrics to automation priorities: use SLIs, SLOs, and error budgets as inputs to prioritize which manual processes get automated first, and to quantify the reliability impact of each automation delivered.
  • Collaborate with SRE teams: partner with SRE and platform engineering teams to align automation initiatives with broader reliability objectives, shared tooling, and incident-response standards.
  • Reduce toil systematically: treat any repeated manual task as a defect to be engineered out, not simply managed.

Monitoring, Insight & Continuous Improvement

  • Track what matters to a 24x7 operation: build and maintain real-time dashboards covering ticket-deflection rate, automation coverage, SLA attainment, MTTR, MTTD, MTTA, and toil reduction.
  • Reduce noise: continuously tune alerting thresholds and noise-reduction logic to improve signal-to-noise ratio and reduce alert fatigue for operations teams.
  • Drive a metrics-led culture: translate operational data into executive-ready insight and recommendations.

24x7 Operations Leadership

  • Run the operational incident lifecycle: detection, triage, escalation, resolution, and structured post-incident review — applying an automation-first approach at every stage.
  • Manage on-call: own on-call rotations and escalation chains via PagerDuty, ensuring incidents are routed, tracked, and resolved within agreed SLAs across the 24x7 model.
  • Build operational readiness: define operational-readiness and automation-governance standards so new services enter operations with automation and monitoring built in from day one.
  • Grow the team: mentor and upskill operations engineers in automation-first ways of working, championing reusable automation over repeated manual effort, and encouraging foundational SRE literacy across the team.


What You'll Bring

Education & Experience

12+ years in IT Operations, Operations Automation, or AIOps roles, including at least 3 years at manager level in a 24x7 operational environment.

Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.

Demonstrated track record of delivering operational automation at scale — measurable ticket reduction, toil elimination, and process automation in complex enterprise environments.

Technical Skills

  • RPA & workflow automation — Automation Anywhere, Microsoft Power Automate, n8n, Camunda, or equivalent enterprise-grade tooling, including end-to-end design of event-driven and scheduled workflows with human-in-the-loop approvals.
  • Operational toolchain — ServiceNow (incident, change, problem, workflow automation) and PagerDuty (on-call scheduling, escalation policies, AIOps integration).
  • Monitoring & observability — Datadog, Splunk, or equivalent — operational dashboards, monitors, and alert policies.
  • Integration — REST, GraphQL, webhooks, and event-driven architecture across the operational toolchain.
  • AI/LLM integration — Claude, OpenAI, or equivalent embedded into operational workflows.
  • Scripting — Python, Bash, PowerShell, or equivalent for automation tooling.
  • Operational & reliability metrics — SLA, SLI, SLO, MTTR, MTTD, MTTA — definition, tracking, and governance, with a working understanding of how these concepts are used in SRE practice to set reliability targets and error budgets.

Operational Competencies

  • Proven ability to operate and lead in 24x7 on-call environments; comfortable with shared on-call ownership.
  • Automation-first mindset — treats every recurring operational task as an automation opportunity, engineering reusable solutions rather than repeating manual effort.
  • Strong stakeholder communication — able to translate operational and technical findings into executive-level narratives.
  • Experience running structured post-incident reviews and converting learnings into systemic operational improvements.

PREFERRED QUALIFICATIONS & ADDED ADVANTAGE

  • Infrastructure automation (strong added advantage): knowledge of automating infrastructure components — compute, network, storage, and cloud resources.
  • Infrastructure-as-code & orchestration: familiarity with Terraform, Ansible, or Kubernetes in an operations-automation context.
  • AIOps platforms: experience with AIOps and predictive-operations tooling.
  • Cloud exposure: operational monitoring and automation patterns across AWS, Azure, or GCP.
  • Command-centre leadership: prior experience building or operating NOC / command-centre teams in a global 24x7 model.
  • Process certification: ITIL v4 Foundation or above; automation/platform certifications (e.g., Automation Anywhere, Power Platform).
  • SRE certification (preferred): an SRE certification — such as Google Cloud Professional Site Reliability Engineer, DevOps Institute SRE Foundation, or equivalent — is preferred, and directly supports the reliability-alignment aspects of this role.
  • SRE & DevOps fluency: solid working knowledge of SRE and DevOps practices — SLIs/SLOs, error budgets, blameless postmortems — sufficient to partner effectively with SRE and platform teams.


Additional info

At BCG, our people and relationships are at the heart of everything we do. We believe that in-person collaboration plays an important role in our culture, mentorship, and professional development.

For this role, you will generally be expected to work from a BCG office or in person with clients on a regular basis (typically around 50% of working time), depending on business needs and in line with local policies and practices.

We aim to provide a dynamic and collaborative environment that fosters connection and teamwork.



Boston Consulting Group is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, age, religion, sex, sexual orientation, gender identity / expression, national origin, disability, protected veteran status, or any other characteristic protected under national, provincial, or local law, where applicable, and those with criminal histories will be considered in a manner consistent with applicable state and local laws.
BCG is an E - Verify Employer. Click here for more information on E-Verify.