Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President
Software Engineering, IT, Data Science
London, UK
What We Do
At Goldman Sachs, our Engineers don’t just make things – we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. Goldman Sachs Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more.
Cloud Engineering & Architecture (CE&A)
As part of Core Engineering at Goldman Sachs, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting-edge platform that enable Goldman Sachs Engineering teams to deploy and manage services in public cloud safely and securely.
The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting-edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer-centric product development. They will work closely with cross-functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast-paced global environment.
Responsibilities And Qualifications:
We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands-on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large-scale platform initiatives from conception to delivery.
Key Responsibilities:
- Design, develop, and operationalize enterprise-grade cloud platform capabilities.
- Architect scalable, resilient, and secure infrastructure solutions on AWS.
- Architect and operationalize autonomous AI based, self-healing infrastructure
- Define technical standards, best practices, and reference architectures for cloud adoption across the firm.
- Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud.
- Drive automation and infrastructure-as-code practices to improve operational efficiency.
- Drive AI-powered FinOps and predictive resource optimization
- Mentor and guide engineers across teams, raising the overall technical bar.
- Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements.
- Evaluate emerging technologies and make recommendations for platform evolution.
- Participate in architecture design reviews and provide technical leadership on complex initiatives.
Complementary AI Based Skills:
- LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks.
- AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering.
- Self-Healing Infrastructure Engineering: Experience designing closed-loop, self-healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built-in verification and safety guardrails.
- AI-Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right-size cloud resources, manage spot instances, and optimize data platform workloads.
- Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over-provisioning while maintaining strict SLAs.
Complementary Behaviours:
- Toil-Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation.
- Risk-Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human-in-the-loop fallbacks for autonomous AI actions.
- Value-Oriented Engineering: Treats cost optimization as a first-class architectural metric, aligning infrastructure spend directly with business value and platform efficiency.
- Impact-Preserving Innovation: Executes large-scale optimization initiatives with a meticulous, risk-mitigated approach, ensuring zero disruption to production environments or developer velocity.
Basic Qualifications:
- 10+ years of experience in software engineering or infrastructure engineering.
- Deep hands-on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.).
- Strong background in platform engineering, building internal developer platforms, or infrastructure tooling.
- Experience designing and operating large-scale distributed systems.
- Proficiency with infrastructure-as-code tools (Terraform, CloudFormation).
- Strong understanding of containerization and orchestration (Docker, Kubernetes).
- Experience with CI/CD pipelines and DevOps practices.
- Knowledge of networking, security, and identity management in cloud environments.
- Excellent communication skills with the ability to influence technical decisions across teams.
- Experience working in regulated industries is a plus.
Preferred Qualifications:
- AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.).
- Experience building self-service platforms for development teams.
- Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch).
- Background in financial services or other highly regulated environments.
About Goldman Sachs
At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world.
We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at GS.com/careers.
We’re committed to finding reasonable accommodation for candidates with special needs or disabilities during our recruiting process. Learn more: https://www.goldmansachs.com/careers/footer/disability-statement.html