Systems Administrator, GPU/AI Infrastructure
Software Engineering, Other Engineering, IT, Data Science
Morrisville, NC, USA
Why Work at Lenovo
Description and Requirements
Position Overview
Lenovo seeks a Systems Administrator to own end-to-end deployment, support, and administration of the consolidated Morrisville AI lab, across the network, compute, storage, operating system, and container orchestration (Docker and Kubernetes) layers, plus the NVIDIA technology stack across 50-plus servers. This is a hands-on infrastructure role: the same person who designs and administers these layers also does the day-to-day admin, support, and engineering work in the lab, including rack-and-stack and smart hands.
You are the only dedicated support resource for the Lab Operations Manager, and you keep the lab’s FY26/27 Net CapEx investment operational at the scale the AI Lab Consolidation requires. You will work alongside the lab’s Network Engineer and Junior Systems Administrator roles, who own adjacent, more specialized slices of network engineering and routine rack-and-stack work respectively; you provide the full-stack administration that connects their work end to end. You will also provide secondary technical support for capital equipment hosted in Bangalore and validate lab configurations against NVIDIA reference architecture certification standards. The role sits within Lenovo’s Hybrid Cloud and AI Infrastructure Services (HCAIS) Lab Operations organization.
Key Responsibilities
Infrastructure Deployment and Administration
Network: deploy, support, and administer lab networking, in coordination with the dedicated Network Engineer role for switch and fabric configuration.
Compute: deploy, support, and administer physical and virtual compute infrastructure across the lab.
Storage: deploy, support, and administer lab storage systems across block and file tiers.
Operating Systems: install, patch, and administer operating systems across lab infrastructure.
Container Orchestration
Docker: deploy and administer Docker container runtime environments.
Kubernetes: own full production Kubernetes administration, including cluster operations, role-based access control (RBAC), cluster networking, and version upgrades.
NVIDIA Platform Administration
NVIDIA Technology Stack: administer NVIDIA drivers, the CUDA toolkit, InfiniBand fabric, Run:AI orchestration, and ClearML across the lab’s GPU infrastructure.
Configuration Validation: validate lab configurations against NVIDIA reference architecture (Lenovo Validated Design) certification requirements.
Lab Operations and Smart Hands
Day-to-Day Support: provide day-to-day administration, support, and engineering functions for the network, compute, and storage layers in the physical lab.
Rack and Stack: perform rack-and-stack, structured cabling, and smart-hands support for lab hardware alongside higher-level administration duties.
Cross-Site Support and Coordination
Bangalore Secondary Support: provide secondary technical support for NVIDIA technology stack capital equipment hosted in Bangalore.
Lab Coordination: coordinate with the adjacent ISG AI COE lab (Tech Marketing, customer proof-of-concepts) and the Morrisville Executive Briefing Center given shared physical proximity.
NVIDIA Relationship: maintain a direct working relationship with NVIDIA field contacts based in Morrisville.
Required Qualifications
Required Experience
- 2+ years of systems administration experience with hands-on breadth across network, compute, storage, and operating system layers, plus a GPU/AI infrastructure specialization.
- Production Kubernetes Administration: demonstrated experience with cluster operations, RBAC, cluster networking, and version upgrades in a production environment.
- Docker: hands-on experience administering Docker container runtime environments.
- Hands-on NVIDIA administration: demonstrated experience administering NVIDIA drivers and the CUDA toolkit.
- InfiniBand Fabric: working familiarity with InfiniBand fabric in a multi-node GPU environment.
- GPU Orchestration: experience with Run:AI, ClearML, or a comparable GPU orchestration and MLOps toolset.
- Physical Infrastructure: comfortable with hands-on hardware work, including rack-and-stack, structured cabling, and smart hands, in addition to higher-level administration.
Core Professional Skills
- Incident Response: able to own after-hours incident response for a shared, multi-team infrastructure environment.
- Cross-Team Coordination: comfortable supporting multiple HC/AI teams across a shared lab environment with competing priorities, and coordinating day to day with the Network Engineer and Junior Systems Administrator roles on adjacent, overlapping infrastructure.
Preferred Qualifications
- Kubernetes Certification: Certified Kubernetes Administrator (CKA) or equivalent.
- Multi-Tenant Lab Experience: experience supporting a multi-tenant shared lab environment serving distributed teams.
- Familiarity with NVIDIA certification and reference architecture (Lenovo Validated Design) processes.
- Experience with DCIM, IPAM, or monitoring tooling comparable to the lab’s stack (Hyperview-class DCIM, BlueCat / Infoblox-class IPAM, NVIDIA DCGM monitoring).