AI/HPC Implementation Technical Consultant
Software Engineering, IT, Data Science
Morrisville, NC, USA
USD 110,600-169,510 / year
Why Work at Lenovo
Description and Requirements
REMOTE ROLE anywhere in the U.S.
Lenovo is a global leader in AI infrastructure, high-performance compute platforms, and data centre technology, helping customers design, deploy, and scale the next generation of accelerated computing environments. Our Lenovo Infrastructure Deployment Services (LIDS) team is committed to fostering a culture of ownership and entrepreneurship, where individuals are empowered to challenge themselves, develop their skills, and be recognised and rewarded for their contributions.
Lenovo is looking for an experienced Technical Consultant to join its LIDS team. The role will combine complex technical delivery leadership with bid management support, ensuring that pre-sales commitments, commercial assumptions, delivery risks, and implementation plans are clearly understood, validated, and transitioned into executable customer projects.
The increasing demand for AI infrastructure projects, including large-scale AI Factory environments and NeoCloud clusters, requires a scalable delivery model that can consistently execute multiple deployments in parallel while maintaining quality and customer satisfaction. This role will join a team of North America deployment POD's, consisting of four Technical Consultants with complementary infrastructure, networking, AI platform, and deployment expertise. These PODs will operate as dedicated delivery teams responsible for the end-to-end deployment, integration, validation, and customer handover of AI Factory and NeoCloud environments.
This role offers the opportunity to deploy large-scale, complex AI infrastructure & HPC deployments, including liquid-cooled GPU clusters, accelerated compute environments, AI-ready data center platforms, and integrated infrastructure solutions for some of Lenovo’s most strategic global customers.
Organization
You will be part of a global HPC team supporting customers and deploying projects primarily in the North American region but on occasion worldwide. You will participate in weekly collaborative planning calls and meetings.
What you’ll bring/Position Requirements
The Technical Consultant will be responsible for the deployment, integration, validation, and support of large-scale HPC and AI infrastructure solutions. The role spans the full technology stack from data center infrastructure, networking, storage, and Kubernetes platforms through to AI orchestration, MLOps, GPU resource management, and workload optimization. The consultant will work closely with customers, partners, and project teams to deliver successful HPC and AI deployments from initial installation through production handover and operational readiness.
Applicant must possess strong customer interaction skills and ability to make technical decisions. Collaborate during projects with several organization verticals, partners, and customers. Develop training and knowledge base documentation for technical skills throughout career.
Required Skills & Experience
- Strong Linux administration experience (RHEL, Rocky Linux, Ubuntu, SLES) in enterprise, HPC and AI environments.
- Experience deploying and supporting HPC and AI clusters
- Knowledge of NVIDIA AI platforms and architectures, including GB300 NVL72, NVLink, NVSwitch, GPU clustering, and AI factory concepts.
- Experience with high-performance networking technologies, including InfiniBand, Ethernet, RoCEv2 and Spectrum-X and fabric management.
- Experience with Kubernetes, OpenShift, containerized workloads, and GPU orchestration platforms.
- Familiarity with HPC workload scheduling and resource management platforms such as Slurm, PBS Pro.
- Experience deploying and integrating high-performance storage solutions for HPC and AI workloads.
- Experience with cluster provisioning, system bring-up, validation, performance benchmarking, and acceptance testing.
- Knowledge of monitoring and observability tools such as Icinga, Grafana, Prometheus, Loki, and NVIDIA DCGM.
- Automation and scripting experience using Python, Bash, Ansible, and Git.
- Strong troubleshooting skills across compute, networking, storage, Kubernetes, and AI software stacks.
- Excellent customer-facing skills with experience delivering workshops, technical documentation, deployment runbooks, and knowledge transfer sessions.
Preferred Skills
- Experience with NVIDIA AI Factory deployments and enterprise AI infrastructure.
- Knowledge of GPFS, VAST, DDN Data storage platforms.
- Experience with NVIDIA Cluster Validation Suite (CVS).
- Familiarity with Kubeflow, MLFlow, JupyterHub, or similar AI/ML platforms.
- Experience with direct liquid-cooled infrastructure.
- Knowledge of AI platform technologies including Run:AI, Rafay, ClearML, MLOps
- NVIDIA, Kubernetes, Red Hat, or AI infrastructure certifications.
What We Will Offer You:
- A multitude of professional and personal opportunities.
- An open and stimulating environment within one of the most forwarding thinking IT companies.
- Flat structures and fast decision-making processes.
- A modern and flexible way of working to combine personal and professional life.
- An international team with a high focus on Diversity.