Senior Staff Consulting Architect - AI Infrastructure
Role Overview
Nutanix is seeking a Sr. Staff Consulting Architect – AI to provide senior technical leadership in the architecture, design, and transformation of enterprise AI platforms based on the Nutanix AI Platform (NAI) and equivalent enterprise AI/ML platforms.
The role is intended for a highly experienced AI platform architect who can lead complex enterprise AI transformation initiatives spanning AI infrastructure, GPU-enabled Kubernetes, MLOps/LLMOps, Generative AI, RAG, model serving, data services, automation, security, and hybrid/multi-cloud environments.
The Sr. Staff Consulting Architect will act as a trusted technical advisor to senior customer stakeholders, helping organizations define their AI platform strategy, target architecture, adoption roadmap, and operating model.
The role requires a combination of deep hands-on expertise and strategic architecture leadership across AI/ML platforms, Kubernetes, virtualization, GPU infrastructure, cloud-native technologies, data platforms, and automation.
The architect will also provide technical leadership to Staff Consulting Architects and Consultants, establish reference architectures and reusable AI solutions, and collaborate with Product Management, Engineering, and other technical organizations to influence the evolution and adoption of Nutanix AI solutions.
Key Responsibilities
1. Enterprise AI Strategy & Architecture
- Lead architecture strategy for enterprise AI/ML and Generative AI platform transformation programs.
- Define target-state AI platform architectures aligned to customer business, data, application, security, and operational requirements.
- Architect enterprise AI platforms across:
- Nutanix AI Platform (NAI)
- Kubernetes-based AI platforms
- Nutanix AHV
- Bare-metal GPU infrastructure
- Hybrid and multi-cloud environments
- Public cloud AI/ML platforms where applicable
- Develop phased AI adoption roadmaps covering:
- AI infrastructure
- Model development
- Model training and fine-tuning
- Inference
- RAG
- MLOps/LLMOps
- Application integration
- Governance and operations
- Advise customer executives, enterprise architects, data scientists, ML engineers, platform teams, and application teams on enterprise AI adoption.
- Translate business use cases into scalable, secure, resilient, and operationally sustainable AI architectures.
- Evaluate architecture trade-offs across performance, scalability, GPU utilization, cost, security, data locality, and operational complexity.
2. AI Infrastructure & GPU Platform Architecture
- Design enterprise AI infrastructure supporting:
- GPU-enabled Kubernetes
- AI/ML workloads
- Model training
- Fine-tuning
- Inference
- Vector search
- RAG workloads
- Define GPU architecture considering:
- GPU selection and sizing
- GPU scheduling
- GPU partitioning and sharing
- CPU/GPU topology
- NUMA
- Memory requirements
- Network bandwidth
- Storage throughput
- Architect high-performance AI infrastructure across virtualized and bare-metal environments.
- Design GPU-enabled Kubernetes platforms using technologies such as:
- NVIDIA GPU Operator
- NVIDIA CUDA
- NVIDIA Container Toolkit
- NVIDIA Triton Inference Server
- KServe
- Evaluate and architect high-performance networking for AI workloads, including where applicable:
- RDMA
- InfiniBand
- RoCE
- High-speed Ethernet
- NVIDIA networking technologies
- Optimize infrastructure for GPU utilization, workload density, performance, availability, and cost.
- Define architectures for scalable AI compute pools and shared enterprise AI platforms.
3. Kubernetes & AI Platform Engineering
- Architect Kubernetes platforms optimized for enterprise AI workloads.
- Define AI platform architecture across:
- Kubernetes control plane
- GPU worker nodes
- Storage
- Networking
- Ingress
- Security
- Observability
- Design multi-cluster and multi-tenant AI environments.
- Define workload isolation across:
- Cluster
- Namespace
- Node
- GPU
- Architect integration with enterprise infrastructure and services including:
- Identity and access management
- LDAP/Active Directory
- DNS
- PKI/certificates
- Load balancers
- Storage
- Backup and DR
- Monitoring and logging
- Drive platform engineering and self-service capabilities for data scientists and AI developers.
- Define standardized patterns for AI workspace provisioning and lifecycle management.
4. Generative AI, RAG & AI Application Architecture
- Lead architecture for enterprise Generative AI solutions and platforms.
- Design end-to-end RAG architectures covering:
- Data ingestion
- Document processing
- Chunking
- Embedding generation
- Vector databases
- Retrieval
- Re-ranking
- Prompt orchestration
- LLM inference
- Architect enterprise AI application patterns using technologies such as:
- Large Language Models (LLMs)
- Embedding models
- Vector databases
- LangChain
- LlamaIndex
- NVIDIA NeMo
- NVIDIA NIM
- Triton
- KServe
- Define architectures for:
- LLM inference
- Model serving
- Fine-tuning
- RAG
- AI agents
- Enterprise copilots
- Evaluate open-source and commercial AI models and frameworks based on customer requirements.
- Design model deployment strategies considering:
- Performance
- Latency
- Throughput
- GPU utilization
- Model size
- Security
- Data privacy
- Cost
- Guide customers in moving AI prototypes and proof-of-concepts into production-grade platforms.
5. MLOps, LLMOps & Automation
- Define enterprise MLOps and LLMOps architectures covering the complete AI lifecycle.
- Design workflows for:
- Data preparation
- Model training
- Fine-tuning
- Model evaluation
- Model registration
- Model deployment
- Model monitoring
- Model retirement
- Integrate AI platforms with CI/CD and GitOps practices.
- Design automated AI platform provisioning using:
- Terraform
- Ansible
- Kubernetes operators
- APIs
- Establish automation for:
- GPU infrastructure provisioning
- Kubernetes clusters
- AI workspaces
- Model deployment
- Inference endpoints
- RAG pipelines
- Integrate automation with:
- Nutanix APIs
- NKP APIs
- Kubernetes APIs
- NVIDIA technologies
- Cloud provider APIs
- Develop repeatable, version-controlled, auditable, and supportable AI platform delivery patterns.
6. Data, Storage & AI Platform Integration
- Architect data platforms supporting enterprise AI workloads.
- Define integration between AI platforms and:
- File services
- Object storage
- Databases
- Data lakes
- Vector databases
- Enterprise data sources
- Design storage architectures optimized for:
- Model datasets
- Training data
- Model artifacts
- Checkpoints
- Embeddings
- Vector data
- Inference workloads
- Address performance requirements across:
- IOPS
- Throughput
- Latency
- Metadata performance
- Network bandwidth
- Define data movement and data locality strategies for hybrid AI environments.
- Establish architecture patterns for data protection, backup, disaster recovery, and lifecycle management.
7. AI Security, Governance & Responsible AI
- Define security architecture across the AI infrastructure, platform, data, model, and application layers.
- Establish controls for:
- Identity and access management
- RBAC
- Secrets management
- Network security
- Data protection
- Model access
- API security
- Design secure multi-tenant AI platforms with appropriate isolation.
- Define governance for:
- Models
- Datasets
- Prompts
- Embeddings
- AI applications
- Inference endpoints
- Establish security and governance patterns for enterprise RAG and GenAI applications.
- Address risks associated with:
- Sensitive data
- Data leakage
- Prompt injection
- Model misuse
- Unauthorized model access
- Integrate policy-as-code and security controls into AI platform lifecycle management.
- Guide customers toward responsible, auditable, and governed enterprise AI adoption.
8. Observability, Performance & Cost Optimization
- Define comprehensive observability for AI platforms covering:
- Infrastructure
- GPUs
- Kubernetes
- Models
- Inference services
- Applications
- Establish monitoring and optimization for:
- GPU utilization
- GPU memory
- CPU utilization
- Network performance
- Storage throughput
- Model latency
- Inference throughput
- Define AI platform SLOs, SLIs, and operational KPIs.
- Troubleshoot complex issues across:
- GPU infrastructure
- Kubernetes
- Networking
- Storage
- AI frameworks
- Model serving
- Application layers
- Drive capacity planning and AI infrastructure optimization.
- Develop strategies for optimizing AI infrastructure and cloud consumption costs.
9. Consulting, Technical Leadership & Customer Enablement
- Lead executive-level architecture workshops for:
- Enterprise AI strategy
- AI platform adoption
- Generative AI
- RAG
- MLOps/LLMOps
- AI infrastructure modernization
- Act as the senior technical advisor to customer:
- CIO/CTO organizations
- Enterprise architects
- Data and AI leaders
- Infrastructure leaders
- Platform engineering teams
- Data science and ML engineering teams
- Lead complex technical discovery and architecture engagements.
- Provide technical direction for strategic AI programs and high-value customer engagements.
- Act as the senior escalation point for complex AI architecture and implementation challenges.
- Review and govern AI architectures developed by Staff Consulting Architects and Consultants.
- Mentor architects and consultants and build organizational AI capability.
10. Reference Architecture, Innovation & Thought Leadership
- Create enterprise-grade AI reference architectures and solution blueprints.
- Develop reusable:
- Design patterns
- Implementation guides
- Automation frameworks
- AI solution templates
- MLOps/LLMOps patterns
- RAG reference implementations
- Identify and evaluate emerging AI technologies and their applicability to Nutanix solutions.
- Drive innovation across:
- Generative AI
- AI agents
- RAG
- GPU infrastructure
- AI platform engineering
- Model serving
- AI observability
- Collaborate with Product Management, Engineering, Support, and other technical teams to provide field feedback and influence product direction.
- Contribute to technical whitepapers, solution briefs, blogs, reference architectures, and internal enablement content.
- Represent Nutanix in strategic customer discussions, technical forums, and industry events where appropriate.
Technical Skills & Experience
Mandatory Skills – Must Have
- 15+ years of experience across infrastructure, cloud, virtualization, Kubernetes, AI/ML platforms, or related technologies.
- Proven experience architecting and delivering enterprise AI/ML platforms and AI infrastructure.
- Deep understanding of AI/ML platform architecture and enterprise AI adoption.
- Strong hands-on experience with:
- Kubernetes
- Nutanix AHV or equivalent virtualization platforms
- GPU-enabled infrastructure
- Linux
- Container technologies
- Strong experience with Nutanix AI Platform (NAI) or equivalent enterprise AI platforms.
- Deep understanding of GPU architecture and AI workload requirements.
- Strong understanding of:
- GPU scheduling
- CUDA
- GPU Operator
- GPU resource management
- GPU performance optimization
- Strong experience with Kubernetes-based AI platforms.
- Strong understanding of AI/ML frameworks and platforms, including several of:
- NVIDIA NeMo
- NVIDIA NIM
- Triton
- KServe
- Kubeflow
- MLflow
- Strong understanding of Generative AI and LLM architectures.
- Proven experience designing RAG architectures and AI application platforms.
- Strong understanding of:
- LLMs
- Embedding models
- Vector databases
- Model serving
- Fine-tuning
- Inference
- Strong experience with Infrastructure as Code:
- Terraform
- Ansible
- Strong experience with CI/CD, GitOps, and platform automation.
- Strong understanding of enterprise storage, networking, security, and data architecture for AI workloads.
- Strong troubleshooting skills across infrastructure, Kubernetes, GPU, networking, storage, and AI platform layers.
- Proven ability to lead complex customer engagements and architecture decisions.
- Excellent customer-facing, consulting, presentation, and executive communication skills.
Added Advantage – Good to Have
- Experience with NVIDIA AI Enterprise.
- Experience with NVIDIA DGX, HGX, or equivalent GPU infrastructure.
- Experience with NVIDIA networking technologies, InfiniBand, or RoCE.
- Experience with GPU partitioning and sharing technologies.
- Experience with Kubeflow, MLflow, KServe, Ray, or equivalent AI/ML platforms.
- Experience with LangChain, LlamaIndex, Flowise, or equivalent GenAI orchestration frameworks.
- Experience with vector databases such as Milvus, Qdrant, Weaviate, pgvector, or equivalent.
- Experience with AI agents and agentic architectures.
- Experience with multimodal AI.
- Experience with AI/ML model fine-tuning and optimization.
- Experience with public cloud AI services and GPU platforms on AWS or Azure.
- Experience with hybrid and multi-cloud AI architectures.
- Experience with AI governance, security, and responsible AI practices.
- Experience with FinOps and AI infrastructure cost optimization.
Trainable / Developable Skills
- Advanced AI agent and agentic workflow architectures.
- Emerging LLM and multimodal model architectures.
- Advanced model optimization and inference technologies.
- Emerging GPU and AI accelerator technologies.
- Advanced AI governance and responsible AI frameworks.
- Advanced AI observability and LLM evaluation frameworks.
- New Nutanix AI platform capabilities and ecosystem integrations.
Leadership & Consulting Expectations
The Sr. Staff Consulting Architect – AI is expected to:
- Operate effectively with C-level executives, AI leaders, enterprise architects, data scientists, ML engineers, infrastructure leaders, and platform engineering teams.
- Lead strategic AI architecture discussions and influence enterprise technology direction.
- Convert business AI opportunities into practical, production-ready architectures.
- Balance AI innovation with enterprise requirements for:
- Security
- Governance
- Reliability
- Performance
- Scalability
- Cost
- Operational sustainability
- Independently lead ambiguous and high-impact AI architecture problems.
- Provide technical authority for complex AI engagements and critical escalations.
- Mentor Staff Consulting Architects and Consultants and raise the organization's AI technical maturity.
- Convert customer and field experience into reusable AI reference architectures, methodologies, and intellectual property.
- Influence product and engineering roadmaps through structured field feedback and technical collaboration.
Education & Certifications
Education
- Bachelor's Degree in Engineering, Computer Science, Artificial Intelligence, Data Science, or equivalent discipline.
- Master's degree in Computer Science, AI/ML, Data Science, or a related discipline is an advantage.
Preferred Certifications
- Nutanix certifications:
- NCP
- NAI / relevant Nutanix AI certification
- NCX preferred
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- NVIDIA certifications relevant to AI/ML or data center technologies
- Terraform Associate
- Relevant AWS/Azure certifications
- Relevant AI/ML or cloud architecture certifications
Location & Work Shifts
- Pune / Bangalore – Hybrid model (3 days a week)
- Travel as required for customer assignments
- Willingness to work US and EMEA shifts based on customer and engagement requirements
--
Nutanix is an equal opportunity employer.
Nutanix is an Equal Employment Opportunity and (in the U.S.) an Affirmative Action employer. Qualified applicants are considered for employment opportunities without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, marital status, protected veteran status, disability status or any other category protected by applicable law. We hire and promote individuals solely on the basis of qualifications for the job to be filled. We strive to foster an inclusive working environment that enables all our Nutants to be themselves and to do great work in a safe and welcoming environment, free of unlawful discrimination, intimidation or harassment. As part of this commitment, we will ensure that persons with disabilities are provided reasonable accommodations. If you need a reasonable accommodation, please let us know by contacting [email protected].