Skip to main content
Documentation
close
Get Started
Get Started with Google Cloud
Product List
Cloud Customer Care
Featured Products
Agent Platform
Apigee API Management
BigQuery
Compute Engine
Cloud CDN
Cloud Run
Cloud Storage
Cloud SQL
Gemini Enterprise
Google Kubernetes Engine
Looker
Cross-product Tools
Access and resources management
Costs and usage management
Infrastructure as code
SDK, languages, frameworks, and tools
Technology Areas
AI and ML
Application development
Application hosting
Compute
Data analytics and pipelines
Databases
Distributed, hybrid, and multicloud
Industry solutions
Migration
Networking
Observability and monitoring
Security
Storage
/
Console
English
Deutsch
Español
Español – América Latina
Français
Indonesia
Italiano
Português
Português – Brasil
עברית
中文 – 简体
中文 – 繁體
日本語
한국어
Sign in
Google Kubernetes Engine (GKE)
GKE AI/ML
Start free
Overview
Guides
Documentation
More
Overview
Guides
Console
Discover
Introduction to AI/ML workloads on GKE
Explore GKE documentation
Overview
Main GKE documentation
GKE AI/ML documentation
GKE networking documentation
GKE security documentation
GKE fleet management documentation
Select how to obtain and consume accelerators on GKE
Design for resource obtainability with Gemini
GKE AI/ML conformance
Get started
Why use GKE for AI/ML inference
Simplified autoscaling concepts for AI/ML workloads in GKE
Quickstart: Serve your first AI model on GKE
Serve AI models for inference
About AI/ML model inference on GKE
Analyze model serving performance and costs with GKE Inference Quickstart
Expose AI applications with GKE Inference Gateway
Best practices for inference
Overview
Choose a load balancing strategy for inference
Autoscale inference workloads on GPUs
Autoscale LLM inference workloads on TPUs
Optimize LLM inference workloads on GPUs
Optimize batch inference workloads
Try inference examples
GPUs
Serve Gemma open models using GPUs with vLLM
Serve LLMs like DeepSeek-R1 671B or Llama 3.1 405B
Serve an LLM with GKE Inference Gateway
Serve an LLM with multiple GPUs
Serve T5 with Torch Serve
TPUs
Serve Llama on TPUs with vLLM
Serve LLMs using multi-host TPUs with JetStream and Pathways
Serve Stable Diffusion XL on TPUs with MaxDiffusion
Serve open models on TPUs with Terraform
Train AI models at scale
Train large-scale models with Multi-tier Checkpointing
Try training examples
Train a model with GPUs on GKE Standard mode
Train a model with GPUs on GKE Autopilot mode
Train a Llama model on GPUs with Megatron-LM
Fine-tune a LLM using TPUs on GKE with JAX
Run reinforcement learning workloads on GKE
Fine-tune and scale reinforcement learning with verl
Fine-tune and scale reinforcement learning with NeMo RL
Monitor reinforcement learning workloads with OpenTelemetry
Deploy and orchestrate AI agents
Deploy AI agents with the ADK and Agent Platform API
Deploy AI agents with the ADK and a self-hosted LLM
Scale AI agentic workloads with Agent Substrate
About Agent Substrate
Install Agent Substrate on GKE
Secure AI agentic workloads with Agent Sandbox
About Agent Sandbox
Choose storage for AI agentic workloads
Enable Agent Sandbox on GKE
Isolate AI code execution with Agent Sandbox
Manage Agent Sandbox storage
Save and restore Agent Sandbox environments
Trigger Agent Sandbox snapshots from inside a cluster
Use Ray for distributed AI/ML applications
About Ray on GKE
Quickstart: Deploy your first Ray application on GKE
Enable managed KubeRay with the Ray Operator add-on
Deploy AI/ML workloads with Ray on GKE
Serve AI models with Ray Serve on GKE
Serve an LLM on GPUs with Ray Serve
Serve an LLM with multi-cluster Ray Serve and Inference Gateway
Serve an LLM on TPUs with Ray Serve
Serve Gemma open models using multi-host TPUs on GKE with Ray
Serve a diffusion model on GPUs with Ray
Serve a diffusion model on TPUs with Ray
Train AI models with Ray on GKE
Train with PyTorch, Ray, and GKE
Train an LLM using Jax and Ray Train on TPUs with GKE
Multislice and Elastic Training on TPUs using Ray Train on GKE
Share compute between teams with Kueue
Start RayJobs faster across multiple compute options
Use GPUs with Ray on GKE
Set up Ray on GKE with A4X and GB200
Use TPUs with Ray on GKE
Set up Ray on GKE with TPU Trillium
Optimize AI training on TPUs with DWS and Kueue
Monitor Ray on GKE
View logs for the Ray Operator on GKE
View logs and metrics for Ray clusters on GKE
Debug completed Ray Jobs with Ray History Server
Use Slurm Operator for HPC and AI/ML workloads
About Slurm on GKE
Quickstart: Deploy a Slurm cluster on GKE
Enable the Slurm Operator add-on
Build custom Slurm Docker images
Autoscale a Slurm cluster using metrics
Configure shared storage for Slurm on GKE
Configure Filestore for Slurm
Managed Lustre for Slurm
Manage GPUs on GKE
About GPUs
Configure GPUs for AI/ML workloads
Configure A3 or A4 VMs with AI Hypercomputer
Deploy GPU workloads in Standard clusters
Deploy GPU workloads in Autopilot clusters
Configure autoscaling for LLM workloads on GPUs
Manage the GPU stack with the NVIDIA GPU Operator
Encrypt GPU workloads in-place
Provision resources for high performance computing (HPC)
Deploy AI Hypercomputer clusters (A3, A4 VMs)
Optimize GPU utilization with GPU sharing strategies
About GPU sharing strategies on GKE
Use multi-instance GPUs
Configure timesharing GPUs
Use NVIDIA MPS
Optimize GPU provisioning with flex-start
Overview
Run a large-scale workload with flex-start
Run a small batch workload with GPUs and flex-start
Manage TPUs on GKE
About TPUs
Versions
Ironwood (TPU7x)
Plan TPUs on GKE
Configure TPUs for AI/ML workloads
Deploy TPU workloads on Standard clusters
Deploy TPU workloads on Autopilot clusters
Deploy high-performance TPU workloads with auto-networking
Deploy TPU Multislices on GKE
Orchestrate TPU Multislice workloads using JobSet and Kueue
Configure autoscaling for LLM workloads on TPUs
Optimize TPU provisioning with flex-start
Overview
Run a small batch workload with TPUs and flex-start
Request TPUs with future reservation in calendar mode
Optimize TPU utilization with dynamic slicing
About TPU All Capacity mode