Scheduling GPUs on Kubernetes, optimizing for inference workloads, and running GPU operators in production.