Articles

Blog post

Optimizing AI Workloads with NVIDIA GPUs, Time Slicing, and Karpenter (Part 2)

Introduction: Overcoming GPU Management Challenges   In Part 1 of this blog series, we explored the challenges of hosting large language models (LLMs) on CPU-based workloads within an EKS cluster. We discussed the inefficiencies associated with using CPUs for such tasks, primarily due to the large model sizes and slower inference speeds. The introduction of GPU […]

January 22, 2025 • 8 min read