🔍

Google Cloud Professional Machine Learning Engineer PMLE — Question 238

Topic 1 · Question 238 of 339

Topic 1 · Question 238

You have deployed a scikit-team model to a Vertex AI endpoint using a custom model server. You enabled autoscaling: however, the deployed model fails to scale beyond one replica, which led to dropped requests. You notice that CPU utilization remains low even during periods of high load. What should you do?