Prefix-aware routing for lower LLM latency on SageMaker
Amazon introduced prefix-aware routing on SageMaker Inference to reduce LLM response latency, improving inference performance for real-time applications.