Quit Emailing Yourself

# optimization → vllm

3 links tagged with all of: optimization + vllm

Click any tag below to further narrow down your results

Links

How prompt caching works - Paged Attention and Automatic Prefix Caching plus practical tips

This article explains how prompt caching works in large language models, focusing on techniques like paged attention and KV-cache reuse. It offers practical tips for improving cache hits to enhance performance and reduce costs in API usage.

Saved by tldr-importer · Last saved February 14, 2026 · 7 min read

+ prompt-caching + kv-cache + inference optimization ✓ vllm ✓

[no-title]

The article provides an in-depth exploration of the process involved in handling inference requests using the VLLM framework. It details the steps from receiving a request to processing it efficiently, emphasizing the benefits of utilizing VLLM for machine learning applications. Key aspects include optimizing performance and resource management during inference tasks.

Saved by tldr-importer · Last saved October 29, 2025 · 1 min read

+ inference vllm ✓ + machine-learning optimization ✓ + performance

Disaggregated Inference at Scale with PyTorch & vLLM

PyTorch and vLLM have been integrated to enhance generative AI applications by implementing Prefill/Decode Disaggregation, which improves inference efficiency at scale. This collaboration has optimized Meta's internal inference stack by allowing independent scaling of prefill and decode processes, resulting in better performance metrics. Key optimizations include enhanced KV cache transfer and load balancing, ultimately leading to reduced latency and increased throughput.

Saved by tldr-importer · Last saved October 29, 2025 · 6 min read

+ pytorch vllm ✓ + generative-ai + inference optimization ✓