LLM Inference Serving: Survey of Recent Advances and Opportunities

July 17, 2024 Β· Declared Dead Β· πŸ› IEEE Conference on High Performance Extreme Computing

πŸ‘» CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Baolin Li, Yankai Jiang, Vijay Gadepally, Devesh Tiwari arXiv ID 2407.12391 Category cs.DC: Distributed Computing Cross-listed cs.AI Citations 47 Venue IEEE Conference on High Performance Extreme Computing Last Checked 6 months ago
Abstract
This survey offers a comprehensive overview of recent advancements in Large Language Model (LLM) serving systems, focusing on research since the year 2023. We specifically examine system-level enhancements that improve performance and efficiency without altering the core LLM decoding mechanisms. By selecting and reviewing high-quality papers from prestigious ML and system venues, we highlight key innovations and practical considerations for deploying and scaling LLMs in real-world production environments. This survey serves as a valuable resource for LLM practitioners seeking to stay abreast of the latest developments in this rapidly evolving field.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Distributed Computing

Died the same way β€” πŸ‘» Ghosted