Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang’s inference runtime where the core request path meets deployment pressure.
This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, distributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.
The book is written for engineers, researchers, and advanced students who want to understand how modern LLM serving systems preserve correctness while changing representation, placement, execution, and measurement.
This is not a command reference or quick-start guide. It is a deep technical reading of the machinery behind high-throughput inference serving. Each chapter explains the algorithmic or systems problem first, then ties SGLang-specific claims to source files, tests, benchmark artifacts, and operational invariants.
Volume II assumes familiarity with the core serving loop from Volume I: admission, prefill and decode scheduling, KV storage, forward execution, sampling, streaming, and cleanup. Readers with GPU systems and transformer inference background can also use the opening preliminaries bridge as a compact orientation before entering the specialized chapters.
"Sinopsis" puede pertenecer a otra edición de este libro.
Librería: California Books, Miami, FL, Estados Unidos de America
Condición: New. Print on Demand. Nº de ref. del artículo: I-9798183580419
Cantidad disponible: Más de 20 disponibles
Librería: PBShop.store US, Wood Dale, IL, Estados Unidos de America
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000. Nº de ref. del artículo: L2-9798183580419
Cantidad disponible: Más de 20 disponibles
Librería: PBShop.store UK, Fairford, GLOS, Reino Unido
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000. Nº de ref. del artículo: L2-9798183580419
Cantidad disponible: Más de 20 disponibles
Librería: CitiRetail, Stevenage, Reino Unido
Paperback. Condición: new. Paperback. Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang's inference runtime where the core request path meets deployment pressure.This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, distributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.The book is written for engineers, researchers, and advanced students who want to understand how modern LLM serving systems preserve correctness while changing representation, placement, execution, and measurement.How rows stay aligned across optimized execution pathsHow KV state moves safely between runtime componentsHow kernels become eligible or ineligible for a request stepHow distributed workers coordinate placement and visibilityHow benchmark claims become comparableHow extensions can be evaluated without breaking hidden runtime contractsThis is not a command reference or quick-start guide. It is a deep technical reading of the machinery behind high-throughput inference serving. Each chapter explains the algorithmic or systems problem first, then ties SGLang-specific claims to source files, tests, benchmark artifacts, and operational invariants.Volume II assumes familiarity with the core serving loop from Volume I: admission, prefill and decode scheduling, KV storage, forward execution, sampling, streaming, and cleanup. Readers with GPU systems and transformer inference background can also use the opening preliminaries bridge as a compact orientation before entering the specialized chapters. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. Nº de ref. del artículo: 9798183580419
Cantidad disponible: 1 disponibles
Librería: AHA-BUCH GmbH, Einbeck, Alemania
Taschenbuch. Condición: Neu. Neuware - Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang's inference runtime where the core request path meets deployment pressure.This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, distributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.The book is written for engineers, researchers, and advanced students who want to understand how modern LLM serving systems preserve correctness while changing representation, placement, execution, and measurement.- How rows stay aligned across optimized execution paths- How KV state moves safely between runtime components- How kernels become eligible or ineligible for a request step- How distributed workers coordinate placement and visibility- How benchmark claims become comparable- How extensions can be evaluated without breaking hidden runtime contractsThis is not a command reference or quick-start guide. It is a deep technical reading of the machinery behind high-throughput inference serving. Each chapter explains the algorithmic or systems problem first, then ties SGLang-specific claims to source files, tests, benchmark artifacts, and operational invariants.Volume II assumes familiarity with the core serving loop from Volume I: admission, prefill and decode scheduling, KV storage, forward execution, sampling, streaming, and cleanup. Readers with GPU systems and transformer inference background can also use the opening preliminaries bridge as a compact orientation before entering the specialized chapters. Nº de ref. del artículo: 9798183580419
Cantidad disponible: 2 disponibles