AI Inference Optimization Engineering (Paperback)

Idioma: inglés

Editorial: Independently Published, 2026

9798199720021

Serie: Libro 6 de 21 - Production AI Engineering Series

  • Tapa blanda
  • Nuevo
Ver todos los detalles

Librería: CitiRetail, Stevenage, Reino UnidoCitiRetail

Vendedor de 5 estrellas

Vendedor de AbeBooks desde 29 de junio de 2022

Ver los artículos de este vendedor
Tapa blanda

Condición: Nuevo

EUR 16,76

Envío por EUR 43,04 
Se envía de Reino Unido a Estados Unidos de America

Cantidad disponible: 1 disponibles

Añadir al carrito
Devoluciones gratuitas de 30 días

Descripción del artículo del vendedor

Paperback. Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.

N° de ref. del artículo 9798199720021

Título
AI Inference Optimization Engineering (Paperback)
Autor
Chatvariety Team
Editorial
Independently Published
Año de publicación
2026
Estado
new
Encuadernación
Paperback
Idioma
inglés
ISBN 13
9798199720021
Serie
Libro 6 de 21: Production AI Engineering Series

CitiRetail

Stevenage, Reino Unido

Vendedor de 5 estrellas

Vendedor de AbeBooks desde 29 de junio de 2022

Tarifas de envío de Reino Unido a Estados Unidos de America

ArtículoDe 7 a 14 días hábilesDe 7 a 60 días hábiles
Primer artículoEUR 43,04EUR 43,04
Los plazos de entrega los establecen los vendedores y varían según el transportista y la ubicación. Los pedidos que pasan por la aduana pueden sufrir retrasos y los compradores son responsables de los aranceles o tarifas asociadas. Los vendedores pueden ponerse en contacto con usted en relación con cargos adicionales para cubrir cualquier aumento en los costes de envío de los artículos.

Métodos de pago

  • Visa
  • Mastercard
  • American Express
  • Carte Bleue
  • Apple Pay
  • Google Pay

Descripción de la tienda

Online business

Información empresarial del vendedor

ABC BOOKS LIMITED

10 John Street
London, Reino Unido WC1N 2EB