LLM Inference in C++ (Paperback)
Idioma: inglés
Editorial: Independently Published, 2026
- Tapa blanda
- Nuevo

Librería: CitiRetail, Stevenage, Reino UnidoCitiRetail
Vendedor de IberLibro desde 29 de junio de 2022
Condición: Nuevo
EUR 34,52
Cantidad disponible: 1 disponible
Añadir al carritoDescripción del artículo del vendedor
Paperback. Stop Wasting GPU Compute. Build the High-Throughput, Low-Latency AI Infrastructure of 2026. The "VRAM Wall" is the biggest bottleneck in modern AI. Standard Python wrappers and out-of-the-box runtimes are fine for prototyping, but at scale, memory fragmentation and Global Interpreter Lock (GIL) overhead will destroy your throughput. LLM Inference in C++ is the definitive engineering manual for bypassing Python entirely and building custom, bare-metal inference engines that maximize hardware utilization. Focusing on the cutting-edge 2026 landscape, this book bridges the gap between high-level AI concepts and low-level GPU execution. You will learn how to implement enterprise-grade features like PagedAttention, FlashAttention-3, and Continuous Batching directly in C++ and CUDA, unlocking massive performance gains for large-scale language models.Inside, you will discover: Hardware-Aware Memory Management: Eliminate memory waste by implementing PagedAttention logic and custom allocators to bypass std:: malloc overhead.Accelerated Tensor Algebra: Master C++23's std:: mdspan and write fused SIMD kernels with AVX-512 to minimize GPU context switching.Custom CUDA Kernels: Write high-speed FlashAttention-3, LayerNorm, and RMSNorm kernels while managing CUDA streams for maximum GPU occupancy.The Cost Killer (Quantization): Slash VRAM requirements with bit-level manipulation for 4-bit (AWQ) and 8-bit (FP8) inference using NVIDIA Tensor Cores.Distributed & Speculative Execution: Scale across clusters using zero-copy NCCL/RDMA interconnects and implement Draft Models to accelerate massive architectures.The Production Serving Layer: Build lock-free C++ request queues for continuous batching and track P99 "Time to First Token" (TTFT) at the systems level.THE IMPLEMENTATION VAULT (Appendix) Built for the infrastructure engineer in the trenches, the Appendix provides immediate, battle-tested utility: The 15-Point Production-Ready Checklist: Your mandatory safety and performance audit before deploying any custom engine.Latency vs. Throughput Reference Table: The ultimate cheat sheet for balancing batch sizes against user wait times.Troubleshooting Guide: Direct solutions for the top 10 most common and devastating CUDA and C++ memory errors.Don't let inefficient software architecture throttle your hardware. Master C++ LLM inference and build the fastest, most cost-effective AI engines in the industry. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.…
N° de ref. del artículo 9798259069299
- Título
- LLM Inference in C++ (Paperback)
- Autor
- Billie S. Lightner
- Editorial
- Independently Published
- Año de publicación
- 2026
- Estado
- new
- Encuadernación
- Paperback
- Idioma
- inglés
- ISBN 13
- 9798259069299
- Serie
- Libro 7 de 11: High-Performance C++ Engineering
The "VRAM Wall" is the biggest bottleneck in modern AI. Standard Python wrappers and out-of-the-box runtimes are fine for prototyping, but at scale, memory fragmentation and Global Interpreter Lock (GIL) overhead will destroy your throughput. LLM Inference in C++ is the definitive engineering manual for bypassing Python entirely and building custom, bare-metal inference engines that maximize hardware utilization.
Focusing on the cutting-edge 2026 landscape, this book bridges the gap between high-level AI concepts and low-level GPU execution. You will learn how to implement enterprise-grade features like PagedAttention, FlashAttention-3, and Continuous Batching directly in C++ and CUDA, unlocking massive performance gains for large-scale language models.
Inside, you will discover:
- Hardware-Aware Memory Management: Eliminate memory waste by implementing PagedAttention logic and custom allocators to bypass std::malloc overhead.
- Accelerated Tensor Algebra: Master C++23's std::mdspan and write fused SIMD kernels with AVX-512 to minimize GPU context switching.
- Custom CUDA Kernels: Write high-speed FlashAttention-3, LayerNorm, and RMSNorm kernels while managing CUDA streams for maximum GPU occupancy.
- The Cost Killer (Quantization): Slash VRAM requirements with bit-level manipulation for 4-bit (AWQ) and 8-bit (FP8) inference using NVIDIA Tensor Cores.
- Distributed & Speculative Execution: Scale across clusters using zero-copy NCCL/RDMA interconnects and implement Draft Models to accelerate massive architectures.
- The Production Serving Layer: Build lock-free C++ request queues for continuous batching and track P99 "Time to First Token" (TTFT) at the systems level.
Built for the infrastructure engineer in the trenches, the Appendix provides immediate, battle-tested utility:
- The 15-Point Production-Ready Checklist: Your mandatory safety and performance audit before deploying any custom engine.
- Latency vs. Throughput Reference Table: The ultimate cheat sheet for balancing batch sizes against user wait times.
- Troubleshooting Guide: Direct solutions for the top 10 most common and devastating CUDA and C++ memory errors.
“Sinopsis” puede pertenecer a otra edición de este título.
CitiRetail
Stevenage, Reino Unido
Vendedor de IberLibro desde 29 de junio de 2022
Tarifas de envío de Reino Unido a Estados Unidos de America
| Artículo | De 7 a 14 días hábiles | De 7 a 60 días hábiles |
|---|---|---|
| Primer artículo | EUR 43,54 | EUR 43,54 |
Métodos de pago
Descripción de la tienda
Online business
Información empresarial del vendedor
ABC BOOKS LIMITED
10 John Street
London, Reino Unido WC1N 2EB
Condiciones de venta
Orders can be returned within 30 days of receipt.
Derecho al desistimiento
Si es un consumidor, puede rescindir el contrato de acuerdo con lo siguiente. Por consumidor se entiende cualquier persona física que actúe con fines ajenos a su actividad comercial, empresarial, oficio o profesión.
Información sobre el derecho de desistimiento
Derecho legal de desistimiento
Tiene derecho a rescindir este contrato en un plazo de 14 días sin dar ningún motivo.
El periodo de desistimiento vencerá a los 14 días desde que usted, o un tercero que no sea el transportista e indicado por usted, adquiera la posesión física del último bien o del último lote o pieza.
Para ejercer el derecho de desistimiento, complete de forma electrónica y envíe una declaración clara en nuestro sitio web, desde "Mis compras" en "Mi cuenta". Le enviaremos sin demora un acuse de recibo de dicho desistimiento a través de un soporte duradero (por ejemplo, por correo electrónico).
Para cumplir con el plazo de desistimiento, basta con que envíe su comunicación relativa al ejercicio del derecho de desistimiento antes de que venza el periodo de desistimiento.
Efectos del desistimiento
Si rescinde este contrato, le reembolsaremos todos los pagos que hayamos recibido de usted, incluidos los gastos de envío (excepto los gastos adicionales que surjan si elige un tipo de envío que no sea el tipo de envío estándar más económico que ofrecemos).
Podemos hacer una deducción del reembolso por la pérdida de valor de cualquier bien suministrado, si la pérdida es el resultado de una manipulación innecesaria por su parte.
Efectuaremos el reembolso sin demoras indebidas y, a más tardar, 14 días después de que se nos informe de su decisión de rescindir este contrato.
Efectuaremos el reembolso utilizando el mismo medio de pago que utilizó para la transacción inicial, a menos que haya acordado expresamente lo contrario; en cualquier caso, no incurrirá en ningún cargo como resultado de dicho reembolso.
Podremos retener el reembolso hasta que hayamos recibido los bienes o hasta que nos haya presentado una prueba de que los ha devuelto, lo que ocurra primero.
Deberá devolver los bienes o entregarlos a CitiRetail, Stevenage, United Kingdom, sin demoras indebidas y, en cualquier caso, en un plazo máximo de 14 días a partir del día en que nos comunique su desistimiento del presente contrato. El plazo se cumple si devuelve la mercancía antes de que venza el periodo de 14 días. Tendrá que asumir los gastos directos de devolución de los bienes. Usted solo es responsable de la disminución del valor de los bienes como resultado de una manipulación distinta a la necesaria para establecer la naturaleza, las características y el funcionamiento de los bienes.
Excepciones al derecho de desistimiento
El derecho de desistimiento no se aplica a lo siguiente:
- La entrega de periódicos, diarios o revistas, con la excepción de los contratos de suscripción; y
- El suministro de contenido digital que no se proporcione en un soporte tangible (por ejemplo, en un CD o DVD) si, al hacer el pedido, aceptó que podíamos empezar a entregarlo y que no podría desistir una vez iniciada la entrega.
Condiciones de envío
Please note that titles are dispatched from our US, Canadian or Australian warehouses. Delivery times specified in shipping terms. Orders ship within 2 business days. Delivery to your door then takes 7-14 days.