Local LLM Inference Optimization (Paperback)

Idioma: inglés

Editorial: Independently Published, 2026

9798258375193

  • Tapa blanda
  • Nuevo
Ver todos los detalles

Librería: Grand Eagle Retail, Bensenville, IL, Estados Unidos de AmericaGrand Eagle Retail

Vendedor de 5 estrellas

Vendedor de AbeBooks desde 12 de octubre de 2005

Ver los artículos de este vendedor
Tapa blanda

Condición: Nuevo

EUR 19,72

 Gastos de envío gratis 
Se envía dentro de Estados Unidos de America

Cantidad disponible: 1 disponibles

Añadir al carrito
Devoluciones gratuitas de 30 días

Descripción del artículo del vendedor

Paperback. Stop Renting Intelligence. Start Optimizing Your Own.Do you want to run 70B parameter models on a single consumer GPU? Are you tired of high API costs, network latency, and the privacy risks of cloud-based AI?The "Local LLM Revolution" is here, but running Large Language Models (LLMs) privately is only half the battle. To make them truly useful, you must master Inference Optimization.In Local LLM Inference Optimization, you will move beyond basic "out-of-the-box" setups and dive into the high-performance engineering required to squeeze every drop of power from your hardware. Whether you are using NVIDIA CUDA, Apple Silicon (MLX), or AMD ROCm, this comprehensive guide provides the technical blueprint for the sovereign engineer. What You Will Master: The Quantization Deep-Dive: Learn to navigate the "Quantization Tax" using GGUF, EXL2, AWQ, and GPTQ. Move from FP32 to 4-bit and even 1.58-bit (BitNet) without losing the model's "mind."Advanced Memory Management: Defeat "Out of Memory" (OOM) errors by mastering KV Cache Management, PagedAttention, and FlashAttention 2 & 3.The Speed Multipliers: Double your Tokens Per Second (TPS) using Speculative Decoding, Continuous Batching, and Lookahead Heuristics.Hardware Architecture: Architect high-performance local servers using Multi-GPU Pipeline Parallelism and CPU/GPU offloading strategies.Context Window Expansion: Use RoPE Scaling, YaRN, and LongRoPE to push 8k models to 128k+ context on consumer hardware.The Full Local Stack: Step-by-step guides for Llama.cpp, Ollama, vLLM, and TGI (Text Generation Inference).Security & Privacy: Deploy Air-Gapped AI environments and secure your infrastructure using Safetensors and local sandboxing.Why This Book?This book focuses on Deployment and Efficiency. It is written for the Lead Engineer, the Privacy-Conscious CTO, and the Prosumer Hobbyist who demands low Time to First Token (TTFT) and maximum Perf/Watt.Stop paying for tokens. Own your weights. Optimize your future. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.

N° de ref. del artículo 9798258375193

Título
Local LLM Inference Optimization (Paperback)
Autor
Thomas O. Greene
Editorial
Independently Published
Año de publicación
2026
Estado
new
Encuadernación
Paperback
Idioma
inglés
ISBN 13
9798258375193

Grand Eagle Retail

Bensenville, IL, Estados Unidos de America

Vendedor de 5 estrellas

Vendedor de AbeBooks desde 12 de octubre de 2005

Tarifas de envío en Estados Unidos de America

ArtículoDe 6 a 14 días hábilesDe 6 a 16 días hábiles
Primer artículoEUR 0,00EUR 0,00
Los plazos de entrega los establecen los vendedores y varían según el transportista y la ubicación. Los pedidos que pasan por la aduana pueden sufrir retrasos y los compradores son responsables de los aranceles o tarifas asociadas. Los vendedores pueden ponerse en contacto con usted en relación con cargos adicionales para cubrir cualquier aumento en los costes de envío de los artículos.

Métodos de pago

  • Visa
  • Mastercard
  • American Express
  • Carte Bleue
  • Apple Pay
  • Google Pay

Información empresarial del vendedor

APOLLO ONLINE CORP.

605 Geddes Street
Wilmington, DE Estados Unidos de America 19805