Isbn: 9798258375193 - local llm inference optimization: a comprehensive guide to quantization, hardware acceleration, and efficient private ai deployment (9 resultados)

ISBN
Refinar con la Búsqueda avanzada

Filtrar la búsqueda

  • Libros (9)

a

Intervalo de precios personalizado (EUR)

a

  • Idioma: Inglés

    Editorial: Independently published, 2026

    9798258375193

    • Tapa blanda

    Librería: GreatBookPrices, Columbia, MD, Estados Unidos de AmericaGreatBookPrices

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 17,32

    Envío por EUR 2,30 
    Se envía dentro de Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    Condición: New.

  • Idioma: Inglés

    Editorial: Independently published, 2026

    9798258375193

    • Tapa blanda

    Librería: GreatBookPrices, Columbia, MD, Estados Unidos de AmericaGreatBookPrices

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Usado - Como Nuevo

    EUR 17,58

    Envío por EUR 2,30 
    Se envía dentro de Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    Condición: As New. Unread book in perfect condition.

  • Idioma: Inglés

    Editorial: Independently Published, 2026

    9798258375193

    • Tapa blanda

    Librería: PBShop.store UK, Fairford, GLOS, Reino UnidoPBShop.store UK

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 19,04

    Envío por EUR 3,84 
    Se envía de Reino Unido a Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.

  • Idioma: Inglés

    Editorial: Independently published, 2026

    9798258375193

    • Tapa blanda

    Librería: GreatBookPricesUK, Woodford Green, Reino UnidoGreatBookPricesUK

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 19,03

    Envío por EUR 17,49 
    Se envía de Reino Unido a Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    Condición: New.

  • Idioma: Inglés

    Editorial: Independently published, 2026

    9798258375193

    • Tapa blanda

    Librería: GreatBookPricesUK, Woodford Green, Reino UnidoGreatBookPricesUK

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Usado - Como Nuevo

    EUR 19,99

    Envío por EUR 17,49 
    Se envía de Reino Unido a Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    Condición: As New. Unread book in perfect condition.

  • Idioma: Inglés

    Editorial: Independently Published, 2026

    9798258375193

    • Tapa blanda
    • Impresión bajo demanda

    Librería: Grand Eagle Retail, Bensenville, IL, Estados Unidos de AmericaGrand Eagle Retail

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 19,69

     Gastos de envío gratis 
    Se envía dentro de Estados Unidos de America

    Cantidad disponible: 1 disponibles

    Paperback. Condición: new. Paperback. Stop Renting Intelligence. Start Optimizing Your Own.Do you want to run 70B parameter models on a single consumer GPU? Are you tired of high API costs, network latency, and the privacy risks of cloud-based AI?The "Local LLM Revolution" is here, but running Large Language Models (LLMs) privately is only half the battle. To make them truly useful, you must master Inference Optimization.In Local LLM Inference Optimization, you will move beyond basic "out-of-the-box" setups and dive into the high-performance engineering required to squeeze every drop of power from your hardware. Whether you are using NVIDIA CUDA, Apple Silicon (MLX), or AMD ROCm, this comprehensive guide provides the technical blueprint for the sovereign engineer. What You Will Master: The Quantization Deep-Dive: Learn to navigate the "Quantization Tax" using GGUF, EXL2, AWQ, and GPTQ. Move from FP32 to 4-bit and even 1.58-bit (BitNet) without losing the model's "mind."Advanced Memory Management: Defeat "Out of Memory" (OOM) errors by mastering KV Cache Management, PagedAttention, and FlashAttention 2 & 3.The Speed Multipliers: Double your Tokens Per Second (TPS) using Speculative Decoding, Continuous Batching, and Lookahead Heuristics.Hardware Architecture: Architect high-performance local servers using Multi-GPU Pipeline Parallelism and CPU/GPU offloading strategies.Context Window Expansion: Use RoPE Scaling, YaRN, and LongRoPE to push 8k models to 128k+ context on consumer hardware.The Full Local Stack: Step-by-step guides for Llama.cpp, Ollama, vLLM, and TGI (Text Generation Inference).Security & Privacy: Deploy Air-Gapped AI environments and secure your infrastructure using Safetensors and local sandboxing.Why This Book?This book focuses on Deployment and Efficiency. It is written for the Lead Engineer, the Privacy-Conscious CTO, and the Prosumer Hobbyist who demands low Time to First Token (TTFT) and maximum Perf/Watt.Stop paying for tokens. Own your weights. Optimize your future. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.

  • Idioma: Inglés

    Editorial: Independently published, 2026

    9798258375193

    • Tapa blanda
    • Impresión bajo demanda

    Librería: California Books, Miami, FL, Estados Unidos de AmericaCalifornia Books

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 19,77

     Gastos de envío gratis 
    Se envía dentro de Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    Condición: New. Print on Demand.

  • Idioma: Inglés

    Editorial: Independently Published, 2026

    9798258375193

    • Tapa blanda

    Librería: PBShop.store US, Wood Dale, IL, Estados Unidos de AmericaPBShop.store US

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 2027,39

     Gastos de envío gratis 
    Se envía dentro de Estados Unidos de America

    Cantidad disponible: Más de 20 disponibles

    PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.

  • Idioma: Inglés

    Editorial: Independently Published, 2026

    9798258375193

    • Tapa blanda
    • Impresión bajo demanda

    Librería: CitiRetail, Stevenage, Reino UnidoCitiRetail

    Vendedor de 5 estrellas
    Contactar con el vendedor

    Condición: Nuevo

    EUR 22,81

    Envío por EUR 43,15 
    Se envía de Reino Unido a Estados Unidos de America

    Cantidad disponible: 1 disponibles

    Paperback. Condición: new. Paperback. Stop Renting Intelligence. Start Optimizing Your Own.Do you want to run 70B parameter models on a single consumer GPU? Are you tired of high API costs, network latency, and the privacy risks of cloud-based AI?The "Local LLM Revolution" is here, but running Large Language Models (LLMs) privately is only half the battle. To make them truly useful, you must master Inference Optimization.In Local LLM Inference Optimization, you will move beyond basic "out-of-the-box" setups and dive into the high-performance engineering required to squeeze every drop of power from your hardware. Whether you are using NVIDIA CUDA, Apple Silicon (MLX), or AMD ROCm, this comprehensive guide provides the technical blueprint for the sovereign engineer. What You Will Master: The Quantization Deep-Dive: Learn to navigate the "Quantization Tax" using GGUF, EXL2, AWQ, and GPTQ. Move from FP32 to 4-bit and even 1.58-bit (BitNet) without losing the model's "mind."Advanced Memory Management: Defeat "Out of Memory" (OOM) errors by mastering KV Cache Management, PagedAttention, and FlashAttention 2 & 3.The Speed Multipliers: Double your Tokens Per Second (TPS) using Speculative Decoding, Continuous Batching, and Lookahead Heuristics.Hardware Architecture: Architect high-performance local servers using Multi-GPU Pipeline Parallelism and CPU/GPU offloading strategies.Context Window Expansion: Use RoPE Scaling, YaRN, and LongRoPE to push 8k models to 128k+ context on consumer hardware.The Full Local Stack: Step-by-step guides for Llama.cpp, Ollama, vLLM, and TGI (Text Generation Inference).Security & Privacy: Deploy Air-Gapped AI environments and secure your infrastructure using Safetensors and local sandboxing.Why This Book?This book focuses on Deployment and Efficiency. It is written for the Lead Engineer, the Privacy-Conscious CTO, and the Prosumer Hobbyist who demands low Time to First Token (TTFT) and maximum Perf/Watt.Stop paying for tokens. Own your weights. Optimize your future. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.