9798195802172 - the local ai performance handbook: optimizing ollama for multi-gpu and hardware acceleration: 4 (architecting enterprise agents series) de tyson, ethan (5 resultados)
Idioma: Inglés
Editorial: Independently Published, 2026
- Tapa blanda
Librería: PBShop.store US, Wood Dale, IL, Estados Unidos de AmericaPBShop.store US
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 22,11
Gastos de envío gratisSe envía dentro de Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.
Idioma: Inglés
Editorial: Independently Published, 2026
- Tapa blanda
Librería: PBShop.store UK, Fairford, GLOS, Reino UnidoPBShop.store UK
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 21,08
Envío por EUR 3,84Se envía de Reino Unido a Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.
- Más imágenes
Idioma: Inglés
Editorial: Createspace Independent Publishing Platform Mai 2026, 2026
- Tapa blanda
Librería: AHA-BUCH GmbH, Einbeck, AlemaniaAHA-BUCH GmbH
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 26,00
Envío por EUR 61,25Se envía de Alemania a Estados Unidos de AmericaCantidad disponible: 2 disponibles
Taschenbuch. Condición: Neu. Neuware - The Local AI Performance Handbook: Optimizing Ollama for Multi-GPU and Hardware AccelerationLocal AI is powerful, but poor configuration can turn expensive hardware into a slow, unstable bottleneck. If your Ollama setup struggles with VRAM limits, weak token throughput, GPU underuse, long c…ontext slowdowns, or unreliable multi-user workloads, this handbook gives you the practical performance playbook you need.The Local AI Performance Handbook is a technical guide to building faster, more private, and more reliable Ollama systems across NVIDIA CUDA, AMD ROCm, Apple Silicon, WSL2, Docker, Kubernetes, and multi-GPU environments. It moves beyond basic local model setup and focuses on the engineering details that determine real-world performance: hardware acceleration, VRAM planning, quantization, request concurrency, private RAG, secure deployment, benchmarking, and production maintenance. The book's scope is reflected in its coverage of hardware-specific runtimes, memory engineering, multi-GPU scheduling, quantization, high-concurrency handling, private RAG, deployment, agentic workflows, and troubleshooting.Inside, readers will learn how to: - Configure Ollama for CUDA, ROCm, Apple Silicon, Vulkan, Docker, and WSL2.- Calculate model memory footprints and avoid out-of-memory failures.- Tune VRAM usage, KV cache behavior, context windows, and quantization choices.- Scale Ollama across multiple GPUs and isolate workloads with resource controls.- Benchmark tokens per second, latency, GPU utilization, and system bottlenecks.- Deploy private AI inference with Docker Compose, Kubernetes, health checks, and secure API access.- Build faster private RAG and local agent workflows without depending on cloud APIs.For developers, AI engineers, homelab builders, and technical teams serious about private AI performance, this book turns Ollama from a simple local model runner into a tuned inference platform.
Idioma: Inglés
Editorial: Independently published, 2026
- Tapa blanda
- Impresión bajo demanda
Librería: California Books, Miami, FL, Estados Unidos de AmericaCalifornia Books
Contactar con el vendedorVendedor de 4 estrellasCondición: Nuevo
EUR 21,43
Gastos de envío gratisSe envía dentro de Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
Condición: New. Print on Demand.
Idioma: Inglés
Editorial: Independently Published, 2026
- Tapa blanda
- Impresión bajo demanda
Librería: CitiRetail, Stevenage, Reino UnidoCitiRetail
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 24,66
Envío por EUR 43,24Se envía de Reino Unido a Estados Unidos de AmericaCantidad disponible: 1 disponibles
Paperback. Condición: new. Paperback. The Local AI Performance Handbook: Optimizing Ollama for Multi-GPU and Hardware AccelerationLocal AI is powerful, but poor configuration can turn expensive hardware into a slow, unstable bottleneck. If your Ollama setup struggles with VRAM limits, weak token throughput, GPU underuse, long co…ntext slowdowns, or unreliable multi-user workloads, this handbook gives you the practical performance playbook you need.The Local AI Performance Handbook is a technical guide to building faster, more private, and more reliable Ollama systems across NVIDIA CUDA, AMD ROCm, Apple Silicon, WSL2, Docker, Kubernetes, and multi-GPU environments. It moves beyond basic local model setup and focuses on the engineering details that determine real-world performance: hardware acceleration, VRAM planning, quantization, request concurrency, private RAG, secure deployment, benchmarking, and production maintenance. The book's scope is reflected in its coverage of hardware-specific runtimes, memory engineering, multi-GPU scheduling, quantization, high-concurrency handling, private RAG, deployment, agentic workflows, and troubleshooting.Inside, readers will learn how to: Configure Ollama for CUDA, ROCm, Apple Silicon, Vulkan, Docker, and WSL2.Calculate model memory footprints and avoid out-of-memory failures.Tune VRAM usage, KV cache behavior, context windows, and quantization choices.Scale Ollama across multiple GPUs and isolate workloads with resource controls.Benchmark tokens per second, latency, GPU utilization, and system bottlenecks.Deploy private AI inference with Docker Compose, Kubernetes, health checks, and secure API access.Build faster private RAG and local agent workflows without depending on cloud APIs.For developers, AI engineers, homelab builders, and technical teams serious about private AI performance, this book turns Ollama from a simple local model runner into a tuned inference platform. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.

