Isbn: 9798189579370 - fast and frugal llm apps: systematic cost and latency optimization with caching, routing, and model right-sizing: 10 (applied llm engineering series) (5 resultados)

Idioma: Inglés
Editorial: Amazon Digital Services LLC - Kdp, 2026
- Tapa blanda
Librería: PBShop.store US, Wood Dale, IL, Estados Unidos de AmericaPBShop.store US
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 14,24
Gastos de envío gratisSe envía dentro de Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.

Idioma: Inglés
Editorial: Amazon Digital Services LLC - Kdp, 2026
- Tapa blanda
Librería: PBShop.store UK, Fairford, GLOS, Reino UnidoPBShop.store UK
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 13,10
Envío por EUR 3,87Se envía de Reino Unido a Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.

Idioma: Inglés
Editorial: Amazon Digital Services LLC - Kdp Jul 2026, 2026
- Tapa blanda
Librería: AHA-BUCH GmbH, Einbeck, AlemaniaAHA-BUCH GmbH
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 13,16
Envío por EUR 35,00Se envía de Alemania a Estados Unidos de AmericaCantidad disponible: 2 disponibles
Taschenbuch. Condición: Neu. Neuware - Stop Letting LLM Inference Bills Drain Your Product's MarginsIn the modern AI landscape, model inference spend has skyrocketed to become the dominant line item for software teams. Many organizations attempt to optimize naively, resulting in a fragile, hard-to-maintain stack of half-finished techniques. Fast and Frugal LLM Apps replaces this chaotic pattern with a single, repeatable engineering framework designed to systematically slash costs and latency without sacrificing output quality.Written specifically for AI engineers, backend developers, engineering managers, and technical founders, this book provides a production-ready playbook to scale your AI systems sustainably. You will learn to treat cost as a first-class engineering metric alongside latency and accuracy.What you will master inside: - The Four-Phase Loop: Run a highly effective six-week optimization cadence with clear ownership.- Smart Caching Architecture: Implement prompt and semantic caching to slash input costs by 70-90%.- Dynamic Routing: Automatically classify requests to send them to the cheapest sufficient model tier.- Open-Source and Quantization: Deploy highly efficient open-weight models in hybrid architectures.- Cost-Aware Evaluations: Establish automated regression gates so optimizations never silently destroy quality.- Batch and Latency Engineering: Transition non-interactive tasks to batch APIs and implement smart streaming.Stop guessing your monthly API invoice. Build a structured cost-engineering practice with named owners, finance partnerships, and solid quarterly reviews. Scroll up, click 'Buy Now,' and transform your model invoice from an unpredictable crisis into a predictable engineering roadmap today. …

Idioma: Inglés
Editorial: Independently published, 2026
- Tapa blanda
- Impresión bajo demanda
Librería: California Books, Miami, FL, Estados Unidos de AmericaCalifornia Books
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 14,64
Gastos de envío gratisSe envía dentro de Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
Condición: New. Print on Demand.

Idioma: Inglés
Editorial: Independently Published, 2026
- Tapa blanda
- Impresión bajo demanda
Librería: CitiRetail, Stevenage, Reino UnidoCitiRetail
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 16,96
Envío por EUR 43,54Se envía de Reino Unido a Estados Unidos de AmericaCantidad disponible: 1 disponible
Paperback. Condición: new. Paperback. Stop Letting LLM Inference Bills Drain Your Product's MarginsIn the modern AI landscape, model inference spend has skyrocketed to become the dominant line item for software teams. Many organizations attempt to optimize naively, resulting in a fragile, hard-to-maintain stack of half-finished techniques. Fast and Frugal LLM Apps replaces this chaotic pattern with a single, repeatable engineering framework designed to systematically slash costs and latency without sacrificing output quality.Written specifically for AI engineers, backend developers, engineering managers, and technical founders, this book provides a production-ready playbook to scale your AI systems sustainably. You will learn to treat cost as a first-class engineering metric alongside latency and accuracy.What you will master inside: The Four-Phase Loop: Run a highly effective six-week optimization cadence with clear ownership.Smart Caching Architecture: Implement prompt and semantic caching to slash input costs by 70-90%.Dynamic Routing: Automatically classify requests to send them to the cheapest sufficient model tier.Open-Source and Quantization: Deploy highly efficient open-weight models in hybrid architectures.Cost-Aware Evaluations: Establish automated regression gates so optimizations never silently destroy quality.Batch and Latency Engineering: Transition non-interactive tasks to batch APIs and implement smart streaming.Stop guessing your monthly API invoice. Build a structured cost-engineering practice with named owners, finance partnerships, and solid quarterly reviews. Scroll up, click "Buy Now," and transform your model invoice from an unpredictable crisis into a predictable engineering roadmap today! This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. …