Make your large language models faster, cheaper, and ready for production.
Training a model happens once. Serving it happens every time someone uses your product, and that is where most AI budgets go. This book explains, step by step, what really happens when an LLM generates text and how engineers make it fast and affordable at scale.
Starting from a single question, "what happens when a model produces one token?", you will build a complete mental model of modern inference, from first principles to production deployment.
Software and ML engineers, platform and MLOps engineers, solution architects, technical product managers, and anyone preparing for AI infrastructure interviews. You need basic Python and a rough idea of what a neural network is. No CUDA or GPU required.
Written by an AI architect and former Amazon Web Services engineer who has built LLM applications serving millions of customers.
"Sinopsis" puede pertenecer a otra edición de este libro.
Librería: California Books, Miami, FL, Estados Unidos de America
Condición: New. Print on Demand. Nº de ref. del artículo: I-9798178522462
Cantidad disponible: Más de 20 disponibles
Librería: PBShop.store US, Wood Dale, IL, Estados Unidos de America
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000. Nº de ref. del artículo: L2-9798178522462
Cantidad disponible: Más de 20 disponibles
Librería: PBShop.store UK, Fairford, GLOS, Reino Unido
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000. Nº de ref. del artículo: L2-9798178522462
Cantidad disponible: Más de 20 disponibles
Librería: AHA-BUCH GmbH, Einbeck, Alemania
Taschenbuch. Condición: Neu. Neuware. Nº de ref. del artículo: 9798178522462
Cantidad disponible: 2 disponibles