Gpu kernel engineering llm de team chatvariety (4 resultados)

Idioma: Inglés
Editorial: Independently published, 2026
Serie: Libro 6 de 6 - AI Infrastructure, Hardware & Compiler Engineering Series
- Tapa blanda
Librería: PBShop.store US, Wood Dale, IL, Estados Unidos de AmericaPBShop.store US
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 13,79
Gastos de envío gratisSe envía dentro de Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.

Idioma: Inglés
Editorial: Independently published, 2026
Serie: Libro 6 de 6 - AI Infrastructure, Hardware & Compiler Engineering Series
- Tapa blanda
Librería: PBShop.store UK, Fairford, GLOS, Reino UnidoPBShop.store UK
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 13,48
Envío por EUR 3,84Se envía de Reino Unido a Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
PAP. Condición: New. New Book. Shipped from UK. Established seller since 2000.

Idioma: Inglés
Editorial: Independently published, 2026
Serie: Libro 6 de 6 - AI Infrastructure, Hardware & Compiler Engineering Series
- Tapa blanda
- Impresión bajo demanda
Librería: California Books, Miami, FL, Estados Unidos de AmericaCalifornia Books
Contactar con el vendedorVendedor de 4 estrellasCondición: Nuevo
EUR 13,22
Gastos de envío gratisSe envía dentro de Estados Unidos de AmericaCantidad disponible: Más de 20 disponibles
Condición: New. Print on Demand.

Idioma: Inglés
Editorial: Independently Published, 2026
Serie: Libro 6 de 6 - AI Infrastructure, Hardware & Compiler Engineering Series
- Tapa blanda
- Impresión bajo demanda
Librería: CitiRetail, Stevenage, Reino UnidoCitiRetail
Contactar con el vendedorVendedor de 5 estrellasCondición: Nuevo
EUR 16,82
Envío por EUR 43,19Se envía de Reino Unido a Estados Unidos de AmericaCantidad disponible: 1 disponibles
Paperback. Condición: new. Paperback. Master GPU Kernel Engineering for Production LLM InferenceIn the high-stakes world of AI production, the bottleneck isn't model architecture-it's execution efficiency. GPU Kernel Engineering for LLM Inference is the definitive engineering guide designed to close the gap between framework-lev…el PyTorch code and hardware-optimal hardware performance.Written specifically for machine learning engineers, systems software developers, and performance engineers, this practical manual provides deep, production-ready blueprints for building high-throughput AI serving infrastructure. You will move past theoretical concepts and dive straight into raw hardware-level optimizations that top AI teams use to slash latency and reduce infrastructure costs.What You Will Master: Custom CUDA Kernels: Write and optimize high-performance CUDA C++ kernels tailored specifically for modern Transformer workloads.Flash Attention 2 & 3: Implement advanced IO-aware attention algorithms with Hopper-specific asynchronous memory and Tensor Core operations.Triton Development: Build optimized fusion kernels for layer normalization, activation functions, and positional encodings using Python-level Triton.Quantization & GEMM: Develop high-throughput INT8, INT4, and weight-only quantized GEMM kernels for optimized memory footprints.PagedAttention & KV-Cache: Design vLLM-style virtual memory management kernels to eliminate memory fragmentation.Multi-GPU Scaling: Coordinate tensor-parallel all-reduce operations with custom NCCL collectives to scale seamlessly across cluster nodes.Nsight Profiling: Locate hardware bottlenecks using Nsight Systems and Nsight Compute.Stop relying on out-of-the-box configurations. Equip yourself with the systems engineering skills required to build the next generation of high-speed, cost-efficient AI infrastructure. Learn to build kernels that run at the absolute physical limits of NVIDIA silicon. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.