Speech AI and Multimodal Models with Nvidia Nemo (Paperback)

Idioma: inglés

Editorial: Independently Published, 2025

9798273025103

  • Tapa blanda
  • Nuevo
Ver todos los detalles

Librería: Grand Eagle Retail, Bensenville, IL, Estados Unidos de AmericaGrand Eagle Retail

Vendedor de 5 estrellas

Vendedor de AbeBooks desde 12 de octubre de 2005

Ver los artículos de este vendedor
Tapa blanda

Condición: Nuevo

EUR 37,98

 Gastos de envío gratis 
Se envía dentro de Estados Unidos de America

Cantidad disponible: 1 disponibles

Añadir al carrito
Devoluciones gratuitas de 30 días

Descripción del artículo del vendedor

Paperback. Build dependable speech and multimodal systems from data to deployment with NeMo, Riva, Triton, and NIM.Shipping ASR, TTS, and vision language features is hard because real traffic, latency budgets, and safety rules punish vague guidance. Teams need a concrete stack, tested workflows, and playbooks that hold up under load.This book gives practitioners a practical path. Train with NeMo, serve with Triton and Riva, package stable APIs with NIM, and wire observability, safety, and rollout controls so your services stay reliable after launch.Map the NVIDIA stack in production, NeMo for training, Riva for runtime, NIM for standard APIs, Triton for serving and metricsSet up containers, GPU drivers, CUDA, and validation checks for a clean starting environmentBuild NeMo manifests, create tarred WebDataset shards, and manage data versions for repeatable trainingApply text processing that works in products, PnC models for punctuation and case, grammar based ITN with SparrowhawkChoose and justify architectures, CTC and RNNT tradeoffs, FastConformer for short and long speech, Parakeet for multilingual, Canary for translation and timestampsDesign streaming with intent, lookahead, chunk size, and padding choices that balance latency and accuracyRun NeMo 2 configs and NeMo Run cleanly, migrate experiments, track ablations, and keep results comparableEvaluate with WER, CER, MER, and slice by accent, SNR, and channel so quality numbers reflect realityAdd diarization that operators can trust, VAD with MarbleNet, embeddings with TitaNet, and MSDD integrationExport for serving the right way, ONNX or TorchScript paths, TensorRT where appropriate, and Triton model repos that scaleTune Riva streaming ASR, chunk and padding settings, punctuation and ITN options, diarization flags and limitsStand up NIM ASR endpoints with an OpenAI compatible surface and autoscale them with Helm on KubernetesBuild TTS that sounds right and runs fast, FastPitch with HiFi GAN or BigVGAN, voice cloning data, lexicons, SSML controlsManage prosody and latency for streaming audio, set clause sizes and playback buffers that feel responsiveProtect your product, content safeguards in TTS, consent gates for data and cloning, redaction and retention policiesMeasure what matters, Triton metrics in Prometheus and Grafana, practical alert rules that catch real issuesLoad test with perf analyzer sweeps, batch and concurrency tuning, sequence batching for conversational trafficEngineer reliability, fault injection and backpressure, graceful degradation under spikes and partial failuresWire NeMo Guardrails around ASR, TTS, and VLM flows so outputs stay on policyWatermark and detect audio with AudioSeal and formalize a detection pipelineUnderstand licenses and terms, NVIDIA AI Enterprise scope, Riva EULA, and NGC usage expectationsUse production playbooks with SLOs, cost caps, and rollback guards that turn operations into repeatable stepsThis is a code heavy guide with working Python, YAML, JSON, and Shell examples that you can adapt directly into real services.Get the guide and build systems your users can rely on. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.

N° de ref. del artículo 9798273025103

Título
Speech AI and Multimodal Models with Nvidia Nemo (Paperback)
Autor
Ansel Corbyn
Editorial
Independently Published
Año de publicación
2025
Estado
new
Encuadernación
Paperback
Idioma
inglés
ISBN 13
9798273025103

Grand Eagle Retail

Bensenville, IL, Estados Unidos de America

Vendedor de 5 estrellas

Vendedor de AbeBooks desde 12 de octubre de 2005

Tarifas de envío en Estados Unidos de America

ArtículoDe 6 a 14 días hábilesDe 6 a 16 días hábiles
Primer artículoEUR 0,00EUR 0,00
Los plazos de entrega los establecen los vendedores y varían según el transportista y la ubicación. Los pedidos que pasan por la aduana pueden sufrir retrasos y los compradores son responsables de los aranceles o tarifas asociadas. Los vendedores pueden ponerse en contacto con usted en relación con cargos adicionales para cubrir cualquier aumento en los costes de envío de los artículos.

Métodos de pago

  • Visa
  • Mastercard
  • American Express
  • Carte Bleue
  • Apple Pay
  • Google Pay

Información empresarial del vendedor

APOLLO ONLINE CORP.

605 Geddes Street
Wilmington, DE Estados Unidos de America 19805