Hands-On Big Data Analytics with PySpark
Idioma: inglés
Editorial: Packt Publishing Limited, GB, 2019
- Tapa blanda
- Nuevo

Librería: Rarewaves.com UK, London, Reino UnidoRarewaves.com UK
Vendedor de AbeBooks desde el 11 de junio de 2025
Condición: Nuevo
EUR 31,09
Cantidad disponible: Más de 20 disponibles
Añadir al carritoDescripción del artículo del vendedor
Use PySpark to easily crush messy data at-scale and discover proven techniques to create testable, immutable, and easily parallelizable Spark jobsKey FeaturesWork with large amounts of agile data using distributed datasets and in-memory cachingSource data from all popular data hosting platforms, such as HDFS, Hive, JSON, and S3Employ the easy-to-use PySpark API to deploy big data Analytics for productionBook DescriptionApache Spark is an open source parallel-processing framework that has been around for quite some time now. One of the many uses of Apache Spark is for data analytics applications across clustered computers. In this book, you will not only learn how to use Spark and the Python API to create high-performance analytics with big data, but also discover techniques for testing, immunizing, and parallelizing Spark jobs.You will learn how to source data from all popular data hosting platforms, including HDFS, Hive, JSON, and S3, and deal with large datasets with PySpark to gain practical big data experience. This book will help you work on prototypes on local machines and subsequently go on to handle messy data in production and at scale. This book covers installing and setting up PySpark, RDD operations, big data cleaning and wrangling, and aggregating and summarizing data into useful reports. You will also learn how to implement some practical and proven techniques to improve certain aspects of programming and administration in Apache Spark.By the end of the book, you will be able to build big data analytical solutions using the various PySpark offerings and also optimize them effectively.What you will learnGet practical big data experience while working on messy datasetsAnalyze patterns with Spark SQL to improve your business intelligenceUse PySpark s interactive shell to speed up development timeCreate highly concurrent Spark programs by leveraging immutabilityDiscover ways to avoid the most expensive operation in the Spark API: the shuffle operationRe-design your jobs to use reduceByKey instead of groupByCreate robust processing pipelines by testing Apache Spark jobsWho this book is forThis book is for developers, data scientists, business analysts, or anyone who needs to reliably analyze large amounts of large-scale, real-world data. Whether you're tasked with creating your company's business intelligence function or creating great data platforms for your machine learning models, or are looking to use code to magnify the impact of your business, this book is for you.…
N° de ref. del artículo LU-9781838644130
- Título
- Hands-On Big Data Analytics with PySpark
- Autor
- Rudy Lai, Bartlomiej Potaczek
- Editorial
- Packt Publishing Limited, GB
- Año de publicación
- 2019
- Estado
- New
- Encuadernación
- Paperback
- Idioma
- inglés
- ISBN 10
- 183864413X
- ISBN 13
- 9781838644130
Use PySpark to easily crush messy data at-scale and discover proven techniques to create testable, immutable, and easily parallelizable Spark jobs
Key Features:
- Work with large amounts of agile data using distributed datasets and in-memory caching
- Source data from all popular data hosting platforms, such as HDFS, Hive, JSON, and S3
- Employ the easy-to-use PySpark API to deploy big data Analytics for production
Book Description:
Apache Spark is an open source parallel-processing framework that has been around for quite some time now. One of the many uses of Apache Spark is for data analytics applications across clustered computers. In this book, you will not only learn how to use Spark and the Python API to create high-performance analytics with big data, but also discover techniques for testing, immunizing, and parallelizing Spark jobs.
You will learn how to source data from all popular data hosting platforms, including HDFS, Hive, JSON, and S3, and deal with large datasets with PySpark to gain practical big data experience. This book will help you work on prototypes on local machines and subsequently go on to handle messy data in production and at scale. This book covers installing and setting up PySpark, RDD operations, big data cleaning and wrangling, and aggregating and summarizing data into useful reports. You will also learn how to implement some practical and proven techniques to improve certain aspects of programming and administration in Apache Spark.
By the end of the book, you will be able to build big data analytical solutions using the various PySpark offerings and also optimize them effectively.
What You Will Learn:
- Get practical big data experience while working on messy datasets
- Analyze patterns with Spark SQL to improve your business intelligence
- Use PySpark s interactive shell to speed up development time
- Create highly concurrent Spark programs by leveraging immutability
- Discover ways to avoid the most expensive operation in the Spark API: the shuffle operation
- Re-design your jobs to use reduceByKey instead of groupBy
- Create robust processing pipelines by testing Apache Spark jobs
Who this book is for:
This book is for developers, data scientists, business analysts, or anyone who needs to reliably analyze large amounts of large-scale, real-world data. Whether you're tasked with creating your company's business intelligence function or creating great data platforms for your machine learning models, or are looking to use code to magnify the impact of your business, this book is for you.
“Sinopsis” puede pertenecer a otra edición de este título.
Acerca del autor
Bartłomiej Potaczek is a software engineer working for Schibsted Tech Polska and programming mostly in JavaScript. He is a big fan of everything related to the react world, functional programming, and data visualization. He founded and created InitLearn, a portal that allows users to learn to program in a pair-programming fashion. He was also involved in InitLearn frontend, which is built on the React-Redux technologies. Besides programming, he enjoys football and crossfit. Currently, he is working on rewriting the frontend for tv.nu―Sweden's most complete TV guide, with over 200 channels. He has also recently worked on technologies including React, React Router, and Redux.
“Acerca de” puede pertenecer a otra edición de este título.
Rarewaves.com UK
London, Reino Unido
Vendedor de AbeBooks desde el 11 de junio de 2025
Tarifas de envío de Reino Unido a Estados Unidos de America
| Artículo | De 60 a 60 días hábiles | De 60 a 60 días hábiles |
|---|---|---|
| Primer artículo | EUR 76,48 | EUR 117,67 |
Métodos de pago
Información empresarial del vendedor
RAREWAVES.COM LIMITED
Elsley Court, 20-22 Great Titchfield Street
London, Reino Unido W1W 8BE
Derecho al desistimiento
Si es un consumidor, puede rescindir el contrato de acuerdo con lo siguiente. Por consumidor se entiende cualquier persona física que actúe con fines ajenos a su actividad comercial, empresarial, oficio o profesión.
Información sobre el derecho de desistimiento
Derecho legal de desistimiento
Tiene derecho a rescindir este contrato en un plazo de 14 días sin dar ningún motivo.
El periodo de desistimiento vencerá a los 14 días desde que usted, o un tercero que no sea el transportista e indicado por usted, adquiera la posesión física del último bien o del último lote o pieza.
Para ejercer el derecho de desistimiento, complete de forma electrónica y envíe una declaración clara en nuestro sitio web, desde "Mis compras" en "Mi cuenta". Le enviaremos sin demora un acuse de recibo de dicho desistimiento a través de un soporte duradero (por ejemplo, por correo electrónico).
Para cumplir con el plazo de desistimiento, basta con que envíe su comunicación relativa al ejercicio del derecho de desistimiento antes de que venza el periodo de desistimiento.
Efectos del desistimiento
Si rescinde este contrato, le reembolsaremos todos los pagos que hayamos recibido de usted, incluidos los gastos de envío (excepto los gastos adicionales que surjan si elige un tipo de envío que no sea el tipo de envío estándar más económico que ofrecemos).
Podemos hacer una deducción del reembolso por la pérdida de valor de cualquier bien suministrado, si la pérdida es el resultado de una manipulación innecesaria por su parte.
Efectuaremos el reembolso sin demoras indebidas y, a más tardar, 14 días después de que se nos informe de su decisión de rescindir este contrato.
Efectuaremos el reembolso utilizando el mismo medio de pago que utilizó para la transacción inicial, a menos que haya acordado expresamente lo contrario; en cualquier caso, no incurrirá en ningún cargo como resultado de dicho reembolso.
Podremos retener el reembolso hasta que hayamos recibido los bienes o hasta que nos haya presentado una prueba de que los ha devuelto, lo que ocurra primero.
Deberá devolver los bienes o entregarlos a Rarewaves.com UK, Unit 144 The Lightbox, 111 Power Road, W4 5PY, London, London, United Kingdom, sin demoras indebidas y, en cualquier caso, en un plazo máximo de 14 días a partir del día en que nos comunique su desistimiento del presente contrato. El plazo se cumple si devuelve la mercancía antes de que venza el periodo de 14 días. Tendrá que asumir los gastos directos de devolución de los bienes. Usted solo es responsable de la disminución del valor de los bienes como resultado de una manipulación distinta a la necesaria para establecer la naturaleza, las características y el funcionamiento de los bienes.
Excepciones al derecho de desistimiento
El derecho de desistimiento no se aplica a lo siguiente:
- La entrega de periódicos, diarios o revistas, con la excepción de los contratos de suscripción; y
- El suministro de contenido digital que no se proporcione en un soporte tangible (por ejemplo, en un CD o DVD) si, al hacer el pedido, aceptó que podíamos empezar a entregarlo y que no podría desistir una vez iniciada la entrega.
Condiciones de envío
Please note that we do not offer Priority shipping to any country.
We currently do not ship to the below countries:
Russia
Belarus
Ukraine
Please do not attempt to place orders with any of these countries as a ship to address - they will be cancelled.