Apache Spark for Distributed Systems Engineers / Najlacnejšie knihy
Apache Spark for Distributed Systems Engineers

Code: 54026513

Apache Spark for Distributed Systems Engineers

by Devlin Nexley

Apache Spark is not just a data processing framework-it is a distributed execution system built on deep principles of cluster computing, DAG-based scheduling, memory-aware computation, and fault-tolerant design.This book provides ... more

22.50 €

RRP: 22.52 €

You save 0.02 €


In stock at our supplier
30.09.2026

Availability alert

Add to wishlist

You might also like

Give this book as a present today
  1. Order book and choose Gift Order.
  2. We will send you book gift voucher at once. You can give it out to anyone.
  3. Book will be send to donee, nothing more to care about.

Book gift voucher sampleRead more

Availability alert

Availability alert


Your agreement - Submiting you agree to the Terms and Condtions.

We will watch availability for you

Enter your e-mail address and once book will be available,
we will send you a message. It's that simple.

More about Apache Spark for Distributed Systems Engineers

You get 55 loyalty points

Book synopsis

Apache Spark is not just a data processing framework-it is a distributed execution system built on deep principles of cluster computing, DAG-based scheduling, memory-aware computation, and fault-tolerant design.

This book provides a rigorous, systems-level examination of Spark as a production-grade distributed execution engine. It moves beyond APIs, tutorials, and surface-level usage patterns to expose the internal mechanisms that govern how Spark actually executes workloads at scale.

Designed for experienced engineers working in distributed systems, backend infrastructure, and data platform engineering, this book dissects Spark as an architectural system rather than a development tool.

Inside, you will explore how Spark transforms high-level computations into distributed execution graphs, how it schedules and coordinates work across clusters, and how it manages the complexity of large-scale data movement in cloud-native environments.

Key areas covered include:

Internal architecture of Spark's driver, executors, and cluster coordination model

DAG construction, stage decomposition, and task scheduling mechanics

Shuffle architecture, data movement patterns, and network bottlenecks

Memory management, execution optimization, and JVM runtime behavior

Fault tolerance through lineage reconstruction and retry semantics

Query execution via Spark SQL, Catalyst optimizer, and Tungsten engine

Structured Streaming and incremental computation models

Performance bottlenecks, skew handling, and production tuning strategies

Cloud-native execution on object storage systems and Kubernetes

Integration with modern lakehouse ecosystems such as Delta Lake, Iceberg, and Hudi

Rather than presenting Spark as a tool to be used, this book treats it as a distributed systems case study-revealing how large-scale data infrastructure is engineered, optimized, and operated under real production constraints.

By the end, readers will understand not only how Spark works, but why its architecture is designed the way it is, what trade-offs shape its execution model, and how it fits into the broader evolution of modern distributed data platforms.

This is a book for engineers who build systems, not just pipelines.

Book details

22.50 €



Collection points Bratislava a 13474 dalších

Copyright ©2008-26 najlacnejsie-knihy.sk All rights reservedPrivacyCookies


Account: Log in
Všetky knihy sveta na jednom mieste. Navyše za skvelé ceny.

Shopping cart ( Empty )

For free shipping
shop for 59,99 € and more

You are here: