• All tags
  • spark

Apache Spark

Tagยท1.3K stories

Apache Spark, a distributed engine for large-scale data processing. Readers can learn about the DataFrame and SQL APIs, execution planning and shuffles, structured streaming, memory and partition tuning, running on Kubernetes or managed services, and alternatives.

Deploy an on-premise data hub with Canonical MAAS, Spark, Kubernetes and CephApache Hadoop and Apache Spark for Big Data AnalysisAmazon EMR Serverless introduces Shuffle-optimized disks delivering improved performance for I/O intensive workloadsAmazon EMR on EKS now supports Apache LivyUnderstanding Distributed ComputingIris - Turning observations into actionable insights for enhanced decision makingCost Optimization Strategies for scalable Data LakehouseEnhancing Data Security with Spark: A Guide to Column-Level Encryption - Part 1Sentiment Analysis of Yelp Restaurants Reviews in Real-TimeEnabling near real-time data analytics on the data lake

Roadmaps

roadmap.sh logo

Comprehensive roadmap for Apache Spark

By roadmap.sh

Recommended Apache Spark stories

Who to follow for Apache Spark

Top sources covering Apache Spark

Most upvoted Apache Spark posts

Best discussed Apache Spark posts

All posts about Apache Spark