Automates PySpark DataFrame transformations including schema inference, partition pruning, and Delta Lake merge operations. Integrates with AWS Glue Data Catalog and Apache Iceberg table formats for lakehouse architectures.
Scanned 6/2/2026
Install to Claude Code
npx -y skills add agentskillexchange/skills --skill spark-dataframe-etl-pipeline --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Spark Dataframe Etl Pipeline?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/agentskillexchange-spark-dataframe-etl-pipeline)More formats (shields.io, HTML) on the badges page.
---
name: "Apache Spark DataFrame ETL Pipeline"
slug: "spark-dataframe-etl-pipeline"
description: "Automates PySpark DataFrame transformations including schema inference, partition pruning, and Delta Lake merge operations. Integrates with AWS Glue Data Catalog and Apache Iceberg table formats for lakehouse architectures."
github_stars: 43117
verification: "security_reviewed"
source: "https://github.com/apache/spark"
category: "Data Extraction & Transformation"
framework: "OpenClaw"
tool_ecosystem:
github_repo: "apache/spark"
github_stars: 43117
---
# Apache Spark DataFrame ETL Pipeline
Automates PySpark DataFrame transformations including schema inference, partition pruning, and Delta Lake merge operations. Integrates with AWS Glue Data Catalog and Apache Iceberg table formats for lakehouse architectures.
## Installation
Requirements and caveats from upstream:
- high-level APIs in Scala, Java, Python, and R (Deprecated), and an optimized engine that
- ## Interactive Python Shell
- Alternatively, if you prefer Python, you can use the Python shell:
Basic usage or getting-started notes:
- To build Spark and its example programs, run:
- And run the following command, which should also return 1,000,000,000:
- ## Example Programs
- Source: https://github.com/apache/spark
- Extracted from upstream docs: https://raw.githubusercontent.com/apache/spark/HEAD/README.md
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/spark-dataframe-etl-pipeline/)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!