A community showcase featuring Christian Casazza's open-source vertical data stack built with Dagster, exploring NYC and NY State public datasets. The project uses Arrow, Parquet, DuckDB, Polars, and DuckLake as core primitives, with Dagster acting as the orchestration backbone. Key highlights include factory-style asset patterns for scalable pipeline creation, AI agent integration for code generation, and a public SQL query interface at QueryStation.app. The post covers why Dagster was chosen, the hardest problems solved (asset factory design, scalable patterns), and practical advice for newcomers.

9m read timeFrom dagster.io
Post cover image
Table of contents
Tell us a little about yourself and what you work onHow did you first discover Dagster?What project have you been building with Dagster?What was the hardest or most interesting problem you solved?Why was Dagster a good fit for this project?What advice would you give someone starting with Dagster?
4.4K Impressions