Artwork

Indhold leveret af The Data Flowcast. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af The Data Flowcast eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.
Player FM - Podcast-app
Gå offline med appen Player FM !

Inside Vinted’s Code-Generated Airflow Pipelines with Oscar Ligthart and Rodrigo Loredo

29:36
 
Del
 

Manage episode 515154788 series 2948506
Indhold leveret af The Data Flowcast. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af The Data Flowcast eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.

The shift from monolithic to decentralized data workflows changes how teams build, connect and scale pipelines.

In this episode, we feature Oscar Ligthart, Lead Data Engineer, and Rodrigo Loredo, Lead Analytics Engineer, both at Vinted, as we unpack their YAML-driven abstraction that generates Airflow DAGs and standardizes cross-team orchestration.

Key Takeaways:

00:00 Introduction.

05:28 Challenges of decentralization.

06:45 YAML-based generator standardizes pipelines and dependencies.

12:28 Declarative assets and sensors align cross-DAG dependencies.

17:29 Task-level callbacks enable auto-recovery and clear ownership.

21:39 Standardized building blocks simplify upgrades and maintenance.

24:52 Platform focus frees domain work.

26:49 Container-only standardization prevents sprawl.

Resources Mentioned:

Oscar Ligthart

https://www.linkedin.com/in/oscar-ligthart/

Rodrigo Loredo

https://www.linkedin.com/in/rodrigo-loredo-410a16134/

Vinted | LinkedIn

https://www.linkedin.com/company/vinted/

Vinted | Website

https://www.vinted.com/?srsltid=AfmBOor87MGR_eLOauCO93V9A-aLDaAhGYx9cnu_oN8s1SAXMlCRuhW7

Apache Airflow

https://airflow.apache.org/

Kubernetes

https://kubernetes.io/

dbt

https://www.getdbt.com/

Google Cloud Vertex AI

https://cloud.google.com/vertex-ai

Airflow Datasets & Assets (concepts)

https://www.astronomer.io/docs/learn/airflow-datasets

Airflow Summit

https://airflowsummit.org/

Thanks for listening to “The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI.” If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations.

#AI #Automation #Airflow #MachineLearning

  continue reading

82 episoder

Artwork
iconDel
 
Manage episode 515154788 series 2948506
Indhold leveret af The Data Flowcast. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af The Data Flowcast eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.

The shift from monolithic to decentralized data workflows changes how teams build, connect and scale pipelines.

In this episode, we feature Oscar Ligthart, Lead Data Engineer, and Rodrigo Loredo, Lead Analytics Engineer, both at Vinted, as we unpack their YAML-driven abstraction that generates Airflow DAGs and standardizes cross-team orchestration.

Key Takeaways:

00:00 Introduction.

05:28 Challenges of decentralization.

06:45 YAML-based generator standardizes pipelines and dependencies.

12:28 Declarative assets and sensors align cross-DAG dependencies.

17:29 Task-level callbacks enable auto-recovery and clear ownership.

21:39 Standardized building blocks simplify upgrades and maintenance.

24:52 Platform focus frees domain work.

26:49 Container-only standardization prevents sprawl.

Resources Mentioned:

Oscar Ligthart

https://www.linkedin.com/in/oscar-ligthart/

Rodrigo Loredo

https://www.linkedin.com/in/rodrigo-loredo-410a16134/

Vinted | LinkedIn

https://www.linkedin.com/company/vinted/

Vinted | Website

https://www.vinted.com/?srsltid=AfmBOor87MGR_eLOauCO93V9A-aLDaAhGYx9cnu_oN8s1SAXMlCRuhW7

Apache Airflow

https://airflow.apache.org/

Kubernetes

https://kubernetes.io/

dbt

https://www.getdbt.com/

Google Cloud Vertex AI

https://cloud.google.com/vertex-ai

Airflow Datasets & Assets (concepts)

https://www.astronomer.io/docs/learn/airflow-datasets

Airflow Summit

https://airflowsummit.org/

Thanks for listening to “The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI.” If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations.

#AI #Automation #Airflow #MachineLearning

  continue reading

82 episoder

모든 에피소드

×
 
Loading …

Velkommen til Player FM!

Player FM is scanning the web for high-quality podcasts for you to enjoy right now. It's the best podcast app and works on Android, iPhone, and the web. Signup to sync subscriptions across devices.

 

Hurtig referencevejledning

Lyt til dette show, mens du udforsker
Afspil