Knowledge Hub
Discover, browse, and search Dataminded's knowledge assets in one place.
Type @ for a person, # for a technology or theme
Agent Skills for Data Practitioners
A half-day workshop on what Agent Skills are, how other data professionals use them, and how to build your own.
AI-Assisted Coding Workshop
Level up your engineering workflow with AI coding tools.
AI Discovery Workshop
A one-day workshop that turns scattered AI experiments into a prioritized, feasibility-checked roadmap leadership can act on.
Building Agentic AI Solutions Workshop
A hands-on, one-day workshop that takes engineers from prompting AI tools to building production-grade agentic AI systems: RAG, MCP, agents, and AI security.
Capstone Project
Apply all your knowledge in a capstone project similar to an actual data product in a modern data platform.
CI/CD with Azure DevOps
Build end-to-end CI/CD pipelines on Azure DevOps. From repos and code reviews to YAML pipelines, multi-environment deployments, and automated testing.
CI/CD with GitHub Actions
Set up CI/CD pipelines with GitHub Actions: workflows, triggers, secrets, environments, and container registries.
Containerization with Docker
Learn how to get your application working on any machine. Reach production faster.
Data Product Workshop
A one-day, cross-functional workshop that gets business and technical teams on the same page about data products, with a shared vocabulary, a scoping framework, quality principles, and clear roles.
DataFrame of Mind
Build data pipelines with Polars and PySpark. Learn when to use single-node processing and when to scale out.
Developer Productivity
Get more out of your dev environment: Python tooling, debugging, code quality automation, and agentic coding harnesses.
Effective PySpark
Build production-grade PySpark pipelines with sound design patterns, performance tuning, data quality checks, and testing.
Infrastructure as Code with Terraform
Learn how to create and maintain infrastructure as code, using Terraform.
Introduction to Cloud Providers (with AWS)
Learn when and how to use some of Amazon Web Services' most commonly used resources.
Introduction to Git for Version Control
Become proficient with the most widely used version control software. Stop sending code as attachments. Collaborate fearlessly.
Introduction to Linux & Bash
Learn bash and Linux basics to make you productive on any data project.
Kafka & Spark Structured Streaming
Build real-time streaming data pipelines by integrating Apache Kafka with Spark Structured Streaming.
Kubernetes
Deploy, scale, and manage containerized applications on Kubernetes, from core components to operators and networking.
Modernize SQL Analytics with DBT
A core component of any good analytics solution is still a relational database. Learn about SQL and build pipelines using DBT.
Principles of Modern Data Platforms
Prepare your data platform to support multiple use cases, at scale.
PySpark Fundamentals
Learn distributed data processing with PySpark. Understand Spark's architecture and DataFrame API, then apply them to build clean, modular, and testable data transformation pipelines.
Python for Data Engineering
Learn how to process data with Python and maintain a good code base.
Snowflake, a Modern Data Warehouse
Learn how to maintain and run analytics in this virtually infinitely scalable warehouse.
Sovereign Data Workshop
A one-day workshop that assesses whether and how to extend your data platform with sovereign EU cloud capabilities, from regulatory exposure to a migration roadmap.
Spark on Kubernetes
Submit, configure, and troubleshoot Spark jobs running on a Kubernetes cluster.
Workflow Orchestration with Airflow
Manage multiple workloads with the orchestration framework of Apache Airflow.
S1:E3 · Nov 2025
Accelerate Data Engineering using MCP Tools
Jonny Daenen, Emil Krause
Oct 2025
AWS Outage: When The Cloud Fails...
Jonny Daenen, Stijn De Haes
S1:E5 · Dec 2025
Build a RAG-based Agent with MindsDB
Jonny Daenen, Tarik Jamoulle
S1:E6 · Dec 2025
Cross-Project AI Code Assistance using Cursor Workspaces
Jonny Daenen, Emil Krause
S1:E4 · Nov 2025
Data Ingestion using PyAirbyte
Jonny Daenen, Tarik Jamoulle
S1:E2 · Oct 2025
MCP 101: The Model Context Protocol for AI Agents
Jonny Daenen, Pierre Crochelet
S1:E1 · Oct 2025
Whisper & SuperWhisper: 3x Faster Prompting with Speech-to-text?
Jonny Daenen, Emil Krause
S2:E7 · Jan 2026
AI Code Reviews with CodeRabbit and Sourcery
Jonny Daenen, Hannes De Smet
S2:E12 · Apr 2026
AI Workflows in Agno: Building Deterministic Agents
Jonny Daenen, Pascal Knapen
S2:E8 · Feb 2026
Azure Log Analytics Costs Are Out of Control - Here's How We Cut Them by 60%
Jonny Daenen, Niels Claeys
S2:E10 · Mar 2026
Building an AI Agent with Subagents and Skills
Jonny Daenen, Jesús García Ramírez
S2:E9 · Feb 2026
From Prompts to Agents: AI Agent Skills in Claude Code
Jonny Daenen, Jesús García Ramírez
S2:E11 · Apr 2026
Managing Airflow at Scale using the Flowrs TUI
Jonny Daenen, Jan Vanbuel
S2:E14 · May 2026
Metric Views in Databricks: The Missing Layer for AI Agents
Jonny Daenen, Stefan Van Raemdonck
S2:E13 · May 2026
Snowflake Intelligence: The End of Dashboards?
Jonny Daenen, Jelle De Vleminck
S1:E1 · Apr 2025
Building Data Architecture from Scratch: Cloud, Data Mesh, and the Real-World Tradeoffs
Kris Peeters, Thorsten Foltz
S1:E10 · Jul 2025
AI Innovation That Delivers: How to Align Strategy, Adoption, and Business Value
Kris Peeters, Joris Renkens
S1:E11 · Jul 2025
The Art of the Data Platform: Reducing Cognitive Load and Driving Adoption
Kris Peeters, Jelle De Vleminck
S1:E12 · Jul 2025
RADAR by publiq: How AI is Reshaping Cultural Discovery, Without Compromising Privacy
Kris Peeters, Sven Houtmeyers, Elia Van Wolputte
S1:E2 · May 2025
What Not to Build with AI: Avoiding the New Technical Debt in Data-Driven Organizations
Kris Peeters, Tim Schröder, Pascal Brokmeier
S1:E3 · May 2025
Building Sustainable Data Products that actually get used
Kris Peeters, Geert Verstraeten
S1:E4 · May 2025
From On-Prem to Cloud (Again): How a Government Agency Made the Big Bang Work
Kris Peeters, Niels Melotte
S1:E5 · Jun 2025
Data Modeling: Why should I care? Unveiling the Sense and Nonsense with Jonas De Keuster
Kris Peeters, Jonas De Keuster
S1:E6 · Jun 2025
Data Mesh Live: How to make it successful in organisations, with Jacek Majchrzak & Andrew Jones
Kris Peeters, Jacek Majchrzak, Andrew Jones
S1:E7 · Jun 2025
A Deep Dive into SQLMesh: Structured Query Validation and Safe Pipeline Testing
Kris Peeters, Michiel De Muynck
S1:E8 · Jun 2025
What It Really Takes to Build a Data-Centric Organization
Kris Peeters, Jonny Daenen
S1:E9 · Jul 2025
You Don’t Need the Latest Stack. You Need Better Questions
Kris Peeters, Rushil Daya
S2:E1 · Oct 2025
Live from the POA Summit: Simon Harrer on Data Contracts, AI, and the Next Phase of Data Mesh
Kris Peeters, Simon Harrer
S2:E10 · Jan 2026
How OBI Built a Lean, High-Impact AI Function That Scales
Kris Peeters, Dr. Ruth Janning
S2:E11 · Jan 2026
How Dataminded Was Built: Kris Peeters on 11 Years of Data Engineering & Culture
Kris Peeters, Pascal Brokmeier
S2:E2 · Oct 2025
Federated Data Governance in Banking: How ABN AMRO Scaled from 300 Data Owners to 15 Data Domains
Kris Peeters, Jan Mark Pleijsant
S2:E3 · Oct 2025
How imec scales research with data platforms: governance, workbenches, and adoption
Kris Peeters, Wim Vancuyck
S2:E4 · Nov 2025
Build vs Buy in the GenAI Era: Inside Belfius’ Data & AI Strategy
Kris Peeters, Hannes Heylen
S2:E5 · Nov 2025
Beyond Hyperscalers: How to Run Modern Data Platforms on European Clouds
Kris Peeters, Niels Claeys
S2:E6 · Nov 2025
5 Years Kate 🎂: Inside KBC’s AI Playbook
Kris Peeters, Dr. Barak Chizi
S2:E7 · Nov 2025
Data Engineering Meets Excel: Building Explainable and Reliable Decision Models with River Solutions
Kris Peeters, Amaury Anciaux
S2:E8 · Dec 2025
A Structured Framework for Building Successful Data Solutions
Kris Peeters, Frederic Vanderveken
S2:E9 · Dec 2025
Data Science vs Data Engineering: Breaking the Wall
Kris Peeters, Jelena Grujic
S3:E1 · Mar 2026
Machine Learning in Energy: Forecasting, MLOps, and Business Impact
Kris Peeters, Jean-Michel Begon
S3:E10 · Sep 2026
Data Contracts and Data Products: What AI Agents Need to Access Enterprise Data
Kris Peeters, Simon Harrer
S3:E2 · Mar 2026
Scaling Data in Aviation: Inside Brussels Airlines’ Data Strategy
Kris Peeters, Tom Holsteens
S3:E3 · Apr 2026
The Data Challenge behind the Einstein Telescope
Kris Peeters, Tjonnie Li
S3:E4 · Jun 2026
Can We Outsource Thinking? AI, Education, and the Future of Knowledge Work
Kris Peeters, Frank Neven
S3:E5 · Jun 2026
Agentic AI in Production: What Data Leaders Need to Know Before Scaling AI
Kris Peeters, Jesús García Ramírez
S3:E6 · Jul 2026
How NMBS/SNCB Uses Data Science to Improve Public Transport
Kris Peeters, Max Hubeau
S3:E7 · Jul 2026
How Tomorrowland Keeps Data Simple While Scaling Globally
Kris Peeters, Wannes Rosiers
S3:E8 · Aug 2026
Why AI Agents Need Knowledge Graphs, Not Just Data
Kris Peeters, Eric Broda
S3:E9 · Aug 2026
Data Mesh After the Hype: Why Data Products Matter More Than Ever
Kris Peeters, Arif Wider
Nov 2017
3 years Data Minded!
This month Data Minded turns 3 years already. Like every startup, we’ve had our ups and downs. But looking back I’m extremely proud of the…
Nov 2017
Cloud is a business strategy, not an IT implementation detail
British Airways is not having their best weekend: A big IT outage, resulting in a lockdown at Heathrow, all apparently caused by a single…
Nov 2017
Order some 21st century management with your big data lunch
Why is it that every company today wants to have an Hadoop-enabled data lake, large-scale data pipelines in Spark, and artificial…
Nov 2017
Our three favourite data analytics tools of 2015
In this blog, we look back at the 3 data analytics tools that shaped 2015 for us. This is by no means an objective blog, and I’m sure your…
May 2018
Connect to AWS Athena using Datagrip
Datagrip is a great database IDE, just like the other IDEs from Jetbrains: Pycharm, IntelliJ, … In this blog, I describe how to connect an…
Jul 2018
Good leaders take ownership of everything in their world
Last week, yet another political fight started in Belgium. You can read all about it here http://www.standaard.be/cnt/dmf20180719_03624520…
Sep 2018
Hell’s Kitchen — IT edition
Imagine you start a new job in the kitchen of a fancy restaurant. You just graduated from your culinary studies and you’re eager to learn…
Nov 2018
Little known Spark DataFrame join types
Probably most of you know the basic join types from SQL: left, right, inner and outer. Since these are supported by most of the…
Nov 2018
Run Spark Jobs on Azure Batch using Azure Container Registry and Blob storage
This is another one of those “how to” blogs that can hopefully help people get up-and-running quickly because it took me a while to figure…
Sep 2019
Avoiding copy paste in Terraform: Two approaches for multi-environment Infra as code setups
Terraform workspaces vs symlinks and overrides
Nov 2019
AWS MSK secure Python Kafka client
How to write a secured python client for AWS MSK using TLS for encryption and authentication
Feb 2019
Capture clickstream data with Azure Application Insights
Raw clickstream data is a valuable data source in almost any analytics project. But it’s not always easy to capture. Free tools like…
Oct 2019
Hooray, I’m an AWS Certified Pro Architect. Now what?
Is it worth it to spend time and money collecting advanced certifications? In this blog I share my opinion and lessons learned. TL/DR: Yes…
Oct 2019
How to build a cleaning pipeline with BigQuery and DataFlow on GCP
I have a small script running on my phone which sends a set of information to the cloud every 5 minutes. I wanted to build a dataset with…
Oct 2019
Organize your data lake using Lighthouse
Lighthouse is an open source library (using Apache Spark and Scala) that we developed at Data Minded, with the aim of providing a way to…
Aug 2019
Test new Kafka application with prod data with Kafka streams applications running on K8S
Or how to pipe messages from one kafka cluster to another cluster through kubectl and your local machine
Jul 2020
A summary of cookies
A short, technical summary of cookie concepts such as “third party cookies” and attributes such as a expiration date, host-only et cetera.
Mar 2020
Cloud story: Spin down unneeded infrastructure quickly with terraform to save OPEX
The benefit of infrastructure as code and cloud: If you don’t need it, spin it down.
Oct 2020
Combine AWS cloud with OVH storage for handling sensitive EU data
Now that the Privacy Shield has been invalidated, there are some legal disputes whether European companies and government agencies can…
Sep 2020
Data Quality Libraries: The Right Fit
A high-level comparison of TensorFlow Data Validation, Great Expectations, and Deequ
Jun 2020
Enrichment and batch processing in Snowplow
A close look at Snowplow’s enrichment component as well as the deprecation of its batch pipeline
Feb 2020
Google Pub/Sub: Putting a number on the lack of order
Evaluating Pub/Sub’s lack of order. Making good decisions about watermarks
Nov 2020
How I prepared for the Azure Data Engineer Associate (DP-200 and DP-201) exams
As the title says, this article is about how I personally prepared for the aforementioned exams and not about how one should best prepare…
Jun 2020
How to conditionally disable modules in terraform
Terraform doesn’t support the count parameter on modules. A proposal was made for a enabled parameter, but this is also not yet present…
Sep 2020
How to deploy analytics workloads
“It works on my machine”. That’s great. But now how do you make sure it runs in production, repeatedly and reliably? Here we share our…
Dec 2020
How to share tabular data in a privacy-preserving way
Adding noise to existing rows, only adding noise to outcomes of tasks performed on that data, or synthetic data generation? An intuition.
Jun 2020
Import SQL Server data in BigQuery
A list of four approaches for a one-off data dump from a RDBMS like SQL Server to BigQuery, and an in-depth look at how to use Apache…
Nov 2020
Ingesting custom event sources with Snowplow
How to use Snowplow to ingest data from an unsupported source, such as Auth0’s log streaming service.
Dec 2020
Learnings from the AWS Data Analytics Speciality
Last year I blogged about how I got my AWS Pro Architect certificate, which you can read about here…
Nov 2020
Running Spark 3 on AKS with Azure AD integration
Do you want to run Spark 3 on AKS in pro mode? Meaning no more “just copy-paste the storage account access key into the source code, and…
Apr 2020
Running through the Google GCP Cloud Foundation Toolkit Setup
An experience report of applying the CFT toolkit foundation steps
Jun 2020
Save money on MSK
Let’s be blunt here for a second: MSK is not a mature managed service. The author of that post may have changed his mind in the meantime…
Jan 2020
The data product lifecycle
Your organisation wants to dive head-first into data and AI but you don’t really know where to start? Data&AI is on the radar of most…
Dec 2020
Why DBT will one day be bigger than Spark
The world of data is moving and shaking again. Ever since Hadoop came around, people were offloading workloads from their data warehouses…
Sep 2021
Authentication on GCP with Docker: Application Default Credentials
How applications magically authenticate themselves with GCP through their environment, and how to make locally running containers magic too
Oct 2021
Backing up and restoring SageMaker Notebook Instances
How to store and restore SageMaker Notebook instances to and from S3, for example for migration to Amazon Linux 2
Mar 2021
Consulting 101 for data engineering
At Data Minded, we are data engineers first and foremost. But in reality, we do a lot of consulting. What is consulting really? I recently…
Apr 2021
CORS and the SOP explained
Introduction to Cross-Origin Resource Sharing (CORS) and the Same-Origin Policy (SOP). Structured as a dialogue, and focused on the why.
Nov 2021
Customizing SageMaker Notebook Instances
How SageMaker’s lifecycle configuration works, a collection of useful startup scripts and how to bundle them
May 2021
How RDS proxy allowed us to run Airflow 200% more efficient
Running multiple Airflow instances on the same RDS means a lot of open database connections. RDS proxy can make this setup more efficient.
Apr 2021
How to Access Key Vaults from Azure Batch Jobs
The cheapest and simplest way of running computational jobs on Azure is by using Azure Batch. This service enables you to launch managed…
Dec 2021
How to access private Git repositories during a Docker image build
A complete guide to building images that require access to SSH keys during the build process.
Jan 2021
Joining Spark Datasets
Ever wanted to do better than joins on Apache Spark DataFrames? Now you can!
Aug 2021
Mastering the Google Cloud Platform SDK tools
A look at some lesser-known GCP SDK settings and features that make your day-to-day interactions with GCP more enjoyable.
Nov 2021
Sending mail from Google Cloud Build
A simple solution for sending messages from a GCP Cloud Build pipeline
May 2021
Storing Snowplow bad row events in BigQuery
How to use Cloud Functions and a BigQuery schema generator to make Snowplow bad row schema violation events easily queryable
Jan 2021
Tracking prevention in the modern browser
An overview of cross-domain tracking prevention in the modern browser, specifically of Intelligent Tracking Prevention (ITP) in WebKit.
Sep 2021
What I wish I knew before going into Data Engineering
Disclaimer: this is my opinion, not necessarily the one of my employer or any organisation.
Sep 2021
What is lakeFS: A Critical Survey
A critical introduction to lakeFS, a new metadata layering solution that brings Git-like operations and versioning to object storage.
Mar 2021
What to consider before choosing Argo Workflow?
To go full Kubernetes-native or not?
Apr 2022
Automatically following insiders transactions on the belgian stock market with Serverless on AWS
Recently I was reading the newspaper and I saw an article in which the owner family of D’Ieteren, a company listed on the Belgian stock…
Aug 2022
Aws security for software engineers
Security breaches are more and more common. This preview shows you what we will talk about during our webinar on AWS security.
Sep 2022
AWS Web Identity Token Authentication in Legacy Applications
How to run make applications without support for AWS web identity token authentication run in environments requiring it
Apr 2022
Batch orchestration on Azure flowchart
Managed solutions vs building it yourself
Sep 2022
CI/CD for data projects: Why manual deployments are not good enough
CI/CD for data projects: Why manual deployments are not good enough Data engineering is slowly but surely adopting many of the best practices from Software engineering. Two of these best practices …
Sep 2022
Containerizing Git credential helpers
How to expose Git credential helpers to containerized processes, allowing the use of Bitbucket, Github and Gitlab inside of Docker on…
Nov 2022
Content is no longer the king … It’s all about experiential marketing!
My honest takeaways on where HubSpot can do better & be better implemented!
Aug 2022
Dagger vs. the current state of CI/CD
How Dagger tries to solve some of the pain points of current CI systems.
Jan 2022
Datafy feature release Q4 2021
Today we want to share with you the features the Datafy team has been working in the last part of 2021.
Oct 2022
Debugging Google Application Default Credentials
Inspecting gcloud application default credentials, Google access tokens, and ID tokens through the refresh token grant & token…
Feb 2022
From notebook hell to container heaven, Part II.
This article is the second chapter of a 3-parts tutorial
Mar 2022
Keep Long-Lived AWS Credentials Out of Untrusted Environments
Generate and use short-lived AWS session credentials, keeping your AWS IAM keys secure
Dec 2022
Make Gitpod Open Sites in the Browser
Gitpod configuration for opening links in the browser instead of in the terminal.
Jul 2022
Make Spark resilient against spot interruptions on kubernetes
Based on our experience of running spark in production at our customers, we discuss 3 ways to improve the resilience of spark on kubernetes
Oct 2022
Mastering Cloud: can we do better than certificates?
A flipped classroom approach can help you get productive in the cloud more quickly.
Oct 2022
ML Pipelines in Azure Machine Learning Studio the right way
An opinionated way to get started quickly
Oct 2022
Not every m5.4xlarge is created equally
You might be surprised to learn that not every instance of the same type has the exact same amount of memory on AWS and this can have some…
Apr 2022
Porting a data platform from AWS to Azure
We recently created an Azure version of our data platform and this blogpost elaborates on our learnings/issues to support a new cloud…
Mar 2022
Running Containers on Windows Subsystem for Linux (WSL 2)
How to install and automatically start Docker Engine on WSL distributions such as Ubuntu.
Mar 2022
Running project-specific CI/CD pipelines for a monorepo in AWS
Implementing a flexible (yet powerful) CI/CD setup for monorepos using AWS CodeBuild, CodePipeline, and Lambda Functions.
Nov 2022
Slash your cloud bill by moving data workloads to cost-effective compute
These are tough times for technology
Apr 2022
The 6 pillars of data maturity
In our previous blog, we already explained that you’ll have to put in some effort if you want to grow your data maturity. Now, let’s take a…
Apr 2022
The different ways to configure AKS connectivity
This post elaborates on the different ways to configure AKS within your network as well as explains the most important configuration…
Nov 2022
The rise of remote development environments
Gitpod and Codespaces are the first remote development environments that we would use ourselves and may also be useful for you.
Mar 2022
Three methods for obtaining GCP access tokens
Using user credentials, service account credentials or the metadata service to obtain access tokens from Google’s identity service
Apr 2022
Using kubebuilder in production
Some extra tips on using kubebuilder in production
Aug 2022
What does it take to build a data platform
On August 9th, I will host a webinar on what it takes to build a data platform and which lessons I learned from helping my clients.
Apr 2022
Why grow in data maturity?
Every company pays lip service to data. Who takes action?
Jun 2022
Why I will not build my next data platform myself
The blogpost explains the difference between can a company build it’s own data platform and should it build it themself.
Feb 2022
Why you should govern data access through Purpose-Based Access Control
PBAC is a powerful data access governance strategy that can make your data access policies more practical and secure.
Apr 2022
Why you should not use IAM users
AWS provides two main options for giving users access to your AWS resources: IAM users and roles. You should avoid the former.
Oct 2022
Will Dagger revolutionize CI/CD?
A hands-on exploration of Dagger CI/CD pipelines
Jul 2023
Automating refactoring across teams and projects
Conquering breaking changes Chaos: Leveraging automatic refactoring tools across teams and projects.
Dec 2023
Connecting to Databases using JDBC from the CLI
A quick guide to using the CLI tool sqlline to connect to any database with JDBC drivers
Jul 2023
Cross-DAG Dependencies in Apache Airflow: A Comprehensive Guide
Exploring four methods to effectively manage and scale your data workflow dependencies with Apache Airflow.
Jun 2023
Enhance Your ETL Ingestion: Unlocking the Power of the Apache Iceberg Table Format
Regardless of all the changes that are happening in the data landscape today, data ingestion and ETL processes still play an important role…
Oct 2023
From notebook hell to container heaven
This article is the first chapter of a 3-parts tutorial
May 2023
How Leading Data Organizations Achieve Success: Prioritize People, Process, and Product
Technology is not the biggest challenge anymore when building data platforms
Dec 2023
Making the Airflow web UI faster
It’s a nice sunny Wednesday afternoon in Belgium, and my colleague Jan and I decide to go for a walk in the cozy city of Leuven. As we…
Aug 2023
Navigating the Iceberg: unit testing iceberg tables with Pyspark
The table format iceberg has gained a lot traction and has caused great excitement across the data-landscape. The architecture allows…
Nov 2023
Polars dataframe’s plugins and extensibility: getting started
Interesting feature of Polars explained.
Sep 2023
Soft Skills in a Tech World: A Psychologist’s Journey to IT Consulting
Five factors that facilitated my career switch to IT
Nov 2023
Terra-Do’s and Terra-Don’ts — a few common issues with Terraform iterables and how to avoid them
One of the most common issues I observe when teaching Terraform is improper iteration over resources, data sources and modules, which in…
Sep 2023
Testing frameworks in dbt
When I was struggling to find the right way to write tests in dbt, I got great inspiration from Mariah Rogers’ talk about testing at…
Oct 2023
Twelve-Factor Python applications using Pydantic Settings
A look at Pydantic Settings and how it can help you reliably deploy applications across environments
Nov 2023
Upserting Data using Spark and Iceberg
Use Spark and Iceberg’s MERGE INTO syntax to efficiently store daily, incremental snapshots of a mutable source table.
Apr 2023
Use dbt and Duckdb instead of Spark in data pipelines
Dbt has become very popular for transformation on top of your data warehouse. We see potential to use dbt with Duckdb on top of a data…
Apr 2023
Write Cookiecutters faster
This one’s about writing cookiecutter templates. 🍪 If you’ve ever done that, you might recognize the following git log
Jan 2024
3 software design patterns that every software data engineer should know
With real-life examples
Jan 2024
7 Lessons Learned migrating dbt code from Snowflake to Trino
Just change the target in profiles.yml, right?
Oct 2024
A glimpse into the life of a data leader
Last week, Dataminded organised a data leadership roundtable. We invited 20 decision takers of large data organisations, from Belgium, the…
May 2024
Age of DataFrames 2: Polars edition
In this publication, I showcase some Polars tricks and features.
Mar 2024
Becoming Clout* certified
Hot takes about my experience with cloud certifications
Jul 2024
Clear signals: Enhancing communication within a data team
Clear, effective communication is as crucial to successful data engineering as technical expertise.
Dec 2024
Data Product Portal Integrations 1: OIDC
How to integrate Open ID Connect with the Data Product Portal
Dec 2024
Data Product Portal Integrations 2: Helm
How to install portal in your production environment using Helm
Dec 2024
Data Product Thinking Rethought
On November 27, 2024, at Startplatz in Düsseldorf, we had a fantastic evening where data professionals gathered for the Data Product…
Aug 2024
Data Stability with Python: How to Catch Even the Smallest Changes
As a data engineer, it is nearly always the safest option to run data pipelines every X minutes. This allows you to sleep well at night…
Jul 2024
Demystifying Device Flow
Implementing OAuth 2.0 Device Authorization Grant with AWS Cognito and FastAPI
Jan 2024
Everyone to the data dance floor: a story of trust
Wandering through the vast realm of the internet, on stumbles upon a countless amount of articles about companies transitioning from…
Oct 2024
From Good AI to Good Data Engineering. Or how Responsible AI interplays with High Data Quality
The intersection of artificial intelligence (AI) and data engineering has become increasingly critical. As AI technologies proliferate, the…
Jan 2024
Growing your data program with a use-case-driven approach
How to Bridge the Gap Between Strategic Vision and Tactical Execution in Data Initiatives
Jan 2024
Harnessing AWS MSK and AWS EMR for Real-time Analytics
A Deep Dive into Building a Scalable Analytics Pipeline, automating the infrastructure with Terraform
Mar 2024
How to organize a data team to get the most value out of data
To state the obvious: a data team is there to bring value to the company. But is it this obvious? Haven’t companies too often created a…
Aug 2024
How To Reduce Pressure On Your Data Teams
In August 2016, BARC published the results of a global survey on Data-Driven Decision-Making in Business.Results are astonishing: only 22%…
Jan 2024
How to run PySpark jobs in an Amazon EMR Serverless Cluster with Terraform
Ignite Your Data Revolution and Harness the Power of Amazon EMR Serverless
Sep 2024
How we democratized data access with Streamlit and Microsoft-powered automation
How we democratized data access with Streamlit and Microsoft-powered automation Within the community of data professionals, the term “data governance” often conjures up an image of a large …
Feb 2024
How we used GenAI to make sense of the government
We built a RAG chatbot with AWS Bedrock and GPT to answer questions about the Flemish government
Feb 2024
Leveraging Pydantic for validation.
Ensuring clean and reliable input is crucial for building robust services. One powerful tool that simplifies this process is Pydantic, a…
Aug 2024
Microsoft Fabric’s Migration Hurdles: My Experience
It has been more than a year since Microsoft announced their new all-mighty Fabric data platform. It offers a wide range of capabilities…
May 2024
Prompt Engineering for a Better SQL Code Generation With LLMs
Picture yourself as a marketing executive tasked with optimising advertising strategies to target different customer segments effectively…
Jan 2024
Pulumi vs. Terraform: Choosing your IaC Tool
Similarities and differences
Oct 2024
Quack, Quack, Ka-Ching: Cut Costs by Querying Snowflake from DuckDB
How to leverage Snowflake’s support for interoperable open lakehouse technology — Iceberg — to save money.
Apr 2024
Querying Hierarchical Data with Postgres
Hierarchical data is prevalent and simple to store, but querying it can be challenging. This post will guide you through the process of…
Jan 2024
SAP CDC with Azure Data Factory
Using a self-hosted Integration Runtime
Jun 2024
Short feedback cycles on AWS Lambda
A Makefile that enables to iterate quickly
Jan 2024
Step-by-Step Guide to deploy a Kafka Cluster with AWS MSK and Terraform
A Comprehensive Walkthrough to Deploying and Managing Kafka Clusters with Amazon MSK and Terraform
Mar 2024
Two Lifecycle Policies Every S3 Bucket Should Have
Abandoned multipart uploads and expired delete markers: what are they, and why you must care about them thanks to bad AWS defaults.
Apr 2024
Unlock Insights & Learnings: Data Minded Newsletter — March/April 2024 Edition
Welcome to the March — April 2024 edition of the Data Minded Newsletter! We’re thrilled to share some exciting updates, valuable insights…
Jul 2024
Unlock Insights & Learnings: Dataminded Newsletter — June/July 2024 Edition
Welcome to the June — July 2024 edition of the Dataminded Newsletter! These months have been full of events and exciting developments for…
Sep 2024
Unlocking the new Power of Advanced Analytics
In recent years advanced analytics has become a cornerstone for businesses aiming to gain deeper insights and make informed decisions. This…
Dec 2024
Why not to build your own data platform
A round-table discussion summary on imec’s approach to their data platform
Oct 2024
Why rising cloud costs are the silent killers of data platforms
Building data platforms in the cloud is changing. Gone are the days that you would manually set up a few EC2 instances and run some modest…
May 2024
A 5-step approach to improve data platform experience
A guide to turning user feedback into continuous platform improvement.
Apr 2025
Are your AKS logging costs too high? Here’s how to reduce them
Explore how to reduce the cost of logging on Azure by analyzing your applications and investigating Basics log analytics tables in Azure.
Nov 2024
Beyond the Buzzwords: Let’s Talk About the Real Challenges in Data
We’re on a mission to cut through the noise and talk about the real challenges that data teams face in our fast-moving industry. And what…
Jun 2025
Cloud Independence: Testing a European Cloud Provider Against the Giants
Somewhere in Europe. You might be running a small online business, managing a mid-sized automotive supplier, or leading a global…
Mar 2025
Debugging Running Pods on Kubernetes
Exploring Kubernetes’s debugging feature, kubectl debug, and extending kubectl debug to support volume mounts
Mar 2025
Head-to-head comparison of dbt SQL engines
Compare usage and performance of dbt against 3 popular open-source SQL engines, namely: Spark, Trino and Duckdb
Aug 2024
How To Conquer The Complexity Of The Modern Data Stack
The more people you add to a team, the more lines of communication you introduce. The same holds for the number of tools in your data…
Oct 2024
How to Effectively Structure Data for Self-Service Data Teams
For years, data platforms — particularly data lakes and lakehouses — have relied on the medallion architecture. This tiered system…
Oct 2025
How to Prevent Crippling Your Infrastructure When AWS US-EAST-1 Fails
What the October 2025 outage reminded us about dependencies, preparedness, and the people behind the cloud.
Oct 2023
How we reduced our docker build times by 40%
This post describes two ways to speed up building your Docker images: caching build info remotely, using the link option when copying files
Mar 2025
Improve the security of pods on kubernetes
In this article we will show you four settings you should enable on your pods to improve the security of your kubernetes cluster.
Mar 2025
Improving Apache Spark performance on k8s
This post talks about adding external disks to your Kubernetes executors in order to speed up your spark jobs using the TPC-DS benchmark.
Apr 2025
Integrating MegaLinter to Automate Linting Across Multiple Codebases. A Technical Description.
If you’re not familiar with linters, or specifically with MegaLinter, please take a look at my previous article on the topic. In contrast…
Nov 2025
Introducing Conveyor: Build, Deploy & Scale Your Data Projects
Today, the team at Data Minded is proud to announce the launch of Conveyor, our managed, cloud-based data platform.
Jun 2024
Introducing Data Product Portal: An open source tool for scaling your data products
In the fast-evolving world of data, companies are discovering that the key to success for scaling their data initiatives is not to rely on…
Sep 2025
Locking down your data: fine-grained data access on EU Clouds
This post investigates how to restrict data access for both SQL and Spark/Python applications when using EU clouds
Feb 2024
My key takeaways for building a data engineering platform
Having been a member of a product team for two years, I aim to share three valuable insights that I have gained.
Jul 2025
Portable by design: Rethinking data platforms in the age of digital sovereignty
Recent geo-political and legal rulings have triggered us to investigate how data platforms can be designed for portability across providers
Jan 2024
Quacking Queries in the Azure Cloud with DuckDB
This post describes 2 Duckdb extensions that enable you to read data from Azure blob storage. It also shows code for both Python and dbt.
Aug 2025
Rethinking the data product workbench in the age of AI
In this blogpost we will explore the challenges related to building data products share our vision for a data product workbench using AI.
Apr 2025
Running thousands of Spark applications without losing your cool
I explain how to troubleshoot and detect problematic Spark applications at scale as well as show how this can be used to reduce your costs.
Mar 2024
Securely use Snowflake from VS Code in the browser
At Conveyor we help you to build, deploy, and scale your data products. One way we do that is by offering IDEs that run in the cloud. A…
Nov 2025
Slaying the Terraform Monostate Beast
You start out building your data platform. You choose Terraform because you want to do it the right way and put your infra in code. There’s…
Mar 2025
Source-Aligned Data Products: The Foundation of a Scalable Data Mesh
Introduction
Apr 2025
Stop loading bad quality data
Rule number one of having good quality data: Stop loading bad quality data. Really, it’s that simple. I see so many companies make this…
Mar 2024
The benefits of a data platform team
Shift your focus from maintenance to value
Apr 2025
The building blocks of successful Data Teams
Based on my experience I will elaborate on key criteria for building successful data teams
Apr 2025
The Data Engineer’s guide to optimizing Kubernetes
By default Kubernetes is not optimal for running batch workloads. I tackle spot instances, tweaks in autoscaling and node efficiency…
Aug 2024
The Data Product Portal Integrates With Your Preferred Data Platform
A couple of weeks ago we announced the release of the Data Product Portal as an open source repository. The Data Product Portal is an…
Jul 2024
The Missing Piece to Data Democratization is More Actionable Than a Catalog
As of the nineties, with the advent of Business Intelligence, organizations do attempt to install data driven decision making and do aim to…
Jul 2025
The Mystery of Folders on AWS S3
The difference between objects and files, and what the AWS Console really does when you press the “create folder” button.
Sep 2025
The ROI Challenge: Why Measuring Data’s Value is Hard, but Crucial
Here’s a story playing out in many data-driven organizations, and it often follows a familiar three-act structure.
Jul 2024
The State of Data Products in 2024
Gartner has released their hype cycle for data management 2024 quite recently and has identified Data Products at the gate of the peak of…
Sep 2025
The State of Data Work in 2025: Insights From 32 In-Depth Conversations
Understanding the real challenges faced by data professionals is essential for strategic planning. Our team recently completed an extensive…
Dec 2025
Using AWS IAM with STS as an identity Provider
How EKS tokens are created, and how we can use the same technique to use AWS IAM as an identity provider.
Mar 2025
What Is Data Product Thinking?
Data product thinking is gaining momentum in the data world right now so we recently organised a live online learning session for Conveyor…
Oct 2025
When writing SQL isn't enough: debugging PostgreSQL in production
While writing efficient SQL queries is essential, it is not enough to operate a database at scale. I will illustrate this using 3 issues
Mar 2025
Why data engineers should be more like software engineers
Data engineers are better when using a product mindset as well as software best practices: cicd pipelines, test code, develop iteratively.
Apr 2025
Why the ‘Private’ API Gateway of AWS Might Not Be as Secure as You Think
Designing secure applications is a challenge for everyone. A big part of this is based on who can access what. In this blog I want to dig…
Mar 2024
You can use a supercomputer to send an email but should you?
Discover the next evolution in data processing with DuckDB and Polars
Apr 2026
Authorizing AWS Principals on Azure
How to delegate trust from Entra to AWS IAM through Cognito, authorizing Azure actions without needing long-lived credentials.
Jun 2026
Beyond the Cron: Reclaiming Your Data Intervals in Hybrid Scheduling
Many data pipelines are either on a time schedule or triggered by events. Events are ideal for maximum data freshness, but are unpredictable.
Aug 2026
Building Idempotent Infrastructure: A Guide to provisioning your infrastructure with Data Product…
As data architecture scales, keeping your real-world infrastructure, like S3 buckets, database schemas, or Snowflake roles, in sync with your data governance platform is a big challenge.
Aug 2026
Can small LLMs match Claude Sonnet for agentic coding?
The cost of AI coding assistants is climbing and not just because prices are increasing.
Dec 2024
Data Modelling In A Data Product World
Many organisations are hitting the limits of data warehousing, especially as they grow in size. They often see adopting data products as a…
May 2026
Data Platforms for humans
At Dataminded we build Data Platforms for a living. We’re good at it and we have all kinds of things to say about the technical stuff.
May 2026
Dezoom! How to pitch a data platform to your leadership & organization.
This article has been co-written with my colleague, Jelle De Vleminck. Thanks for all the ideas coming from your gigantic reading list as…
Mar 2026
DuckLake Wants to Fix the Lakehouse. Can It?
While I was on the bench between projects at Dataminded, I had a chance to explore DuckLake and lakehouse architecture in more depth. A…
Mar 2026
From Idea to Implementation: Building an MCP Server for our Data Product Portal
Over the past weeks I’ve been experimenting with building an MCP (Model Context Protocol) server for our Data Product Portal. The portal is…
Mar 2026
From knowledge graphs, to star schemas, to data products
How three powerful concepts come together. A Data Engineer’s reflection.
Mar 2026
How to set up Databricks for data products?
Companies increasingly rely on data products for decision-making, insights, and AI/ML. data products separate data in meaningful entities…
Aug 2026
Making your self-hosted LLM production-ready
I wanted to host Gemma 431B and had an H100 GPU with 80 GB of VRAM. The math looked straightforward: 31 billion parameters × 2 bytes per parameter ≈ 62 GB, making the H100 a great fit for my model.
Jul 2026
Migrating a Data Platform: 20% technology, 80% everything else
At some point, when the bills become too high, or the platform is not scalable anymore, almost every organisation with a data platform reaches the same crossroads: “We need to migrate”.
Mar 2026
Musings from the Second Summit on Data Product Oriented Architectures
What is the state of the art in data products? Last week we joined the Second Summit on Data Product Oriented Architectures to find out…
Jun 2026
Not every data question deserves a full data product (and that's okay)
When we started building the Data Product Portal, we had a clear conviction: data product thinking is the right approach for getting value out of data in organisations. A data product has an owner.
May 2026
Portable S3 security for EU clouds
At the core of most modern data platforms sits a lakehouse on S3-compatible storage. It is the central store for all organisational data, from HR to sales and beyond.
Apr 2026
Scaling your Data Platform with Reference Data Products
A well-designed data platform makes it safe and easy to launch the infrastructure required to build successful data products. The platform…
Apr 2026
The Boring Stuff That Keeps You in the 5%
95% of GenAI pilots fail. What do you actually need to take GenAI from POC to production?
May 2026
uv scripts: micro-production situationship
Sometimes you are not trying to start a Python project at all.
Jul 2024
Why You Should Build A User Interface To Your Data Platform
Modern data platforms are complex. If you look at reference architectures, like the one from A16Z below, it contains 30+ boxes. Each box…
Mar 2026
Why Your Next Data Catalog Should Be a Marketplace
Explore why traditional data catalogs are failing to meet business needs and our vision for the solution: a data product marketplace.
Apr 2026
You Built a Data Mesh, But Your Metrics Are Still a Mess. Here’s Why.
At Dataminded, we work with organizations at various stages of their data journey. One pattern we keep running into: teams that have…
Mar 2026
You Don’t Have a Data Platform Without Excel
The Most Used Feature is “Export”: Why Your Data Stack Needs a Spreadsheet Strategy
Apr 2026
Your cloud strategy after the hyperscaler era
For years, choosing a cloud provider was mostly a question of features, price, and convenience.
Aug 13, 2020
Deploy analytics workloads
Kris Peeters and Stijn De Haes on where to deploy analytics workloads: VMs, serverless, or containers.
Oct 13, 2020
ML Industrialisation frameworks
Kris Peeters on ML industrialisation, comparing Azure ML, AWS SageMaker, and Google AI.
Jun 30, 2020
Save costs on cloud data analytics workloads
Pascal Brokmeier and Kris Peeters on optimizing cloud analytics costs across AWS, Azure, and GCP.
Oct 12, 2022
CI/CD for data projects: Why manual deployments are not good enough
Pierre Borckmans on bringing CI/CD confidence to the data project lifecycle with Conveyor.
Feb 4, 2022
Ensure smooth sailing by orchestrating your containerised data pipelines
Pierre Borckmans on doing MLOps Level 2 in real life with containers and an Airflow orchestrator.
Sep 7, 2022
How to Protect your AWS account against intruders
Stijn De Haes on improving and automating AWS account security.
Jun 23, 2022
Spark on Kubernetes in the real world
Stijn De Haes on writing fast, efficient Spark pipelines on Kubernetes.
Aug 10, 2022
What does it take to build a data platform
Niels Claeys on the real effort behind building a data platform, and common build-vs-buy misconceptions.
Jun 1, 2023
dbt + DuckDB vs Spark: What's the best way to build data pipelines in 2023?
Kris Peeters and Niels Claeys compare dbt + DuckDB against Spark for data lake transformations.
Dec 7, 2023
Future-Proofing Your Data Engineering Career: Essential Skills for 2024 and Beyond
Jelena Grujic, Cedric Mingneau, and Oliver Willekens on the essential skills for data engineers in 2024 and beyond.
Oct 16, 2023
Head-to-head comparison of 3 dbt SQL engines
Kris Peeters and Niels Claeys benchmark Trino, Spark, and DuckDB as SQL engines for dbt pipelines.
Mar 5, 2024
Decentralized vs Centralized: How to organize your data teams?
Kris Peeters, Kristof Martens, and Jelle De Vleminck on choosing between centralized and decentralized data team structures.
Sep 26, 2024
Empowering Every User: How Data Product Portal Accelerates Collaboration
Wannes Rosiers and Amy Raygada (Thoughtworks) on how the Data Product Portal accelerates collaboration across data roles.
Nov 14, 2024
Getting Data Done in Healthcare: Lessons from the Frontlines
Thomas Kranzkowski and Kris Peeters recap "The Doctor's Data Knight" on data and AI in healthcare.
Aug 28, 2024
Initial Reactions: Market Insights on the Data Product Portal
Wannes Rosiers and Niels Cautaerts (VITO) on the Data Product Portal's initial market reception.
Jul 30, 2024
Simplifying Data Management: Introducing the Data Product Portal
Kris Peeters and Wannes Rosiers introduce the open-source Data Product Portal.
Jun 27, 2024
The Impact of Product Thinking for Data
Kinda El Maarry and Wannes Rosiers on how product thinking transforms data governance.
Aug 18, 2025
European Clouds vs. Hyperscalers: Can You Really Build a Sovereign Data Platform?
Kris Peeters, Gergely Soti, and Niels Claeys on whether European cloud providers can replace US hyperscalers for sovereign data platforms.
Aug 27, 2026
From Scattered Data To Trustworthy Science (with VITO)
Katarina Milosevic and VITO's Tom Kuppens on turning scattered research data into a FAIR, AI-ready foundation.
Summer/Winter Bootcamp
The academy à la carte courses taught as part of the winter/summer school bootcamp track.
No matches found. Try a different search.




































































