Setting up Airflow (and other OSS) - Past vs Today

From my experience working at GreyOrange. Refactored my article a bit with help of GPT.

I had the fortune of working in data teams of different companies, and one thing that I noticed is the evolution of setting up OSS (open-source software) on cloud.

Tldr; From VMs to Docker to Helm-Kubernetes

For instance, let’s talk about Airflow. Apache Airflow is an open-source workflow orchestration platform. You define data pipelines as DAGs (Directed Acyclic Graphs) in Python — each DAG describes what tasks to run, in what order, and on what schedule. Airflow handles scheduling, execution, retries, logging, and a web UI for visibility.

It has a few core components like scheduler, workers, webserver/apiserver, etc.

When setting up for the first time on cloud - the deployment model for all these components have changed over years. I would divide it as:

1) VM era:

Teams would spin up one or more virtual machines and install Airflow directly on the OS.

# Typical setup steps on Ubuntu/Debian
sudo apt-get install python3-pip
pip install apache-airflow
export AIRFLOW_HOME=~/airflow
airflow db init
airflow users create --username admin --role Admin ...
airflow webserver --port 8080 &
airflow scheduler &

Workers would be started similarly, either on the same machine or additional VMs. The metadata database (Postgres) was either installed on the same VM or pointed to an external managed instance.

Managing looked like:

Pros:

Cons:

2) Container era:

The era of Docker came. Instead of installing Airflow on a host OS, you’d pull (or build) a Docker image containing Airflow and all its dependencies, then run it as a container.

# Custom Airflow image with extra packages
FROM apache/airflow:3.x.x
RUN pip install pandas boto3 google-cloud-bigquery
COPY dags/ /opt/airflow/dags/
# Running Airflow webserver in Docker
docker run -d \
  -p 8080:8080 \
  -e AIRFLOW__CORE__EXECUTOR=LocalExecutor \
  -e AIRFLOW__DATABASE__SQL_ALCHEMY_CONN=postgresql+psycopg2://airflow:airflow@postgres/airflow \
  -v $(pwd)/dags:/opt/airflow/dags \
  apache/airflow:3.x.x webserver

Airflow project provided an official docker-compose.yaml as well to run entire stack.

Managing looked like:

Pros:

Cons:

3) K8s era:

We can now setup Airflow using helm charts on K8s. K8s provides:

But deploying a complex multi-component application like Airflow on Kubernetes by hand (writing Deployments, Services, PVCs, ConfigMaps, Secrets for each component) is repetitive and error-prone. That’s where Helm comes in. Helm is the package manager for Kubernetes. A Helm chart is a collection of templated Kubernetes manifests packaged together with configurable values.

We can think of it like apt or pip but for Kubernetes applications.

# Install Airflow on Kubernetes with Helm
helm repo add apache-airflow https://airflow.apache.org
helm install airflow apache-airflow/airflow \
  --namespace airflow \
  --create-namespace \
  -f values.yaml

One command which you could run from a bastion machine that has access to the K8s namespace and Airflow’s entire stack — scheduler, webserver, workers, triggerer, metadata DB connection, ingress, RBAC — deployed on that namespace.

Notice in the command, we are passing a values.yaml file, it basically has our complete Airflow deployment config, it may look like below for example:

# Airflow image — use official or custom
images:
  airflow:
    repository: your-registry/custom-airflow
    tag: "3.x.x"
    pullPolicy: IfNotPresent

# Executor choice
executor: KubernetesExecutor   # or CeleryExecutor, LocalExecutor

# Webserver
webserver:
  replicas: 2
  resources:
    requests:
      memory: "1Gi"
      cpu: "500m"
    limits:
      memory: "2Gi"
      cpu: "1000m"

# Scheduler
scheduler:
  replicas: 1
  resources:
    requests:
      memory: "2Gi"
      cpu: "1000m"

# Workers (CeleryExecutor)
workers:
  replicas: 3
  resources:
    requests:
      memory: "4Gi"
      cpu: "2000m"

...Similarly database, env variables, and other details... 

There are 3 common patterns to deploy/update DAGs via helm:

Similarly, upgrading airflow via helm is simple as well.

Pros:

Cons:

Hence, for teams setting up OSS Airflow rather than using any managed service, helm is preffered way. Note that a lot of other OSS like say: Superset, Trino, Metabase, their helm charts exist as well, and can be setup similarly.