Skip to main content

Airflow Logs & Log Retention Policies

From Debugging to Governance πŸ“œπŸ”β€‹

The Story: When Something Breaks at 2 AM​

A DAG fails at 2:03 AM.

The scheduler looks healthy.
The task retried twice.
The UI just says β€œFailed”.

Now comes the only question that matters:

What actually happened?

That answer lives in Airflow logs.

Logs are not just debugging tools β€” they are:

  • Your audit trail
  • Your compliance evidence
  • Your root-cause analysis system

In this article, we’ll explain how Airflow logging works, where logs are stored, how retention works, and how to manage logs professionally at scale.


What Are Airflow Logs?​

Airflow logs capture everything that happens during task execution and system operation, including:

  • Task start and end
  • Errors and stack traces
  • Retry attempts
  • Environment details
  • Operator-level output
  • Scheduler and worker decisions

πŸ“Œ Logs answer why something happened β€” metadata only answers what happened.


Types of Logs in Apache Airflow​

Airflow generates multiple categories of logs, each serving a different purpose.


1️⃣ Task Logs (Most Important)​

Generated for every task execution.

Includes:

  • Python print() output
  • Exceptions and tracebacks
  • Operator logs
  • Retry information

πŸ“Œ This is what you see when you click β€œView Log” in the UI.


2️⃣ Scheduler Logs​

Generated by the Scheduler process.

Captures:

  • DAG parsing issues
  • Scheduling delays
  • Dependency resolution
  • Heartbeat failures

πŸ“Œ First place to check when DAGs don’t trigger.


3️⃣ Webserver Logs​

Generated by the Airflow Web UI.

Captures:

  • UI errors
  • Authentication issues
  • API requests

4️⃣ Worker Logs​

Generated by Celery / Kubernetes / Local workers.

Captures:

  • Task execution environment
  • Resource failures
  • Worker crashes

How Airflow Stores Logs (Architecture)​

By default, Airflow writes logs like this:

logs/
└── dag_id=example_dag/
└── task_id=extract_data/
└── execution_date=2024-01-10/
└── attempt=1.log

Each retry creates a new log file.

πŸ“Œ This structure enables:

  • Per-task isolation
  • Retry-level visibility
  • Easy debugging

Local Logging vs Remote Logging​

Local Logging (Default)​

Logs stored on:

  • Scheduler disk
  • Worker disk

❌ Problems:

  • Logs disappear if workers restart
  • Not scalable
  • Not suitable for distributed systems

Logs stored in:

  • Amazon S3
  • Google Cloud Storage
  • Azure Blob Storage
  • Elasticsearch (advanced)

βœ… Benefits:

  • Centralized logs
  • Durable storage
  • Multi-worker support
  • Compliance-friendly

πŸ“Œ Remote logging is mandatory for production.

(Remote logging is covered deeply in the next article.)


Log Lifecycle: From Creation to Cleanup​

Task Starts
β†’ Log File Created
β†’ Logs Written Incrementally
β†’ Task Ends
β†’ Log Archived (Optional)
β†’ Log Retention Policy Applies
β†’ Log Deleted

πŸ“Œ Without retention policies, logs grow endlessly.


What Is Log Retention in Airflow?​

Log retention defines:

  • How long logs are kept
  • When old logs are deleted
  • How storage costs are controlled

Retention applies to:

  • Task logs
  • Scheduler logs
  • Worker logs
  • Webserver logs

Why Log Retention Is Critical​

Without retention:

  • Disk fills up
  • Airflow crashes
  • Cloud storage costs explode
  • Compliance risks increase

With retention:

  • Predictable storage usage
  • Faster UI
  • Cleaner debugging
  • Compliance alignment

Log Retention Strategies​

Strategy 1: Time-Based Retention (Most Common)​

Example:

  • Keep logs for 30 days
  • Delete anything older
find airflow/logs -type f -mtime +30 -delete

πŸ“Œ Works well for most teams.


Strategy 2: Compliance-Based Retention​

Example:

  • Finance logs: 1 year
  • PII-related logs: 90 days
  • Dev logs: 14 days

πŸ“Œ Often implemented using remote storage lifecycle rules.


Strategy 3: Size-Based Retention​

Example:

  • Keep total logs under 500 GB
  • Delete oldest first

πŸ“Œ More complex, less common.


Log Retention with Remote Storage (Best Practice)​

If using S3 / GCS / Azure Blob, retention is usually handled by:

  • Lifecycle policies
  • Not Airflow itself

Example: S3 Lifecycle Rule (Conceptual)​

RuleAction
After 30 daysMove to Glacier
After 90 daysDelete

πŸ“Œ This is cheap, reliable, and scalable.


Input & Output Example (Realistic)​

Input​

  • 100 DAGs
  • Each DAG runs daily
  • 20 tasks per DAG
  • 2 retries
  • Logs retained for 30 days

Output​

MetricValue
Log files/day~4,000
Log files/month~120,000
Storage required~50–100 GB

πŸ“Œ Without retention β†’ unbounded growth.


Common Logging Mistakes​

❌ Storing logs only locally
❌ Never cleaning old logs
❌ Logging sensitive secrets
❌ Excessive debug logging
❌ Ignoring worker disk limits


Logging Best Practices (Production-Grade)​

βœ… Enable remote logging
βœ… Apply storage lifecycle policies
βœ… Mask secrets in logs
βœ… Use INFO level by default
βœ… Monitor log storage growth
βœ… Separate system logs from task logs


Relationship to Metadata Database​

LogsMetadata DB
Stores why something happenedStores what happened
Large filesStructured rows
Stored in storage backendStored in relational DB
Retention-criticalCleanup-critical

πŸ“Œ Both must be managed together.


Summary πŸ§ β€‹

  • Airflow logs are the foundation of debugging and observability
  • Task logs are the most valuable
  • Local logging doesn’t scale
  • Remote logging is production-standard
  • Log retention prevents failures and cost explosions
  • Storage lifecycle policies are the safest approach

Key Takeaways​

  • βœ… Logs explain failures β€” metadata does not
  • βœ… Remote logging is mandatory at scale
  • βœ… Retention policies protect stability and cost
  • βœ… Lifecycle rules outperform manual cleanup
  • βœ… Logging is part of governance, not just debugging

What’s Next? πŸš€β€‹

➑️ Remote Logging in Airflow (S3, GCS, Azure Blob)
Deep dive into configuration, architecture, and real production setups