astronomer-data
Data engineering plugin - warehouse exploration, pipeline authoring, Airflow integration
LABSCO SUMMARY
A 27-skill toolkit from Astronomer covering Airflow DAG development, dbt integration, data lineage, and warehouse queries — most of it needs a running Airflow instance or a warehouse connection before it does anything.
Twenty of the 27 skills are grouped in the README into five areas: data discovery and analysis (warehouse-init, analyzing-data, checking-freshness, profiling-tables), data lineage (the tracing skills plus custom OpenLineage extractors), DAG development (the airflow entrypoint plus authoring, testing, debugging, deploying, blueprint templates, and human-in-the-loop workflows), dbt integration through Astronomer's own Cosmos, and Airflow 2-to-3 migration. The other seven — including dag-factory, an Airflow-plugin builder, and delegating-to-otto, which hands a task to Astronomer's own hosted Otto agent — aren't in that table at all, which says the README has fallen behind the skill list rather than the reverse.
This is not something to sample for general coding help: every skill assumes a project already built on Apache Airflow, open-source or Astronomer's managed Astro, and the warehouse-focused skills need a configured connection to Snowflake, Postgres, BigQuery, or another supported source before they answer a single question. If your team already runs Airflow, this saves re-explaining DAG conventions and warehouse schema on every task; if it doesn't, none of it applies.
What each skill needs before it runs. 18 of the 27 need a local tool already on your machine — the af CLI, a Jupyter kernel for warehouse queries, or the Astro CLI itself — 6 need an account key such as warehouse or Airflow API credentials, only 2 are ready with nothing configured, and one, dag-factory, we could not verify independently of the others.
The MCP server and CLI aren't part of the skill count above. astro-airflow-mcp, the Airflow REST API server, and its bundled af command-line tool ship in the same repository but install separately from the skills; the af CLI also collects anonymous command-name telemetry by default, opt out with af telemetry disable.
Data engineering plugin - warehouse exploration, pipeline authoring, Airflow integration
WHAT'S INSIDE
Nothing else to set up — install it and go.
GROUPS 6
35 of 35
Showing all 35 skills
Airflow 2 and Airflow 3 answer requests differently, and this is the translation layer inside this codebase that hides the difference and picks the right one by itself.
Puts a human decision in the middle of an automated pipeline: it stops and waits for someone to approve, pick an option, or fill in a form before carrying on.
Extends the Airflow dashboard itself — your own screens and your own code run inside it, instead of in a separate app you have to deploy and keep running.
Lets an Apache Airflow pipeline remember where it got to, so a step that fails and retries carries on from its last checkpoint instead of starting over or launching the same outside job twice.
Writes and extends Airflow data pipelines in the shape this project already uses, and catches import errors before you run them.
An engineer defines the building blocks once, in Python; from then on a pipeline is described in a short settings file rather than written as code.
Takes a dbt project — a set of SQL transformations — and runs it on Airflow's schedule, with each transformation appearing as its own step you can watch and retry.
Running a dbt project on Airflow when that project uses dbt's newer Fusion engine, which has tighter rules than dbt Core about where and how it can run.
Describes a whole Airflow pipeline in a configuration file rather than in Python — every step, its settings and what runs after what, listed in one file.
The command line for Apache Airflow, the system that runs a company's scheduled data pipelines — list what pipelines exist, start one, or find out why last night's run failed.
Works out why a data pipeline failed and writes it up: what actually broke, what it held up, the fix to run now, and how to stop it happening again.
Hands a job off to Otto, Astronomer's own Airflow specialist, which knows the version-by-version upgrade history that a general-purpose assistant would have to guess at.
Gets your Airflow pipelines off your laptop and running for real — on Astronomer's managed service, or on your own servers with Docker or Kubernetes.
Astronomer runs your data pipelines on its own servers; this is how you set those environments up from a terminal and push your code to them.
Airflow is the software that runs data pipelines on a schedule, and this is the part that keeps a working copy of it on your own laptop instead of a server.
Creates the starting folder for a new data-pipeline project and fills in what it needs to run — which packages to install, which databases to connect to.
The run-it-and-see loop for a data pipeline: start it, and if it fails, work out what broke, fix that, and start it again until it passes.
When pipelines are failing on your Astronomer deployment and nobody knows why, this works through its logs and settings in the order that usually finds the cause.
Ask a question about the company's data in plain English and get the answer back — it works out which tables hold it and runs the query for you.
Tells you whether the data in a table is current enough to trust right now, or whether something has stopped refreshing it.
Someone hands you a database table you have never seen before, and this writes up what is in it — every column, how many rows, how fresh it is, and what looks wrong with the data.
Builds a one-page map of your company's data warehouse — which table actually holds the customers, which one holds the orders — so nobody has to go hunting for the right table again.
Writes the Go code behind an Apache Airflow pipeline step, compiled into one native program, while the pipeline itself stays defined in Python. The Go SDK is experimental and not yet production-ready.
Writes the Java code behind an Apache Airflow pipeline step, with each run starting a short-lived Java process, while the pipeline itself is still defined in Python. The Java SDK is in preview.
The groundwork shared by every Airflow language SDK: the pipeline's schedule, order and retries stay in Python, while chosen steps run code written in another language such as Java or Go.
Sets up Airflow so that steps written in Java or Go reach the right runner, using two settings: one that names each runner and one that sends a queue to it.
Builds a Go-written Airflow step into one self-contained program, puts it where Airflow looks for it, and ships it on Docker, Kubernetes or Astro.
Packages the Java code behind Airflow steps into JAR files with Gradle or Maven and gets them onto the workers, on plain Docker or Kubernetes or on Astro.
An old add-on for calling AI models inside data pipelines has an official replacement now, and this does the swap across a whole project.
Moving to Airflow 3 is not just a version bump — this is the code side of it: everything in your own pipelines that version 3 no longer accepts, found and fixed.
Moves a data-pipeline project from Dagster to Airflow 3 on Astronomer's Astro one area at a time, with Dagster kept in charge until the new version matches it, and writes down plainly whatever is lost on the way.
Labels each pipeline step with the tables and files it reads and writes, so Airflow can draw a map of where the data came from even when a step reports nothing itself.
Teaches a pipeline step to report its own inputs and outputs in code, so a step from someone else's library still shows up on the data-flow map.
Answers the question you ask before changing a table or a pipeline — who else is reading it, and how much damage a change would do.
Follows a table back to where its data actually came from — which pipeline wrote it, and which system or file that pipeline read.
HOW TO GET IT
SINGLE SKILL
npx skills add astronomer/agents --skill <name> --full-depthPick the skill name from the Skills tab — each entry there installs independently.
PLUGIN MARKETPLACE
/plugin marketplace add astronomer/agents/plugin install astronomer-data@astronomerTyped inside the agent's own prompt, not in a terminal. The marketplace is called astronomer, which is the part after the @.