LabscoConnect MCP ↗
PLUGIN / ASTRONOMEROFFICIAL

Astronomer Data Engineering

From Astronomer: skills for writing Apache Airflow pipelines, diagnosing the ones that broke, and querying the warehouse they load.

What gets installed: 34 skills, 2 hooks

Skills

Questions about the data itself 4

These answer a question instead of building something: which table holds what, whether it is fresh enough to rely on this morning, and what a column really contains once somebody looks inside it.

analyzing-data

Ask a question about the company's data in plain English and get the answer back — it works out which tables hold it and runs the query for you.

checking-freshness

Tells you whether the data in a table is current enough to trust right now, or whether something has stopped refreshing it.

profiling-tables

Someone hands you a database table you have never seen before, and this writes up what is in it — every column, how many rows, how fresh it is, and what looks wrong with the data.

warehouse-init

Builds a one-page map of your company's data warehouse — which table actually holds the customers, which one holds the orders — so nobody has to go hunting for the right table again.

Code you hand to Airflow 8

The things you write and Airflow then runs: pipelines in Python, pipelines described in a settings file instead of code, a step that stops and waits for a person to approve it, a step that picks up where it left off after a crash, a dbt project put on Airflow's schedule, and your own screens inside the Airflow dashboard.

airflow-hitl

Puts a human decision in the middle of an automated pipeline: it stops and waits for someone to approve, pick an option, or fill in a form before carrying on.

airflow-plugins

Extends the Airflow dashboard itself — your own screens and your own code run inside it, instead of in a separate app you have to deploy and keep running.

airflow-state-store

Lets an Apache Airflow pipeline remember where it got to, so a step that fails and retries carries on from its last checkpoint instead of starting over or launching the same outside job twice.

authoring-dags

Writes and extends Airflow data pipelines in the shape this project already uses, and catches import errors before you run them.

blueprint

An engineer defines the building blocks once, in Python; from then on a pipeline is described in a short settings file rather than written as code.

cosmos-dbt-core

Takes a dbt project — a set of SQL transformations — and runs it on Airflow's schedule, with each transformation appearing as its own step you can watch and retry.

cosmos-dbt-fusion

Running a dbt project on Airflow when that project uses dbt's newer Fusion engine, which has tighter rules than dbt Core about where and how it can run.

dag-factory

Describes a whole Airflow pipeline in a configuration file rather than in Python — every step, its settings and what runs after what, listed in one file.

Steps written in Go or Java 6

For a step whose work is done in Go, Java or Kotlin while the pipeline itself stays in Python: writing that code, packing it into a bundle the workers can run, and setting Airflow up to hand those steps to it.

authoring-go-sdk-tasks

Writes the Go code behind an Apache Airflow pipeline step, compiled into one native program, while the pipeline itself stays defined in Python. The Go SDK is experimental and not yet production-ready.

authoring-java-sdk-tasks

Writes the Java code behind an Apache Airflow pipeline step, with each run starting a short-lived Java process, while the pipeline itself is still defined in Python. The Java SDK is in preview.

authoring-language-sdk-tasks

The groundwork shared by every Airflow language SDK: the pipeline's schedule, order and retries stay in Python, while chosen steps run code written in another language such as Java or Go.

configuring-airflow-language-sdks

Sets up Airflow so that steps written in Java or Go reach the right runner, using two settings: one that names each runner and one that sends a queue to it.

deploying-go-sdk-bundles

Builds a Go-written Airflow step into one self-contained program, puts it where Airflow looks for it, and ships it on Docker, Kubernetes or Astro.

deploying-java-sdk-bundles

Packages the Java code behind Airflow steps into JAR files with Gradle or Maven and gets them onto the workers, on plain Docker or Kubernetes or on Astro.

Getting it running, and finding out why it stopped 9

From a bare folder to a pipeline that died at three in the morning on a production deployment: project setup, an Airflow running on your own machine, the test-then-fix loop, release to Astronomer's service or to your own servers, and the two that go reading a deployment's logs.

airflow

The command line for Apache Airflow, the system that runs a company's scheduled data pipelines — list what pipelines exist, start one, or find out why last night's run failed.

debugging-dags

Works out why a data pipeline failed and writes it up: what actually broke, what it held up, the fix to run now, and how to stop it happening again.

delegating-to-otto

Hands a job off to Otto, Astronomer's own Airflow specialist, which knows the version-by-version upgrade history that a general-purpose assistant would have to guess at.

deploying-airflow

Gets your Airflow pipelines off your laptop and running for real — on Astronomer's managed service, or on your own servers with Docker or Kubernetes.

managing-astro-deployments

Astronomer runs your data pipelines on its own servers; this is how you set those environments up from a terminal and push your code to them.

managing-astro-local-env

Airflow is the software that runs data pipelines on a schedule, and this is the part that keeps a working copy of it on your own laptop instead of a server.

setting-up-astro-project

Creates the starting folder for a new data-pipeline project and fills in what it needs to run — which packages to install, which databases to connect to.

testing-dags

The run-it-and-see loop for a data pipeline: start it, and if it fails, work out what broke, fix that, and start it again until it passes.

troubleshooting-astro-deployments

When pipelines are failing on your Astronomer deployment and nobody knows why, this works through its logs and settings in the order that usually finds the cause.

Where the data came from 4

About the map rather than the pipeline: follow a table back to the system that fed it, see what downstream would break before changing it, and two for making a step declare what it read and wrote when it reports nothing by itself.

annotating-task-lineage

Labels each pipeline step with the tables and files it reads and writes, so Airflow can draw a map of where the data came from even when a step reports nothing itself.

creating-openlineage-extractors

Teaches a pipeline step to report its own inputs and outputs in code, so a step from someone else's library still shows up on the data-flow map.

tracing-downstream-lineage

Answers the question you ask before changing a table or a pipeline — who else is reading it, and how much damage a change would do.

tracing-upstream-lineage

Follows a table back to where its data actually came from — which pipeline wrote it, and which system or file that pipeline read.

Surviving a version change, or a move 3

For code that has to come through a move intact: a project walked up from an older Airflow to the current one, a Dagster project carried over to Airflow, and an AI add-on swapped for the provider that replaced it.

migrating-ai-sdk-to-common-ai

An old add-on for calling AI models inside data pipelines has an official replacement now, and this does the swap across a whole project.

migrating-airflow-2-to-3

Moving to Airflow 3 is not just a version bump — this is the code side of it: everything in your own pipelines that version 3 no longer accepts, found and fixed.

migrating-dagster-to-airflow

Moves a data-pipeline project from Dagster to Airflow 3 on Astronomer's Astro one area at a time, with Dagster kept in charge until the new version matches it, and writes down plainly whatever is lost on the way.

A warm-up when a session opens, a shut-off after each reply.

Installing the plugin also registers two scripts that run by themselves at fixed moments; neither asks before it runs.

When a session opens

Runs the af command once in the background, output discarded, so the Airflow package behind it is already fetched when the first real call comes. If af is not installed, it does nothing.

Each time the assistant stops

Shuts down the background Python session the warehouse skill keeps open for queries. The next query starts a fresh one and connects to the warehouse again.

Read more

Both run on your own machine with your own permissions. Each serves another part of the set: one the af Airflow command line, the other the warehouse skill's background Python session.

The second is started through uv, so it fails each time it fires on a machine without uv. Install uv to clear it.

Labsco Summary

Astronomer's skills help an assistant write Airflow pipelines, work out why one failed, and query the warehouse tables they fill. Show

Airflow, and the tables behind it.

Astronomer builds the managed version of Apache Airflow, the scheduler a company's overnight data loads run on, and these skills are its own answer for working on one.

Two different questions come out of that work, and the set takes both: why did last night's run fail and what should the pipeline do differently — and, separately, what the table it wrote actually holds.

What debugging and querying need running.

Writing a pipeline can start from an empty folder, but the skills that debug a failed run or query its data expect something of yours to be up already:

  • an Airflow it can reach — the open-source project, or Astronomer's hosted Astro
  • a warehouse it is allowed to query: Snowflake, Postgres, BigQuery, or anything SQLAlchemy can open

The warehouse half reads that login from a file of its own, ~/.astro/agents/warehouse.yml, and answers nothing until someone has put it there.

Astronomer's own service, or plain Airflow.

Four of the skills are about Astro, Astronomer's hosted Airflow, and nothing else: starting a project, keeping one alive on your laptop, pushing code up to a deployment, and reading that deployment's logs when it misbehaves.

On open-source Airflow those four are the wrong instructions — the deploying skill carries the Docker Compose and Helm route instead, and the rest of the set never asks whose Airflow it is.

One skill hands the job to Otto instead.

One entry in the set does no Airflow work at all. delegating-to-otto passes the request to Otto, Astronomer's own Airflow agent, and its instructions tell the assistant to offer Otto for a major Airflow version upgrade even when nobody asked for Otto — on the grounds that Otto holds compatibility knowledge the local upgrade skill does not.

That route arrives with the rest, so it is worth deciding up front whether you want the assistant taking it.

02 INSTALL

One plugin holds the whole set:
the skills, and two scripts the plugin runs by itself — one as a session opens, one each time the assistant finishes a reply.

PLUGIN MARKETPLACE

Add the marketplace, then install the plugin

Typed inside the agent's own prompt, not in a terminal. The marketplace is called astronomer, which is the part after the @.

  1. /plugin marketplace add astronomer/agents
  2. /plugin install astronomer-data@astronomer