firecrawl ai-research-skills
COMMUNITYLABSCO SUMMARY
The 83 skills span 20 categories: 8 in post-training (TRL, GRPO, OpenRLHF, verl, slime among them), 6 each in distributed training (DeepSpeed, FSDP2, Megatron-Core, Ray Train) and optimization (Flash Attention, GPTQ, AWQ, bitsandbytes), 5 in RAG (Chroma, FAISS, Pinecone, Qdrant), and smaller sets covering model architecture, tokenization, fine-tuning, mechanistic interpretability, safety and alignment, inference serving, evaluation, agents, multimodal models, prompt engineering, observability, infrastructure, MLOps, emerging techniques, and one dedicated skill for writing an ML paper.
This is for someone doing hands-on ML research or engineering — training, fine-tuning, evaluating, or deploying models — who wants their coding agent to already know a given framework's API and its rough edges, instead of guessing; it has nothing to do with scraping or extracting web content, despite sitting in an organization whose main product does exactly that.
READ THE FULL ANALYSIS
Who actually wrote this. The README's own byline is a separate company, Orchestra Research: its own npm package (@orchestra-research/ai-research-skills), its own Slack community, blog, and social accounts, and the word Firecrawl does not appear anywhere in the 12 KB of README we read.
What it costs to run. 66 of the 83 skills need only a local tool already on the machine — the framework itself, such as PyTorch or a CLI — 16 need an account key (GPU cloud vendors such as Modal and Lambda Labs, or services like LangSmith and Pinecone), and 1 talks to an MCP service directly.
We only read a third of the README. The full document is 37 KB; our material cuts off at 12 KB, partway through the category list. If a later section changes any of this — a credit to Firecrawl, a different license, a deprecation notice — we have not seen it.
ALSO IN THIS PACKAGE
model-architecture
LLM architectures and implementations including LitGPT, Mamba, NanoGPT, RWKV, and TorchTitan. Use when implementing, training, or understanding transformer and alternative architectures.
tokenization
Text tokenization for LLMs including HuggingFace Tokenizers and SentencePiece. Use when training custom tokenizers or handling multilingual text.
fine-tuning
LLM fine-tuning frameworks including Axolotl, LLaMA-Factory, PEFT, and Unsloth. Use when fine-tuning models with LoRA, QLoRA, or full fine-tuning.
mechanistic-interpretability
Neural network interpretability tools including TransformerLens, SAELens, NNSight, and pyvene. Use when analyzing model internals, finding circuits, or understanding how models compute.
data-processing
Data curation and processing at scale including NeMo Curator and Ray Data. Use when preparing training datasets or processing large-scale data.
post-training
RLHF and preference alignment including TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge. Use when aligning models with human preferences, training reward models, or large-scale RL training.
safety-alignment
AI safety and content moderation including Constitutional AI, LlamaGuard, NeMo Guardrails, and Prompt Guard. Use when implementing safety filters, content moderation, or prompt injection detection.
distributed-training
Multi-GPU and multi-node training including DeepSpeed, PyTorch FSDP, Accelerate, Megatron-Core, PyTorch Lightning, and Ray Train. Use when training large models across GPUs.
infrastructure
GPU cloud and compute orchestration including Modal, Lambda Labs, and SkyPilot. Use when deploying training jobs or managing GPU resources.
optimization
Model optimization and quantization including Flash Attention, bitsandbytes, GPTQ, AWQ, GGUF, and HQQ. Use when reducing memory, accelerating inference, or quantizing models.
evaluation
LLM benchmarking and evaluation including lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator. Use when benchmarking models or measuring performance.
inference-serving
Production LLM inference including vLLM, TensorRT-LLM, llama.cpp, and SGLang. Use when deploying models for production inference.
mlops
ML experiment tracking and lifecycle including Weights & Biases, MLflow, and TensorBoard. Use when tracking experiments or managing models.
agents
LLM agent frameworks including LangChain, LlamaIndex, CrewAI, and AutoGPT. Use when building chatbots, autonomous agents, or tool-using systems.
rag
Retrieval-Augmented Generation including Chroma, FAISS, Pinecone, Qdrant, and Sentence Transformers. Use when building semantic search or document retrieval systems.
prompt-engineering
Structured LLM outputs including DSPy, Instructor, Guidance, and Outlines. Use when extracting structured data or constraining LLM outputs.
observability
LLM application monitoring including LangSmith and Phoenix. Use when debugging LLM apps or monitoring production systems.
multimodal
Vision, audio, and multimodal models including CLIP, Whisper, LLaVA, BLIP-2, Segment Anything, Stable Diffusion, and AudioCraft. Use when working with images, audio, or multimodal tasks.
emerging-techniques
Advanced ML techniques including MoE Training, Model Merging, Long Context, Speculative Decoding, Knowledge Distillation, and Model Pruning. Use when implementing cutting-edge optimization or architecture techniques.
ml-paper-writing
Write publication-ready ML/AI papers for NeurIPS, ICML, ICLR, ACL, AAAI, COLM. Includes LaTeX templates, citation verification, reviewer guidelines, and writing best practices from top researchers.
WHAT'S INSIDE
83 showing · 83 totalaudiocraft-audio-generation
Generates music and sound effects from text descriptions using Meta's AudioCraft models, with options to shape the result by melody, style reference, or stereo output.
autogpt-agents
A self-hosted platform for building AI agents as drag-and-drop workflows, or with a Python developer toolkit, that keep running continuously and respond to triggers like a schedule or an incoming webhook.
awq-quantization
Shrinks a large AI model to about a quarter of the memory so it fits on a smaller graphics card and runs roughly three times faster, protecting the few numbers inside that matter most so quality barely drops.
axolotl
Retrains an open-source AI model on your own data from a single small config file instead of code, including the cheaper methods that only adjust part of the model and training spread across many GPUs.
blip-2-vision-language
Describes what is in a photo and answers questions about it in plain language — it joins an existing image model to an existing language model with a small bridge, so neither of them has to be retrained.
chroma
The store an AI app keeps its documents in so it can pull back the ones closest in meaning to a question — a free, open-source vector database that runs in a notebook or on a server.
clip
OpenAI's CLIP scores how well a picture and a piece of text go together — enough to sort photos into any categories you name, search a library by description, or flag unwanted images, with no training of your own.
constitutional-ai
Anthropic's method for making a model safer without people having to label harmful content: the model judges its own answers against a short written set of principles and rewrites them, then is trained on its own judgements.
crewai-multi-agent
Assembles a team of role-based AI agents — each with a role, goal, and backstory — that hand tasks off to each other in sequence or under a manager agent, without depending on LangChain.
deepspeed
A model too big for one graphics card, trained anyway — Microsoft's DeepSpeed splits it and everything training needs across many GPUs, and spills to ordinary memory or disk when even that is not enough.
distributed-llm-pretraining-torchtitan
PyTorch's official framework for pretraining large language models from scratch, combining four kinds of parallelism — FSDP2, tensor, pipeline, and context — to scale one training run from a single node to hundreds of GPUs.
dspy
Instead of hand-writing prompts, you declare what goes in and what should come out — then a search over your own examples and your own scoring finds the wording that actually performs best.
evaluating-code-models
Measures how good a model is at writing code that actually runs, against the standard problem sets everyone quotes — and in 18 languages, not just Python.
evaluating-llms-harness
Runs a language model against 60+ standardized academic benchmarks — MMLU, GSM8K, HellaSwag, TruthfulQA, and more — using the same prompts everyone else uses, so scores are directly comparable to published results.
faiss
Facebook's library for finding, among millions or billions of items, the handful most like the one you are holding — with knobs to trade accuracy for speed or memory, and an option to run on the graphics card.
fine-tuning-with-trl
Teaches a model which answers people prefer — show it good examples, train a scorer on pairs of better-and-worse answers, then let the model practise against that scorer, alone or as the whole chain.
gguf-quantization
Turns a downloaded AI model into a single self-contained file, shrunk by as much or as little as you choose, that ordinary desktop apps can open and run on your own machine.
gptq
A very large AI model normally wants more graphics-card memory than anyone has spare — this cuts what it needs to roughly a quarter and makes it answer three to four times faster, with the answers themselves barely changed.
grpo-rl-training
Instead of showing a model the right answers, you write rules that score its attempts — it takes several shots at every question and learns to produce more of whatever your rules rate highly.
guidance
You normally ask a model for output in a particular format and hope it obeys; this locks the format in while the text is being written, so a reply that does not fit simply cannot come out.
hqq-quantization
Other ways of shrinking an AI model want a pile of sample text first and hours of processing; this one needs neither — point it at the model and a much smaller version comes back in minutes.
huggingface-accelerate
You write the training script once and choose afterwards where it runs — one graphics card, eight of them, several machines, or a TPU — instead of rewriting it for each.
huggingface-tokenizers
Text has to be chopped into small pieces and turned into numbers before a model can read it; this does the chopping at about a gigabyte every twenty seconds, and can learn a set of pieces tailored to your own kind of writing.
implementing-llms-litgpt
Twenty-odd well-known open AI models — Llama, Gemma, Phi, Qwen, Mistral — each written out as one readable file instead of being buried under layers of framework, so you can follow what the model actually does and change it.
instructor
A layer that sits over an AI model and turns its answer into a filled-in form rather than a paragraph — and when the answer doesn't fit the form, it sends it back to the model to try again.
knowledge-distillation
Trains a smaller "student" model to imitate a larger "teacher" model's behavior, so you end up with a model that keeps most of the teacher's quality at a fraction of its size and inference cost.
lambda-labs-gpu-cloud
Rents dedicated GPU machines by the minute, from a single card up to a multi-hundred-GPU cluster, with full SSH access, a pre-installed ML software stack, and storage that survives a restart.
langchain
The pieces an AI app usually needs — a model, your own documents, outside tools it can call, a memory of the conversation — already exist here and snap together, so a chatbot or a question-answering app gets assembled rather than written from scratch.
langsmith-observability
When an AI feature gives a bad answer there is usually no way to see why it did — this records everything that happened behind each request, and lets you collect the bad cases into a test set you re-run after every change.
llama-cpp
The program that actually runs an AI chat model on hardware people already own, with no Nvidia graphics card required and no stack of other software to install underneath it.
llama-factory
llamaguard
A second model from Meta that sits between a chatbot and its users, reading each message and each reply and marking anything that falls into one of six risk areas — violence and hate, sexual content, weapons, drugs, self-harm, or crime.
llamaindex
Reads a pile of your own documents — files, web pages, a database — into a form an AI model can search, so its answers come from what is written there instead of from whatever it half-remembers.
llava
An open model you can run on your own machine that looks at a picture and talks about it — you ask what is in the photo, and then keep asking follow-up questions about the same one.
long-context
A model that could only take in a short report at a time can be stretched to swallow a whole book — this is the four established ways of doing that stretching, and how much retraining each one costs.
mamba-architecture
A different design for the inside of an AI model: the usual kind gets slower and hungrier for memory the longer the text runs, while this one holds a steady cost per word and generates about five times faster.
miles-rl-training
The very largest open AI models break the ordinary reward-based training tools — the run goes unstable and the model no longer fits the hardware — and this is the industrial-grade version built for exactly that scale.
ml-paper-writing
Turns a research repository's code and results into a submission-ready paper draft for a top AI conference, verifying every citation against a live search instead of ever inventing one from memory.
mlflow
Logs the parameters, metrics, and model file from each training run, then tracks that model's versions through a shared registry as it moves from staging to production.
modal-serverless-gpu
Runs Python functions on demand in the cloud with GPUs attached, scaling containers from zero to over a hundred instances automatically and billing only for the seconds they actually run.
model-merging
Combines the weights of several fine-tuned models into one, blending their different skills — like math, coding, and chat — without any further training.
model-pruning
Zeroes out the least important weights in a trained model, cutting its size 40-60% with under 1% accuracy loss and no retraining required.
moe-training
Instead of one big network where every part works on every word, the model is built as a panel of specialists with a router that hands each word to just a couple of them — so it can hold far more knowledge without costing much more to run.
nanogpt
A roughly 300-line, from-scratch GPT implementation that trains in minutes on a CPU for small experiments and scales up to reproduce GPT-2 on multiple GPUs, built for understanding transformers rather than production use.
nemo-curator
A model is only as good as the pile of text, images, video or audio it learns from — this is the cleaning step before training: duplicates thrown out, junk filtered, personal details and explicit material stripped, all of it running across graphics cards rather than ordinary processors.
nemo-evaluator-sdk
Runs a model against 100 or more standard benchmarks — academic, coding, safety, and vision-language — from inside reproducible containers, executing locally, on a Slurm cluster, or in the cloud.
nemo-guardrails
Every message going in and every reply coming out passes a set of checks first — for attempts to talk the bot out of its rules, for made-up facts, for personal details, and for abusive content.
nnsight-remote-interpretability
Opens up the inside of an AI model so you can watch what each layer is doing while it works, and change a layer partway through to see how the answer shifts.
openrlhf-training
Teaching a model what people prefer takes a whole cluster running in lockstep — one part writes answers, another scores them, another retrains on the scores — and this is the machinery that coordinates it, at about twice the speed of the older option.
optimizing-attention-flash
Inside every modern AI model, one step compares every word against every other word, and that step is what devours the memory; this swaps it for a version that needs ten to twenty times less and returns exactly the same answers.
outlines
A model running on your own hardware can be pinned to a shape you define in advance, so the text it produces always fits — nothing left to parse afterwards, and no retry loop for when it does not.
peft-fine-tuning
Fine-tunes a large language model by training small adapter weights instead of the full model, cutting GPU memory and storage needs to a fraction of full fine-tuning.
phoenix-observability
Everything an AI app does gets recorded step by step, on servers you run yourself rather than a vendor's — the traces, the test sets built out of them, and the scores, all in one place you own.
pinecone
A hosted database that finds things by meaning rather than by matching words: each document goes in as a numeric fingerprint, and a query comes back with the closest ones — scaling to billions of them without you running a single server.
prompt-guard
A small, fast checker for one specific danger: text that is trying to hijack a model's instructions, whether a user typed it in or it was hidden inside a web page or document the model is about to read.
pytorch-fsdp2
When a model is too big to fit on one graphics card, its parts have to be cut up and spread across several — this is how to make that split correctly in PyTorch, including the order of steps it has to be done in.
pytorch-lightning
Training code written by hand is mostly the same plumbing every time; put yours into this standard shape and the plumbing comes written for you, whether the run ends up on a laptop or on a supercomputer.
pyvene-interventions
Watching one part of a model light up does not prove it is responsible for anything; this changes that part on purpose and checks whether the model's behaviour changes with it.
qdrant-vector-search
A search engine that matches on meaning rather than on the words themselves, and that you run on your own machines — kept on your premises, under your control, instead of handed to a hosted service.
quantizing-models-bitsandbytes
Loads a large language model in 8-bit or 4-bit precision instead of full precision, cutting GPU memory use by 50 to 75 percent with little accuracy loss.
ray-data
The step that gets data to a model — reading it, cleaning it, resizing it, handing it over — spread across a whole cluster instead of one machine, and taken a piece at a time rather than all at once.
ray-train
Training across a room full of separate machines is more an organising problem than a coding one — who runs what, who reports to whom, what happens when one of them dies — and this is the part that handles all of that.
rwkv-architecture
A model design with two modes: it learns all at once, the fast modern way, then writes text one word at a time on a fixed amount of memory, no matter how long the piece runs.
segment-anything-model
Trained on a billion outlines of every kind of thing, it can trace the exact shape of an object it has never been shown before — a cell under a microscope, a roof in a satellite photo — from one mark you put on the picture.
sentence-transformers
Text and images get turned into lists of numbers that land close together when the meaning is close — which is the trick behind "find me something like this" — and all of it runs on your own machine.
sentencepiece
Chopping text into pieces normally starts by splitting on spaces, which falls apart for Chinese, Japanese and Korean, where there are none; this works straight off the raw characters instead, so one method covers every language.
serving-llms-vllm
Answering one request at a time leaves most of an expensive graphics card sitting idle; this keeps many conversations in flight at once and gets roughly twenty-four times the work out of the same hardware.
sglang
Nearly every request to a chatbot or an agent opens with the same block of text, and most servers redo that work every single time — this one remembers it, so only the genuinely new part of a message costs anything.
simpo-training
Teaching a model which of two answers people prefer normally means keeping a second, untouched copy of it alongside as a yardstick; this method throws that copy away and still comes out ahead of the one that keeps it.
skypilot-multi-cloud-orchestration
Renting computing power from a cloud company normally ties your work to that one company's way of doing things; here you write the job down once and it will run on any of twenty-odd providers, picking itself back up if a cheap machine gets taken away mid-run.
slime-rl-training
The framework Tsinghua's team actually used to build the GLM models, released openly: it runs the stage after pre-training where a model improves by being rewarded, and leaves the logic for producing practice material up to you.
sparse-autoencoder-training
A single unit inside a model fires for several unrelated things at once, which is why nobody can read it directly; this pulls those tangled signals apart into a much larger set where each one stands for one recognisable idea.
speculative-decoding
A chatbot writes its reply one word at a time, and every single word costs a full trip through a huge model; letting a small quick model guess the next few words and having the big one merely check them gives back exactly the same reply, up to three times faster.
stable-diffusion-image-generation
Describe a picture in words and get it back as an image, made on your own machine rather than someone's website — and the same models will redraw a chosen patch of a photo you already have, or hold a pose or an outline while changing everything around it.
tensorboard
While a model is being trained, all you otherwise get is numbers scrolling past in a terminal; this collects them as they go and draws them as charts in a browser tab, so you can see whether the thing is actually learning.
tensorrt-llm
Rather than running a model as it comes, this compiles it ahead of time into a build tuned for the exact NVIDIA card it will sit on — the fastest option available, at the price of that build step and of being tied to one make of hardware.
torchforge-rl-training
You write the rule for what counts as a good answer, and nothing else — the machines, the generating, and the keeping-everything-in-step are all handled underneath.
training-llms-megatron
Building a model from scratch at the largest sizes means splitting the work across thousands of graphics cards in several directions at once — this is the machinery that does the splitting, tuned so as little of that hardware as possible sits idle.
transformer-lens-interpretability
The standard workbench for mapping the circuitry inside a language model: which parts feed into which, in what order, and what each of them is carrying while the model works out its answer.
unsloth
verl-rl-training
A training rig for teaching a language model by rewarding its good answers, built to run across a whole cluster of GPUs rather than one machine.
weights-and-biases
Training a model means dozens of attempts, and by week two nobody remembers which one worked — this keeps a live record of every run, its settings and its results.
whisper
A speech-to-text model you run yourself, on your own machine, covering 99 languages and able to turn foreign-language audio straight into English.
HOW TO GET IT
npx skills add firecrawl/ai-research-skillsnpx skills add firecrawl/ai-research-skills --skill <name> --full-depthPick the skill name from the Skills tab — each entry there installs independently.
/plugin marketplace add firecrawl/ai-research-skills/plugin install model-architecture@ai-research-skills/plugin install tokenization@ai-research-skills/plugin install fine-tuning@ai-research-skillsTyped inside the agent's own prompt, not in a terminal. The marketplace lists 20 plugins; three are shown. Every one installs the same way, with its own name before @ai-research-skills.