huggingface kernels
OFFICIALLABSCO SUMMARY
Each of the four skills owns one hardware target and stays there: cpu-kernels covers C++ SIMD kernels for AVX2 and AVX512 with a two-phase correctness-then-performance workflow, cuda-kernels targets NVIDIA H100, A100, and T4 GPUs with kernel-builder/ABI3-compliant bindings (no pybind11, no setup.py) for models like LTX-Video, Stable Diffusion, LLaMA, Mistral, and Qwen, rocm-kernels writes Triton kernels for AMD's MI355X and R9700 covering RMSNorm, RoPE, GEGLU, and AdaLN, and xpu-kernels does the same for Intel's Battlemage and Arc Pro B50 GPUs through the Xe-Forge framework. All four are classed as needing only local tools — a compiler and the matching hardware or its emulator, not an account or API key.
This is for people building or optimizing a kernel to publish into kernels-community on the Hub, using the separate kernel-builder tool the skills assume is already set up. It is not for the much larger audience the kernels package itself targets: pip install kernels, then get_kernel("kernels-community/activation") to pull a precompiled kernel and run it, with no C++, CUDA, Triton, or hardware-specific knowledge required and none of these four skills relevant.
READ THE FULL ANALYSIS
The README never mentions any of this. huggingface/kernels' own README documents the Python package — portable, unique, and compatible kernel loading from the Hub — and points kernel authors toward a separate kernel-builder directory. Skills, Claude, or agent workflows do not appear anywhere in it; the four kernel-writing skills exist as files in the repository with no public documentation connecting them to the project's own stated purpose.
Checked 19 September 2026 from all four skill descriptions and the full README; we have not run any of the four hardware-specific build-and-benchmark workflows ourselves, since each needs the matching GPU family or CPU instruction set to test.
WHAT'S INSIDE
4 showing · 4 totalcpu-kernels
Hand-writing the heavy math inside an AI model so it runs fast on an ordinary server processor — correctness first, speed second, in that order.
cuda-kernels
Replaces one slow layer inside an image, video or language model with hand-tuned code for NVIDIA graphics cards, and proves it is actually faster.
rocm-kernels
Most fast GPU code assumes an NVIDIA card; this is how the same handful of hot operations get written for AMD hardware instead, and proved faster on it.
xpu-kernels
An automatic trial-and-error loop that rewrites a piece of PyTorch code over and over, benchmarking each attempt on an Intel graphics card until one is clearly faster.
HOW TO GET IT
npx skills add huggingface/kernelsnpx skills add huggingface/kernels --skill <name> --full-depthPick the skill name from the Skills tab — each entry there installs independently.