LANCER Lab
Official@lancerlab · China
Language And Compilation Optimization for Next-gen High Performance Computing Research (Joint Lab of Shanghai Enflame Technology and SJTU)
Agent Skills by LANCER Lab
Showing 19 vetted skills indexed across 1 GitHub repositories.
testing
Run unit tests and simulated sessions for GPU tuning scripts.
cursor-croq-tune
Orchestrate GPU kernel autotuning cycles for CUDA and DSL-targeted kernels on NVIDIA GPUs.
choreo-kernel-examples
Provide GPU kernel design patterns for warp specialization and pipeline staging in Choreo DSL.
cursor-croq-dsl-croqtile
Automate GPU kernel tuning for CroqTile and Choreo using DSL constructs.
cursor-croq-tune-iterative
Wrap GPU kernel tuning workflows with checkpoints for resume and manual boundary control.
croq-dsl-croqtile
Guides GPU kernel developers through CroqTile/Choreo build, test, and profiling workflows.
croq-dsl-helion
Validate environments, build, run, and profile Helion kernels with NVIDIA Nsight Compute.
croq-dsl-triton
Tunes and optimizes Triton GPU kernels with build, run, and profiling templates.
croq-tune
Automate GPU kernel tuning through iterative profiling, ideation, and implementation.
perf-nsight-compute-analysis
Analyze NVIDIA Nsight Compute reports to classify GPU kernel bottlenecks.
skill-creator
Design, test, and refine AI skills using structured prompts and eval datasets.
boost-harness
Analyze AI agent skill architectures to detect design flaws and propose improvements.
croq-dsl-cuda
Generate CUDA kernel build scripts and profile performance with NVIDIA Nsight Compute.
croq-dsl-cute-cpp
Build, run, and profile CuTe/CUTLASS C++ GPU kernels with nvcc.
croq-dsl-tilelang
Automate performance tuning and validation of TileLang GPU kernels.
croq-dsl-cute-dsl
Tune and validate Python JIT GPU kernels using CuTe DSL.
base-tune
Automate GPU kernel tuning through continuous profiling and structural optimization.
croq-tune-iterative
Checkpoint GPU kernel tuning sessions with human validation between runs.
choreo-syntax
Validate Choreo `.co` file syntax, primitives, and patterns for GPU kernels.
Frequently Asked Questions About LANCER Lab
FAQPage SchemaWhat specific tasks can engineers perform using LANCER Lab resources?▼
Engineers can perform iterative GPU kernel tuning, profile performance bottlenecks using Nsight Compute, validate Choreo and TileLang syntax, and optimize C++ CuTe or Triton implementations for high-performance hardware environments.
Which technical personas benefit from these kernel optimization capabilities?▼
These resources are designed for GPU kernel developers, high-performance computing researchers, and systems engineers focused on low-level hardware acceleration and domain-specific language compilation for NVIDIA architectures.
What are the primary dependencies for running these kernel tuning environments?▼
Execution requires a compatible NVIDIA GPU environment, the NVIDIA Nsight Compute suite for profiling, and standard C++ or JIT-based build environments for compiling CUDA, CuTe, or Triton kernel targets.