唯品会
Official@vipshop · Guangzhou China
全球精选,正品特卖 (NASDAQ: VIPS)
Agent Skills by 唯品会
Showing 7 vetted skills indexed across 1 GitHub repositories.
triton-kernel
Write optimized Triton GPU kernels for deep learning operations.
cuda-cpp-kernel
Implements, debugs, and optimizes CUDA C++ and PTX kernels across NVIDIA GPU architectures.
cache-dit-model-integration
Integrates new DiT models into cache-dit with cache, parallelism, and CLI support.
ptq-workflow-integration
Integrate a new PTQ workflow into cache-dit with a public API and validation.
cute-dsl-kernel
Implement and optimize CuTe DSL GPU kernels across NVIDIA architectures.
operator-migration
Migrate native CUDA, Triton, or custom operators into cache-dit with import safety.
cutlass-cpp-kernel
Implement and optimize CUTLASS and CuTe C++ kernels with validation.
Frequently Asked Questions About 唯品会
FAQPage SchemaWhat specific tasks are enabled by these kernel optimization skills?▼
These skills enable the implementation and optimization of high-performance GPU kernels using CuTe DSL and CUTLASS. They facilitate the migration of native CUDA or Triton operators into the cache-dit environment, ensuring rigorous validation and performance parity across NVIDIA hardware architectures.
Which engineering personas benefit from these technical capabilities?▼
These capabilities are designed for GPU performance engineers, machine learning infrastructure developers, and systems architects focused on low-level hardware acceleration. They are specifically intended for engineers tasked with optimizing tensor computation layers and managing operator portability within high-performance computing environments.
What are the primary prerequisites for implementing these kernels?▼
Implementation requires a deep understanding of NVIDIA GPU architectures, proficiency in C++ for high-performance computing, and familiarity with the CUTLASS and CuTe programming models. Users must also have an existing cache-dit environment to integrate and validate the migrated operators.