Together avatar

Together

Official

@togethercomputer

0Followers
|
109Public Repos
|
6Published Skills

Offers high-performance kernel optimization and model integration services for SGLang-based diffusion and generative architectures.

Skills Distribution
DomainAI Models & ...Kernel Engineering (40%)Model Integration (30%)Performance Optimi.. (30%)

Agent Skills by Together

Showing 6 vetted skills indexed across 1 GitHub repositories.

Frequently Asked Questions About Together

FAQPage Schema
What specific tasks can be performed using Together's kernel expertise?β–Ό

Together enables the integration of new diffusion models into SGLang, the development of JIT CUDA and Triton kernels, and the implementation of AOT C++ kernels with Torch bindings. These capabilities focus on reducing latency and optimizing VRAM usage for high-performance generative model serving.

Which technical personas benefit from these SGLang-focused capabilities?β–Ό

These skills are designed for machine learning engineers, kernel developers, and infrastructure architects working on high-throughput generative model deployment. Professionals focused on low-level performance tuning and custom model integration within the SGLang ecosystem will find these technical resources essential for production-grade inference.

What are the prerequisites for implementing these kernel and model integrations?β–Ό

Implementation requires an existing SGLang environment, proficiency in CUDA and Triton programming, and familiarity with CMake and Torch bindings. Users must have a baseline understanding of diffusion model architectures and the specific requirements for JIT and AOT kernel compilation within the SGLang framework.