vllm/csrc/quantization/cutlass_w8a8 at 5eda2ea02a01b2457f4d6ac2a217f2fa8a2e5d5f - vllm

mirror of https://github.com/wassname/vllm.git synced 2026-07-05 10:35:32 +08:00

Files

T

Tyler Michael Smith 8674f9880e [Kernel] Fixup for CUTLASS kernels in CUDA graphs (#4954 )

Pass the CUDA stream into the CUTLASS GEMMs, to avoid future issues with CUDA graphs

2024-05-22 14:10:43 +00:00

common.hpp

2024-05-16 18:32:50 -04:00

cutlass_visitor_2x_broadcast_epilogue.hpp

2024-05-16 18:32:50 -04:00

scaled_mm_dq_c2x.cu

2024-05-22 14:10:43 +00:00

scaled_mm_dq_c3x.cu

2024-05-22 14:10:43 +00:00

scaled_mm_dq_entry.cu

2024-05-22 07:18:41 +00:00