vllm/csrc/quantization at f081c3ce4b020fb094e33575d178345c477ab0c6 - vllm

mirror of https://github.com/wassname/vllm.git synced 2026-07-03 21:32:23 +08:00

Files

T

Varun Sundar Rabindranath f081c3ce4b [Kernel] Update Cutlass fp8 configs (#5144 )

Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-neuralmagic@users.noreply.github.com>

2024-06-01 08:46:07 +00:00

aqlm

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

awq

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

compressed_tensors

[Kernel] Initial Activation Quantization Support (#4525 )

2024-05-23 21:29:18 +00:00

cutlass_w8a8

[Kernel] Update Cutlass fp8 configs (#5144 )

2024-06-01 08:46:07 +00:00

fp8

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

gptq

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

gptq_marlin

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

marlin

Revert "[Kernel] Marlin_24: Ensure the mma.sp instruction is using the ::ordered_metadata modifier (introduced with PTX 8.5)" (#5149 )

2024-05-30 22:00:26 -07:00

squeezellm

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00