This website requires JavaScript.
Explore
Help
Sign In
wassname
/
vllm
Watch
1
Star
0
Fork
0
mirror of
https://github.com/wassname/vllm.git
synced
2026-10-04 13:10:10 +08:00
Code
Issues
Packages
Projects
Releases
Wiki
Activity
Files
e2fbaee7258810bcd43725e9ca7f1444a88f91f3
vllm
/
csrc
T
History
Varun Sundar Rabindranath
and
Varun Sundar Rabindranath
b5241e41d9
[ Kernel ] FP8 Dynamic-Per-Token Quant Kernel (
#6511
)
...
Co-authored-by: Varun Sundar Rabindranath <
varun@neuralmagic.com
>
2024-07-18 01:38:35 +00:00
..
attention
[Kernel][Attention] Separate
Attention.kv_scale
into
k_scale
and
v_scale
(
#6081
)
2024-07-16 15:31:32 -07:00
cpu
[Kernel][Attention] Separate
Attention.kv_scale
into
k_scale
and
v_scale
(
#6081
)
2024-07-16 15:31:32 -07:00
moe
…
prepare_inputs
[Core] draft_model_runner: Implement prepare_inputs on GPU for advance_step (
#6338
)
2024-07-17 14:30:28 -07:00
punica
[Kernel] Add punica dimensions for Granite 3b and 8b (
#5930
)
2024-06-29 10:48:25 +08:00
quantization
[ Kernel ] FP8 Dynamic-Per-Token Quant Kernel (
#6511
)
2024-07-18 01:38:35 +00:00
activation_kernels.cu
[Model] Port over CLIPVisionModel for VLMs (
#5591
)
2024-06-20 11:52:09 +00:00
cache_kernels.cu
[Kernel][Attention] Separate
Attention.kv_scale
into
k_scale
and
v_scale
(
#6081
)
2024-07-16 15:31:32 -07:00
cache.h
[Kernel][Attention] Separate
Attention.kv_scale
into
k_scale
and
v_scale
(
#6081
)
2024-07-16 15:31:32 -07:00
cuda_compat.h
…
cuda_utils_kernels.cu
…
cuda_utils.h
…
custom_all_reduce_test.cu
…
custom_all_reduce.cu
…
custom_all_reduce.cuh
…
dispatch_utils.h
…
layernorm_kernels.cu
…
moe_align_block_size_kernels.cu
…
ops.h
[ Kernel ] FP8 Dynamic-Per-Token Quant Kernel (
#6511
)
2024-07-18 01:38:35 +00:00
pos_encoding_kernels.cu
…
reduction_utils.cuh
…
registration.h
…
torch_bindings.cpp
[ Kernel ] FP8 Dynamic-Per-Token Quant Kernel (
#6511
)
2024-07-18 01:38:35 +00:00