mirror of
https://github.com/wassname/vllm.git
synced 2026-10-02 12:50:41 +08:00
custom allreduce + torch.compile (#10121)
Signed-off-by: youkaichao <youkaichao@gmail.com> Co-authored-by: youkaichao <youkaichao@gmail.com>
This commit is contained in:
1 parent
519e8e4182
commit
9a88f89799
6 files changed
+62
-104
No files matched your search
@@ -86,7 +86,6 @@ If GPU/CPU communication cannot be established, you can use the following Python
|
||||
from vllm.distributed.device_communicators.pynccl import PyNcclCommunicator
|
||||
|
||||
pynccl = PyNcclCommunicator(group=gloo_group, device=local_rank)
|
||||
pynccl.disabled = False
|
||||
|
||||
s = torch.cuda.Stream()
|
||||
with torch.cuda.stream(s):
|
||||
|
||||
Reference in new issue
Block a user