  
|
2f4e108f75
|
[Bugfix] Clean up MiniCPM-V (#6939)
Co-authored-by: hezhihui <hzh7269@modelbest.cn>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
|
2024-07-31 14:39:19 +00:00 |
|
 HandH1998andGitHub
|
6512937de1
|
Support W4A8 quantization for vllm (#5218)
|
2024-07-31 07:55:21 -06:00 |
|
 FeiandGitHub
|
c0644cf9ce
|
[Bugfix] fix logit processor excceed vocab size issue (#6927)
|
2024-07-31 16:16:01 +08:00 |
|
 Woosuk KwonandGitHub
|
533d1932d2
|
[Bugfix][TPU] Set readonly=True for non-root devices (#6980)
|
2024-07-31 00:19:28 -07:00 |
|
 Cyrus LeungandGitHub
|
9f0e69b653
|
[CI/Build] Fix mypy errors (#6968)
|
2024-07-30 19:49:48 -07:00 |
|
 Cyrus LeungandGitHub
|
f230cc2ca6
|
[Bugfix] Fix broadcasting logic for multi_modal_kwargs (#6836)
|
2024-07-31 10:38:45 +08:00 |
|
 Cyrus LeungandGitHub
|
da1f7cc12a
|
[mypy] Enable following imports for some directories (#6681)
|
2024-07-31 10:38:03 +08:00 |
|
 
|
6ca8031e71
|
[core][misc] improve free_finished_seq_groups (#6865)
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>
|
2024-07-30 14:32:12 -07:00 |
|
 Tyler Michael SmithandGitHub
|
d7a299edaa
|
[Kernel] Remove scaled_fp8_quant kernel padding footgun (#6842)
|
2024-07-30 16:37:01 -04:00 |
|
 Nick HillandGitHub
|
5cf9254a9c
|
[BugFix] Fix use of per-request seed with pipeline parallel (#6698)
|
2024-07-30 10:40:08 -07:00 |
|
 fzyzcjyandGitHub
|
f058403683
|
[Doc] Super tiny fix doc typo (#6949)
|
2024-07-30 09:14:03 -07:00 |
|
 Roger WangandGitHub
|
c66c7f86ac
|
[Bugfix] Fix PaliGemma MMP (#6930)
|
2024-07-30 02:20:57 -07:00 |
|
 Woosuk KwonandGitHub
|
6e063ea35b
|
[TPU] Fix greedy decoding (#6933)
|
2024-07-30 02:06:29 -07:00 |
|
 Nick HillandGitHub
|
9f69d8245a
|
[Frontend] New allowed_token_ids decoding request parameter (#6753)
|
2024-07-29 23:37:27 +00:00 |
|
 Thomas ParnellandGitHub
|
9a7e2d0534
|
[Bugfix] Allow vllm to still work if triton is not installed. (#6786)
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com>
|
2024-07-29 14:51:27 -07:00 |
|
 EarthwalkerandGitHub
|
7f8d612d24
|
[TPU] Support tensor parallelism in async llm engine (#6891)
|
2024-07-29 12:42:21 -07:00 |
|
 Peng GuanwenandGitHub
|
db9e5708a9
|
[Core] Reduce unnecessary compute when logprobs=None (#6532)
|
2024-07-29 16:47:31 +00:00 |
|
 
|
7cbd9ec7a9
|
[Model] Initialize support for InternVL2 series models (#6514)
Co-authored-by: Roger Wang <ywang@roblox.com>
|
2024-07-29 10:16:30 +00:00 |
|
 Elsa GrangerandGitHub
|
3eeb148f46
|
[Misc] Pass cutlass_fp8_supported correctly in fbgemm_fp8 (#6871)
|
2024-07-28 11:13:49 -04:00 |
|
 Michael GoinandGitHub
|
b1366a9534
|
Add Nemotron to PP_SUPPORTED_MODELS (#6863)
|
2024-07-27 15:05:17 -07:00 |
|
 Alexander MatveevandGitHub
|
75acdaa4b6
|
[Kernel] Increase precision of GPTQ/AWQ Marlin kernel (#6795)
|
2024-07-27 17:52:33 -04:00 |
|
 Woosuk KwonandGitHub
|
fad5576c58
|
[TPU] Reduce compilation time & Upgrade PyTorch XLA version (#6856)
|
2024-07-27 10:28:33 -07:00 |
|
 
|
1ad86acf17
|
[Model] Initial support for BLIP-2 (#5920)
Co-authored-by: ywang96 <ywang@roblox.com>
|
2024-07-27 11:53:07 +00:00 |
|
 Travis JohnsonandGitHub
|
593e79e733
|
[Bugfix] torch.set_num_threads() in multiproc_gpu_executor (#6802)
[Bugfix] Use torch.set_num_threads() to configure parallelism in multiproc_gpu_executor (#6802)
Signed-off-by: Travis Johnson <tsjohnso@us.ibm.com>
|
2024-07-26 22:15:20 -07:00 |
|
 Woosuk KwonandGitHub
|
52f07e3dec
|
[Hardware][TPU] Implement tensor parallelism with Ray (#5871)
|
2024-07-26 20:54:27 -07:00 |
|
 JoeandGitHub
|
14dbd5a767
|
[Model] H2O Danube3-4b (#6451)
|
2024-07-26 20:47:50 -07:00 |
|
 tomeras91andGitHub
|
ed94e4f427
|
[Bugfix][Model] Jamba assertions and no chunked prefill by default for Jamba (#6784)
|
2024-07-26 20:45:31 -07:00 |
|
 Cyrus LeungandGitHub
|
981b0d5673
|
[Frontend] Factor out code for running uvicorn (#6828)
|
2024-07-27 09:58:25 +08:00 |
|
 Woosuk KwonandGitHub
|
d09b94ca58
|
[TPU] Support collective communications in XLA devices (#6813)
|
2024-07-27 01:45:57 +00:00 |
|
 chenqianfzhandGitHub
|
bb5494676f
|
enforce eager mode with bnb quantization temporarily (#6846)
|
2024-07-27 01:32:20 +00:00 |
|
 Michael GoinandGitHub
|
281977bd6e
|
[Doc] Add Nemotron to supported model docs (#6843)
|
2024-07-26 17:32:44 -04:00 |
|
 Li, JiangandGitHub
|
3bbb4936dc
|
[Hardware] [Intel] Enable Multiprocessing and tensor parallel in CPU backend and update documentation (#6125)
|
2024-07-26 13:50:10 -07:00 |
|
 Woosuk KwonandGitHub
|
aa4867791e
|
[Misc][TPU] Support TPU in initialize_ray_cluster (#6812)
|
2024-07-26 19:39:49 +00:00 |
|
 Michael GoinandGitHub
|
07278c37dd
|
[Model] Support Nemotron models (Nemotron-3, Nemotron-4, Minitron) (#6611)
|
2024-07-26 14:33:42 -04:00 |
|
 Peng GuanwenandGitHub
|
89a84b0bb7
|
[Core] Use array to speedup padding (#6779)
|
2024-07-25 21:31:31 -07:00 |
|
 Anthony PlataniosandGitHub
|
084a01fd35
|
[Bugfix] [Easy] Fixed a bug in the multiprocessing GPU executor. (#6770)
|
2024-07-25 21:25:35 -07:00 |
|
 QQSongandGitHub
|
062a1d0fab
|
Fix ReplicatedLinear weight loading (#6793)
|
2024-07-25 19:24:58 -07:00 |
|
 SangBin ChoandGitHub
|
1adddb14bf
|
[Core] Fix ray forward_dag error mssg (#6792)
|
2024-07-25 16:53:25 -07:00 |
|
 Lucas WilkinsonandGitHub
|
cd7edc4e87
|
[Bugfix] Fix empty (nullptr) channelwise scales when loading wNa16 using compressed tensors (#6798)
|
2024-07-25 15:05:09 -07:00 |
|
 
|
95db75de64
|
[Bugfix] Add synchronize to prevent possible data race (#6788)
Co-authored-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
|
2024-07-25 10:40:01 -07:00 |
|
 Michael GoinandGitHub
|
65b1f121c8
|
[Bugfix] Fix kv_cache_dtype=fp8 without scales for FP8 checkpoints (#6761)
|
2024-07-25 09:46:15 -07:00 |
|
 
|
889da130e7
|
[ Misc ] fp8-marlin channelwise via compressed-tensors (#6524)
Co-authored-by: mgoin <michael@neuralmagic.com>
|
2024-07-25 09:46:04 -07:00 |
|
  
|
b75e314fff
|
[Bugfix] Add image placeholder for OpenAI Compatible Server of MiniCPM-V (#6787)
Co-authored-by: hezhihui <hzh7269@modelbest.cn>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
|
2024-07-25 09:42:49 -07:00 |
|
 Alexander MatveevandGitHub
|
0310029a2f
|
[Bugfix] Fix awq_marlin and gptq_marlin flags (#6745)
|
2024-07-24 22:34:11 -07:00 |
|
 Cody YuandGitHub
|
309aaef825
|
[Bugfix] Fix decode tokens w. CUDA graph (#6757)
|
2024-07-24 22:33:56 -07:00 |
|
 AlphiandGitHub
|
9e169a4c61
|
[Model] Adding support for MiniCPM-V (#4087)
|
2024-07-24 20:59:30 -07:00 |
|
 Evan Z. LiuandGitHub
|
5689e256ba
|
[Frontend] Represent tokens with identifiable strings (#6626)
|
2024-07-25 09:51:00 +08:00 |
|
 youkaichaoandGitHub
|
740374d456
|
[core][distributed] fix zmq hang (#6759)
|
2024-07-24 17:37:12 -07:00 |
|
 Antoni BaumandGitHub
|
5448f67635
|
[Core] Tweaks to model runner/input builder developer APIs (#6712)
|
2024-07-24 12:17:12 -07:00 |
|
 Antoni BaumandGitHub
|
0e63494cf3
|
Add fp8 support to reshape_and_cache_flash (#6667)
|
2024-07-24 18:36:52 +00:00 |
|