Allen Wang
|
c6cf9295e1
|
[Bugfix] Sets is_first_step_output for TPUModelRunner (#9202)
|
2024-10-11 13:28:10 -07:00 |
|
Wallas Henrique
|
8baf85e4e9
|
[Doc] Compatibility matrix for mutual exclusive features (#8512)
Signed-off-by: Wallas Santos <wallashss@ibm.com>
|
2024-10-11 11:18:50 -07:00 |
|
Tyler Michael Smith
|
7342a7d7f8
|
[Model] Support Mamba (#6484)
|
2024-10-11 15:40:06 +00:00 |
|
 youkaichaoandBrendan Wong
|
cbc2ef5529
|
[misc] hide best_of from engine (#9261)
Co-authored-by: Brendan Wong <bjwpokemon@gmail.com>
|
2024-10-10 21:30:44 -07:00 |
|
youkaichao
|
e4d652ea3e
|
[torch.compile] integration with compilation control (#9058)
|
2024-10-10 12:39:36 -07:00 |
|
Li, Jiang
|
ca77dd7a44
|
[Hardware][CPU] Support AWQ for CPU backend (#7515)
|
2024-10-09 10:28:08 -06:00 |
|
Alex Brooks
|
a3691b6b5e
|
[Core][Frontend] Add Support for Inference Time mm_processor_kwargs (#9131)
Signed-off-by: Alex-Brooks <Alex.Brooks@ibm.com>
|
2024-10-08 14:12:56 +00:00 |
|
Kunshang Ji
|
80b57f00d5
|
[Intel GPU] Fix xpu decode input (#9145)
|
2024-10-08 03:51:14 +00:00 |
|
Isotr0py
|
4f95ffee6f
|
[Hardware][CPU] Cross-attention and Encoder-Decoder models support on CPU backend (#9089)
|
2024-10-07 06:50:35 +00:00 |
|
youkaichao
|
18b296fdb2
|
[core] remove beam search from the core (#9105)
|
2024-10-07 05:47:04 +00:00 |
|
Isotr0py
|
487678d046
|
[Bugfix][Hardware][CPU] Fix CPU model input for decode (#9044)
|
2024-10-06 19:14:27 -07:00 |
|
 Cyrus LeungandRoger Wang
|
b22b798471
|
[Model] PP support for embedding models and update docs (#9090)
Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>
|
2024-10-06 16:35:27 +08:00 |
|
 Chongming NiandAshraf Mahgoub
|
cc90419e89
|
[Hardware][Neuron] Add on-device sampling support for Neuron (#8746)
Co-authored-by: Ashraf Mahgoub <ashymahg@amazon.com>
|
2024-10-04 16:42:20 -07:00 |
|
Cyrus Leung
|
0e36fd4909
|
[Misc] Move registry to its own file (#9064)
|
2024-10-04 10:01:37 +00:00 |
|
youkaichao
|
9aaf14c62e
|
[misc] add forward context for attention (#9029)
|
2024-10-03 12:09:42 -07:00 |
|
Sergey Shlyapnikov
|
f58d4fccc9
|
[OpenVINO] Enable GPU support for OpenVINO vLLM backend (#8192)
|
2024-10-02 17:50:01 -04:00 |
|
 Varun Sundar RabindranathandVarun Sundar Rabindranath
|
afb050b29d
|
[Core] CUDA Graphs for Multi-Step + Chunked-Prefill (#8645)
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
|
2024-10-02 19:44:39 +00:00 |
|
Lily Liu
|
1570203864
|
[Spec Decode] (1/2) Remove batch expansion (#8839)
|
2024-10-01 16:04:42 -07:00 |
|
youkaichao
|
7da2487591
|
[torch.compile] fix tensor alias (#8982)
|
2024-10-01 03:40:48 +00:00 |
|
Jee Jee Li
|
1cabfcefb6
|
[Misc] Adjust max_position_embeddings for LoRA compatibility (#8957)
|
2024-09-30 12:57:39 +00:00 |
|
 Nick HillandRoger Wang
|
31f46a0d35
|
[BugFix] Fix seeded random sampling with encoder-decoder models (#8870)
Co-authored-by: Roger Wang <ywang@roblox.com>
|
2024-09-29 09:43:14 +00:00 |
|
Jee Jee Li
|
3d49776bbb
|
[Model][LoRA]LoRA support added for MiniCPMV2.5 (#7199)
|
2024-09-29 06:59:45 +00:00 |
|
Varun Sundar Rabindranath
|
19d02ff938
|
[Bugfix] Fix PP for Multi-Step (#8887)
|
2024-09-28 08:52:46 -07:00 |
|
 Varun Sundar RabindranathandVarun Sundar Rabindranath
|
c2ec430ab5
|
[Core] Multi-Step + Single Step Prefills via Chunked Prefill code path (#8378)
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
|
2024-09-27 13:32:07 -07:00 |
|
youkaichao
|
a9b15c606f
|
[torch.compile] use empty tensor instead of None for profiling (#8875)
|
2024-09-27 08:11:32 -07:00 |
|
    
|
770ec6024f
|
[Model] Add support for the multi-modal Llama 3.2 model (#8811)
Co-authored-by: simon-mo <xmo@berkeley.edu>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simon Mo <simon.mo@hey.com>
Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>
Co-authored-by: Roger Wang <ywang@roblox.com>
|
2024-09-25 13:29:32 -07:00 |
|
Isotr0py
|
c23953675f
|
[Hardware][CPU] Enable mrope and support Qwen2-VL on CPU backend (#8770)
|
2024-09-24 23:16:11 -07:00 |
|
Cody Yu
|
b8747e8a7c
|
[MISC] Skip dumping inputs when unpicklable (#8744)
|
2024-09-24 06:10:03 +00:00 |
|
Li, Jiang
|
3e83c12b5c
|
[Bugfix][CPU] fix missing input intermediate_tensors in the cpu_model_runner (#8733)
|
2024-09-23 13:15:16 +00:00 |
|
Isotr0py
|
e551ca1555
|
[Hardware][CPU] Refactor CPU model runner (#8729)
|
2024-09-23 20:12:20 +08:00 |
|
Huazhong Ji
|
ca2b628b3c
|
[MISC] rename CudaMemoryProfiler to DeviceMemoryProfiler (#8703)
|
2024-09-22 10:44:09 -07:00 |
|
William Lin
|
9e5ec35b1f
|
[bugfix] [AMD] add multi-step advance_step to ROCmFlashAttentionMetadata (#8474)
|
2024-09-19 20:49:54 -07:00 |
|
sroy745
|
3118f63385
|
[Bugfix] [Encoder-Decoder] Bugfix for encoder specific metadata construction during decode of encoder-decoder models. (#8545)
|
2024-09-19 02:24:15 +00:00 |
|
afeldman-nm
|
a8c1d161a7
|
[Core] *Prompt* logprobs support in Multi-step (#8199)
|
2024-09-18 08:38:43 -07:00 |
|
Cyrus Leung
|
6ffa3f314c
|
[CI/Build] Avoid CUDA initialization (#8534)
|
2024-09-18 10:38:11 +00:00 |
|
Nick Hill
|
56c3de018c
|
[Misc] Don't dump contents of kvcache tensors on errors (#8527)
|
2024-09-17 12:24:29 -07:00 |
|
sroy745
|
1009e93c5d
|
[Encoder decoder] Add cuda graph support during decoding for encoder-decoder models (#7631)
|
2024-09-17 07:35:01 -07:00 |
|
Woosuk Kwon
|
50e9ec41fc
|
[TPU] Implement multi-step scheduling (#8489)
|
2024-09-14 16:58:31 -07:00 |
|
youkaichao
|
a36e070dad
|
[torch.compile] fix functionalization (#8480)
|
2024-09-14 09:46:04 -07:00 |
|
youkaichao
|
0a4806f0a9
|
[plugin][torch.compile] allow to add custom compile backend (#8445)
|
2024-09-13 09:32:42 -07:00 |
|
William Lin
|
a6c0f3658d
|
[multi-step] add flashinfer backend (#7928)
|
2024-09-12 11:16:22 -07:00 |
|
WANGWEI
|
8a23e93302
|
[BugFix] lazy init _copy_stream to avoid torch init wrong gpu instance (#8403)
|
2024-09-12 10:47:42 -07:00 |
|
Kevin Lin
|
295c4730a8
|
[Misc] Raise error when using encoder/decoder model with cpu backend (#8355)
|
2024-09-12 05:45:24 +00:00 |
|
Cody Yu
|
a65cb16067
|
[MISC] Dump model runner inputs when crashing (#8305)
|
2024-09-12 01:12:25 +00:00 |
|
 bnellnmandSage Moore
|
73202dbe77
|
[Kernel][Misc] register ops to prevent graph breaks (#6917)
Co-authored-by: Sage Moore <sage@neuralmagic.com>
|
2024-09-11 12:52:19 -07:00 |
|
Li, Jiang
|
0b952af458
|
[Hardware][Intel] Support compressed-tensor W8A8 for CPU backend (#7257)
|
2024-09-11 09:46:46 -07:00 |
|
 
|
3b7fea770f
|
[Model][VLM] Add Qwen2-VL model support (#7905)
Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>
Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2024-09-11 09:31:19 -07:00 |
|
Kevin Lin
|
5faedf1b62
|
[Spec Decode] Move ops.advance_step to flash attn advance_step (#8224)
|
2024-09-10 13:18:14 -07:00 |
|
Alexander Matveev
|
4ef41b8476
|
[Bugfix] Fix async postprocessor in case of preemption (#8267)
|
2024-09-07 21:01:51 -07:00 |
|
youkaichao
|
ce2702a923
|
[tpu][misc] fix typo (#8260)
|
2024-09-06 22:40:46 -07:00 |
|