 ChaunceyandGitHub
|
ac6b8f19b9
|
[Frontend] Multi-Modality Support for Loading Local Image Files (#9915)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2024-11-04 15:34:57 +00:00 |
|
 
|
54597724f4
|
[Model] Add support for H2OVL-Mississippi models (#9747)
Signed-off-by: Shanshan Wang <shanshan.wang@h2o.ai>
Signed-off-by: Roger Wang <ywang@roblox.com>
Co-authored-by: Roger Wang <ywang@roblox.com>
|
2024-11-04 00:15:36 +00:00 |
|
 Cyrus LeungandGitHub
|
ba0d892074
|
[Frontend] Use a proper chat template for VLM2Vec (#9912)
|
2024-11-01 14:09:07 +00:00 |
|
 Cyrus LeungandGitHub
|
06386a64dd
|
[Frontend] Chat-based Embeddings API (#9759)
|
2024-11-01 08:13:35 +00:00 |
|
 Roger WangandGitHub
|
3ea2dc2ec4
|
[Misc] Remove deprecated arg for cuda graph capture (#9864)
Signed-off-by: Roger Wang <ywang@roblox.com>
|
2024-10-31 07:22:07 +00:00 |
|
 Kevin H. LuuandGitHub
|
890ca36072
|
Revert "[Bugfix] Use host argument to bind to interface (#9798)" (#9852)
|
2024-10-31 01:44:51 +00:00 |
|
 Guillaume CalmettesandGitHub
|
abbfb6134d
|
[Misc][OpenAI] deprecate max_tokens in favor of new max_completion_tokens field for chat completion endpoint (#9837)
|
2024-10-30 18:15:56 -07:00 |
|
 Joe RundeandGitHub
|
3b3f1e7436
|
[Bugfix][core] replace heartbeat with pid check (#9818)
Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>
|
2024-10-30 09:34:07 -07:00 |
|
 Went-LiangandGitHub
|
81f09cfd80
|
[Model] Support math-shepherd-mistral-7b-prm model (#9697)
Signed-off-by: Went-Liang <wenteng_liang@163.com>
|
2024-10-30 09:33:42 -07:00 |
|
  
|
882a1ad0de
|
[Model] tool calling support for ibm-granite/granite-20b-functioncalling (#8339)
Signed-off-by: Max de Bayser <mbayser@br.ibm.com>
Co-authored-by: Max de Bayser <mbayser@br.ibm.com>
Co-authored-by: Maximilien de Bayser <maxdebayser@gmail.com>
|
2024-10-29 15:07:37 -07:00 |
|
 Sven SeebergandGitHub
|
0f43387157
|
[Bugfix] Use host argument to bind to interface (#9798)
|
2024-10-29 10:37:59 -07:00 |
|
 Zhong QishuaiandGitHub
|
ef7865b4f9
|
[Frontend] re-enable multi-modality input in the new beam search implementation (#9427)
Signed-off-by: Qishuai Ferdinandzhong@gmail.com
|
2024-10-29 11:49:47 +00:00 |
|
 Sam StoelingaandGitHub
|
067e77f9a8
|
[Bugfix] Steaming continuous_usage_stats default to False (#9709)
Signed-off-by: Sam Stoelinga <sammiestoel@gmail.com>
|
2024-10-26 05:05:47 +00:00 |
|
 Michael GoinandGitHub
|
e26d37a185
|
[Log][Bugfix] Fix default value check for image_url.detail (#9663)
|
2024-10-24 10:44:38 -07:00 |
|
 Vinay R DamodaranandGitHub
|
33bab41060
|
[Bugfix]: Make chat content text allow type content (#9358)
Signed-off-by: Vinay Damodaran <vrdn@hey.com>
|
2024-10-24 05:05:49 +00:00 |
|
 
|
fc6c274626
|
[Model] Add Qwen2-Audio model support (#9248)
Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2024-10-23 17:54:22 +00:00 |
|
 
|
150b779081
|
[Frontend] Enable Online Multi-image Support for MLlama (#9393)
Signed-off-by: Alex-Brooks <Alex.Brooks@ibm.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
|
2024-10-23 17:28:57 +00:00 |
|
 
|
434984e665
|
[Frontend] Support custom request_id from request (#9550)
Co-authored-by: Yuhong Guo <yuhong.gyh@antgroup.com>
|
2024-10-22 18:07:30 +00:00 |
|
 Woosuk KwonandGitHub
|
6c5af09b39
|
[V1] Implement vLLM V1 [1/N] (#9289)
|
2024-10-22 01:24:07 -07:00 |
|
 Chen ZhangandGitHub
|
5b59fe0f08
|
[Bugfix] Pass json-schema to GuidedDecodingParams and make test stronger (#9530)
|
2024-10-20 00:05:02 +00:00 |
|
 Yue ZhangandGitHub
|
c5eea3c8ba
|
[Frontend] Support simpler image input format (#9478)
|
2024-10-18 23:17:07 -07:00 |
|
 Cyrus LeungandGitHub
|
051eaf6db3
|
[Model] Add user-configurable task for models that support both generation and embedding (#9424)
|
2024-10-18 11:31:58 -07:00 |
|
 Nick HillandGitHub
|
25aeb7d4c9
|
[BugFix] Fix and simplify completion API usage streaming (#9475)
|
2024-10-18 14:10:26 +00:00 |
|
 tomeras91andGitHub
|
d2b1bf55ec
|
[Frontend][Feature] Add jamba tool parser (#9154)
|
2024-10-18 10:27:48 +00:00 |
|
 Nick HillandGitHub
|
1ffc8a7362
|
[BugFix] Typing fixes to RequestOutput.prompt and beam search (#9473)
|
2024-10-18 07:19:53 +00:00 |
|
 sasha0552andGitHub
|
d615b5c9f8
|
[Bugfix] Print warnings related to mistral_common tokenizer only once (#9468)
|
2024-10-17 21:44:20 +00:00 |
|
 Cyrus LeungandGitHub
|
390be74649
|
[Misc] Print stack trace using logger.exception (#9461)
|
2024-10-17 13:55:48 +00:00 |
|
 
|
ba30942240
|
[Bugfix] Fix vLLM UsageInfo and logprobs None AssertionError with empty token_ids (#9034)
Co-authored-by: Nick Hill <nickhill@us.ibm.com>
|
2024-10-15 15:40:43 -07:00 |
|
 Nick HillandGitHub
|
e9d517f276
|
[BugFix] Fix chat API continuous usage stats (#9357)
|
2024-10-14 23:19:48 -07:00 |
|
 Steve GrubbandGitHub
|
44eaa5a5d9
|
[Frontend] Clarify model_type error messages (#9345)
|
2024-10-14 21:29:01 -07:00 |
|
 Brendan WongandGitHub
|
4d31cd424b
|
[Frontend] merge beam search implementations (#9296)
|
2024-10-14 15:05:52 -07:00 |
|
   
|
dfe43a2071
|
[Model] Molmo vLLM Integration (#9016)
Co-authored-by: sanghol <sanghol@allenai.org>
Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>
Co-authored-by: Roger Wang <ywang@roblox.com>
|
2024-10-14 07:56:24 -07:00 |
|
 Maximilien de BayserandGitHub
|
ec10cb8511
|
[BugFix] Fix tool call finish reason in streaming case (#9209)
Signed-off-by: Max de Bayser <mbayser@br.ibm.com>
|
2024-10-11 18:24:26 -07:00 |
|
 Russell BryantandGitHub
|
cdca8994bd
|
[CI/Build] mypy: check vllm/entrypoints (#9194)
Signed-off-by: Russell Bryant <rbryant@redhat.com>
|
2024-10-09 17:15:28 +00:00 |
|
 Cyrus LeungandGitHub
|
cfaa6008e6
|
[Bugfix] Access get_vocab instead of vocab in tool parsers (#9188)
|
2024-10-09 08:59:57 -06:00 |
|
 DanieleandGitHub
|
9a94ca4a5d
|
[Bugfix] fix OpenAI API server startup with --disable-frontend-multiprocessing (#8537)
|
2024-10-08 09:38:40 -07:00 |
|
 Alex BrooksandGitHub
|
069d3bd8d0
|
[Frontend] Add Early Validation For Chat Template / Tool Call Parser (#9151)
Signed-off-by: Alex-Brooks <Alex.Brooks@ibm.com>
|
2024-10-08 14:31:26 +00:00 |
|
 Alex BrooksandGitHub
|
a3691b6b5e
|
[Core][Frontend] Add Support for Inference Time mm_processor_kwargs (#9131)
Signed-off-by: Alex-Brooks <Alex.Brooks@ibm.com>
|
2024-10-08 14:12:56 +00:00 |
|
 Brendan WongandGitHub
|
8c746226c9
|
[Frontend] API support for beam search for MQLLMEngine (#9117)
|
2024-10-08 05:51:43 +00:00 |
|
  
|
151ef4efd2
|
[Model] Support NVLM-D and fix QK Norm in InternViT (#9045)
Co-authored-by: Roger Wang <ywang@roblox.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2024-10-07 11:55:12 +00:00 |
|
 youkaichaoandGitHub
|
18b296fdb2
|
[core] remove beam search from the core (#9105)
|
2024-10-07 05:47:04 +00:00 |
|
 Yanyi LiuandGitHub
|
fdf59d30ea
|
[Bugfix] fix tool_parser error handling when serve a model not support it (#8709)
|
2024-10-06 12:51:08 +00:00 |
|
 Cyrus LeungandGitHub
|
f22619fe96
|
[Misc] Remove user-facing error for removed VLM args (#9104)
|
2024-10-06 01:33:52 -07:00 |
|
 
|
168cab6bbf
|
[Frontend] API support for beam search (#9087)
Co-authored-by: youkaichao <youkaichao@126.com>
|
2024-10-05 23:39:03 -07:00 |
|
 Flávia BéoandGitHub
|
0dcc8cbe5a
|
Adds truncate_prompt_tokens param for embeddings creation (#8999)
Signed-off-by: Flavia Beo <flavia.beo@ibm.com>
|
2024-10-04 18:31:40 +00:00 |
|
 代君andGitHub
|
3dbb215b38
|
[Frontend][Feature] support tool calling for internlm/internlm2_5-7b-chat model (#8405)
|
2024-10-04 10:36:39 +08:00 |
|
 Guillaume CalmettesandGitHub
|
83caf35e08
|
[BugFix] Enforce Mistral ToolCall id constraint when using the Mistral tool call parser (#9020)
|
2024-10-03 16:44:52 +08:00 |
|
 Sebastian SchoennenbeckandGitHub
|
35bd215168
|
[Core] [Frontend] Priority scheduling for embeddings and in the OpenAI-API (#8965)
|
2024-10-01 09:58:06 +00:00 |
|
 
|
062c89e7c9
|
[Frontend][Core] Move guided decoding params into sampling params (#8252)
Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>
Co-authored-by: Nick Hill <nickhill@us.ibm.com>
|
2024-10-01 09:34:25 +08:00 |
|
 danieljannai21andGitHub
|
6c9ba48fde
|
[Frontend] Added support for HF's new continue_final_message parameter (#8942)
|
2024-09-29 17:59:47 +00:00 |
|