Logo
Explore Help
Register Sign In
wassname/vllm
Watch 1
Star 0
Fork 0
mirror of https://github.com/wassname/vllm.git synced 2026-08-17 11:28:47 +08:00
Code Issues Packages Projects Releases Wiki Activity
112 Commits 1 Branch 0 Tags
27f1410d065ceca53a07abd2518082eb25228e4f
Commit Graph
14 Commits
Author SHA1 Message Date
Woosuk Kwon a96d63c21d Add support for GPT-NeoX (Pythia) (#50) 2023-04-28 00:32:10 -07:00
Woosuk Kwon 897cb2ae28 Optimize data movement (#20) 2023-04-02 00:30:17 -07:00
Woosuk Kwon 88c0268a18 Implement custom kernel for LLaMA rotary embedding (#14) 2023-03-30 11:04:21 -07:00
Zhuohan Li 2f49f15585 Support tensor parallel (#2) 2023-03-21 13:45:42 -07:00
Woosuk Kwon cfae35b861 Add miscellaneous updates (#8) 2023-03-13 13:48:38 -07:00
Woosuk Kwon 04e5acc08e Fix a bug in 1D input shape (#5) 2023-03-06 10:05:27 -08:00
Woosuk Kwon 3e9f991d6a Use FlashAttention for multi_query_kv_attention (#4) 2023-03-01 21:13:08 -08:00
Woosuk Kwon 0deacbce6e Implement single_query_cached_kv_attention kernel (#3) 2023-03-01 15:02:19 -08:00
Woosuk Kwon 762fd1c3fa Refactor and annotate types for attention 2023-02-24 08:58:46 +00:00
Woosuk Kwon 7f22f90e8c Remove xformers 2023-02-24 08:36:16 +00:00
Woosuk Kwon 932844f1cd Fix attention 2023-02-23 23:02:25 +00:00
Woosuk Kwon ba84b8728a Fix attention 2023-02-23 22:29:46 +00:00
Woosuk Kwon 87e0bcd426 Fix attention 2023-02-23 21:32:02 +00:00
Woosuk Kwon d4bc1a4d24 Add unoptimized OPT Attention 2023-02-23 09:31:55 +00:00
Powered by Gitea Version: 1.27.2 Page: 17ms Template: 3ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API