Xiaoyu Zhang
|
1ff51555f2
|
[Diffusion] modelopt diffusion fp8 support for flux1/flux2 and wan2.2 (#22365)
|
2026-04-10 20:56:57 +08:00 |
|
 
|
5ba7d4e523
|
[HiSparse]: Update HiSparse's user-guide (#22499)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-04-10 15:06:43 +08:00 |
|
billishyahao
|
1df9f4e2f6
|
[AMD] Add prealloc token env for mori-ep (#22329)
|
2026-04-09 09:34:35 -07:00 |
|
amote-i
|
7965573eb4
|
fix issues for npu docs (#22307)
|
2026-04-09 16:27:34 +08:00 |
|
Liwansi
|
8ec0934f8f
|
[NPU]add Qwen3-32b and Qwen3-8b low latency md (#22429)
|
2026-04-09 16:18:34 +08:00 |
|
Mick
|
9709192ce9
|
[diffusion] feat: support FLUX.2-small-decoder (#22414)
|
2026-04-09 15:53:14 +08:00 |
|
Nicolas Castet
|
e379befbac
|
Add symmetric debug mode to print stack trace of comm ops with unregistered tensors (#18569)
|
2026-04-08 22:34:58 -07:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
Zhangheng
|
3d3a32c0b9
|
[HiSparse]: Add readme docs for HiSparse Feature (#22238)
|
2026-04-07 00:39:24 -07:00 |
|
 Aditya SharmaandXinyuan Tong
|
f6e85676b5
|
model: support qwen3-asr (#22073)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-07 13:27:05 +08:00 |
|
Khoa Pham
|
12272b6791
|
[Spec][Ngram] 6/N: Load an external corpus and construct a Suffix Automaton (#21425)
|
2026-04-06 00:11:14 -07:00 |
|
Mohammad Miadh Angkad
|
b311db2e49
|
[Doc] Fix and improve DeepSeek V3.2/GLM-5 documentation (#22179)
|
2026-04-05 23:26:42 -07:00 |
|
 YAMYandShangming Cai
|
dc125afffb
|
Add staging buffer CI test and documentation for heterogeneous TP (#21921)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-06 14:00:20 +08:00 |
|
Mick
|
82c41a2d9e
|
[diffusion] model: support LTX2.3 (#22111)
|
2026-04-06 12:26:30 +08:00 |
|
Baizhou Zhang
|
106baedbfb
|
[Doc] Update GLM-5 instructions in sglang documentation (#21716)
|
2026-04-05 03:13:07 -07:00 |
|
 narutolhyandluhongyu.4869
|
24763256b9
|
[Speculative Decoding] Add FA4-based Spec Support (#21080)
Co-authored-by: luhongyu.4869 <luhongyu.4869@bytedance.com>
|
2026-04-04 02:09:45 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
b5e8c4b9e3
|
model: support LFM2-VL (Liquid Foundation Model 2 Vision-Language) (#21230)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-04 16:36:04 +08:00 |
|
  
|
db3d4f4b76
|
[diffusion] model: support two stage pipeline of LTX-2 (#20707)
Co-authored-by: daiweitao <dwti614707404@163.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: GMI Xiao Jin <xiao.j@gmicloud.ai>
|
2026-04-04 09:37:28 +08:00 |
|
Brayden Zhong
|
6aafe756b9
|
Revert "[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+… (#22047)
|
2026-04-03 13:12:30 -07:00 |
|
amote-i
|
81efcc353a
|
[NPU] Optimized the wording in the npu docs (#21998)
|
2026-04-03 11:51:40 +08:00 |
|
Mook
|
991f3aa5b3
|
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) (#19652)
|
2026-04-03 10:48:15 +08:00 |
|
Liangsheng Yin
|
f25bf86065
|
Fix ngram doc for speculative_num_draft_tokens default (#21910)
|
2026-04-01 22:18:24 -07:00 |
|
Khoa Pham
|
f836658077
|
[Spec][Ngram] 4/N: Remove max_match_window_size and min_match_window_size, matching all suffixes of the Trie (#21225)
|
2026-04-01 22:09:46 -07:00 |
|
David Cheung
|
ed427e1299
|
Migrate all callers from /get_server_info to /server_info (#21463)
|
2026-04-01 21:17:50 -07:00 |
|
Noa Neria
|
8d9145d97e
|
Direct model loading from object storage with Runai Model Streamer (#17948)
Signed-off-by: Noa Neria <noa@run.ai>
|
2026-04-01 18:41:22 -07:00 |
|
yuefeng Wu
|
c9f5d1d502
|
[Diffusion][NPU] add ring sp performance benchmark page in npu (#21811)
|
2026-04-01 18:53:10 +03:00 |
|
amote-i
|
80b1bc5f56
|
[NPU] update ascend docs (#21807)
|
2026-04-01 17:14:26 +08:00 |
|
Brayden Zhong
|
6a9b09847c
|
CUTLASS NVFP4 GEMM improvement of SM120 (#21314)
|
2026-04-01 09:04:34 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) yuefeng Wuandgemini-code-assist[bot]
|
a20d12ae96
|
[diffusion][doc]: add ring sp performance benchmark page (#20998)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-30 20:26:05 +03:00 |
|
Makcum888e
|
f4b0e9c64a
|
[diffusion] [NPU] support ring attention on NPU with FA (#21383)
|
2026-03-30 20:10:55 +03:00 |
|
Mick
|
b76730701b
|
[diffusion] feat: enhance overlay mechanism (#21648)
|
2026-03-30 19:45:34 +08:00 |
|
 Michelle Wuandwuxue
|
965f03cdc2
|
[NPU] Update DeepSeek-V3.2 model deployment instructions in documentation (#21468)
Co-authored-by: wuxue (C) <w00964934@china.huawei.com>
|
2026-03-30 15:51:42 +08:00 |
|
Aishwarya Ramasethu
|
c32ee48886
|
MFU metrics in Prometheus (#19395)
|
2026-03-29 23:40:06 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Baizhou Zhangandgemini-code-assist[bot]
|
5b19c9a05d
|
[Doc] Update tips for developer new-comers (#21659)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-29 22:40:36 -07:00 |
|
Shu Wang
|
efebcab43e
|
Support skip-softmax attention (#19089)
|
2026-03-28 15:55:48 -07:00 |
|
 Артем СавкинandTamir Baydasov
|
27071e0a43
|
[NPU] Update quantization&CI documentation (#21100)
Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com>
|
2026-03-28 21:42:21 +03:00 |
|
Mick
|
fc9de157f9
|
[diffusion] feat: support overlay model materialization (#21600)
|
2026-03-28 23:02:38 +08:00 |
|
Baizhou Zhang
|
edd4d54023
|
[Clean] Remove deprecated environs (#21536)
|
2026-03-28 00:35:44 -07:00 |
|
Lianmin Zheng
|
83997080a6
|
docs: flesh out MAINTAINER.md oncall lists and link GitHub profiles (#21575)
|
2026-03-27 17:39:16 -07:00 |
|
zwang86
|
5fc5c18bed
|
fix(security): replace unsafe pickle.loads with SafeUnpickler for CVE-2026-3989 (#20904)
|
2026-03-27 00:43:41 -07:00 |
|
SevenJ
|
2e65c27b29
|
Api add flush cache timeout (#21413)
Signed-off-by: root <wenjun7j@gmail.com>
|
2026-03-26 14:44:37 -07:00 |
|
Nave Assaf
|
77872a8d55
|
Update Nemotron Example docs to include Super v3 and Nano 4B (#21416)
Signed-off-by: Nave Assaf <nassaf@nvidia.com>
|
2026-03-25 12:03:19 -04:00 |
|
Mick
|
6425df5c8a
|
[diffusion] doc: consolidate documentation (#21373)
|
2026-03-25 16:01:32 +08:00 |
|
amote-i
|
2d583799eb
|
Update ascend docs (#20846)
|
2026-03-25 09:58:44 +03:00 |
|
Mick
|
6cc5717e8a
|
[diffusion] doc: update quantization.md (#21356)
|
2026-03-25 14:48:38 +08:00 |
|
Duyi-Wang
|
61a902ce88
|
[AMD][MoRI] Auto-select dispatch quantization type from MoE weight dtype. (#21040)
|
2026-03-24 22:53:57 -07:00 |
|
  
|
c4db64c16b
|
Add Lychee Doc Links Check to Local and CI (#19742)
Co-authored-by: Zijie Xia <zijie_xia@icloud.com>
Co-authored-by: Zijie Xia <zijiexia@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-03-24 13:48:26 -07:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
27ac831a84
|
docs: improve CI and testing documentation (#21202)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-23 10:48:50 -07:00 |
|
 kpham-sglandClaude Opus 4.6
|
bc4aaab6a1
|
[Spec][Ngram] 2/N: Rename branch length to max trie depth (#21181)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-22 23:35:25 -07:00 |
|
Xiaoyu Zhang
|
766d225fcc
|
Add SGLang CUDA crash API logging inspired by FlashInfer (#20910)
|
2026-03-22 16:39:40 +08:00 |
|