Xiaoyu Zhang
|
695ab705cb
|
[diffusion] quant: update modelopt quantization docs and CI coverage (#22772)
|
2026-04-15 21:30:28 +08:00 |
|
 jianzhao-xuandJianzhao Xu
|
45a83ffbe3
|
[NPU] Offloading docs update (#22860)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-15 15:04:41 +08:00 |
|
 Po-Han Huang (NVIDIA)andClaude Opus 4.6
|
ada52e5972
|
[Docs] Move ptxas sm_103a workaround into For CUDA 13 section (#22852)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-14 22:30:21 -07:00 |
|
chx96642264
|
680bd4b429
|
[NPU] Modify the parameter name and optional values, and add the parameter restrictions. Modify some parameters supported type. (#22804)
|
2026-04-14 21:34:07 +08:00 |
|
 McZyWuandroot
|
1588856e9b
|
[NPU] qwen3next low latency best practice docs. (#22808)
Co-authored-by: root <root@localhost.localdomain>
|
2026-04-14 21:21:37 +08:00 |
|
amote-i
|
ddc7daaf89
|
[NPU] [DOC] Update NPU docs to match latest code (#22796)
|
2026-04-14 21:10:28 +08:00 |
|
loading66
|
074c2a476d
|
fix:[NPU]correct the full name of then Kimi model (#22799)
|
2026-04-14 20:15:22 +08:00 |
|
 jianzhao-xuandJianzhao Xu
|
68dfffaaa3
|
Offloading docs update (#22795)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-14 20:03:29 +08:00 |
|
xdtbynd
|
88253c39b0
|
[Docs] Fix formatting of tool-call-parser options (#22793)
|
2026-04-14 19:21:31 +08:00 |
|
amote-i
|
368cdfbe2f
|
[NPU] [DOC] Fix outdated descriptions in the NPU documentation (#22707)
|
2026-04-14 19:21:15 +08:00 |
|
Xiaoyu Zhang
|
f97c608caa
|
[diffusion] quant: add FLUX.1-dev modelopt nvfp4 support (#22672)
|
2026-04-14 15:00:59 +08:00 |
|
看海的人
|
13a4aafdbe
|
[NPU] update glm5 running guide (#22712)
|
2026-04-13 22:53:24 +08:00 |
|
chx96642264
|
c6403a11cb
|
Modify the optional values and constraints of parameter. (#22705)
|
2026-04-13 22:50:48 +08:00 |
|
 jianzhao-xuandJianzhao Xu
|
b6a91b1afe
|
[NPU] --attn-cp-size --init-expert-location --eplb-algorithm parameter docs update (#22704)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-13 22:42:34 +08:00 |
|
Liwansi
|
8d904e50f2
|
[NPU]qwen3-8b and 32b md bugfix (#22687)
|
2026-04-13 22:20:17 +08:00 |
|
 loading66andh30064329
|
2089ac86a7
|
Improve parameters usage constraints for npu deployment (#22700)
Co-authored-by: h30064329 <hanbing45@h-partners.com>
|
2026-04-13 22:02:56 +08:00 |
|
 看海的人andzhsurpass
|
56c97c7738
|
[NPU] update npu doc (#22697)
Co-authored-by: zhsurpass <zhsurpass@users.noreply.github.com>
|
2026-04-13 21:55:38 +08:00 |
|
 xdtbyndandxdtbynd
|
d01b2bf257
|
[Docs] Fix default values and options in Ascend server arguments documentation (#22698)
Co-authored-by: xdtbynd <supercluster@vip.qq.com>
|
2026-04-13 21:22:37 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
7d2c11970c
|
[Intel GPU] Upgrade pytorch xpu version to 2.11 (#21908)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-13 13:16:24 +08:00 |
|
Mick
|
bf022e177c
|
Revert "[Diffusion] Add FLUX.1-dev ModelOpt NVFP4 support (#22574)" (#22649)
|
2026-04-13 11:17:32 +08:00 |
|
Xiaoyu Zhang
|
37fc47c645
|
diffusion: fix layerwise offload for ModelOpt quantized DiTs (#22594)
|
2026-04-13 08:01:54 +08:00 |
|
Xiaoyu Zhang
|
03a1a7b81c
|
[Diffusion] Add FLUX.1-dev ModelOpt NVFP4 support (#22574)
|
2026-04-13 07:57:41 +08:00 |
|
Mick
|
495ef8ec64
|
[diffusion] model: support LTX2.3 two stage (#22182)
|
2026-04-12 22:15:57 +08:00 |
|
Mohammad Miadh Angkad
|
bcc0c65aa8
|
[DSA] Hopper FP8 FlashMLA KV padding (#22372)
|
2026-04-12 02:19:17 -07:00 |
|
Wenyao Gao
|
4dfc8e1c3f
|
VLM: support passing --mm-process-config for all models (#18467)
|
2026-04-12 17:08:05 +08:00 |
|
Baizhou Zhang
|
d14d368191
|
[Kernel] Set sgl_per_token_group_quant_8bit_v2 as default choice (#22467)
|
2026-04-11 01:59:57 -07:00 |
|
heziiop
|
4f45472f34
|
[NPU][Doc] add qwen3-30b-a3b low latency example (#22446)
|
2026-04-11 15:52:47 +08:00 |
|
  
|
f855a0bde6
|
Introduce CUDA graph debug mode with breakable CUDA graph (#19102)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-11 00:36:56 -07:00 |
|
Xiaoyu Zhang
|
1ff51555f2
|
[Diffusion] modelopt diffusion fp8 support for flux1/flux2 and wan2.2 (#22365)
|
2026-04-10 20:56:57 +08:00 |
|
 
|
5ba7d4e523
|
[HiSparse]: Update HiSparse's user-guide (#22499)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-04-10 15:06:43 +08:00 |
|
billishyahao
|
1df9f4e2f6
|
[AMD] Add prealloc token env for mori-ep (#22329)
|
2026-04-09 09:34:35 -07:00 |
|
amote-i
|
7965573eb4
|
fix issues for npu docs (#22307)
|
2026-04-09 16:27:34 +08:00 |
|
Liwansi
|
8ec0934f8f
|
[NPU]add Qwen3-32b and Qwen3-8b low latency md (#22429)
|
2026-04-09 16:18:34 +08:00 |
|
Mick
|
9709192ce9
|
[diffusion] feat: support FLUX.2-small-decoder (#22414)
|
2026-04-09 15:53:14 +08:00 |
|
Nicolas Castet
|
e379befbac
|
Add symmetric debug mode to print stack trace of comm ops with unregistered tensors (#18569)
|
2026-04-08 22:34:58 -07:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
Zhangheng
|
3d3a32c0b9
|
[HiSparse]: Add readme docs for HiSparse Feature (#22238)
|
2026-04-07 00:39:24 -07:00 |
|
 Aditya SharmaandXinyuan Tong
|
f6e85676b5
|
model: support qwen3-asr (#22073)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-07 13:27:05 +08:00 |
|
Khoa Pham
|
12272b6791
|
[Spec][Ngram] 6/N: Load an external corpus and construct a Suffix Automaton (#21425)
|
2026-04-06 00:11:14 -07:00 |
|
Mohammad Miadh Angkad
|
b311db2e49
|
[Doc] Fix and improve DeepSeek V3.2/GLM-5 documentation (#22179)
|
2026-04-05 23:26:42 -07:00 |
|
 YAMYandShangming Cai
|
dc125afffb
|
Add staging buffer CI test and documentation for heterogeneous TP (#21921)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-06 14:00:20 +08:00 |
|
Mick
|
82c41a2d9e
|
[diffusion] model: support LTX2.3 (#22111)
|
2026-04-06 12:26:30 +08:00 |
|
Baizhou Zhang
|
106baedbfb
|
[Doc] Update GLM-5 instructions in sglang documentation (#21716)
|
2026-04-05 03:13:07 -07:00 |
|
 narutolhyandluhongyu.4869
|
24763256b9
|
[Speculative Decoding] Add FA4-based Spec Support (#21080)
Co-authored-by: luhongyu.4869 <luhongyu.4869@bytedance.com>
|
2026-04-04 02:09:45 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
b5e8c4b9e3
|
model: support LFM2-VL (Liquid Foundation Model 2 Vision-Language) (#21230)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-04 16:36:04 +08:00 |
|
  
|
db3d4f4b76
|
[diffusion] model: support two stage pipeline of LTX-2 (#20707)
Co-authored-by: daiweitao <dwti614707404@163.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: GMI Xiao Jin <xiao.j@gmicloud.ai>
|
2026-04-04 09:37:28 +08:00 |
|
Brayden Zhong
|
6aafe756b9
|
Revert "[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+… (#22047)
|
2026-04-03 13:12:30 -07:00 |
|
amote-i
|
81efcc353a
|
[NPU] Optimized the wording in the npu docs (#21998)
|
2026-04-03 11:51:40 +08:00 |
|
Mook
|
991f3aa5b3
|
[Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) (#19652)
|
2026-04-03 10:48:15 +08:00 |
|
Liangsheng Yin
|
f25bf86065
|
Fix ngram doc for speculative_num_draft_tokens default (#21910)
|
2026-04-01 22:18:24 -07:00 |
|