Commit Graph
388 Commits
Author SHA1 Message Date
amote-i 69ddbf9ef6 [NPU] [DOC] fix model name error on Ascend NPU (#30577) 2026-07-09 11:28:29 +08:00
HuangJi 3d96bb9721 [diffusion] chore: rename lingbot world v2 (#30518) 2026-07-08 19:06:44 +08:00
amote-i 042228a195 [NPU] [DOC] Remove unsupported options of features on Ascend NPU (#30504) 2026-07-08 17:35:06 +08:00
HuangJi db40fd83d2 [diffusion] model: support LingBot-World 2.0 (#30361) 2026-07-08 10:42:50 +08:00
Yihao WangandClaude Opus 4.8 68901ba387 [diffusion] Support SP for Krea-2 (#29777)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 19:39:00 -07:00
631213c3bf Add DeepReinforce Ornith-1.0 to cookbook (#29404)
Co-authored-by: averyjones4 <averyjones4@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 15:14:37 -07:00
sglang-botandsglang-bot d88644b430 docs: sync LMSYS SGLang blog cards (#30395)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-07 12:30:47 -07:00
zijiexia 0bf7ddb481 docs(install): add nightly install + docker tag guidance, and auto-bump version on release tag (#30308) 2026-07-07 12:10:05 -07:00
Xiaoyu ZhangandZijie Xia ead1e490b5 [Doc] Add LongCat 2.0 FP8 cookbook (#30320)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-07-07 11:48:13 -07:00
amote-i cfd3fdc54f [NPU] [DOC] Update features and mainstream models on ascend npu (#30370) 2026-07-07 19:03:43 +08:00
Mohammad Miadh Angkad f32b4ecd26 [Docs] Use trtllm_mha for Qwen3.6 B300 (#29964) 2026-07-07 01:44:01 -07:00
qinsir5522 5e9032c527 [NPU]Modify LoRA heading in ascend_npu_support_features.mdx to specify Qwen model limitations. (#30358) 2026-07-07 16:01:51 +08:00
ZeyuanChen2000 2d9f0b3317 [NPU] [DOC] Update arguments detail to NPU support features page (#30328) 2026-07-07 14:09:46 +08:00
loading66 998acf7df8 [DOCS][NPU]update npu support features (#30324) 2026-07-07 11:37:19 +08:00
3a679459e5 [bench] Add agentic-trace multi-turn dataset to bench_serving (#29215)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:45:44 -07:00
sglang-botandsglang-bot cf4edda956 docs: sync LMSYS SGLang blog cards (#30311)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-07 01:04:21 +00:00
Junlin Wuandronnie_zheng 3abdbab9bb [llm][npu][quant] Add W4A8 MXFP quantization support for Qwen3 Dense on Ascend NPU (#23650)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-06 19:23:26 +03:00
zijiexiaandClaude Opus 4.8 7c9bb316cf docs(cookbook): total (input+output) throughput per GPU + percentile latency labels (#30214)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 00:39:48 -07:00
Baizhou Zhang 5eb1b6a7ba Remove retired DSA env paths (#29912) 2026-07-05 22:58:02 -07:00
Xinyuan Tong 6f22790943 cookbook: add Hunyuan 3 (Hy3) Day-0 page (#30201) 2026-07-06 13:30:47 +08:00
addffd7489 [Diffusion] Diffusion model support log-requests (#23049)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-05 12:01:26 +03:00
Cao E 9df16b5ba9 [XPU] Remove redundant xpu graph backend and make xpu graph opt-in by default (#29911) 2026-07-03 15:58:56 +08:00
amote-i 2fe7182e75 [DOC] [NPU] update supported features on ascend npu (#30011) 2026-07-03 15:36:41 +08:00
ming_wang d8a4f7a7aa add mimo-v2-flash model tutorial (#29932) 2026-07-03 11:02:24 +08:00
Peng Xingchen a6bc7fef90 glm5.2 on ascend doc (new version) (#29828) 2026-07-03 10:19:44 +08:00
amote-i 70a813493f [NPU] [DOC] add missing DEEP_NORMAL_MODE_USE_INT8_QUANT for w8a8+deepep scenarios (#29937) 2026-07-03 10:18:22 +08:00
sglang-bot 91d7645aca docs: sync LMSYS SGLang blog cards (#29990) 2026-07-03 01:01:32 +00:00
Jimmy Shong 85e71b7e13 [Doc] Cookbook Laguna-XS-2.1: add AIME25 accuracy (B300 + GB300) (#29974) 2026-07-02 13:15:01 -07:00
zijiexiaandClaude Opus 4.8 cba3801f52 docs: add PD disaggregation to GLM-5.2 cookbook playground (#29544)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 09:57:38 -07:00
Jimmy Shong 476c946543 [Doc] Cookbook: Laguna-XS-2.1 (DFlash low-latency + high-throughput) (#29884) 2026-07-02 20:05:33 +08:00
zijiexiaandClaude Opus 4.8 1c75243f5e docs: add Qwen3.6-27B-NVFP4 variant to cookbook (#29905)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 00:05:29 -07:00
Cao E 926140d789 [XPU] Enable XPU graph support (decode full-graph + prefill tc_piecewise) (#29053) 2026-07-02 13:24:35 +08:00
Zaili Wang cb06c4e6ce [CPU] Fix model failures on Xeon (#29497) 2026-07-02 13:20:18 +08:00
Cheng Wanandlch1475369 4a8e76805c feat(mem_cache): unified memory pool for hybrid Mamba / SWA models (#29678)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-07-01 13:21:59 -07:00
Baizhou Zhang 677a11bfa9 [Doc] Tiny update dsv4 doc (#29827) 2026-07-01 01:47:35 -07:00
qinsir5522 a7390b17f8 [NPU]Modify --lora-backend & --moe-runner-backend description. (#29793) 2026-07-01 11:15:46 +08:00
bb98629157 docs(cookbook): add AMD MI300X/MI325X/MI355X support for GLM-5.2 (#28471)
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-30 21:08:51 +00:00
sglang-botandsglang-bot 081a01c37d docs: sync LMSYS SGLang blog cards (#29307)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-30 20:48:10 +00:00
97fc4dfd73 [Doc]Checking and modifying Markdown formatting issues and link validity (#28586)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-30 20:27:38 +03:00
Yihao WangandClaude Opus 4.8 a531d81c19 [diffusion][cache-dit] support Krea-2 + run-driven has_separate_cfg (#29688)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 08:22:05 -07:00
Trevor Morris b8c25bfaa7 [NVIDIA] Support flashinfer a2a with flashinfer_trtllm_routed moe (#22394) 2026-06-29 16:23:58 -07:00
Cheng Wanandlch1475369 fc96edd297 feat(mem_cache): page-major (layer-major within a page) KV/state layout (#29533)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-06-29 14:49:54 -07:00
zijiexiaandClaude Opus 4.8 5106b42cbd [cookbook] GLM-5.2 NVFP4 B300: TP8 recipe + 3 strategies (#29557)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:42:19 -07:00
zijiexiaandClaude Opus 4.8 11b7ed7c9e [Docs] Add --prerelease=allow so uv installs the latest sglang (#29676)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:13:34 -07:00
zijiexiaandClaude Opus 4.8 74a197af9d docs: add B200 NVFP4 recipes + benchmarks to GLM-5.2 cookbook (#29674)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:06:30 -07:00
Jyothirmai KottuandXinyuan Tong 473a278dd1 model: support nvidia/LocateAnything-3B (#28958)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-30 00:16:42 +08:00
amote-i 489017b3d6 [NPU] [DOC] Update deterministic inference feature support status to A2, A3 (#29632) 2026-06-29 19:17:06 +08:00
danielafrimiandDaniel Afrimi a2b5ce2ed1 Add stochastic rounding for FP16 Mamba SSM cache (#26929)
Signed-off-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
2026-06-29 01:47:09 -07:00
Xinyuan Tong 38d4ffcd86 [cookbook] drop redundant serve flags (GLM-5.2) + fix M3 page-size note (#28731) 2026-06-29 13:34:13 +08:00
jianzhao-xu 2260e612f6 [NPU] update best practicce docs from testcase (#29492) 2026-06-29 11:30:27 +08:00