21 Commits
Author SHA1 Message Date
zijiecandZijie Chen e634ba78a4 [AMD] gfx950 assembly attention for EAGLE verify, draft extend and decode (#37465)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
2026-09-08 03:10:09 -07:00
02236fa38c Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
2026-07-19 22:57:37 -07:00
Liangsheng Yin c53559ba10 [misc] Remove unit test cases that fail the admission criteria (#30690) 2026-07-09 15:31:28 -07:00
9bd02dc5b9 feat(sgl-kernel): add InfLLM v2 attention kernels (#29383)
Co-authored-by: Size Wang <paulgeorge13hhhhh@gmail.com>
Co-authored-by: lijiayi <lijiayi@modelbest.cn>
Co-authored-by: suhmily10 <suhmily@gmail.com>
Co-authored-by: Xiaoyue Xu <xiaoyue.xu.me@gmail.com>
Co-authored-by: hansjohn <74091612+hansjohn@users.noreply.github.com>
Co-authored-by: zhangyan <1762895426@qq.com>
2026-07-06 22:46:54 -07:00
ashwini rathiandSinghal, Shubham 5134dcdcab [Intel XPU] Initially add nightly GSM8K accuracy tests for Llama-3.1-8B (TP=2) and Qwen3-32B (TP=4) (#28908)
Co-authored-by: Singhal, Shubham <shubham.singhal@intel.com>
2026-07-01 16:27:24 +08:00
Qiaolin Yu 3cecc77ccb [perf] Fuse NVFP4 gate_up_gemm + swiglu + output FP4 quant (#26626) 2026-05-29 13:16:24 -07:00
Kangyan-ZhouandClaude Opus 4.7 6e8fe176be sgl-router: experimental Rust HTTP router for SGLang worker pools (#25851)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 15:34:05 +08:00
jianzhao-xu f66881f03c [NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#25384) 2026-05-21 17:32:25 +08:00
jianzhao-xu edb1b3f8f5 [NPU] add Ascend NPU Accuracy Evaluation and Faq docs (#24777) 2026-05-14 11:28:16 +08:00
Lianmin Zheng 1ae3218d03 Add jasonjk-park and charlotte12l to CI_PERMISSIONS.json (#25104) 2026-05-13 01:51:05 -07:00
YC Yen-Ching Tseng 6b6963f426 [AMD] Update scripts/ci/amd/ensure_vram_clear.sh (#24586) 2026-05-11 18:47:09 +08:00
Xinyuan Tong d8f9d32a05 feat(reasoning): auto-detect reasoning/tool-call parser from chat template (#23952) 2026-05-07 14:19:16 -07:00
Shenxiu LiuandXinyuan Tong a3fc982ba7 [Whisper] Automatic language detection via structured generation (#22997)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-04-27 15:54:41 +08:00
Lianmin Zheng 6a3c070ee3 Add 'allready' to ignore words list in .codespellrc (#23465) 2026-04-22 02:39:04 -07:00
Mick 29f56cb230 CI: fix lint (#22991) 2026-04-17 02:09:04 +08:00
Lianmin Zheng ccff59254c Update .codespellrc (#22912) 2026-04-15 16:25:55 -07:00
Yuhao Yang 16f306fd85 [VLM] GPU Image Preprocessing for Kimi-K2.5 (#22368) 2026-04-11 11:13:30 +08:00
Sundara Raman Ramachandran 712c8c5051 [Score API] Add SequenceClassification Model support (#22118) 2026-04-08 01:30:58 -07:00
2813cb6d9a [New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
2026-04-06 20:24:44 -07:00
Jacob Gordon 858f317f13 ci(codespell): centralizes list of ignorable words (#17524) 2026-01-21 12:29:14 -08:00
Jacob Gordon cda43ffa4d ci: avoids duplication of codespell config (#17519) 2026-01-21 12:02:37 -08:00