49 Commits
Author SHA1 Message Date
Jimmy ShongandYangmin Li e97614d10c [Qwen4-Exp] Build the offloaded PLE table on the meta device so --ple-offload-embedding never materialises it on the accelerator (#39928)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-09-20 08:28:25 -07:00
cebca698e2 [Qwen3.8] Enable NVIDIA NVFP4 on DGX Spark with file-backed PLE and PDL router fix (#39126)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: rdxa <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Manrique <nanomlm@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
2026-09-13 16:23:41 +08:00
f4b75b5c36 docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-07 15:33:03 -07:00
Jimmy ShongandClaude Fable 5.1 2da5802bfa [Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-03 11:43:01 -07:00
Jimmy ShongandClaude Fable 5.1 ed82bea146 [Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-01 15:42:39 -07:00
Jimmy Shong 5030637c65 [docs] Split the Qwen3.8-27B NVFP4 cells by lm_head precision (#36020) 2026-08-25 01:50:33 +08:00
Jimmy ShongandClaude Fable 5 4cb5aebfe0 [docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spark grids on 1cf2b8c (#35825)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 13:03:18 +08:00
Jimmy Shong 3efa057449 [docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c (#35786) 2026-08-20 21:54:51 -07:00
Jimmy Shong 14795dcb1a [docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev image tag (#35767) 2026-08-20 23:43:55 +00:00
Jimmy Shong 1a138e13b9 [docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753) 2026-08-20 23:34:44 +00:00
Jimmy Shong d9f6861359 [docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663) 2026-08-20 13:26:55 -07:00
Jimmy Shong 710267dc4c [Quant] Load compressed-tensors kv_cache_scheme scales (#35455) 2026-08-20 19:17:59 +08:00
Jimmy ShongandLING ZHI 1cf2b8c54d [Spec] Support quantized target lm_head in the DFlash2 selector (#35496)
Co-authored-by: LING ZHI <1747985437lz@gmail.com>
2026-08-19 18:06:41 -07:00
Jimmy Shong 5375babbac [Quant] Load compressed-tensors quantized lm_head instead of value-casting it (#35228) 2026-08-19 15:37:45 -07:00
Jimmy ShongandClaude Fable 5 c863760ae1 [Fix] DCP: advertise the logical KV-event block size (#35298)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 20:55:11 -07:00
Jimmy ShongandClaude Opus 5 b956e916ae docs(cookbook): add Qwen3.8-27B DGX Spark configs (#35121)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:04:17 -07:00
Jimmy ShongandClaude Fable 5 e03c53fc13 docs(cookbook): Qwen3.8-27B deployment grid rework (#35065)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 00:56:08 -07:00
Jimmy ShongandClaude Fable 5 07a28ec5cf docs: fix Qwen3.8-27B mamba ratio calculator for speculative decoding (#35064)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 22:54:57 -07:00
Jimmy Shong 198b7e9240 [Docs] Use Meta's canonical Muse Glimmer GGUF filename (#34626) 2026-08-12 16:07:55 -07:00
Jimmy ShongandClaude Opus 5 e510dc58ba Add @Jiminator as codeowner for Laguna model and config (#33472)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 14:10:19 -07:00
Jimmy Shong ed361ae7f0 Fix attention backends for models with per-layer head counts (num_attention_heads_per_layer) (#32625) 2026-07-29 20:03:00 -07:00
Jimmy Shong 410ab4fde5 Add return_token_ids support to completions and chat completions APIs (#30917) 2026-07-23 14:41:52 -07:00
9a6c96083f [Cookbook] Add Laguna-S-2.1 (#31918)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-21 17:01:50 +00:00
Jimmy ShongandClaude Opus 4.8 4c5fe42be4 [DSA] Fix IMA in fused top-k v2: write all output slots on tie overflow (#30512)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 19:49:26 +08:00
Jimmy Shong 85e71b7e13 [Doc] Cookbook Laguna-XS-2.1: add AIME25 accuracy (B300 + GB300) (#29974) 2026-07-02 13:15:01 -07:00
Jimmy Shong 476c946543 [Doc] Cookbook: Laguna-XS-2.1 (DFlash low-latency + high-throughput) (#29884) 2026-07-02 20:05:33 +08:00
Jimmy Shong e745b3af22 [Fix] compressed-tensors block FP8: requantize weight scales to UE8M0 for DeepGEMM on Blackwell (#28662) 2026-06-26 21:41:18 +00:00
be774d0acd [docs][cookbook] Laguna-M.1 playground: add HiCache; refresh EP / DP-Attention notes (#28774)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-22 03:58:42 +00:00
Jimmy Shong 7516f0db9f [cookbook] Laguna-M.1: add PD disaggregation section (#28737) 2026-06-19 19:49:49 -07:00
Jimmy Shong d962d18f15 docs: add --trust-remote-code to Laguna-M.1 / XS.2 cookbook configs (#28693) 2026-06-19 10:21:39 +08:00
Jimmy Shong f7632ef860 [Cookbook] Laguna-M.1: enable FP8 on Blackwell + drop provisional AIME numbers (#28664) 2026-06-18 09:23:53 -07:00
Jimmy Shong 0eded9e208 Add Laguna-M.1 cookbook (#28661) 2026-06-18 23:23:53 +08:00
Jimmy Shong 97e3b8998d Pass quant_config to attention gate projection (#28649) 2026-06-18 20:04:25 +08:00
Jimmy Shong d2539980b6 [Fix] don't force hybrid-SWA when sliding_window is disabled (#28604) 2026-06-17 22:11:27 -07:00
Jimmy Shongandgithub-actions[bot] 54acffc864 Eval accuracy gpqa aime25 mixins (#27102)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-14 00:49:12 -07:00
Jimmy Shong 0ef39784ef [Bugfix] Gate DP-attention even-token padding to CP-enabled configs (#26911) 2026-06-03 02:06:52 -04:00
Jimmy Shong 716e670d3d [bugfix]: size CuteDSL MoE allgather buffers for the worst-case forward (#26696) 2026-05-30 00:27:20 -07:00
Jimmy Shong f838adb7d4 bench_serving: add Zipfian shared-prefix sampling to generated-shared-prefix (#26378) 2026-05-28 14:39:46 -07:00
Jimmy Shong 1a85586738 [Fix]: Restrict Kimi-K2.5 shared-experts fusion to Quark MXFP4 checkpoints (#25974) 2026-05-21 13:07:45 -07:00
Jimmy Shong daade9cc00 [Fix] Probe speculative draft config via sglang get_config (#25428) 2026-05-15 22:01:44 -07:00
Jimmy Shong a741d0cc56 [CI] Lower mem-fraction-static for GLM-5.1 FP8 8-GPU test to 0.85 (#25453) 2026-05-15 20:14:47 -07:00
Jimmy Shongandgithub-actions[bot] fd3eb77d45 [Cookbook]: add Laguna-XS.2 (Poolside) (#24730)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-12 16:06:26 +01:00
Jimmy Shong e9a15b95da [Fix] Disable FlashInfer allreduce fusion under deterministic inference (#24629) 2026-05-10 20:04:52 -05:00
Jimmy Shongandgemini-code-assist[bot] fa8985486e [test/fix]: isolate VLM MMMU eval output dirs to fix nightly-4-gpu cross-test pollution (#24623)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-08 15:01:53 -07:00
Jimmy Shong 096ad02b06 [Model] Laguna-XS.2 Model Support (#24204) 2026-05-09 05:43:13 +08:00
Jimmy Shong 3d31ac2672 [Fix] FP8 Qwen3-Next quant error by removing fallback fused shards (#23973) 2026-04-29 17:33:47 -04:00
Jimmy ShongandSGLang CI 68a8ed9b11 [Fix/Kernel] Add JIT rmsnorm_hf kernel to fix transformers backend MMLU accuracy regression (#22931)
Co-authored-by: SGLang CI <ci@sglang.ai>
2026-04-23 12:00:31 +08:00
Jimmy Shong 28e915b474 [Bugfix] Preserve auto-detected quant_config for GLM NextN draft model (#22823) 2026-04-15 13:25:36 -07:00
Jimmy Shong e83560562b Update CI Permissions (#22826) 2026-04-14 15:13:31 -07:00