Commit Graph
25 Commits
Author SHA1 Message Date
McZyWu 6bb2918938 Bugfix qwen prefix cache circumstances (#30047) 2026-07-06 14:39:30 +08:00
McZyWu 8416544ab0 [NPU] bugfix for Base class add mamba_track_indices parameter (#29999) 2026-07-03 15:10:41 +08:00
McZyWu c05c48b35e bugfix for npu Grok2 model --detokenizer without all special ids (#29853) 2026-07-02 22:12:48 +08:00
McZyWu bdd3515389 NPU case rl update weights for tensor load_format == None and flatten bucket (#29503) 2026-07-02 22:12:05 +08:00
McZyWu bb7d3440b5 bugfix revise interface get cpu copy for npu mem pool to align with gpu (#29146) 2026-06-29 19:30:14 +08:00
McZyWu 8bdb007e58 [NPU] Docs op performance optimize (#28277) 2026-06-15 20:28:18 +08:00
McZyWu bf186cf8fc bugfix revise interface get cpu copy for npu mem pool to align with gpu (#27802) 2026-06-15 17:24:10 +08:00
McZyWu f7041c9dee step3.5 flash revise for graph mode and use triton activation (#27739) 2026-06-13 16:34:20 +08:00
McZyWu c6be251c5b [NPU] RL update_weights_from_disk/ tensor /distributed (#26717) 2026-06-09 16:52:36 +08:00
McZyWu d8487bad06 Update best practice for qwen3-next-80b-a3b-instruct (#27353) 2026-06-05 17:02:38 +08:00
McZyWu b1173c8c14 [NPU] Enhance accuracy for model Step3_5 from 0 to 88% (#24582) 2026-05-29 11:29:30 +08:00
McZyWu 1c2857b064 bugfix: --decrypted-draft-config-file not applied (#25960) 2026-05-29 09:15:05 +08:00
McZyWu b2631a9a4d [NPU] Docs op performance optimize (#25830) 2026-05-22 09:20:13 +08:00
McZyWu 9d0be860a4 [NPU] recover accuracy for gemma3-4b-it from 54% to 72% (reduced by transformer5.3) (#21537) 2026-05-13 16:46:04 +08:00
McZyWu 4435a23a51 [NPU]adapt multibatch fia ops (#20177) 2026-05-11 09:44:14 +08:00
McZyWuandsglang-npu-bot 7d397ad23d [NPU]Support model Trinity-mini for Npu, accuracy 90% (#18172)
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-05-07 20:58:18 +03:00
McZyWuandroot 1588856e9b [NPU] qwen3next low latency best practice docs. (#22808)
Co-authored-by: root <root@localhost.localdomain>
2026-04-14 21:21:37 +08:00
McZyWu 0906e45cec bugfix for weight loading for qwen3-next (#21313) 2026-03-26 21:21:00 +08:00
McZyWu 8662ba7db4 [NPU] bugfix for import sgl-kernel error (#21200) 2026-03-23 19:52:36 +08:00
McZyWu 4641e5a3d2 [NPU] enhance accuracy for model minimaxm2 from 16.5% to 95.5% (#17695) 2026-03-23 19:06:38 +08:00
McZyWuandcy 4f7422f7ba [NPU] support model skywork-reward-gemma2-2-27B-v0.2 (#16947)
Co-authored-by: cy <chenyang08056032@163.com>
2026-02-11 15:34:53 +08:00
McZyWuandcy 70db3398d1 [NPU] enhance accuracy for model kimi-vl-a3b-instruct (#17480)
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-30 15:19:42 +08:00
McZyWuandcy 2734b23481 accuracy enhancement for baichuan2-13B for npu (#16868)
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-26 16:14:35 +08:00
McZyWu b4a611fb33 [NPU] solve accuracy problem for stablelm-2-1-6b for npu (#17470) 2026-01-24 08:27:38 +08:00
McZyWu 8a5ed2434f [NPU]support model MiniCPM3-4B for npu (#16866) 2026-01-24 08:25:12 +08:00