Yuhong Guo
|
4d84f886e7
|
Refactor --debug-tensor-dump-layers to list (#12691)
|
2025-11-05 03:30:01 -08:00 |
|
Yuhong Guo
|
2f6af1a3de
|
Enable bailing_moe to support TP=16 (#12369)
|
2025-10-31 19:32:49 +08:00 |
|
Yuhong Guo
|
ab95d35fcb
|
feat: Add Non-intrusive Tensor Dumping for Model Inference (#10566)
|
2025-10-31 12:04:48 +08:00 |
|
Yuhong Guo
|
caa5d2967c
|
feat: return partial generation results when aborting requests in waiting queue (#11673)
|
2025-10-29 22:03:00 +08:00 |
|
Yuhong Guo
|
d279d4990c
|
Fix aiohttp 'Chunk too big' in bench_serving (#6737)
|
2025-05-30 00:50:36 -07:00 |
|
Yuhong Guo
|
f87a6ab359
|
Resolves the 404 Not Found error when running compile_deep_gemm.py in multi-node setups (#5720)
|
2025-04-26 18:13:13 -07:00 |
|
Yuhong Guo
|
5d93a950ee
|
[BugFix] Fix combination of MTP and --n-share-experts-fusionwith R1 (#5707)
|
2025-04-24 21:13:51 +08:00 |
|
Yuhong Guo
|
3dfc6023ce
|
Fix bench_serving with random-ids (#5214)
|
2025-04-15 01:34:35 -07:00 |
|
Yuhong Guo
|
7d8c0ce7ce
|
[Build] Support build sgl-kernel with ccache (#5020)
|
2025-04-03 00:22:37 -07:00 |
|
Yuhong Guo
|
87fafa0105
|
Revert PR 4764 & 4813 related to R1 RoPE (#4959)
|
2025-03-31 20:56:58 -07:00 |
|
Yuhong Guo
|
ee47a6c1c3
|
[Build] Fix cuda12.8 build error in nvfp4_scaled_mm_kernels.cu (#4953)
|
2025-03-31 12:00:34 -07:00 |
|
Yuhong Guo
|
64edeb798f
|
Support dynamic version name in sglang's pyproject.toml (#4720)
|
2025-03-24 08:56:31 -07:00 |
|
Yuhong Guo
|
417fc72f6f
|
Align completion and chat_completion response to OpenAI API (#4637)
|
2025-03-20 22:59:04 -07:00 |
|