Yuhao Yang
|
1240ac13b8
|
vlm: fix tiny multimodal cache bug (#12984)
|
2025-11-10 21:30:19 +08:00 |
|
 Yuhao YangandMick
|
d8736c756a
|
fix multimodal gen issues (#12765)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-07 13:19:49 +08:00 |
|
yhyang201
|
48e9e71930
|
Add --max-new-tokens CLI flag for MMMU evaluation (#11217)
|
2025-10-04 17:35:53 -07:00 |
|
yhyang201
|
6130529143
|
Quick Fix: fix Qwen3-VL launch failure caused by MRotaryEmbedding arg (#10985)
|
2025-09-30 22:17:05 -07:00 |
|
yhyang201
|
388c05d544
|
Fix bias handling in TritonMoeQuantInfo within quantization/mxfp4.py (#10579)
|
2025-09-18 11:44:43 -07:00 |
|
yhyang201
|
c377923304
|
[feat] Reduce GPU memory overhead by using weakref (#9673)
|
2025-08-28 01:09:06 -07:00 |
|
    
|
a85363c199
|
[docs] Instructions for bench_serving.py (#9071)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-08-26 18:30:57 -07:00 |
|
yhyang201
|
00da906584
|
feat: Support DP Attention for step3_vl (#8699)
|
2025-08-03 19:35:26 +08:00 |
|
 yhyang201andChayenne
|
0dfe2491ac
|
Preliminary Support for Qwen3XMLDetector (#8260)
Co-authored-by: Chayenne <zhaochen20@outlook.com>
|
2025-07-23 06:49:38 +08:00 |
|
 yhyang201andChang Su
|
dea2b84bc3
|
[OAI Server Refactor] [ChatCompletions & Completions] Implement UsageInfo Processor (#7360)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-06-20 14:51:21 -07:00 |
|
yhyang201
|
1dffee31ac
|
OAI Server Skeleton & Core Utility Endpoints (#7179)
|
2025-06-16 20:45:55 -07:00 |
|
yhyang201
|
cec98f1034
|
[Fix] Incorrect Memory Allocation on CUDA:0 by Non-Zero CUDA Processes in TP/DP (#5745)
|
2025-05-08 17:52:26 -07:00 |
|
yhyang201
|
92ab0a2055
|
feat: Add fused moe triton config for qwen3bf16 moe on h20 (#5839)
|
2025-04-28 09:30:59 -07:00 |
|
yhyang201
|
4db463b1ad
|
[Model] Adding Qwen3 and Qwen3MoE (#4693)
|
2025-04-18 09:51:29 -07:00 |
|
yhyang201
|
072df75354
|
Support for Qwen2.5-VL Model in bitsandbytes Format (#5003)
|
2025-04-14 02:03:40 -07:00 |
|