Yingchun Lai
|
875a25ddef
|
refactor: remove duplicate function _get_bootstrap_info_from_server (#13277)
|
2025-11-15 02:48:17 +08:00 |
|
Yingchun Lai
|
b419e20c5b
|
[Dockerfile] Speed up docker image building (#8784)
|
2025-11-04 20:26:18 -08:00 |
|
Yingchun Lai
|
ec92b0cefe
|
EPLB: prefer to use physical experts in the same gpu or node (#10874)
|
2025-10-28 21:01:11 -07:00 |
|
Yingchun Lai
|
5e36a0b455
|
[metrics][EPLB]: Support selected count of physical experts on each GPU (#9825)
|
2025-10-28 20:56:19 -07:00 |
|
Yingchun Lai
|
0fe87213bb
|
fix: fix gpu-proc affinity set incorrectly when pp_size > 1 (#11389)
|
2025-10-09 18:40:05 -07:00 |
|
Yingchun Lai
|
9d7e82a0ab
|
EPLB: prefer to use physical experts in the same node (#9849)
|
2025-09-22 00:34:30 -07:00 |
|
Yingchun Lai
|
b1721edbac
|
[PD metrics] Add latency Histogram metrics of each stage for generate requests (#8710)
|
2025-09-16 01:52:49 +08:00 |
|
Yingchun Lai
|
fc2c3a3d8e
|
metrics: support customer labels specified in request header (#10143)
|
2025-09-14 20:00:08 -07:00 |
|
Yingchun Lai
|
21ca4c3afa
|
[PD metrics] Fix some uncompleted PD related metrics (#8627)
|
2025-09-14 02:26:58 -07:00 |
|
Yingchun Lai
|
b32ab0705e
|
metrics: support customer buckets for prompt/generation_tokens_histogram (#9634)
|
2025-09-04 22:22:08 +08:00 |
|
Yingchun Lai
|
ed6f7597b3
|
Fix the missing 'lof' choice of --schedule-policy server args (#7114)
|
2025-08-03 12:29:42 -07:00 |
|
Yingchun Lai
|
36d6f0ba5b
|
fix: fix the missing metrics on non-rank0 nodes (#7720)
|
2025-07-27 00:55:25 -07:00 |
|
Yingchun Lai
|
610381b75e
|
[health_generate] fix: fix the /health_generate always success bug (#8028)
|
2025-07-18 22:08:46 -07:00 |
|
 Yingchun LaiandStefan He
|
795668dc73
|
feat: add tp_rank, pp_rank and dp_rank labels for scheduler metrics (#7597)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-07-16 17:55:59 -07:00 |
|