This website requires JavaScript.
294ff71d18
[Diffusion] Avoid cpu2gpu sync in flashinfer rope and apply flashinfer rope to wanvideo (#16668 )
Xiaoyu Zhang and Mick
2026-01-08 22:44:38 +08:00
f52ae586b6
Remove migrated e2e_grpc/basic tests (#16738 )
Simo Lin
2026-01-08 06:30:45 -08:00
2f8a36347c
[smg][ci] migrate chat completions tests to new infrastructure and build wheel once and share via artifact (#16709 )
Simo Lin
2026-01-08 06:29:23 -08:00
b6e8a0d851
Add file size hints in bitwise model file verifier (#16735 )
fzyzcjy
2026-01-08 22:28:52 +08:00
fb7609f1dd
Fix FP8 MoE NaN with DeepGEMM on Blackwell (#16622 )
luoyuyan
2026-01-08 22:24:12 +08:00
d2ea44f775
VLM: enhance VL embedding model with video input support and revise warm-up strategy (#16635 )
yuhao and Mick
2026-01-08 22:12:01 +08:00
20ca2c6e1e
[NPU] update model and features supported (#16733 )
Hexq0210
2026-01-08 21:50:42 +08:00
83abecd0c2
Support pre-generating and using expected checksums (#16730 )
fzyzcjy
2026-01-08 20:26:13 +08:00
d54f0a10b4
Support bitwise weight checksum verifier (#16729 )
fzyzcjy
2026-01-08 20:19:25 +08:00
fb04e7e3c8
[Auto Sync] Update schedule_batch.py, common.py, eagle_info... (20260105) (#16519 )
2026-01-08 02:35:46 -08:00
1e5de05e35
fix: adding timeout for install dependencies (#16706 )
Douglas Yang
2026-01-08 00:30:41 -08:00
7dd679cbb9
[NPU][Bugfix] Fix qwen3 error when enable-dp-lm-head (#16115 )
chenxu214
2026-01-08 15:15:43 +08:00
3d51ae18a1
[AMD] Turn on AMD CI if rocm.Dockerfile changed (#16634 )
YC Tseng
2026-01-08 14:51:53 +08:00
f9c0426692
[AMD CI] re-enable testcases missed when migrating ci test files (#16535 )
2026-01-08 14:43:48 +08:00
48b8dcd42e
[jit kernel] support dtype as a cpp template parameter (#16452 )
陈一涵
2026-01-08 13:54:33 +08:00
41b434a7e6
[diffusion] endpoint: add API endpoint to query loaded LoRA adapters information (#16533 )
Fan Lin
2026-01-08 13:51:26 +08:00
4935344fcd
[AMD] Fix aiter page-size handling, DeepSeek MLA tuple inputs, and HiCache/FA3 decode-backend override (#16531 )
Hubert Lu
2026-01-07 21:14:32 -08:00
63cc97f4ef
ci: migrate 2-GPU tests to test/registered/ (#16529 )
Alison Shao
2026-01-07 20:28:16 -08:00
ab7d5829cd
[AMD] Add pip install / wheel build support for ROCm sgl-kernel (#15627 )
Alan Kao
2026-01-08 12:18:29 +08:00
261860e17b
[NPU][Bugfix] move free_page logics to cpu (#16608 )
hw-csong
2026-01-08 12:00:18 +08:00
154740bd4d
Disable PCG TP Unittest (#16693 )
Yuwei An
2026-01-07 19:59:47 -08:00
1c09cbe3ed
[Build] Enable full kernel in aarch64 wheel (#16155 )
MarcoDWei
2026-01-08 11:40:03 +08:00
e14f5ec8a8
[diffusion] refactor: eliminate redundant parameters in req (#16505 )
Yuhao Yang and Mick
2026-01-08 11:14:03 +08:00
8867d24879
Tiny adjust cancel PR workflow. (#16697 )
Liangsheng Yin
2026-01-08 11:13:46 +08:00
6b3f93c4dd
vlm: support SGLANG_MM_SKIP_COMPUTE_HASH for bypassing multimodal feature hashing (#16555 )
siyu
2026-01-08 11:10:00 +08:00
5a5cece561
[Diffusion] clean useless and buggy set_seq_parallel_pg in yunchang (#16669 )
Xiaoyu Zhang
2026-01-08 11:08:34 +08:00
d566739b65
Add sufeng-buaa into CI_PERMISSION (#16625 )
Teng Ma
2026-01-08 11:01:32 +08:00
4c46ecde80
[smg][ci] delete old responses api ci (#16695 )
Simo Lin
2026-01-07 18:24:12 -08:00
12a0292bfd
Revert "[sgl-kernel] Update flashmla to include fp8 sparse_mla optimizations" (#16678 )
hlu1
2026-01-07 18:23:06 -08:00
a08dc5aa10
[smg][ci] rename 3rd models from cloud backend and delete dead code (#16692 )
Simo Lin
2026-01-07 18:19:44 -08:00
eec7dbd31e
remove redundant max_running_reqs calculation in r3 (#16629 )
Junrong Lin
2026-01-08 09:45:21 +08:00
109fe03ad1
[smg][ci] Migrate Response API e2e tests to shared infrastructure (#16680 )
2026-01-07 17:40:03 -08:00
bb798a1c26
[diffusion] fix: reduce default text length for Qwen-Image from 1024 to 512 (#16445 )
Changyi Yang
2026-01-07 17:32:21 -08:00
38dc5839dd
[1/n]deepseek_v2.py Refactor: attention backend handlers and forward method definition (#16306 )
Baizhou Zhang
2026-01-08 09:22:31 +08:00
5e867f60cf
[NPU] Update model and features supported (#16652 )
Hexq0210
2026-01-08 09:13:30 +08:00
65bed8382b
Add google-cloud-storage into Dockerfile (#15343 )
gongwei-130
2026-01-07 16:54:45 -08:00
156d97b219
Fix KeyError when logprobs=false in completions endpoint (#16095 )
Harish
2026-01-07 15:49:02 -08:00
24b30f7757
MoE Refactor: Refactor fp8.py -> flashinfer_trllm.py (#15151 )
b8zhong and Brayden Zhong
2026-01-07 15:35:00 -08:00
3a4767daa3
Fix pytest tests to exit with proper exit code (#16681 )
Alison Shao
2026-01-07 15:20:46 -08:00
6037267f5b
[smg][ci] Add thread safety to ModelPool and GPUAllocator (#16674 )
Simo Lin
2026-01-07 13:25:41 -08:00
0241e0460f
Migrate tokenizer tests to test/registered/tokenizer/ (#16457 )
Alison Shao
2026-01-07 13:21:16 -08:00
0c474273c5
Fix gpt_oss_common import path and migrate core tests (#16426 )
Alison Shao
2026-01-07 12:58:32 -08:00
3e73e12458
Revert "Add SwapAB Optimization for triton fused_moe_kernel on SM90." (#16676 )
Michael
2026-01-07 11:24:47 -08:00
f4742558ac
fix: 8-gpu-b200 increase timeout length (#16658 )
Douglas Yang
2026-01-07 09:45:55 -08:00
7385834c8d
Add reference counting to ModelInstance for parallel test safety (#16672 )
Simo Lin
2026-01-07 08:28:15 -08:00
b5a94f8a8e
[model-gateway] Fix IGW routing for external OpenAI workers (#16633 )
Ziwen Zhao
2026-01-07 08:15:23 -08:00
c356ed03dd
refactor(e2e): unify RouterInstance into Gateway class, split conftest.py into modular fixtures (#16671 )
Simo Lin
2026-01-07 07:50:28 -08:00
ee4d2287ab
Add SwapAB Optimization for triton fused_moe_kernel on SM90. (#15712 )
Insideyyy
2026-01-07 23:45:35 +08:00
153c69f63d
[CI] Enable dpsk v31 test on nightly H200 (#16660 )
Baizhou Zhang
2026-01-07 23:21:19 +08:00
55b7936582
refactor(e2e_test): fix smg ci e2e test code quality (#16664 )
Simo Lin
2026-01-07 06:59:49 -08:00
7fc12e0bfa
support page size large than 64 for mamba radix cache (#16657 )
Yi Zhang and Hanming Lu
2026-01-07 22:52:24 +08:00
8729ad5e6c
fix(e2e_test): remove dead code and fix type annotations (#16661 )
Simo Lin
2026-01-07 06:35:15 -08:00
e432057381
[smg][ci] preserve model launch order with test collected (#16618 )
Simo Lin
2026-01-07 06:16:59 -08:00
4d902c8211
[diffusion] bench: upgrade multimodal benchmarks for diverse applications and create a prettier, more intuitive logger. (#16179 )
Li Jinliang
2026-01-07 22:10:23 +08:00
2ff872311b
ci: adding llama4 placeholder test to nightly (#16599 )
Douglas Yang
2026-01-07 05:51:30 -08:00
fd16c91cb8
Handle Marlin weight restoration and shape recording (QAT INT4 Rollout Part1) (#15238 )
2026-01-07 21:11:24 +08:00
32a6540afc
[Diffusion] Fix Ulysses/Ring process group construction under TP to enable correct Wan2.2 tensor parallelism (#16532 )
Xiaoyu Zhang and Mick
2026-01-07 20:52:36 +08:00
62d0280f62
Tiny fix readme (#16654 )
Xiaoyu Zhang
2026-01-07 20:41:07 +08:00
d4b717c01e
[NPU] update docs (#16651 )
Even Zhou
2026-01-07 20:20:01 +08:00
98a107d491
Re-enable temp_prefill_info assertion after pairing fix (#16203 )
Hudson Xing
2026-01-07 18:05:17 +08:00
b86bbf841e
[AMD] Add 8-GPU MX35X test running DSR1-MXFP4 model for AMD CI (#13602 )
Hubert Lu
2026-01-07 01:43:11 -08:00
48381c3b6d
[AMD] suppress warning for amd (#16620 )
YC Tseng
2026-01-07 17:37:40 +08:00
8bce085321
[AMD] CI - add 2 pp test cases to performance-test-2-gpu-amd (#16514 )
YC Tseng
2026-01-07 15:52:46 +08:00
6b8a9d7058
fix: update AMD CI estimated time for test_torch_compile (#16631 )
Alison Shao
2026-01-06 23:42:18 -08:00
973116e6bb
[Doc] Optimize pipeline parallelism doc (#16630 )
Shangming Cai
2026-01-07 14:52:42 +08:00
820e97d6c9
Upgrade aiter version (#16619 )
Thomas Wang
2026-01-07 14:20:42 +08:00
4c85f9d039
Only allocate encoder metadata for encoder-decoder models (#16527 )
Minglei Zhu
2026-01-06 22:17:04 -08:00
7d757d6f17
Clean Some Environment Variables for DeepSeek V32 (#15938 )
Baizhou Zhang
2026-01-07 14:00:16 +08:00
5c04088b3a
fix: remove performance testing from nightly dpsk v32 cp single node (#16582 )
Douglas Yang
2026-01-06 21:38:31 -08:00
f066036c8b
[diffusion] fix: fix ZImage SP sharding for 5D latents and unpad frames (#16418 )
Chi McIsaac
2026-01-07 00:24:04 -05:00
52de807dd7
Migrate profiling tests to test/registered/profiling/ (#16459 )
Alison Shao
2026-01-06 21:04:48 -08:00
70933f34f1
Migrate attention unit tests to test/registered/attention/ (#16465 )
Alison Shao
2026-01-06 20:50:54 -08:00
ce453fa43b
Migrate backends tests to CI registration system (#16468 )
Alison Shao
2026-01-06 20:31:52 -08:00
3be1e734ee
[model-gateway] extract header extraction in policy and add (#16566 )
fzyzcjy
2026-01-07 12:18:44 +08:00
38895a0064
ci: adjust partition counts for stage-b and unit tests (#16617 )
Alison Shao
2026-01-06 20:18:10 -08:00
d8b8198192
[smg][ci]: migrate benchmarks to e2e_test/benchmarks/, use parent conftest (#16597 )
Simo Lin
2026-01-06 20:15:20 -08:00
913b688f21
fix: fill a meaningful tool_index (#16504 )
Yingchun Lai
2026-01-07 12:00:09 +08:00
951d16c890
fix: adjusting vlm accuracy thresholds (#16593 )
Douglas Yang
2026-01-06 19:41:05 -08:00
53846746bf
[VLM] Fix CUDA IPC OOM (#16118 )
Yuan Luo and luoyuan.luo
2026-01-07 11:30:35 +08:00
534ac384db
[HiCacheStorage & PD] fix prefill bootstrap request host memory leaks (#15439 )
MOHENOO and yangjia1
2026-01-07 11:03:57 +08:00
2a8d5493f3
PCG Unit Test Adjustment (#16609 )
Yuwei An
2026-01-06 18:57:15 -08:00
90eac38a12
Migrate FP8/TorchAO tests to test/registered/quant/ (#16453 )
Alison Shao
2026-01-06 18:27:43 -08:00
badcd02896
[diffusion] chore: automatically enable dit_layerwise_offload for Wan (#16499 )
Mick
2026-01-07 10:22:08 +08:00
d874c8bba4
Tiny support http headers in bench serving (#16606 )
fzyzcjy
2026-01-07 10:15:17 +08:00
9a21d89c5b
Tiny add metrics for prefill delayer (#16603 )
fzyzcjy
2026-01-07 09:53:52 +08:00
4c9ac8566c
[NPU] fix command in npu best practice (#16576 )
Hexq0210
2026-01-07 09:37:27 +08:00
fb5b71d015
[router][openai] Rename prepare_mcp_payload_for_streaming and patch_streaming_response_json (#16596 )
Chang Su
2026-01-06 16:33:29 -08:00
05b54b6d7b
[router][grpc] Replace Vec<(String, String, String)> with ExtractedToolCall (#16598 )
Chang Su
2026-01-06 16:32:59 -08:00
4f443f445a
[model-gateway][cleanup] Fix wrong comment in manager.rs (#16601 )
Chang Su
2026-01-06 16:32:37 -08:00
dce8b0606c
refactor(e2e): keep only benchmark tests in e2e_http, remove redundant tests (#16594 )
Simo Lin
2026-01-06 15:05:08 -08:00
399d5283f8
ci: migrate scheduler tests to test/registered/scheduler/ (#16442 )
Alison Shao
2026-01-06 14:55:52 -08:00
2e0527dd74
ci: migrate Debug Utils, Ops, and Rotary Embedding tests to test/registered/ (#16422 )
Alison Shao
2026-01-06 14:35:57 -08:00
6beb50d612
feat: add .dockerignore to ignore files when build images (#16223 )
Yingchun Lai
2026-01-07 06:34:16 +08:00
18e2ef09d7
Add v1/models endpoint to diffusion model APIs so that they can be discovered by model gateway (#16425 )
Kangyan-Zhou
2026-01-06 14:28:24 -08:00
3271e0e76d
Remove dllm-test-1-gpu-amd job (followup to DLLM migration) (#16589 )
Alison Shao
2026-01-06 14:27:37 -08:00
d415d22daa
refactor(e2e): remove old embedding tests migrated to e2e_test/embeddings (#16592 )
Simo Lin
2026-01-06 14:07:45 -08:00
a49b9a6420
[model-gateway] add embedding tests (#16583 )
Simo Lin
2026-01-06 13:24:37 -08:00
5349764298
fix: kimi k2 thinking accuracy threshold change (#16585 )
Douglas Yang
2026-01-06 13:18:43 -08:00
d57d8e7e5f
[smg] clean up logs in mcp (should be info instead warn) (#16591 )
Simo Lin
2026-01-06 12:44:32 -08:00
0cbd8f3247
ci(test): migrate OpenAI server tests to registered CI system (#16326 )
Alison Shao and Kangyan-Zhou
2026-01-06 11:09:37 -08:00