forked from sgl-project/sglang
-
Notifications
You must be signed in to change notification settings - Fork 21
Pull requests: kvcache-ai/sglang
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix: five small guards (NextN post-load, MLA draft warmup hook, zero-expert Marlin, deep_gemm)
deepseek
#71
opened Jul 23, 2026 by
architehc
Loading…
fix(nsa): tilelang sparse kernel hardcoded head offset + sm_120 smem-fit geometry
#70
opened Jul 23, 2026 by
architehc
Loading…
feat(nsa): sm_120 support — MQA logits torch/Triton replacements + NSA backend fallbacks
#69
opened Jul 23, 2026 by
architehc
Loading…
fix(moe): sm_120 topk_sigmoid writes only 1408 rows of topk_weights — pure-torch routing fallback
#68
opened Jul 23, 2026 by
architehc
Loading…
fix: bump flashinfer to 0.6.9 to match the V4-Flash MXFP4 minimum
dependencies
#65
opened Jul 12, 2026 by
gvr13n
Loading…
feat(kt): add --kt-skip-gpu-expert-cpu-copy to reclaim host RAM for GPU experts
#64
opened Jul 12, 2026 by
gvr13n
Loading…
fix(swa): port tombstone-aware eviction + dec_swa_lock_only from upsream #23882
#49
opened May 11, 2026 by
yyj6666667
Loading…
fix(v4-flash): make ExpertDistributionRecorder cuda-graph-safe + per-layer for DeepSeek-V4-Flash
deepseek
#46
opened May 8, 2026 by
yiqiliu2
Loading…
3 tasks done
fix(v4-flash): make ExpertDistributionRecorder work for DeepSeek-V4-Flash
deepseek
#45
opened May 8, 2026 by
yiqiliu2
Loading…
3 tasks done
Making sgl-kernel work by chaging torch back to 2.9.1
deepseek
dependencies
#33
opened Apr 28, 2026 by
login256
Loading…
[feat]: add --kt-numa-nodes for explicit NUMA node mapping
#28
opened Mar 18, 2026 by
ErvinXie
Loading…
4 tasks done
[fix](kt): Fix barrier call to use cpu_group instead of device_group
#27
opened Mar 12, 2026 by
SCDESPERTATE
Loading…
Automatically detects RDMA devices, eliminating complex manual setup for mooncake
#5
opened Apr 11, 2025 by
whybeyoung
Collaborator
Loading…
ProTip!
Filter pull requests by the default branch with base:main.